mcp-policy-gateway
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-policy-gatewayGet the account summary for account 9876543210"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-policy-gateway
A policy-first MCP server reference implementation. It shows how to give an agent useful tools without allowing the model to decide authorization, cross tenant boundaries, leak direct PII, or execute high-impact mutations on its own.
This is a portfolio/reference implementation. Every record is synthetic; it is not production software and does not claim to integrate with a bank, payment processor, or customer system.
Why this exists
A common MCP demo exposes a raw API to a model and treats a successful tool call as a security model. This project takes the opposite position:
Policy is deterministic. The model cannot authorize itself.
Tenant boundary is enforced server-side. A tool argument cannot switch tenant.
Tool output is minimized. The account summary intentionally excludes owner name and email.
High-impact actions stop at an approval gate. A freeze request produces no state change.
Every decision is auditable. Allow, deny and approval-required outcomes create structured audit events.
Mutation requests require a reason and idempotency key.
Related MCP server: mcp-permission-server
Architecture
MCP client / agent
│ tool call
▼
FastMCP stdio adapter
│ verified principal (demo fixture here)
▼
GatewayService ──► PolicyEngine ──► allow / deny / require approval
│ │
▼ ▼
synthetic domain data structured audit eventIn a production deployment, the principal would be derived from a verified identity/session outside the model context. Never accept tenant_id, roles, or scopes from a tool argument or model output.
Tools
Tool | Risk | Behaviour |
| Read | Allows only cases in the caller's tenant with the right scope. |
| Read | Returns a minimized summary; never direct PII. |
| Mutation | Requires a scope, reason and idempotency key, then always requires human approval. |
Quickstart
Requires Python 3.11+.
python -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/pytest
.venv/bin/ruff check .Run as an MCP stdio server:
.venv/bin/mcp-policy-gatewayThe adapter uses a deliberately fixed demo principal for local exploration. Real identity propagation is intentionally documented as a production integration concern rather than faked here.
Demonstrated abuse controls
The tests prove the gateway:
denies an agent from
tenant_redtrying to readtenant_bluedata;denies a tool call without its required scope;
does not return owner name or email through the summary tool;
returns
require_approvalrather than mutating account state; andrejects state-changing requests that omit a reason or idempotency key.
Threat model and non-goals
See THREAT_MODEL.md. The design is intentionally narrow: it demonstrates a policy boundary around MCP tools, not full identity infrastructure, durable audit storage, secrets management, or a payment/ledger system.
Interview walkthrough
A concise way to explain the project:
I designed the MCP boundary so the model can reason over safe, scoped data but cannot decide access or execute high-impact actions. Authorization is deterministic, tenant isolation is enforced server-side, the tool surface minimizes PII, and sensitive mutations terminate in an auditable human approval gate.
Development
.venv/bin/pytest
.venv/bin/ruff check .CI runs both checks on pushes and pull requests.
License
MIT. See LICENSE.
Available Tools
3 toolsget_account_summaryB
Return a minimized account summary; this tool never exposes direct PII.
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden. It discloses a key behavioral trait: 'this tool never exposes direct PII', which is valuable privacy context. However, it does not mention other behaviors such as whether it requires specific permissions, handles missing accounts, or returns errors. The single disclosed trait is important but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the primary action and object, then adds a meaningful privacy qualifier. It is appropriately short and contains no filler, though it could include a bit more context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema, so the description need not be extensive. Still, it leaves gaps: no mention of what the summary contains, how to interpret the response, or typical usage scenarios. While not critical for a trivial read operation, the lack of any usage guidance and return format means an agent may not fully understand the tool's output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameter descriptions (coverage 0%), and the description does not mention the 'account_id' parameter at all. The parameter name is self-explanatory, which partially compensates, but the description adds no semantic detail about its format, scope, or expected values. Since schema coverage is low, the description should have compensated but did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Return') and the resource ('account summary'), and adds the qualifier 'minimized' to distinguish it from a full account detail. While it does not explicitly name sibling tools, the siblings (get_case, request_account_freeze) are sufficiently different in name and purpose that an agent can distinguish without explicit differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool vs alternatives, nor any exclusions or prerequisites. It only says it returns a summary, implying it is for summary needs, but does not mention when to prefer it over get_case or request_account_freeze, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_caseA
Retrieve a support case inside the caller's tenant boundary.
| Name | Required | Description | Default |
|---|---|---|---|
| case_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It does disclose a meaningful constraint—the tenant-boundary scoping and read-only retrieval semantics of 'Retrieve.' However, it does not disclose error behavior for missing/non-accessible cases, auth requirements, or return expectations, which are meaningful gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the verb and resource front-loaded and zero filler. The tenant-boundary qualifier earns its place as a scope constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple retrieve-by-ID tool with a single parameter and no output schema, the description covers the essentials: what it retrieves and under what scope boundary. Missing output-format and error details are minor for this complexity tier.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate. It connects case_id to 'support case' and the tenant-boundary constraint, adding some meaning beyond the bare name. Still, it provides no format, example, or uniqueness detail for case_id—though the single self-explanatory parameter name reduces the need for heavy elaboration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
'Retrieve a support case' states a specific verb and resource, and the tenant-boundary clause adds a precise scope constraint. The resource (case) is clearly distinct from the sibling tools (account summary, account freeze), though the description stops short of explicitly naming those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the verb and resource—call this when you have a case_id and need case details. The tenant-boundary scoping hints at context, but there is no explicit when-not-to-use guidance, no mention of alternatives, and no condition that routes the agent to get_account_summary or request_account_freeze.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_account_freezeB
Request a high-impact account freeze. Policy always requires human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| account_id | Yes | ||
| idempotency_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does reveal a key trait: human approval is always required, implying the request is not immediately executed. However, it does not disclose what happens after approval, whether it is asynchronous, what the response will be, or any side effects. Partial disclosure is better than nothing, but gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The first sentence states the action and impact; the second adds the critical policy constraint. Every word earns its place and the structure is front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three required parameters, no annotations, no output schema, and a high-impact action, the description is insufficiently complete. It does not explain parameter semantics, expected outcomes, or how to handle approval delays. An agent would need additional documentation or inference to use this tool correctly in a real flow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and the description provides no parameter explanations. The parameter names (account_id, reason, idempotency_key) are somewhat self-explanatory, but the description adds zero meaning beyond the raw names. For a tool with three required parameters, the description should at least hint at their purpose, especially the idempotency key's role in preventing duplicate requests.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('request') and resource ('account freeze') and adds 'high-impact' to convey significance. It inherently distinguishes itself from the read-only sibling tools (get_case, get_account_summary) by being an action that modifies state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites. It mentions human approval is required, but that is a policy constraint, not a usage guide. An agent must infer that this tool is for freezing accounts, but there is no explicit comparison to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
get_account_summary - First observed
get_case - First observed
request_account_freeze
TDQS
Scored across 3 tools
Each tool targets a distinct resource/action: case retrieval, account summary (with PII restriction), and a freeze request. There is no overlap in purpose, and the descriptions clearly differentiate the operations.
All tool names follow a consistent verb_noun pattern (get_case, get_account_summary, request_account_freeze). The verbs are clear and the nouns identify the target, making the set predictable and easy to navigate.
With only 3 tools, the server is tightly scoped and each tool serves a distinct, essential function within the policy gateway context. This is within the typical 3-15 range and feels appropriately lean rather than sparse.
The tools cover core read operations and a high-impact action, but there are notable gaps: no listing of cases or accounts, no update/modification paths, and no way to track the status of a freeze request. Agents would need to work around these missing operations.
Maintenance
Related MCP Connectors
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
- gatewayOAuthai.sealgate
MCP gateway with runtime security policy, tool-call-level control, and audit of agent actions.
Human-in-the-loop review and approval for AI agents. Audit trail, approval policies, native MCP.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables approval-gated incident response workflows that gather evidence through read-only MCP tools, perform idempotent writes, and preserve a durable audit trail.-
- AlicenseNot gradedqualityBmaintenanceEnforces fine-grained, context-aware access control on MCP tool calls, with a tamper-evident, replayable audit log that records denials and verifies every decision.MIT
- FlicenseNot gradedqualityBmaintenanceEnables governing tenant-aware MCP tools with policy enforcement, scoped access, human approval workflows, and tamper-evident audit logging.-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents and MCP clients to safely access upstream MCP servers through centralized policy enforcement, including allow/deny/approval decisions, schema pinning, circuit breakers, and audit logging.MIT