mcp-gatehouse
You can manage orders through a permission-gated MCP server with audit logging.
Look up an order's status (
lookup_order, READ tier).Attach a note to an order (
add_note, WRITE tier); theapi_keyargument is redacted from logs and approver prompts.Cancel an order (
cancel_order, DESTRUCTIVE tier); requires approval and fails closed if no approver is configured.All allowed, denied, and failed calls are written to an append-only JSONL audit log.
Tools emit honest spec annotations (
readOnlyHint/destructiveHint) and can be blocked outright via denylist.
mcp-gatehouse
Permission tiers, approval gates, and audit logging for MCP servers. The server is the gatekeeper: you decide what an AI can read, what it can write, and what's off-limits — and every action gets logged.
Most MCP servers hand the model every tool at full strength and keep no
record of what it did. That's fine for a demo. It's not fine the day an
agent has write access to your CRM, your books, or your order system.
mcp-gatehouse is the missing gate, enforced inside the server — no
proxy, no external policy service, no dependencies beyond the official
mcp SDK.
pip install mcp-gatehouseWhat you get
Permission tiers | Every tool is declared |
Approval gates | Tiers you choose require a sign-off before the tool runs. Your approver is any callable — a terminal prompt, a Slack ping, a ticket. Fails closed: a gated tool with no approver configured is denied, not waved through. |
Audit log | Append-only JSONL, one line per call — allowed, denied, or failed — with UTC timestamps and durations. The answer to "what did the AI actually do?" six months later. |
Redaction | Argument keys you name ( |
Denylist | Block a tool outright, whatever its tier. |
Related MCP server: TheeDiscordMCP
Quickstart
from mcp.server.mcpserver import MCPServer
from mcp_gatehouse import AccessTier, AuditLog, Gatehouse, Policy, terminal_approver
mcp = MCPServer("order-desk")
gatehouse = Gatehouse(
mcp,
policy=Policy(approver=terminal_approver),
audit=AuditLog(path="~/order-desk/audit.jsonl"),
)
@gatehouse.tool(tier=AccessTier.READ)
def lookup_order(order_id: str) -> str:
"""Look up an order's status."""
...
@gatehouse.tool(tier=AccessTier.DESTRUCTIVE)
def cancel_order(order_id: str) -> str:
"""Cancel an order. Runs only if the approver says yes."""
...
mcp.run()That's the whole integration: build your MCPServer exactly as the
SDK docs show, but register tools through the gatehouse. Schema generation
(including Annotated[..., Field(description=...)] parameter docs),
transports, and everything else work unchanged — the guard preserves the
function's signature. Sync tools still run in a worker thread, exactly as
they would on a bare MCPServer.
terminal_approver asks on the controlling terminal (/dev/tty), never
stdin — over stdio, stdin is the protocol pipe, so a plain input()
approver would corrupt it. An approver is any callable, sync or async,
that takes an ApprovalRequest and returns True to allow; sync ones run
in a worker thread, so a human taking their time doesn't stall other calls.
If the approver raises (Slack down, ticket API timing out), the call is
denied and the failure is logged.
Under the default policy, DESTRUCTIVE requires approval and everything
is audited. Gate writes too with one line:
Policy(require_approval=frozenset({AccessTier.WRITE, AccessTier.DESTRUCTIVE}), ...)Add your own secret-bearing keys without losing the defaults:
from mcp_gatehouse import DEFAULT_REDACT
Policy(redact=DEFAULT_REDACT | {"card_pin"}, ...)What the audit trail looks like:
{"ts": "2026-07-16T14:02:11+00:00", "tool": "lookup_order", "tier": "read", "outcome": "ok", "arguments": {"order_id": "4417"}, "duration_ms": 0.42}
{"ts": "2026-07-16T14:02:38+00:00", "tool": "add_note", "tier": "write", "outcome": "ok", "arguments": {"order_id": "4417", "note": "call back", "api_key": "«redacted»"}, "duration_ms": 1.08}
{"ts": "2026-07-16T14:03:05+00:00", "tool": "cancel_order", "tier": "destructive", "outcome": "denied", "reason": "approver refused", "arguments": {"order_id": "4417"}}Try the demo
The package ships a runnable order-desk server with all three tiers wired
up and a terminal-prompt approver. Add it to any MCP client that speaks
stdio — for Claude Desktop, in claude_desktop_config.json:
{
"mcpServers": {
"gatehouse-demo": {
"command": "uvx",
"args": ["mcp-gatehouse", "--audit-log", "/tmp/gatehouse-audit.jsonl"]
}
}
}Or run mcp-gatehouse-demo yourself after pip install mcp-gatehouse.
Ask the model to list the orders and cancel one, then watch the verdict
land in the audit log either way. The audit log goes to --audit-log, else
$MCP_GATEHOUSE_AUDIT_LOG, else ./audit.jsonl — falling back to
~/.mcp-gatehouse/audit.jsonl when the working directory isn't writable
(desktop clients often launch servers from /). The server prints the
path it chose to stderr.
The approval prompt needs a terminal: a client that launches the server in
the background has none, so cancel_order is denied — the gate failing
closed, as designed. Run the server from a terminal to approve
interactively. examples/orders_server.py is the same server as a copyable
template.
Design notes
Enforcement lives inside the server, at the tool boundary. A proxy can't see your tools' semantics, and a policy service is one more thing to deploy. A 40-person plant doesn't have a platform team; this is a few small classes and a JSONL file.
Fail closed. Security defaults that quietly allow are worse than none. That includes redaction: argument values the scrubber can't take apart (arbitrary objects, bytes) are replaced with an opaque placeholder rather than passed through, and exception messages stay out of the log — only the exception type is recorded, because error text loves to embed the very values you just redacted.
The audit log records denials and errors, not just successes — the calls that didn't happen are half the story.
A blocking terminal approver and the stdio transport don't mix — stdout/stdin are the protocol pipe.
terminal_approverprompts on/dev/ttyfor exactly that reason (and denies when no terminal exists). Real deployments should approve out-of-band: Slack, a ticket, a queue.What this is not: authentication, transport encryption, or a sandbox. It's a gate inside your server, not a perimeter around it. See SECURITY.md.
Compatibility
Targets the official mcp Python SDK
v2.x (mcp>=2,<3) and Python 3.10–3.14. Ships type information
(py.typed). Release notes: CHANGELOG.md.
| SDK | Server class |
|
|
|
|
|
|
|
|
|
The public API (Gatehouse, Policy, AuditLog, AccessTier) is
unchanged across that line, as promised. Porting a v1 server is two
import edits — FastMCP became MCPServer and moved to
mcp.server.mcpserver; see the SDK's
migration guide.
Staying on SDK v1 needs no action: 0.1.x pins mcp<2, so pip keeps
resolving it. That line is closed to features but still gets security
fixes.
Who built this
Nick George — I design and run MCP servers in production for a mid-market reverse logistics-tech company, and build them for businesses at nickgeorgeai.com. This library is the permission-and-audit discipline from those builds, extracted.
If you're an owner or operator wondering what MCP even is, start with the plain-English guide: What is an MCP server?
License
Available Tools
4 toolsadd_noteA
Append a note to an order. Existing notes are kept — this never overwrites or cancels anything. Runs without approval, and every call is recorded in the audit log. To actually stop an order, use cancel_order.
| Name | Required | Description | Default |
|---|---|---|---|
| note | Yes | Free-text note to append, e.g. 'customer asked for a callback'. | |
| api_key | No | Optional credential for the order system. Its value is masked in the audit log. | |
| order_id | Yes | The order's ID, e.g. '4417'. Use list_orders to find one. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, destructiveHint=false), it discloses non-obvious traits: existing notes are preserved and nothing is overwritten or cancelled, no approval workflow is triggered, and every call is audit-logged. These are exactly the behaviors an agent cannot infer from the annotation flags alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the safety guarantee, then the alternative. Every sentence carries distinct information with no repetition of the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and annotations plus the schema cover the structured contract. The description supplies the remaining behavioral context (non-destructive append, no approval, audited) so nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — note, order_id and api_key are all documented in the schema, including the masking behavior of api_key and a pointer to list_orders for order_id. The description adds no parameter detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Append a note to an order') and immediately scopes the operation against the sibling 'cancel_order'. An agent can distinguish it from cancel_order/lookup_order without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly routes the agent away from this tool when the goal is to stop an order ('To actually stop an order, use cancel_order'), and notes the operation needs no approval. It does not state when this tool is the right choice versus merely adding a note through another path, but the boundary against the closest sibling is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_orderADestructive
Cancel an order. Irreversible, so it is gated: a human must approve on the server's terminal before it runs. If they refuse, or there is no terminal, the call is denied and the refusal is logged. To record information without cancelling, use add_note.
| Name | Required | Description | Default |
|---|---|---|---|
| order_id | Yes | The order's ID, e.g. '4417'. Use list_orders to find one. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description goes well beyond them: it discloses irreversibility, a server-side human approval gate, denial behavior on refusal, and that refusals are logged. This is exactly the operational context an agent needs before invoking a destructive, gated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler; the core action is front-loaded, followed by the risk/gating context, then the routing alternative. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the description covers the remaining gaps an agent would care about: irreversibility, the approval gate, denial semantics, and the sibling alternative.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single order_id parameter is already documented with an example ('4417') and a pointer to list_orders, so the description adds nothing further. Baseline 3 for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Cancel an order') and immediately differentiates itself from the sibling add_note, which is the nearest-looking alternative. An agent can distinguish it from list_orders/lookup_order/add_note without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the when-not case ('To record information without cancelling, use add_note') and the condition under which the call fails (refusal or no terminal). Alternatives and failure conditions are named rather than left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_ordersARead-only
List every order on the desk as id · customer · status, one per
line. Read-only. Start here to find the order ID that lookup_order,
add_note and cancel_order need.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the 'Read-only' line merely confirms that safety profile. However, the description adds real value by disclosing the exact rendered output format, which is behavioral information beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero waste. The purpose and output format are front-loaded, followed by the safety hint and the routing guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be explained, yet the description still notes the rendering format. For a zero-param read tool, nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There are no parameter semantics to clarify, and the description correctly says nothing misleading about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (orders), plus the exact output shape `id · customer · status`, one per line. Clearly distinguishable from the sibling operations that consume the IDs it produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Start here to find the order ID that lookup_order, add_note and cancel_order need' explicitly positions this as the entry point and names the alternatives that depend on its output. No inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_orderARead-only
Look up one order's customer, status and notes. Read-only; changes
nothing. Returns no order <id> if the ID doesn't exist.
| Name | Required | Description | Default |
|---|---|---|---|
| order_id | Yes | The order's ID, e.g. '4417'. Use list_orders to find one. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so 'Read-only; changes nothing' is largely redundant. What genuinely adds value is the disclosed failure behavior: it returns `no order <id>` for a missing ID, which tells the agent how to interpret an empty result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler, with the primary purpose front-loaded and the error behavior follow-up. Every clause carries information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-field explanation is unnecessary; the description still covers the one return nuance an agent must know (the 'no order <id>' sentinel). Nothing needed to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single order_id parameter is documented with an example ('4417') and a pointer to list_orders. The description adds no parameter semantics beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up) and resource (one order) and enumerates the returned fields (customer, status, notes). The word 'one' cleanly separates it from list_orders, though the description never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the read-only framing and the singular 'one order', and the schema description points to list_orders for ID discovery. But the description itself gives no explicit when-to-use vs alternatives and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.3.0- Changed
add_note3 fields changed- added
Input schema / properties / api_key / descriptionAdded value: +"Optional credential for the order system. Its value is masked in the audit log." - added
Input schema / properties / note / descriptionAdded value: +"Free-text note to append, e.g. 'customer asked for a callback'." - added
Input schema / properties / order_id / descriptionAdded value: +"The order's ID, e.g. '4417'. Use list_orders to find one."
- Changed
cancel_order1 field changed- added
Input schema / properties / order_id / descriptionAdded value: +"The order's ID, e.g. '4417'. Use list_orders to find one."
- Added
list_orders - Changed
lookup_order1 field changed- added
Input schema / properties / order_id / descriptionAdded value: +"The order's ID, e.g. '4417'. Use list_orders to find one."
3 tool updates
v1.0.0- First observed
add_note - First observed
cancel_order - First observed
lookup_order
TDQS
Scored across 4 tools
Each tool targets a distinct action and the descriptions explicitly cross-reference each other (add_note vs cancel_order, list_orders vs lookup_order). There is no realistic way to misselect among these four.
All four names follow a clean verb_noun snake_case pattern: add_note, cancel_order, list_orders, lookup_order. The verbs (add/cancel/list/lookup) are distinct and predictable.
Four tools is a bit lean but well-matched to a gatehouse control surface (observe, annotate, cancel). Each tool earns its place with no redundant or filler operations.
The surface covers the read (list_orders, lookup_order) and controlled-write (add_note, cancel_order) lifecycle cleanly. Order creation and non-cancellation status updates are absent, but these are plausibly out of scope for a gatehouse desk.
Maintenance
Related MCP Connectors
Create, deploy, and operate MCP servers directly from your GitHub repositories.
A MCP server built for developers enabling Git based project management with project and personal…
MCP server for mandates, delegation, policy-gated execution, credential grants, and audit.
Independent A-F trust grade for any MCP server, watched for drift. Free, never for sale.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceA least-privilege enforcement proxy for MCP servers. It sits between MCP clients and upstream servers, enforcing tool policies, hiding denied tools, requiring human approval for risky actions, and providing a structured audit trail.MIT
- AlicenseBqualityBmaintenanceLocal MCP server for inspecting and managing an allowlisted Discord server via Discord's REST API, with safety modes, idempotent JSON blueprints, and destructive-operation safeguards.271MIT
- AlicenseAqualityBmaintenanceA security-first MCP server for managing Whatbox slots with structured read-only inspection and approval-gated mutations, including storage, services, website deployment, and torrent control.30MIT
- AlicenseNot gradedqualityBmaintenanceEnforces identity-based access control and audit logging for MCP servers, letting you grant fine-grained tool permissions to users and systems while failing secure by default.MIT