Skip to main content
Glama
nickgeorgeseo

mcp-gatehouse

mcp-gatehouse

CI PyPI Python License: MIT mcp-gatehouse MCP server

Permission tiers, approval gates, and audit logging for MCP servers. The server is the gatekeeper: you decide what an AI can read, what it can write, and what's off-limits — and every action gets logged.

Most MCP servers hand the model every tool at full strength and keep no record of what it did. That's fine for a demo. It's not fine the day an agent has write access to your CRM, your books, or your order system. mcp-gatehouse is the missing gate, enforced inside the server — no proxy, no external policy service, no dependencies beyond the official mcp SDK.

pip install mcp-gatehouse

What you get

Permission tiers

Every tool is declared READ, WRITE, or DESTRUCTIVE — and the tier also emits honest spec ToolAnnotations (readOnlyHint / destructiveHint), which the wrapper won't let you override to lie.

Approval gates

Tiers you choose require a sign-off before the tool runs. Your approver is any callable — a terminal prompt, a Slack ping, a ticket. Fails closed: a gated tool with no approver configured is denied, not waved through.

Audit log

Append-only JSONL, one line per call — allowed, denied, or failed — with UTC timestamps and durations. The answer to "what did the AI actually do?" six months later.

Redaction

Argument keys you name (api_key, password, token, client_secret, … by default) are masked before they reach the log or the approver — however they're spelled: apiKey, API-Key and api_key are the same key.

Denylist

Block a tool outright, whatever its tier.

Related MCP server: TheeDiscordMCP

Quickstart

from mcp.server.mcpserver import MCPServer
from mcp_gatehouse import AccessTier, AuditLog, Gatehouse, Policy, terminal_approver

mcp = MCPServer("order-desk")
gatehouse = Gatehouse(
    mcp,
    policy=Policy(approver=terminal_approver),
    audit=AuditLog(path="~/order-desk/audit.jsonl"),
)

@gatehouse.tool(tier=AccessTier.READ)
def lookup_order(order_id: str) -> str:
    """Look up an order's status."""
    ...

@gatehouse.tool(tier=AccessTier.DESTRUCTIVE)
def cancel_order(order_id: str) -> str:
    """Cancel an order. Runs only if the approver says yes."""
    ...

mcp.run()

That's the whole integration: build your MCPServer exactly as the SDK docs show, but register tools through the gatehouse. Schema generation (including Annotated[..., Field(description=...)] parameter docs), transports, and everything else work unchanged — the guard preserves the function's signature. Sync tools still run in a worker thread, exactly as they would on a bare MCPServer.

terminal_approver asks on the controlling terminal (/dev/tty), never stdin — over stdio, stdin is the protocol pipe, so a plain input() approver would corrupt it. An approver is any callable, sync or async, that takes an ApprovalRequest and returns True to allow; sync ones run in a worker thread, so a human taking their time doesn't stall other calls. If the approver raises (Slack down, ticket API timing out), the call is denied and the failure is logged.

Under the default policy, DESTRUCTIVE requires approval and everything is audited. Gate writes too with one line:

Policy(require_approval=frozenset({AccessTier.WRITE, AccessTier.DESTRUCTIVE}), ...)

Add your own secret-bearing keys without losing the defaults:

from mcp_gatehouse import DEFAULT_REDACT
Policy(redact=DEFAULT_REDACT | {"card_pin"}, ...)

What the audit trail looks like:

{"ts": "2026-07-16T14:02:11+00:00", "tool": "lookup_order", "tier": "read", "outcome": "ok", "arguments": {"order_id": "4417"}, "duration_ms": 0.42}
{"ts": "2026-07-16T14:02:38+00:00", "tool": "add_note", "tier": "write", "outcome": "ok", "arguments": {"order_id": "4417", "note": "call back", "api_key": "«redacted»"}, "duration_ms": 1.08}
{"ts": "2026-07-16T14:03:05+00:00", "tool": "cancel_order", "tier": "destructive", "outcome": "denied", "reason": "approver refused", "arguments": {"order_id": "4417"}}

Try the demo

The package ships a runnable order-desk server with all three tiers wired up and a terminal-prompt approver. Add it to any MCP client that speaks stdio — for Claude Desktop, in claude_desktop_config.json:

{
  "mcpServers": {
    "gatehouse-demo": {
      "command": "uvx",
      "args": ["mcp-gatehouse", "--audit-log", "/tmp/gatehouse-audit.jsonl"]
    }
  }
}

Or run mcp-gatehouse-demo yourself after pip install mcp-gatehouse. Ask the model to list the orders and cancel one, then watch the verdict land in the audit log either way. The audit log goes to --audit-log, else $MCP_GATEHOUSE_AUDIT_LOG, else ./audit.jsonl — falling back to ~/.mcp-gatehouse/audit.jsonl when the working directory isn't writable (desktop clients often launch servers from /). The server prints the path it chose to stderr.

The approval prompt needs a terminal: a client that launches the server in the background has none, so cancel_order is denied — the gate failing closed, as designed. Run the server from a terminal to approve interactively. examples/orders_server.py is the same server as a copyable template.

Design notes

  • Enforcement lives inside the server, at the tool boundary. A proxy can't see your tools' semantics, and a policy service is one more thing to deploy. A 40-person plant doesn't have a platform team; this is a few small classes and a JSONL file.

  • Fail closed. Security defaults that quietly allow are worse than none. That includes redaction: argument values the scrubber can't take apart (arbitrary objects, bytes) are replaced with an opaque placeholder rather than passed through, and exception messages stay out of the log — only the exception type is recorded, because error text loves to embed the very values you just redacted.

  • The audit log records denials and errors, not just successes — the calls that didn't happen are half the story.

  • A blocking terminal approver and the stdio transport don't mix — stdout/stdin are the protocol pipe. terminal_approver prompts on /dev/tty for exactly that reason (and denies when no terminal exists). Real deployments should approve out-of-band: Slack, a ticket, a queue.

  • What this is not: authentication, transport encryption, or a sandbox. It's a gate inside your server, not a perimeter around it. See SECURITY.md.

Compatibility

Targets the official mcp Python SDK v2.x (mcp>=2,<3) and Python 3.10–3.14. Ships type information (py.typed). Release notes: CHANGELOG.md.

mcp-gatehouse

SDK

Server class

0.3.x

mcp>=2,<3

MCPServer

0.2.x

mcp>=2,<3

MCPServer

0.1.x

mcp>=1.27,<2

FastMCP

The public API (Gatehouse, Policy, AuditLog, AccessTier) is unchanged across that line, as promised. Porting a v1 server is two import edits — FastMCP became MCPServer and moved to mcp.server.mcpserver; see the SDK's migration guide.

Staying on SDK v1 needs no action: 0.1.x pins mcp<2, so pip keeps resolving it. That line is closed to features but still gets security fixes.

Who built this

Nick George — I design and run MCP servers in production for a mid-market reverse logistics-tech company, and build them for businesses at nickgeorgeai.com. This library is the permission-and-audit discipline from those builds, extracted.

If you're an owner or operator wondering what MCP even is, start with the plain-English guide: What is an MCP server?

License

MIT

Available Tools

4 tools
add_noteA

Append a note to an order. Existing notes are kept — this never overwrites or cancels anything. Runs without approval, and every call is recorded in the audit log. To actually stop an order, use cancel_order.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteYesFree-text note to append, e.g. 'customer asked for a callback'.
api_keyNoOptional credential for the order system. Its value is masked in the audit log.
order_idYesThe order's ID, e.g. '4417'. Use list_orders to find one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=false), it discloses non-obvious traits: existing notes are preserved and nothing is overwritten or cancelled, no approval workflow is triggered, and every call is audit-logged. These are exactly the behaviors an agent cannot infer from the annotation flags alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then the safety guarantee, then the alternative. Every sentence carries distinct information with no repetition of the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and annotations plus the schema cover the structured contract. The description supplies the remaining behavioral context (non-destructive append, no approval, audited) so nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — note, order_id and api_key are all documented in the schema, including the masking behavior of api_key and a pointer to list_orders for order_id. The description adds no parameter detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Append a note to an order') and immediately scopes the operation against the sibling 'cancel_order'. An agent can distinguish it from cancel_order/lookup_order without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the agent away from this tool when the goal is to stop an order ('To actually stop an order, use cancel_order'), and notes the operation needs no approval. It does not state when this tool is the right choice versus merely adding a note through another path, but the boundary against the closest sibling is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_orderA
Destructive

Cancel an order. Irreversible, so it is gated: a human must approve on the server's terminal before it runs. If they refuse, or there is no terminal, the call is denied and the refusal is logged. To record information without cancelling, use add_note.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesThe order's ID, e.g. '4417'. Use list_orders to find one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description goes well beyond them: it discloses irreversibility, a server-side human approval gate, denial behavior on refusal, and that refusals are logged. This is exactly the operational context an agent needs before invoking a destructive, gated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler; the core action is front-loaded, followed by the risk/gating context, then the routing alternative. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description covers the remaining gaps an agent would care about: irreversibility, the approval gate, denial semantics, and the sibling alternative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single order_id parameter is already documented with an example ('4417') and a pointer to list_orders, so the description adds nothing further. Baseline 3 for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Cancel an order') and immediately differentiates itself from the sibling add_note, which is the nearest-looking alternative. An agent can distinguish it from list_orders/lookup_order/add_note without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the when-not case ('To record information without cancelling, use add_note') and the condition under which the call fails (refusal or no terminal). Alternatives and failure conditions are named rather than left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_ordersA
Read-only

List every order on the desk as id · customer · status, one per line. Read-only. Start here to find the order ID that lookup_order, add_note and cancel_order need.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the 'Read-only' line merely confirms that safety profile. However, the description adds real value by disclosing the exact rendered output format, which is behavioral information beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste. The purpose and output format are front-loaded, followed by the safety hint and the routing guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be explained, yet the description still notes the rendering format. For a zero-param read tool, nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There are no parameter semantics to clarify, and the description correctly says nothing misleading about inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (orders), plus the exact output shape `id · customer · status`, one per line. Clearly distinguishable from the sibling operations that consume the IDs it produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Start here to find the order ID that lookup_order, add_note and cancel_order need' explicitly positions this as the entry point and names the alternatives that depend on its output. No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_orderA
Read-only

Look up one order's customer, status and notes. Read-only; changes nothing. Returns no order <id> if the ID doesn't exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesThe order's ID, e.g. '4417'. Use list_orders to find one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so 'Read-only; changes nothing' is largely redundant. What genuinely adds value is the disclosed failure behavior: it returns `no order <id>` for a missing ID, which tells the agent how to interpret an empty result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler, with the primary purpose front-loaded and the error behavior follow-up. Every clause carries information the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-field explanation is unnecessary; the description still covers the one return nuance an agent must know (the 'no order <id>' sentinel). Nothing needed to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single order_id parameter is documented with an example ('4417') and a pointer to list_orders. The description adds no parameter semantics beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (look up) and resource (one order) and enumerates the returned fields (customer, status, notes). The word 'one' cleanly separates it from list_orders, though the description never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the read-only framing and the singular 'one order', and the schema description points to list_orders for ID discovery. But the description itself gives no explicit when-to-use vs alternatives and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.3.0
    • Changedadd_note3 fields changed
      • addedInput schema / properties / api_key / description
        Added value: +"Optional credential for the order system. Its value is masked in the audit log."
      • addedInput schema / properties / note / description
        Added value: +"Free-text note to append, e.g. 'customer asked for a callback'."
      • addedInput schema / properties / order_id / description
        Added value: +"The order's ID, e.g. '4417'. Use list_orders to find one."
    • Changedcancel_order1 field changed
      • addedInput schema / properties / order_id / description
        Added value: +"The order's ID, e.g. '4417'. Use list_orders to find one."
    • Addedlist_orders
    • Changedlookup_order1 field changed
      • addedInput schema / properties / order_id / description
        Added value: +"The order's ID, e.g. '4417'. Use list_orders to find one."
  2. 3 tool updatesv1.0.0
    • First observedadd_note
    • First observedcancel_order
    • First observedlookup_order

TDQS

A4.4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a distinct action and the descriptions explicitly cross-reference each other (add_note vs cancel_order, list_orders vs lookup_order). There is no realistic way to misselect among these four.

Naming Consistency5/5

All four names follow a clean verb_noun snake_case pattern: add_note, cancel_order, list_orders, lookup_order. The verbs (add/cancel/list/lookup) are distinct and predictable.

Tool Count4/5

Four tools is a bit lean but well-matched to a gatehouse control surface (observe, annotate, cancel). Each tool earns its place with no redundant or filler operations.

Completeness4/5

The surface covers the read (list_orders, lookup_order) and controlled-write (add_note, cancel_order) lifecycle cleanly. Order creation and non-cancellation status updates are absent, but these are plausibly out of scope for a gatehouse desk.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    A least-privilege enforcement proxy for MCP servers. It sits between MCP clients and upstream servers, enforcing tool policies, hiding denied tools, requiring human approval for risky actions, and providing a structured audit trail.
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Local MCP server for inspecting and managing an allowlisted Discord server via Discord's REST API, with safety modes, idempotent JSON blueprints, and destructive-operation safeguards.
    27
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    A security-first MCP server for managing Whatbox slots with structured read-only inspection and approval-gated mutations, including storage, services, website deployment, and torrent control.
    30
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enforces identity-based access control and audit logging for MCP servers, letting you grant fine-grained tool permissions to users and systems while failing secure by default.
    MIT