MCP Airlock
MCP Airlock is a zero-trust security gateway that enforces capability-based access control for MCP tool calls, with tamper-evident audit logging.
Issue Capability Leases (
airlock_issue_capability): Request short-lived, HMAC-signed leases bound to a specific session, subject (agent), intent, and set of tools — with optional constraints (e.g., allowed domains, max risk) and a custom TTL (default 900s). These leases must be presented with subsequent tool calls to authorize execution.Audit Trail Inspection (
airlock_audit_tail): Read recent signed provenance events from a tamper-evident, hash-chained ledger to review allow/deny decisions (configurable limit, default 20 events).Fetch External JSON (
http_get_json): Make SSRF-resistant HTTP GET requests to public HTTPS endpoints with optional query parameters, subject to domain allowlisting and capability/policy checks.Hourly Weather Lookup (
weather_hourly): Resolve a US city/state to coordinates and fetch hourly forecasts from Open-Meteo (1–168 hours ahead, default 24), gated behind capability and policy enforcement.Dynamic Policy Enforcement: Every tool call is validated against risk thresholds, session/intent continuity, and tool-specific constraints — preventing prompt injection and privilege escalation.
Tamper-Evident Provenance: All authorization decisions are recorded in an append-only hash chain, ensuring auditability and integrity of the security log.
Supports OpenTelemetry traces and SIEM sinks for monitoring and observability of security decisions and tool calls.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Airlockissue a short-lived capability for weather lookup"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Airlock
Zero-trust security gateway for MCP tools.
MCP Airlock turns every tool call into a short-lived, context-bound capability decision with tamper-evident provenance.
Why this exists
Agent tool ecosystems are failing at one painful boundary: the jump from untrusted prompt text to privileged tool execution.
Current patterns are usually one of:
static allowlists (
agent can call tool X)weak regex filtering
post-hoc logs with no integrity guarantees
They fail when prompt injection mutates intent mid-session, causing silent privilege escalation or data exfiltration.
MCP Airlock solves this with a missing primitive for MCP:
Capability Leases: short-lived, signed, context-bound rights (
session + intent + tool scope + constraints)Context-Aware Policy: dynamic authorization on every call (risk score + tool constraints + lease checks)
Tamper-Evident Provenance: append-only hash chain across all allow/deny decisions
Related MCP server: AgentsGate
Core innovation
Context-Bound Capability Leases (CBCL)
Each tool call is authorized against a signed lease:
Bound to
session_idBound to
intent_hashScoped to specific tools
Time-limited
Optional constraints (e.g. allowed domains, max risk)
If a prompt injection tries to change intent or jump tool scope, execution is denied.
Architecture
flowchart LR
A[Agent / MCP Client] -->|tools/call| B[MCP Airlock Server]
B --> C[Risk Engine]
B --> D[Capability Verifier]
B --> E[Policy Engine]
E -->|allow| F[Tool Adapter Layer]
E -->|deny| G[Policy Deny Response]
F --> H[External APIs / Internal Services]
B --> I[Provenance Ledger Hash Chain]Trust boundaries
flowchart TB
subgraph Untrusted
U1[Prompt Content]
U2[Agent Reasoning Trace]
end
subgraph Trusted Control Plane
T1[MCP Airlock]
T2[Policy + Lease Validation]
T3[Signed Provenance Ledger]
end
subgraph External Targets
X1[Public APIs]
X2[Internal APIs]
end
U1 --> T1
U2 --> T1
T1 --> T2
T2 --> X1
T2 --> X2
T1 --> T3Key features
MCP stdio server compatible with
initialize,tools/list,tools/callCapability issuance tool:
airlock_issue_capabilityAgent/API usage analytics tool:
airlock_usage_statsAPI exposure measurement tool:
airlock_exposure_reportPolicy enforcement middleware with per-tool risk thresholds
Prompt-injection signature scoring
SSRF-resistant HTTP tool adapter (
http_get_json)Real API integration example (
weather_hourly)Tamper-evident provenance log + verification command
Sandbox hardening guide for agentic API security
CLI for serve/demo/issue/verify/stats/exposure
2-minute quickstart
git clone https://github.com/lara-muhanna/mcp-airlock
cd mcp-airlock
python -m pip install -e .
python -m mcp_airlock --config examples/airlock.config.json demoWhat you will see:
a human-friendly summary (handshake, capability, allow/deny, audit integrity)
malicious call denied with plain-English reasons
signed provenance evidence
One-command local demo (no install)
python -m mcp_airlock --config examples/airlock.config.json demo --city Austin --state TexasFor full JSON payloads during demo:
python -m mcp_airlock --config examples/airlock.config.json demo --rawRun as MCP server
python -m mcp_airlock --config examples/airlock.config.json serveMCP client setup examples:
CLI
# Issue a capability directly
python -m mcp_airlock --config examples/airlock.config.json issue \
--session-id sess-123 \
--subject agent:planner \
--tools weather_hourly,http_get_json \
--intent "Plan safe outdoor activities" \
--ttl-seconds 900 \
--constraints '{"allowed_domains":["api.open-meteo.com"],"max_risk":0.6}'
# Verify audit integrity
python -m mcp_airlock --config examples/airlock.config.json verify-log
# API usage stats by agent
python -m mcp_airlock --config examples/airlock.config.json stats --lookback-hours 24
# API exposure measurement
python -m mcp_airlock --config examples/airlock.config.json exposure --lookback-hours 24Example agent integration
Run:
python examples/agent_integration.pyThis script:
starts Airlock over stdio
negotiates MCP initialize/list
issues a lease
runs a normal tool call
runs an injected call that gets blocked
Config template
examples/airlock.config.json
{
"secret_key": "dev-secret-change-this-before-production",
"provenance_log": "./airlock-provenance.log",
"max_ttl_seconds": 1800,
"default_risk_threshold": 0.55,
"tools": {
"weather_hourly": {
"require_capability": true,
"risk_threshold": 0.7
},
"http_get_json": {
"require_capability": true,
"risk_threshold": 0.45,
"allowed_domains": ["api.open-meteo.com", "geocoding-api.open-meteo.com"]
}
}
}Security model summary
Agent requests lease via
airlock_issue_capability.Lease is HMAC-signed and includes
session,intent_hash,tool_scope,expiry.Every
tools/callrequest includes_capabilityand_context.Airlock enforces:
lease validity + signature
session and intent continuity
risk threshold
tool-specific constraints (e.g., domain allowlist)
Decision + evidence is hash-chained to provenance log.
Project structure
mcp-airlock/
mcp_airlock/
cli.py
server.py
policy.py
capability.py
risk.py
provenance.py
config.py
tool_ids.py
tools/
http_json.py
weather.py
examples/
airlock.config.json
agent_integration.py
docs/
CLIENT_SETUP.md
SANDBOXING_AGENTIC_APIS.mdRoadmap
Upstream MCP proxy mode (wrap existing MCP servers transparently)
OPA/Rego policy backend
OpenTelemetry traces + SIEM sinks
Managed capability broker + key rotation
Signed replay package for incident response
Community
Contribution guide:
CONTRIBUTING.mdSecurity policy:
SECURITY.mdCode of conduct:
CODE_OF_CONDUCT.mdRelease checklist:
docs/RELEASE_CHECKLIST.md
License
MIT
Available Tools
4 toolsairlock_audit_tailC
Read recent signed provenance events.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. While it notes events are 'recent' and 'signed', it fails to specify the time window for 'recent', result ordering, return format, pagination behavior, or the implications of 'signed' (verification requirements?). This leaves critical behavioral gaps for an audit tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at five words, immediately front-loading the verb and object. While efficient, this brevity is arguably inappropriate given the complete absence of annotations and schema descriptions, leaving the description too terse to stand alone as documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a security/audit tool handling cryptographically signed provenance data, the description is inadequate. With no output schema, no parameter descriptions, and no annotations, the description should explain what data structure is returned and what 'airlock' provenance tracks, but it provides none of this context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage for the 'limit' parameter. The description completely fails to compensate by explaining the parameter's purpose, valid ranges, or default behavior (20). The agent has no textual guidance on how to use the only available parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Read') and the specific resource ('signed provenance events'), distinguishing it from the sibling 'airlock_issue_capability' (which issues/writes) and the unrelated 'weather_hourly' and 'http_get_json' tools. However, it could better clarify what constitutes a 'provenance event' in this specific 'airlock' domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor does it mention prerequisites (e.g., specific permissions needed to read signed audit trails) or when not to use it. The agent receives no signals about appropriate usage contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
airlock_issue_capabilityC
Issue a short-lived capability lease bound to session+intent.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | ||
| subject | Yes | ||
| tools | Yes | ||
| ttl_seconds | No | ||
| intent | Yes | ||
| constraints | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Lacking annotations, the description only discloses the 'short-lived' nature (relating to ttl_seconds) but omits critical behavioral details: what authorization the lease grants, how the constraints object limits usage, side effects of issuance, or security implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single dense sentence with no filler words; information is front-loaded. However, extreme brevity becomes a liability given the complete absence of schema documentation and annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Inadequate for a security-sensitive tool with 6 parameters (including a nested constraints object). Fails to explain the capability model, what the lease authorizes, or the purpose of required fields like 'subject', creating operational risk.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description partially compensates by implying session_id, intent, and ttl_seconds via 'bound to session+intent' and 'short-lived', but leaves 'subject', 'tools', and the nested 'constraints' object completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the core action (Issue) and resource (capability lease) with binding context (session+intent), but uses domain jargon without explanation and fails to differentiate from sibling 'airlock_audit_tail' or explain what the lease enables.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no guidance on when to issue capabilities versus using other tools, no security prerequisites, and no warnings about the sensitivity of granting tool access via leases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
http_get_jsonA
Fetch JSON from a public HTTPS API endpoint.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | HTTPS URL to fetch. | |
| query | No | Optional query parameters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It adds valuable behavioral context by specifying 'public' (implying no auth headers required) and 'JSON' (setting expectation for response parsing), but omits operational details like timeout behavior, redirect handling, or error responses for non-JSON content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence with zero waste. It front-loads the action and precisely qualifies the target resource type and protocol without filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 parameters, no output schema), the description adequately covers the essential contract. The 'public' qualifier is crucial for setting correct expectations, though mentioning error handling for non-JSON responses would strengthen completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the schema fully documenting both 'url' and 'query' parameters. The description adds no additional parameter semantics beyond what the schema provides, meeting the baseline expectation for well-documented schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description provides a specific verb ('Fetch'), resource type ('JSON'), and scope ('public HTTPS API endpoint'), clearly positioning this as a generic external HTTP client distinct from domain-specific siblings like weather_hourly and airlock_audit_tail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While it lacks explicit 'when-to-use' statements, the description implies usage through the 'public HTTPS API' scope, suggesting external/unauthenticated endpoints versus the internal/domain-specific siblings. However, it does not explicitly direct users to alternatives like weather_hourly for weather data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
weather_hourlyA
Resolve US city/state and fetch hourly weather from Open-Meteo.
| Name | Required | Description | Default |
|---|---|---|---|
| city | Yes | ||
| state | Yes | ||
| hours | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. It successfully indicates the geocoding behavior ('Resolve') and data source ('Open-Meteo'), but omits critical behavioral details like error handling for invalid locations, output format, units (imperial/metric), or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence with zero redundancy. Every word contributes essential information (action, scope, data source), making it appropriately sized and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple 3-parameter schema with primitive types and no output schema, the description adequately covers the core function. However, gaps remain: the 'hours' parameter is undocumented, and the absence of annotations or output schema leaves the return structure and units unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, requiring the description to compensate. It implicitly documents the 'city' and 'state' parameters by specifying they are 'US city/state', adding geographic context not present in the raw parameter names. However, it fails to mention the 'hours' parameter or its constraints (1-168, default 24).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('Resolve', 'fetch') and identifies the exact resource ('hourly weather'). It clearly distinguishes the tool from siblings (airlock_audit_tail, http_get_json) by specifying the weather domain and Open-Meteo data source.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies geographic constraints by specifying 'US city/state', hinting at usage boundaries. However, it lacks explicit guidance on when to use versus alternatives (e.g., for non-US locations) or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v0.1.0- First observed
airlock_audit_tail - First observed
airlock_issue_capability - First observed
http_get_json - First observed
weather_hourly
TDQS
Each tool targets a distinct function (provenance auditing, capability issuance, generic HTTP fetching, and weather retrieval) with minimal overlap. An agent can easily distinguish when to use each based on the task requirements.
Mixed naming patterns: two tools use 'airlock_' prefix with verbs (issue, audit_tail), while others use 'http_' prefix or no prefix (weather_hourly). The weather tool breaks the verb-first convention used by others, using a noun-adjective pattern instead.
Four tools is acceptably compact but the set suffers from scope confusion, mixing core Airlock security functions with unrelated utility tools (weather, HTTP). This feels like two different servers merged together rather than a cohesive toolset.
The Airlock domain (capability management) is severely underdeveloped, offering only issuance and audit tail reading without revocation, validation, or capability listing. The utility tools (weather, HTTP) are also minimal, offering only hourly forecasts and GET requests without parameter customization.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
- gatewayOAuthai.sealgate
MCP gateway with runtime security policy, tool-call-level control, and audit of agent actions.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Identity, authorization, audit trails, and revocable permissions for AI agents accessing MCP tools.
Security & DLP proxy for MCP: tool-poisoning scans, PII redaction on tool args/results. Beta.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables secure interaction between LLMs and MCP tools by applying zero-trust security controls, including sensitive data masking, file system protection, and policy enforcement.-

AgentsGateofficial
AlicenseNot gradedqualityAmaintenanceEnables AI agents to securely call MCP tools with risk scoring, checkpoints, rollback, and approval workflows.17MIT
evav-gatewayofficial
AlicenseNot gradedqualityBmaintenanceGoverned MCP gateway that lets AI agents call tools with policy enforcement, prompt-injection screening, a kill-switch, and tamper-evident signed audit logs.Apache 2.0- AlicenseNot gradedqualityBmaintenanceEnforces fine-grained, context-aware access control on MCP tool calls, with a tamper-evident, replayable audit log that records denials and verifies every decision.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/lara-muhanna/mcp-airlock'
If you have feedback or need assistance with the MCP directory API, please join our Discord server