context-firewall
{
"answer": "The Context Firewall MCP server is a local proxy that sits in front of downstream MCP servers (like memory) and exposes just 4 meta-tools to your AI agent. It reduces context window bloat via progressive tool disclosure, output compression, and tool-definition collapsing.\n\n### The 4 Meta-Tools\n\n- list_tool_categories — Lists all connected downstream servers with their status, tool counts, and capability categories. Start here to see what's available.\n- search_tools — Searches downstream tools by keyword and returns their full input schemas, enabling progressive disclosure without loading all schemas upfront.\n- invoke_tool — Invokes a tool on a downstream server (specify server, tool, and args). Large outputs are automatically compressed (base64 stripping, HTML→Markdown conversion, JSON summarization, token-budget truncation), and a handle is returned for full retrieval.\n- read_more — Retrieves a slice of a stored full output by handle and character offsets, letting you page through anything that was compressed.\n\n### Key Benefits\n\n- Output Compression — HTML, JSON, and base64 blobs are compressed by 60–95% before reaching the model's context, with configurable settings (maxOutputTokens, htmlToMarkdown, stripBase64, jsonSummary).\n- Progressive Tool Disclosure — Downstream tool schemas are only loaded when searched or invoked, reducing initial token costs.\n- Full Data Retrieval — Compressed outputs are never permanently lost; read_more provides paginated access to the complete original.\n- Access Control — Per-server and per-tool allow/deny policies can be configured.\n- Session Reporting — Generates reports on token savings from definition collapsing and output compression.\n- Timeout Handling — Configurable timeouts on invoke_tool calls prevent blocking on unresponsive downstream servers.\n- Universal Compatibility — Works with any MCP client (Claude Code, Claude Desktop, Cursor, Cline, etc.)."
}
Context Firewall
Turn 50+ MCP tools into 4, and shrink large tool outputs by 60–95% (real HTML/JSON, measured — see benchmark) — for any MCP client, any model. Any output still over your configured token budget after compression is hard-truncated to that budget, with the full original retrievable via read_more.
Context Firewall is a local MCP proxy that sits between your AI agent (Claude Code, Claude Desktop, Cursor, Cline, ...) and every downstream MCP server you've configured. It exposes exactly 4 tools to the client, no matter how many tools the downstream servers actually have, and compresses large tool outputs (raw HTML, base64 blobs, giant JSON) before they ever reach the model's context window.
Measured results
Metric | Result |
Tool collapse | 122 → 4 exposed meta-tools (5 real downstream servers incl. official GitHub |
Tool-definition savings | ~28,600 tokens (estimated, chars ÷ 3.5) — 102,158 raw definition chars vs. 2,146 exposed |
Output compression | 60–95%+ on real-world HTML/JSON, e.g. a live Wikipedia page via the |
All figures measured against real downstream MCP servers, not synthetic data — see STATE.md ("P1-1 — github server portion" section) for full methodology.
Related MCP server: fastmcp-gateway
What it does
Progressive tool disclosure — instead of loading every downstream tool's full schema at startup, the client sees 4 meta-tools (
list_tool_categories,search_tools,invoke_tool,read_more) and only pays the token cost of a tool's full schema when it actually searches for it.Output compression pipeline — large tool results are run through base64 stripping, HTML→Markdown conversion, JSON structure-aware summarization, and finally character-budget truncation, in that order, before being returned.
Full output retrievable via
read_more— nothing is silently thrown away. Every compressed output is stored in full (in memory, opaque handle) and can be paged back withread_more(handle, offset, length).Session savings report — on shutdown, prints a shareable terminal card (and optional Markdown file) showing tool-definition and output-token savings for the session, plus a breakdown of which tools saved the most.
Quickstart
npx context-firewall --config context-firewall.jsonMinimal context-firewall.json (the downstreams block mirrors the mcpServers format you already know):
{
"downstreams": {
"filesystem": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-filesystem", "/path/to/project"]
},
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}" }
}
}
}(${GITHUB_TOKEN} above is just this example's environment variable name — pick whatever's already set in your shell; it gets expanded into GITHUB_PERSONAL_ACCESS_TOKEN, the env var name the downstream server itself actually reads.)
Which GitHub server? Two options, different tool counts:
@modelcontextprotocol/server-github(used above) — the original npm package, 26 tools, onenpx -yline, zero extra setup. Archived/no longer maintained upstream, but still functional.github/github-mcp-server— the actively-maintained official server, 44 tools (default toolset) to 85 (GITHUB_TOOLSETS=all). Ships as a Go binary or Docker image, not an npm package:"github": { "command": "docker", "args": ["run", "-i", "--rm", "-e", "GITHUB_PERSONAL_ACCESS_TOKEN", "ghcr.io/github/github-mcp-server"], "env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}" } }or, running a locally-built/downloaded binary directly:
"command": "/path/to/github-mcp-server", "args": ["stdio"](sameenvblock; addGITHUB_TOOLSETSto scope which of the 85 tools are exposed).
Per-tool allow/deny policy. Add allowTools/denyTools (array of exact names or * globs) to any downstream entry to restrict which of its tools can be invoked:
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"denyTools": ["delete_*"]
}Deny always wins over allow. When allowTools is set, only matching tools are permitted; everything else on that server is blocked. An empty allowTools: [] is treated the same as omitting it (allow everything), not "deny everything". Blocked tools are hidden from search_tools results, and invoke_tool rejects them before dispatching to the downstream server. Tool counts in list_tool_categories and in the meta-tool descriptions are unfiltered totals — the policy is only enforced at search_tools/invoke_tool time.
Client setup
Set Context Firewall as your only MCP server — move every downstream server you currently configure directly (filesystem, github, everything, ...) into context-firewall.json's downstreams block instead. Your agent then sees 4 tools instead of the sum of every downstream server's tool count; pointing the client at Context Firewall alongside your existing servers doesn't give you the tool-collapse or compression benefit.
Each client below takes the same server entry:
{
"mcpServers": {
"context-firewall": {
"command": "npx",
"args": ["-y", "context-firewall", "--config", "/absolute/path/to/context-firewall.json"]
}
}
}Claude Code
Project-scoped .mcp.json in your repo root (shown above), or via the CLI:
claude mcp add --transport stdio context-firewall -- npx -y context-firewall --config /absolute/path/to/context-firewall.jsonClaude Desktop
claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json; Windows: %APPDATA%\Claude\claude_desktop_config.json) — same mcpServers block as above.
Cursor
.cursor/mcp.json (project-scoped) or ~/.cursor/mcp.json (global) — same mcpServers block as above.
Cline
cline_mcp_settings.json (VS Code extension storage; macOS: ~/Library/Application Support/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json) — same mcpServers block as above.
Compatibility status
Client | Status |
Claude Code | tested in real agent sessions — autonomous list → search → invoke → read_more workflow verified end-to-end |
Claude Desktop | protocol-verified* |
Cursor | config format documented, community testing welcome |
Cline | config format documented, community testing welcome |
* Verified via MCP protocol integration tests (176 automated tests, including full stdio protocol round-trips against real downstream servers). Real-client reports welcome.
Configuration
downstreams
Each entry is either a stdio server (same shape as mcpServers) or a Streamable HTTP server:
{
"downstreams": {
"local-tool": { "command": "npx", "args": ["-y", "some-mcp-server"], "env": { "TOKEN": "${TOKEN}" } },
"remote-tool": { "url": "https://mcp.example.com/mcp", "transport": "streamable-http" }
}
}${VAR_NAME} in any string value is expanded from the environment; a missing variable fails config load with a readable error.
compression
Policy resolution order is default < perServer < perTool (later overrides earlier, field by field):
{
"compression": {
"default": {
"maxOutputTokens": 2000,
"htmlToMarkdown": true,
"stripBase64": true,
"jsonSummary": true,
"bypass": false
},
"perServer": { "github": { "maxOutputTokens": 4000 } },
"perTool": { "filesystem/read_file": { "maxOutputTokens": 8000 } }
}
}Field | Type | Default | Meaning |
| number |
| Soft budget (chars ≈ tokens × 3.5) a compressed output is truncated to as a last resort. |
| boolean |
| Convert detected HTML markup to Markdown. |
| boolean |
| Replace base64 blobs (data URIs and bare blocks) with a |
| boolean |
| Collapse homogeneous JSON arrays and trim long string fields, keeping valid JSON. |
| boolean |
| Skip the whole pipeline for this server/tool - output passes through untouched. |
report
{
"report": {
"enabled": true,
"markdownPath": "./context-firewall-report.md"
}
}Field | Type | Default | Meaning |
| boolean |
| Print the session report to stderr on shutdown. |
| string | (none) | If set, also write the report as a Markdown file at this path. |
callToolTimeoutMs
Top-level (not nested under compression). Per-invoke_tool timeout in milliseconds passed to the downstream MCP SDK client; a downstream that hangs without responding causes invoke_tool to return an isError result once this elapses, instead of blocking. Defaults to the SDK's own default (60,000ms) when unset.
{ "callToolTimeoutMs": 30000 }How it works
The client calls list_tool_categories() to see what's connected and what it's roughly capable of, search_tools(query) to pull the full input schema for candidate tools, invoke_tool(server, tool, args) to actually run one (compressed on the way back), and read_more(handle, offset, length) to page through anything that got compressed. Compression, when it runs, always applies in the same order: strip base64 → HTML to Markdown → JSON structure summary → truncate to budget. Security-relevant outputs (errors, permission/warning/confirmation messages) are never silently compressed - they pass straight through, only hard-capped at 50,000 characters to prevent a single runaway error dump from blowing out the caller's context.
Positioning
Context Firewall complements Anthropic's Tool Search Tool, it doesn't compete with it. Tool Search solves tool definition bloat at startup (the schemas loaded into context before any tool is even called), and is Claude-specific. Context Firewall compresses tool outputs at call time - the half of the problem Tool Search doesn't touch - and works with any MCP client and any model, not just Claude.
A note on token counts
Every token count in this project (truncation budgets, the session report) is estimated as chars / 3.5, never an exact model-specific count. There is no code path that sends your tool output content to an external API to get an exact count - that would conflict with the safety posture below. The session report is always labeled "(estimated)" for this reason.
Safety
Tool arguments and output content are never written to logs or the session report - only server/tool names and character/token counts.
Security-relevant outputs (errors, permission denials, warnings, confirmations) are never silently compressed.
Downstream tool descriptions are treated as untrusted input and only ever displayed, never executed.
Downstream tool descriptions are passed through verbatim, unsanitized -
search_toolsdoes not strip or filter prompt-injection text a malicious downstream might put there. The trust boundary is which downstream servers you choose to configure, not this gateway.Progressive disclosure has a real tradeoff: tool descriptions arrive on demand, mid-session, right when the calling model actively asks for them via
search_tools- which is also when a model is least likely to scrutinize an embedded instruction, compared to tools all being presented up front at session start. As of v0.3.0 this is mitigated two ways:search_toolsresults are wrapped in<untrusted-tool-descriptions nonce="...">...</untrusted-tool-descriptions nonce="...">delimiters carrying a random nonce generated once per process at startup (crypto.randomBytes(8).toString('hex'), fixed for the process's whole lifetime), with a note telling the model that only a closing tag carrying the matching nonce ends the block; and the CLI prints a human-readable digest to stderr on startup (server names, tool counts, top categories) so an operator can see at a glance what actually got connected. The nonce specifically defeats the literal bypass where a downstream embeds its own</untrusted-tool-descriptions>string followed by forged "trusted system" instructions in its description - it can't predict the nonce, so it can't forge a matching closing tag. Residual risk: this is still a text-level convention, not a sandbox - it depends on the calling model actually reading the note and honoring the nonce match; nothing stops a model from ignoring the framing altogether. Neither mitigation sanitizes the description content itself - see the point above. The delimiter framing now costs about 80-85 tokens persearch_toolscall (~289 characters, at this project's chars/3.5 estimate).The categories shown by
list_tool_categoriesare also derived from downstream-supplied data (tool names, via a crude verb-prefix heuristic inregistry.ts) - but a tool name is a far lower-bandwidth channel for smuggling instructions than a free-text description, and this output isn't wrapped in the delimiters above. Treat it as lower-risk thansearch_toolsoutput, not risk-free.
License
MIT
Maintenance
Related MCP Servers
- Alicense-qualityAmaintenanceA proxy server that wraps existing MCP servers to significantly reduce token consumption by compressing tool descriptions into a two-step interface. It enables users to integrate extensive toolsets without exceeding context limits or incurring high API costs.106Apache 2.0
- Alicense-qualityAmaintenanceAggregates tools from multiple upstream MCP servers and exposes them through 4 meta-tools, enabling LLMs to discover and use hundreds of tools without loading all schemas upfront.2Apache 2.0
- Flicense-qualityBmaintenanceA proxy that intercepts MCP responses to reduce token consumption by compressing them via a pipeline and exposing only two meta-tools to the LLM.
Related MCP Connectors
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Markdown-first MCP server for Notion API with 8 composite tools and 39 actions.
Operator-as-agent MCP hub. 6 tools. First $5 free, then $0.001/call.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Alepha188838884/context-firewall'
If you have feedback or need assistance with the MCP directory API, please join our Discord server