headroom_agent_mcp
It is a read-only discovery MCP server that explores noisy docs, logs, codebases, or web sources and returns condensed evidence for a parent agent, without editing files.
Run a single
run_discoverytool for scoped investigationResearch documentation and large README files
Triage logs and terminal output, with severity ordering and deduplication
Discover codebases broadly: collect candidate files, symbols, and snippets
Perform web research with search providers (Brave, Tavily, DuckDuckGo), readability extraction, dedup, BM25 reranking, concurrent fetches, and caching
Run safe read-only terminal commands like
git statusorpytest -qunder allowlisted profilesScope work with
scope_pathscovering files, directories, or URLsBias discovery with
query_hintsand configurable limits (max_files,max_commands,raw_read_budget)Return structured findings:
relevant_findings,candidate_files,candidate_symbols,small_snippets,uncertainties,commands_run, andraw_reads_needed_by_parentOptionally enrich summaries via an LLM, including through a Headroom proxy with compression and CCR retrieval
Control output language and choose optional LLM model profiles per request
Enables web research discovery using Brave Search as a search provider, returning relevant pages and readable content for the parent agent.
Enables web research discovery using DuckDuckGo as a keyless fallback search provider, returning relevant pages and readable content for the parent agent.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@headroom_agent_mcpExplore the codebase to locate the billing logic and suggest which files the parent should read first."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
headroom_agent_mcp
Read-only discovery MCP for OpenClaw and other agent hosts, with optional Headroom-proxied LLM enrichment.
It is designed for the cases where Headroom actually helps:
large docs / README files
noisy logs and terminal output
broad codebase discovery before the parent agent reads raw files
web research, where the fetched page text stays behind the Headroom proxy instead of eating the parent context
It is not a final code-editing agent. The parent agent still reads raw files and patches them directly.
Status
New standalone repository
Intended to be published separately from upstream
headroomLicense:
Apache-2.0Upstream compatibility target:
headroom+OpenClaw
Related MCP server: agy-mcp
What It Exposes
One MCP tool:
run_discovery
Use it when the parent agent needs:
docs_researchlogs_triagecodebase_discoveryweb_research
Do not use it when:
you already know the exact 1-3 files to edit
you need final patch generation
the input is already small and precise
Tool Contract
Input highlights:
objective: concrete goal for this runobjective_type:docs_research|logs_triage|codebase_discovery|web_researchscope_paths: files, directories, or URLs URL fetches are bounded and only the first 20,000 characters of each source are inspected.query_hints: extra terms to bias searchterminal_commands: optional tokenized safe commands like["git", "status"]command_allowlist_profile:safe_readonlyorsafe_terminalresponse_language: optional output language for LLM enrichment; defaultensearch_results_limit:web_researchonly, 1-10 results per query; default5search_provider:web_researchonly,brave|tavily|duckduckgo| auto; picks which provider is tried first, the others stay as fallbacks
The tool accepts the fields either flat (recommended) or bundled inside a legacy params object.
Both forms are valid, so existing callers keep working:
{ "objective": "Find the auth check", "objective_type": "codebase_discovery", "scope_paths": ["src/"] }{ "params": { "objective": "Find the auth check", "objective_type": "codebase_discovery" } }logs_triage extra behavior:
test/fixture directories (
tests,test,__tests__,spec,specs) are ignored when real logs existfindings are deduplicated and ordered by severity (error > warning > info)
if the caller explicitly scopes a single file inside a test directory, it is still honored
web_research extra behavior:
the whole search is delegated to this tool: it queries the backend, then fetches every result
fetched HTML is reduced to readable text (trafilatura when installed, a stdlib HTML parser otherwise) before any preview is built
providers are tried in order and the first one that returns results wins:
tavily->brave->duckduckgoTAVILY_API_KEY/BRAVE_API_KEYgate the first two; DuckDuckGo needs no key and is always the last fallback (it parses the no-JShtml.duckduckgo.comendpoint, so it can be rate-limited)if every provider errors, the run reports an uncertainty naming them; if a provider just returns nothing, the response states which one was tried
results are deduped before fetching: the URL key ignores scheme, fragment (
#...) and tracking parameters (utm_*,fbclid, ...), and no single domain may contribute more than2results (MAX_RESULTS_PER_DOMAINinsrc/headroom_agent_mcp/websearch.py)pages whose readable text repeats an earlier source are skipped (mirrors/syndications), keyed on a normalized 400-char prefix
survivors are reranked with Okapi BM25 over title + body (zero dependencies): term frequency saturates, length is normalized, and each query term is IDF-weighted — so a short on-topic page outranks a long padded one
drops are always declared in
uncertainties, never silentthe surviving pages are fetched concurrently (up to
5in flight), so five sources cost roughly one round-trip instead of fivethe search call and each fetched page are cached on disk with a TTL, keyed on the canonical URL / normalized query: repeating an objective inside the window costs no search credit and no re-download. Only successes are cached — failures and empty result sets always go back to the network
search and fetch calls retry transient failures (
429,5xx, transport errors) with backoff, honouringRetry-After; a page that never resolves falls back to its search snippet instead of failing the run
Measured on the DGX (scripts/measure_websearch_dgx.sh), 5 sources:
search 4.82s -> 0.00s (cache) · fetch 1.71s -> 0.36s (concurrent, 4.7x) -> 0.05s (cache), with byte-identical readable text across all three paths.
Output highlights:
relevant_findingscandidate_filescandidate_symbolssmall_snippetsuncertaintiescommands_run(blocked commands are returned withexit_code=-1,blocked=true)raw_reads_needed_by_parentrecommended_next_actionllm_enrichedllm_errorllm_profile_used
The parent agent should treat raw_reads_needed_by_parent as the handoff for precise next reads before any edit.
Architecture
Parent Agent
-> run_discovery (this MCP)
-> scoped file/url collection
-> safe terminal commands
-> web search + readability extraction (web_research)
-> optional LLM summarization
-> optionally routed through Headroom proxy
<- structured discovery output
Parent Agent
-> reads raw target files itself
-> edits code itselfWhy Headroom Is Optional Here
This repo does not reimplement Headroom compression logic.
Instead, if you configure the subagent model to talk to a Headroom proxy, the subagent gets automatic compression on its own model traffic while it explores noisy inputs. That keeps the parent agent precise and uncompressed for final edits.
How the evidence is framed matters more than the proxy itself. The delegated LLM receives the bulky raw evidence as one JSON item per document section. Measured against a live proxy on the DGX (scripts/probe_proxy_framings_dgx.sh, deepseek/deepseek-v4-flash):
Evidence framing | Prompt tokens before | after | saved |
prose block in the user message | 2,525 | 2,411 | 4.5% |
JSON, 5 items | 1,547 | 1,547 | 0% |
JSON, 10 items | 3,067 | 808 | 73.7% |
JSON, 30 items | 9,187 | 2,368 | 74.2% |
Headroom routes JSON arrays to SmartCrusher, which keeps first/last, error and query-relevant items and drops the rest. Prose in the same position barely compresses, and a payload with fewer than ~10 items has nothing to discard. That is why the service emits sections, not a text blob.
Compression is reversible, and the client does not have to implement the retrieval. With CCR enabled the proxy injects the headroom_retrieve tool, the delegated model calls it when it needs a block back, and the proxy resolves the call server-side before returning a final answer:
"When the LLM calls
headroom_retrieve[...] Response Handler intercepts the tool call [...] Continues the API call automatically. The client never sees CCR tool calls — they're handled transparently."
Source: upstream wiki/ccr.md (CCR Phase 3, shipped with headroom-ai).
Verified end to end on the DGX (scripts/measure_ccr_retrieval_dgx.sh): during a single web_research call the proxy reported toin.total_retrievals going 0 -> 1 and ccr_retrievals 42 -> 48, while the delegated prompt shrank 13,991 -> 3,499 tokens (75.0%) and the returned summary stayed fully grounded.
Two knobs matter for that result:
Do not force
response_format. Forcing{"type": "json_object"}made the model echo the compressed table instead of answering. The service now asks for a JSON object through the prompt and parses a fenced or prose-wrapped body (extract_json_object). SetHEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT=trueonly for providers you have verified.Set
max_tokens. DeepSeek documents that JSON Output "may occasionally return empty content" and asks callers to "set themax_tokensparameter reasonably to prevent the JSON string from being truncated midway" (https://api-docs.deepseek.com/guides/json_mode).
--no-ccr remains available for callers that cannot tolerate a retrieval round trip or want the conservative single-shot mode (scripts/headroom_proxy_service_dgx.sh start-no-ccr).
Quick Start
Copy
.env.templateto.envInstall the package in a Python 3.11+ environment
Run the smoke check or wire the stdio launcher into your MCP host
Windows:
python -m venv C:\Users\giova\.venvs\headroom_agent_mcp
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python -m pip install -e Z:\Repositories\headroom_agent_mcp[dev]
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python Z:\Repositories\headroom_agent_mcp\scripts\headroom_agent_stdio_windows.py --checkLinux / DGX:
python3 -m venv ~/.venvs/headroom_agent_mcp
~/.venvs/headroom_agent_mcp/bin/python -m pip install -e ~/Repositories/headroom_agent_mcp[dev]
~/Repositories/headroom_agent_mcp/scripts/headroom_agent_stdio_unix.sh --checkConfiguration
Copy .env.template to .env or export the variables in your runtime:
HEADROOM_AGENT_PYTHON(optional explicit interpreter path for launchers)HEADROOM_AGENT_MODEL_PROVIDERHEADROOM_AGENT_MODEL_NAMEHEADROOM_AGENT_BASE_URLHEADROOM_AGENT_API_KEY_ENV(optional alternate env var name for auth)HEADROOM_AGENT_API_KEYHEADROOM_AGENT_REQUIRE_API_KEY(trueby default; setfalsefor local keyless endpoints)HEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT(falseby default; forcingjson_objectmade the delegated model echo compressed evidence instead of answering — see the proxy section)HEADROOM_AGENT_MAX_TOKENS(default2048; DeepSeek asks for a reasonable cap to avoid a JSON body truncated midway)HEADROOM_PROXY_URL(optional)HEADROOM_AGENT_TIMEOUT_SECONDS(default45;120whenHEADROOM_PROXY_URLis set, because compression adds latency)HEADROOM_AGENT_LLM_EVIDENCE_CHARS(raw evidence characters handed to the delegated LLM; default12000,40000whenHEADROOM_PROXY_URLis set)BRAVE_API_KEY/TAVILY_API_KEY(keyed web search backends;web_researchalways keeps the keylessduckduckgofallback)HEADROOM_AGENT_SEARCH_PROVIDER(optional; pins which search provider is tried first)HEADROOM_AGENT_TAVILY_SEARCH_DEPTH(advancedby default; setbasicto trade recall for credits)HEADROOM_AGENT_CACHE_TTL_SECONDS(search + page cache TTL, default3600;0disables the cache entirely)HEADROOM_AGENT_CACHE_DIR(cache location; defaults to~/.headroom/agent_cache, i.e. outside the repository)
If HEADROOM_PROXY_URL is set, the configured LLM profile routes through it, the evidence budget and the request timeout grow, and the proxy compresses the subagent's own prompt before it reaches the provider.
For local OpenAI-compatible endpoints:
point
HEADROOM_AGENT_BASE_URLto your local/v1serverset
HEADROOM_AGENT_REQUIRE_API_KEY=falseif the server is keylessset
HEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT=falseif the server does not support JSON response formattingoptionally set
HEADROOM_AGENT_API_KEY_ENVto a different variable name if the server still wants auth
config/config.yaml is live:
defaultsare applied when the caller omits optional request fields likemax_files,max_commands,raw_read_budget,return_snippets, andcommand_allowlist_profileprofilesdefine the allowed tokenized command prefixes for each command profiledirectory scans also collect common config files without standard suffixes, such as
.env,.env.template,Dockerfile,Makefile, andProcfile
LLM enrichment behavior:
if an LLM profile exists and the caller does not pass
model_profile, the server falls back to the configured default profilethe caller can force the enrichment output language with
response_language; default isenthe response body is parsed tolerantly (markdown fence or surrounding prose), so a provider that wraps its JSON still works
LLM failures are exposed in the response via
llm_errorand logged tostderrwithout corrupting the stdio MCP streama response with no text content reports the upstream
finish_reasonand any tool calls, instead of an opaqueNoneTypeerrorzero-score candidates are now labeled as fallback candidates instead of claiming keyword overlap that did not happen
Provider selection:
the active default LLM backend is chosen by
HEADROOM_AGENT_MODEL_PROVIDERthe provider-specific settings come from the matching env values such as
HEADROOM_AGENT_MODEL_NAMEandHEADROOM_AGENT_BASE_URLthe MCP caller can override the default per request by passing
model_profilethe current
config/config.yamlonly defines command profiles and request defaults; the default repo setup builds the active LLM profile from env
Current retrieval behavior:
directory scans collect standard source/docs/log files by suffix
directory scans also collect common config files by name, including
.env,.env.template,Dockerfile,Makefile, andProcfilesnippet budget is now distributed across top documents in rounds, so one dense file does not starve the rest of the evidence set
snippet excerpts are centered on the match column, so long single-line sources (minified JS/JSON, long CSV rows, single-line logs) still include the matching term instead of a mute prefix
local file reads are bounded to the first 20,000 characters per source to cap memory usage and latency
URL fetches strip boilerplate to readable text first, then bound the result to the LLM evidence budget (
20,000characters minimum,40,000when the proxy is configured)when a source is truncated by that cap, the response adds an
uncertaintieswarning instead of treating missing later matches as evidence of absence
Launchers:
Windows stable launcher:
scripts/headroom_agent_stdio_windows.pyWindows convenience shim:
scripts/headroom_agent_stdio_windows.cmdUnix/Linux launcher:
scripts/headroom_agent_stdio_unix.shDGX compatibility shim:
scripts/openclaw_stdio_dgx.sh
MCP JSON Templates
Trae / Windows, use the repo .env as the source of truth:
{
"mcpServers": {
"headroom_agent_discovery": {
"type": "STDIO",
"description": "Headroom Agent MCP discovery server",
"command": "python",
"args": [
"Z:\\Repositories\\headroom_agent_mcp\\scripts\\headroom_agent_stdio_windows.py"
],
"env": {}
}
}
}Trae / Windows, force the fast Windows IQ4 backend directly from MCP config:
{
"mcpServers": {
"headroom_agent_discovery": {
"type": "STDIO",
"description": "Headroom Agent MCP discovery server (Windows IQ4 backend)",
"command": "python",
"args": [
"Z:\\Repositories\\headroom_agent_mcp\\scripts\\headroom_agent_stdio_windows.py"
],
"env": {
"HEADROOM_AGENT_MODEL_PROVIDER": "local",
"HEADROOM_AGENT_MODEL_NAME": "nex-n2.5-mini-uncensored-iq4xs",
"HEADROOM_AGENT_BASE_URL": "http://192.168.1.11:8080/v1",
"HEADROOM_AGENT_REQUIRE_API_KEY": "false",
"HEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT": "false",
"HEADROOM_AGENT_TIMEOUT_SECONDS": "45"
}
}
}
}If env is empty, the launcher loads the repo .env and that file decides the active provider.
If env contains provider variables, the MCP host overrides the repo defaults for that process.
OpenClaw Example
Add a server entry like the example in config/openclaw.headroom_agent_mcp.example.json.
OpenClaw / Linux:
{
"mcp": {
"servers": {
"headroom_agent_discovery": {
"enabled": true,
"command": "/home/jagones/Repositories/headroom_agent_mcp/scripts/headroom_agent_stdio_unix.sh",
"args": [],
"cwd": "/home/jagones/Repositories/headroom_agent_mcp",
"connectionTimeoutMs": 120000
}
}
}
}For Windows hosts, use config/windows.stdio.headroom_agent_mcp.example.json.
The MCP description is intentionally explicit so the parent agent knows:
when to call it
what to pass
what not to expect from it
Development
Windows local test venv:
python -m venv C:\Users\giova\.venvs\headroom_agent_mcp
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python -m pip install -e Z:\Repositories\headroom_agent_mcp[dev]
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python -m pytest Z:\Repositories\headroom_agent_mcp\tests -qDGX smoke scripts:
scripts/run_tests_dgx.shscripts/smoke_check_dgx.shscripts/smoke_openrouter_headroom_dgx.shscripts/smoke_web_research_dgx.sh(web_researchend to end, prints proxy savings)scripts/probe_proxy_framings_dgx.sh(payload-framing compression benchmark)scripts/probe_ccr_retrieval_dgx.sh(does the proxy resolveheadroom_retrieveunder a given response format?)scripts/inspect_proxy_jsonl_dgx.py(reads the official--log-fileJSONL and the/statsCCR counters)scripts/measure_ccr_retrieval_dgx.sh(CCR counters before/after oneweb_researchcall)scripts/measure_websearch_dgx.sh(cold vs warm cache and sequential vs concurrent fetch timings)
DGX Headroom proxy runtime:
scripts/setup_headroom_runtime_dgx.sh(Linux venv for theheadroomrepo +headroom-ai[proxy])scripts/setup_headroom_ml_dgx.sh(addsheadroom-ai[ml], the Kompress ML compressor)scripts/headroom_proxy_service_dgx.sh{start|start-trace|start-no-ccr|stop|status}(proxy on127.0.0.1:8788;startkeeps CCR on,start-tracealso logs full messages)scripts/headroom_proxy_stats.py(one-line savings summary)scripts/wire_openclaw_headroom_dgx.py(registers the officialheadroomMCP +HEADROOM_PROXY_URLin~/.openclaw/openclaw.json)
License And Attribution
This repository is licensed under Apache-2.0, matching the upstream Headroom project.
Why this shape:
upstream
headroomis Apache-2.0 licensedthis repo is a separate overlay/companion project, not a fork that modifies upstream in place
Apache-2.0 allows separate derivative or companion works as long as the license text is included and attribution/trademark rules are respected
Files added for that:
LICENSENOTICE
Upstream reference:
Headroom: headroomlabs-ai/headroom
This project references Headroom for interoperability and architectural patterns, but does not claim affiliation or endorsement.
Current Scope
Implemented:
contract validation
safe terminal policy
deterministic discovery service
optional OpenAI-compatible LLM enrichment
web search (
tavily/brave/ keylessduckduckgo) with readability extraction, provider fallback, URL/content dedup, BM25 reranking, concurrent fetches, disk TTL cache and retry-with-backoffMCP server and CLI smoke check
Not implemented:
write/edit tools
automatic child-process orchestration inside OpenClaw
Available Tools
1 toolrun_discoveryARead-only
Explore noisy docs, logs, or codebases and return only the evidence a parent agent needs.
Use this tool when the parent agent needs discovery or triage before reading raw files itself. Best cases: docs research, logs/output triage, or codebase discovery over broad scopes. Do not use it for final file edits or precise patch generation.
Inputs:
objective: concrete question or goal for this run
objective_type: docs_research, logs_triage, or codebase_discovery
scope_paths: files, directories, or URLs to inspect
query_hints: optional extra terms to bias search/scoring
terminal_commands: optional tokenized safe commands, e.g. [["git","status"],["pytest","-q"]]
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds behavioral context: the tool filters noisy inputs, returns only evidence, handles broad scopes, and accepts 'safe commands.' It does not contradict the annotations and provides useful extra context beyond them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized and front-loaded: purpose first, usage guidance second, and a compact bulleted parameter list. Every sentence contributes useful information, and the example for terminal_commands is valuable without adding unnecessary length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, usage, and key parameters, but the tool has no output schema and the description gives only a vague sense of the return value ('return only the evidence'). It also does not explain how budgets, snippets, or command allowlist profiles affect behavior, so an agent may not know how to tune or interpret the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the core parameters well: objective, objective_type, scope_paths, query_hints, and terminal_commands, including a concrete example. However, it omits several meaningful parameters such as raw_read_budget, max_files, return_snippets, and command_allowlist_profile, leaving gaps an agent must infer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific action verb and resource ('Explore noisy docs, logs, or codebases') and clearly states the output ('return only the evidence a parent agent needs'). It also enumerates the supported objective types, making the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool ('discovery or triage before reading raw files itself'), lists best cases, and gives a concrete exclusion ('Do not use it for final file edits or precise patch generation'). This gives an agent clear routing guidance even without sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
run_discovery
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap between tools. run_discovery has a clearly defined purpose and inputs, making it unambiguous for an agent to select.
The single tool name run_discovery follows a clear verb_noun convention. There are no other tools to create naming inconsistencies or mixed conventions.
One tool is minimal, but it is appropriate for a narrowly scoped discovery/triage subagent. The count feels slightly thin compared to typical multi-tool servers, but the server's purpose is focused enough that a single tool can reasonably fulfill it.
The tool covers the stated discovery domain across docs research, logs triage, and codebase discovery, with support for scope paths, query hints, and safe terminal commands. Minor gaps exist around iterative refinement or returning raw context, but the core triage workflow is well covered.
Maintenance
Related MCP Connectors
AI agent infrastructure for discovery, authorization, execution, identity, and signed receipts.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Outcome-first agent fallback: free discovery, minimal routing, declared costs, verified execution.
MCP delegation fallback for AI agents to discover capabilities, knowledge, tools, and collaborators.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables coding agents to pre-compute repository structure and access structured intelligence briefs, including dependency graphs, hotspots, and blast radius, reducing token usage and improving code understanding.101 npm1-
- AlicenseAqualityAmaintenanceEnables Claude Code and Claude Desktop to delegate token-heavy tasks to Antigravity headless subagents, offloading file edits, test runs, and exploration while preserving Claude's context window.1225 npm1MIT
- FlicenseAqualityCmaintenanceEnables main agents to delegate memory retrieval, web research, and multi-step tasks to internal sub-agents, returning concise conclusions while keeping detailed tool calls and raw content out of the main context.3-
- AlicenseNot gradedqualityCmaintenanceEnables coding agents to delegate file reads, command output triage, page fetching, and image inspection to cheap flash models, returning concise answers and verified pointers while keeping raw dumps out of the main model's context.MIT