Skip to main content
Glama

headroom_agent_mcp

Read-only discovery MCP for OpenClaw and other agent hosts, with optional Headroom-proxied LLM enrichment.

It is designed for the cases where Headroom actually helps:

  • large docs / README files

  • noisy logs and terminal output

  • broad codebase discovery before the parent agent reads raw files

  • web research, where the fetched page text stays behind the Headroom proxy instead of eating the parent context

It is not a final code-editing agent. The parent agent still reads raw files and patches them directly.

Status

  • New standalone repository

  • Intended to be published separately from upstream headroom

  • License: Apache-2.0

  • Upstream compatibility target: headroom + OpenClaw

Related MCP server: agy-mcp

What It Exposes

One MCP tool:

  • run_discovery

Use it when the parent agent needs:

  • docs_research

  • logs_triage

  • codebase_discovery

  • web_research

Do not use it when:

  • you already know the exact 1-3 files to edit

  • you need final patch generation

  • the input is already small and precise

Tool Contract

Input highlights:

  • objective: concrete goal for this run

  • objective_type: docs_research | logs_triage | codebase_discovery | web_research

  • scope_paths: files, directories, or URLs URL fetches are bounded and only the first 20,000 characters of each source are inspected.

  • query_hints: extra terms to bias search

  • terminal_commands: optional tokenized safe commands like ["git", "status"]

  • command_allowlist_profile: safe_readonly or safe_terminal

  • response_language: optional output language for LLM enrichment; default en

  • search_results_limit: web_research only, 1-10 results per query; default 5

  • search_provider: web_research only, brave | tavily | duckduckgo | auto; picks which provider is tried first, the others stay as fallbacks

The tool accepts the fields either flat (recommended) or bundled inside a legacy params object. Both forms are valid, so existing callers keep working:

{ "objective": "Find the auth check", "objective_type": "codebase_discovery", "scope_paths": ["src/"] }
{ "params": { "objective": "Find the auth check", "objective_type": "codebase_discovery" } }

logs_triage extra behavior:

  • test/fixture directories (tests, test, __tests__, spec, specs) are ignored when real logs exist

  • findings are deduplicated and ordered by severity (error > warning > info)

  • if the caller explicitly scopes a single file inside a test directory, it is still honored

web_research extra behavior:

  • the whole search is delegated to this tool: it queries the backend, then fetches every result

  • fetched HTML is reduced to readable text (trafilatura when installed, a stdlib HTML parser otherwise) before any preview is built

  • providers are tried in order and the first one that returns results wins: tavily -> brave -> duckduckgo

  • TAVILY_API_KEY / BRAVE_API_KEY gate the first two; DuckDuckGo needs no key and is always the last fallback (it parses the no-JS html.duckduckgo.com endpoint, so it can be rate-limited)

  • if every provider errors, the run reports an uncertainty naming them; if a provider just returns nothing, the response states which one was tried

  • results are deduped before fetching: the URL key ignores scheme, fragment (#...) and tracking parameters (utm_*, fbclid, ...), and no single domain may contribute more than 2 results (MAX_RESULTS_PER_DOMAIN in src/headroom_agent_mcp/websearch.py)

  • pages whose readable text repeats an earlier source are skipped (mirrors/syndications), keyed on a normalized 400-char prefix

  • survivors are reranked with Okapi BM25 over title + body (zero dependencies): term frequency saturates, length is normalized, and each query term is IDF-weighted — so a short on-topic page outranks a long padded one

  • drops are always declared in uncertainties, never silent

  • the surviving pages are fetched concurrently (up to 5 in flight), so five sources cost roughly one round-trip instead of five

  • the search call and each fetched page are cached on disk with a TTL, keyed on the canonical URL / normalized query: repeating an objective inside the window costs no search credit and no re-download. Only successes are cached — failures and empty result sets always go back to the network

  • search and fetch calls retry transient failures (429, 5xx, transport errors) with backoff, honouring Retry-After; a page that never resolves falls back to its search snippet instead of failing the run

Measured on the DGX (scripts/measure_websearch_dgx.sh), 5 sources: search 4.82s -> 0.00s (cache) · fetch 1.71s -> 0.36s (concurrent, 4.7x) -> 0.05s (cache), with byte-identical readable text across all three paths.

Output highlights:

  • relevant_findings

  • candidate_files

  • candidate_symbols

  • small_snippets

  • uncertainties

  • commands_run (blocked commands are returned with exit_code=-1, blocked=true)

  • raw_reads_needed_by_parent

  • recommended_next_action

  • llm_enriched

  • llm_error

  • llm_profile_used

The parent agent should treat raw_reads_needed_by_parent as the handoff for precise next reads before any edit.

Architecture

Parent Agent
  -> run_discovery (this MCP)
      -> scoped file/url collection
      -> safe terminal commands
      -> web search + readability extraction (web_research)
      -> optional LLM summarization
          -> optionally routed through Headroom proxy
  <- structured discovery output

Parent Agent
  -> reads raw target files itself
  -> edits code itself

Why Headroom Is Optional Here

This repo does not reimplement Headroom compression logic.

Instead, if you configure the subagent model to talk to a Headroom proxy, the subagent gets automatic compression on its own model traffic while it explores noisy inputs. That keeps the parent agent precise and uncompressed for final edits.

How the evidence is framed matters more than the proxy itself. The delegated LLM receives the bulky raw evidence as one JSON item per document section. Measured against a live proxy on the DGX (scripts/probe_proxy_framings_dgx.sh, deepseek/deepseek-v4-flash):

Evidence framing

Prompt tokens before

after

saved

prose block in the user message

2,525

2,411

4.5%

JSON, 5 items

1,547

1,547

0%

JSON, 10 items

3,067

808

73.7%

JSON, 30 items

9,187

2,368

74.2%

Headroom routes JSON arrays to SmartCrusher, which keeps first/last, error and query-relevant items and drops the rest. Prose in the same position barely compresses, and a payload with fewer than ~10 items has nothing to discard. That is why the service emits sections, not a text blob.

Compression is reversible, and the client does not have to implement the retrieval. With CCR enabled the proxy injects the headroom_retrieve tool, the delegated model calls it when it needs a block back, and the proxy resolves the call server-side before returning a final answer:

"When the LLM calls headroom_retrieve [...] Response Handler intercepts the tool call [...] Continues the API call automatically. The client never sees CCR tool calls — they're handled transparently."

Source: upstream wiki/ccr.md (CCR Phase 3, shipped with headroom-ai).

Verified end to end on the DGX (scripts/measure_ccr_retrieval_dgx.sh): during a single web_research call the proxy reported toin.total_retrievals going 0 -> 1 and ccr_retrievals 42 -> 48, while the delegated prompt shrank 13,991 -> 3,499 tokens (75.0%) and the returned summary stayed fully grounded.

Two knobs matter for that result:

  • Do not force response_format. Forcing {"type": "json_object"} made the model echo the compressed table instead of answering. The service now asks for a JSON object through the prompt and parses a fenced or prose-wrapped body (extract_json_object). Set HEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT=true only for providers you have verified.

  • Set max_tokens. DeepSeek documents that JSON Output "may occasionally return empty content" and asks callers to "set the max_tokens parameter reasonably to prevent the JSON string from being truncated midway" (https://api-docs.deepseek.com/guides/json_mode).

--no-ccr remains available for callers that cannot tolerate a retrieval round trip or want the conservative single-shot mode (scripts/headroom_proxy_service_dgx.sh start-no-ccr).

Quick Start

  1. Copy .env.template to .env

  2. Install the package in a Python 3.11+ environment

  3. Run the smoke check or wire the stdio launcher into your MCP host

Windows:

python -m venv C:\Users\giova\.venvs\headroom_agent_mcp
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python -m pip install -e Z:\Repositories\headroom_agent_mcp[dev]
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python Z:\Repositories\headroom_agent_mcp\scripts\headroom_agent_stdio_windows.py --check

Linux / DGX:

python3 -m venv ~/.venvs/headroom_agent_mcp
~/.venvs/headroom_agent_mcp/bin/python -m pip install -e ~/Repositories/headroom_agent_mcp[dev]
~/Repositories/headroom_agent_mcp/scripts/headroom_agent_stdio_unix.sh --check

Configuration

Copy .env.template to .env or export the variables in your runtime:

  • HEADROOM_AGENT_PYTHON (optional explicit interpreter path for launchers)

  • HEADROOM_AGENT_MODEL_PROVIDER

  • HEADROOM_AGENT_MODEL_NAME

  • HEADROOM_AGENT_BASE_URL

  • HEADROOM_AGENT_API_KEY_ENV (optional alternate env var name for auth)

  • HEADROOM_AGENT_API_KEY

  • HEADROOM_AGENT_REQUIRE_API_KEY (true by default; set false for local keyless endpoints)

  • HEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT (false by default; forcing json_object made the delegated model echo compressed evidence instead of answering — see the proxy section)

  • HEADROOM_AGENT_MAX_TOKENS (default 2048; DeepSeek asks for a reasonable cap to avoid a JSON body truncated midway)

  • HEADROOM_PROXY_URL (optional)

  • HEADROOM_AGENT_TIMEOUT_SECONDS (default 45; 120 when HEADROOM_PROXY_URL is set, because compression adds latency)

  • HEADROOM_AGENT_LLM_EVIDENCE_CHARS (raw evidence characters handed to the delegated LLM; default 12000, 40000 when HEADROOM_PROXY_URL is set)

  • BRAVE_API_KEY / TAVILY_API_KEY (keyed web search backends; web_research always keeps the keyless duckduckgo fallback)

  • HEADROOM_AGENT_SEARCH_PROVIDER (optional; pins which search provider is tried first)

  • HEADROOM_AGENT_TAVILY_SEARCH_DEPTH (advanced by default; set basic to trade recall for credits)

  • HEADROOM_AGENT_CACHE_TTL_SECONDS (search + page cache TTL, default 3600; 0 disables the cache entirely)

  • HEADROOM_AGENT_CACHE_DIR (cache location; defaults to ~/.headroom/agent_cache, i.e. outside the repository)

If HEADROOM_PROXY_URL is set, the configured LLM profile routes through it, the evidence budget and the request timeout grow, and the proxy compresses the subagent's own prompt before it reaches the provider.

For local OpenAI-compatible endpoints:

  • point HEADROOM_AGENT_BASE_URL to your local /v1 server

  • set HEADROOM_AGENT_REQUIRE_API_KEY=false if the server is keyless

  • set HEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT=false if the server does not support JSON response formatting

  • optionally set HEADROOM_AGENT_API_KEY_ENV to a different variable name if the server still wants auth

config/config.yaml is live:

  • defaults are applied when the caller omits optional request fields like max_files, max_commands, raw_read_budget, return_snippets, and command_allowlist_profile

  • profiles define the allowed tokenized command prefixes for each command profile

  • directory scans also collect common config files without standard suffixes, such as .env, .env.template, Dockerfile, Makefile, and Procfile

LLM enrichment behavior:

  • if an LLM profile exists and the caller does not pass model_profile, the server falls back to the configured default profile

  • the caller can force the enrichment output language with response_language; default is en

  • the response body is parsed tolerantly (markdown fence or surrounding prose), so a provider that wraps its JSON still works

  • LLM failures are exposed in the response via llm_error and logged to stderr without corrupting the stdio MCP stream

  • a response with no text content reports the upstream finish_reason and any tool calls, instead of an opaque NoneType error

  • zero-score candidates are now labeled as fallback candidates instead of claiming keyword overlap that did not happen

Provider selection:

  • the active default LLM backend is chosen by HEADROOM_AGENT_MODEL_PROVIDER

  • the provider-specific settings come from the matching env values such as HEADROOM_AGENT_MODEL_NAME and HEADROOM_AGENT_BASE_URL

  • the MCP caller can override the default per request by passing model_profile

  • the current config/config.yaml only defines command profiles and request defaults; the default repo setup builds the active LLM profile from env

Current retrieval behavior:

  • directory scans collect standard source/docs/log files by suffix

  • directory scans also collect common config files by name, including .env, .env.template, Dockerfile, Makefile, and Procfile

  • snippet budget is now distributed across top documents in rounds, so one dense file does not starve the rest of the evidence set

  • snippet excerpts are centered on the match column, so long single-line sources (minified JS/JSON, long CSV rows, single-line logs) still include the matching term instead of a mute prefix

  • local file reads are bounded to the first 20,000 characters per source to cap memory usage and latency

  • URL fetches strip boilerplate to readable text first, then bound the result to the LLM evidence budget (20,000 characters minimum, 40,000 when the proxy is configured)

  • when a source is truncated by that cap, the response adds an uncertainties warning instead of treating missing later matches as evidence of absence

Launchers:

  • Windows stable launcher: scripts/headroom_agent_stdio_windows.py

  • Windows convenience shim: scripts/headroom_agent_stdio_windows.cmd

  • Unix/Linux launcher: scripts/headroom_agent_stdio_unix.sh

  • DGX compatibility shim: scripts/openclaw_stdio_dgx.sh

MCP JSON Templates

Trae / Windows, use the repo .env as the source of truth:

{
  "mcpServers": {
    "headroom_agent_discovery": {
      "type": "STDIO",
      "description": "Headroom Agent MCP discovery server",
      "command": "python",
      "args": [
        "Z:\\Repositories\\headroom_agent_mcp\\scripts\\headroom_agent_stdio_windows.py"
      ],
      "env": {}
    }
  }
}

Trae / Windows, force the fast Windows IQ4 backend directly from MCP config:

{
  "mcpServers": {
    "headroom_agent_discovery": {
      "type": "STDIO",
      "description": "Headroom Agent MCP discovery server (Windows IQ4 backend)",
      "command": "python",
      "args": [
        "Z:\\Repositories\\headroom_agent_mcp\\scripts\\headroom_agent_stdio_windows.py"
      ],
      "env": {
        "HEADROOM_AGENT_MODEL_PROVIDER": "local",
        "HEADROOM_AGENT_MODEL_NAME": "nex-n2.5-mini-uncensored-iq4xs",
        "HEADROOM_AGENT_BASE_URL": "http://192.168.1.11:8080/v1",
        "HEADROOM_AGENT_REQUIRE_API_KEY": "false",
        "HEADROOM_AGENT_USE_JSON_RESPONSE_FORMAT": "false",
        "HEADROOM_AGENT_TIMEOUT_SECONDS": "45"
      }
    }
  }
}

If env is empty, the launcher loads the repo .env and that file decides the active provider. If env contains provider variables, the MCP host overrides the repo defaults for that process.

OpenClaw Example

Add a server entry like the example in config/openclaw.headroom_agent_mcp.example.json.

OpenClaw / Linux:

{
  "mcp": {
    "servers": {
      "headroom_agent_discovery": {
        "enabled": true,
        "command": "/home/jagones/Repositories/headroom_agent_mcp/scripts/headroom_agent_stdio_unix.sh",
        "args": [],
        "cwd": "/home/jagones/Repositories/headroom_agent_mcp",
        "connectionTimeoutMs": 120000
      }
    }
  }
}

For Windows hosts, use config/windows.stdio.headroom_agent_mcp.example.json.

The MCP description is intentionally explicit so the parent agent knows:

  • when to call it

  • what to pass

  • what not to expect from it

Development

Windows local test venv:

python -m venv C:\Users\giova\.venvs\headroom_agent_mcp
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python -m pip install -e Z:\Repositories\headroom_agent_mcp[dev]
C:\Users\giova\.venvs\headroom_agent_mcp\Scripts\python -m pytest Z:\Repositories\headroom_agent_mcp\tests -q

DGX smoke scripts:

  • scripts/run_tests_dgx.sh

  • scripts/smoke_check_dgx.sh

  • scripts/smoke_openrouter_headroom_dgx.sh

  • scripts/smoke_web_research_dgx.sh (web_research end to end, prints proxy savings)

  • scripts/probe_proxy_framings_dgx.sh (payload-framing compression benchmark)

  • scripts/probe_ccr_retrieval_dgx.sh (does the proxy resolve headroom_retrieve under a given response format?)

  • scripts/inspect_proxy_jsonl_dgx.py (reads the official --log-file JSONL and the /stats CCR counters)

  • scripts/measure_ccr_retrieval_dgx.sh (CCR counters before/after one web_research call)

  • scripts/measure_websearch_dgx.sh (cold vs warm cache and sequential vs concurrent fetch timings)

DGX Headroom proxy runtime:

  • scripts/setup_headroom_runtime_dgx.sh (Linux venv for the headroom repo + headroom-ai[proxy])

  • scripts/setup_headroom_ml_dgx.sh (adds headroom-ai[ml], the Kompress ML compressor)

  • scripts/headroom_proxy_service_dgx.sh {start|start-trace|start-no-ccr|stop|status} (proxy on 127.0.0.1:8788; start keeps CCR on, start-trace also logs full messages)

  • scripts/headroom_proxy_stats.py (one-line savings summary)

  • scripts/wire_openclaw_headroom_dgx.py (registers the official headroom MCP + HEADROOM_PROXY_URL in ~/.openclaw/openclaw.json)

License And Attribution

This repository is licensed under Apache-2.0, matching the upstream Headroom project.

Why this shape:

  • upstream headroom is Apache-2.0 licensed

  • this repo is a separate overlay/companion project, not a fork that modifies upstream in place

  • Apache-2.0 allows separate derivative or companion works as long as the license text is included and attribution/trademark rules are respected

Files added for that:

  • LICENSE

  • NOTICE

Upstream reference:

This project references Headroom for interoperability and architectural patterns, but does not claim affiliation or endorsement.

Current Scope

Implemented:

  • contract validation

  • safe terminal policy

  • deterministic discovery service

  • optional OpenAI-compatible LLM enrichment

  • web search (tavily / brave / keyless duckduckgo) with readability extraction, provider fallback, URL/content dedup, BM25 reranking, concurrent fetches, disk TTL cache and retry-with-backoff

  • MCP server and CLI smoke check

Not implemented:

  • write/edit tools

  • automatic child-process orchestration inside OpenClaw

Available Tools

1 tool
run_discoveryA
Read-only

Explore noisy docs, logs, or codebases and return only the evidence a parent agent needs.

Use this tool when the parent agent needs discovery or triage before reading raw files itself. Best cases: docs research, logs/output triage, or codebase discovery over broad scopes. Do not use it for final file edits or precise patch generation.

Inputs:

  • objective: concrete question or goal for this run

  • objective_type: docs_research, logs_triage, or codebase_discovery

  • scope_paths: files, directories, or URLs to inspect

  • query_hints: optional extra terms to bias search/scoring

  • terminal_commands: optional tokenized safe commands, e.g. [["git","status"],["pytest","-q"]]

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds behavioral context: the tool filters noisy inputs, returns only evidence, handles broad scopes, and accepts 'safe commands.' It does not contradict the annotations and provides useful extra context beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized and front-loaded: purpose first, usage guidance second, and a compact bulleted parameter list. Every sentence contributes useful information, and the example for terminal_commands is valuable without adding unnecessary length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage, and key parameters, but the tool has no output schema and the description gives only a vague sense of the return value ('return only the evidence'). It also does not explain how budgets, snippets, or command allowlist profiles affect behavior, so an agent may not know how to tune or interpret the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the core parameters well: objective, objective_type, scope_paths, query_hints, and terminal_commands, including a concrete example. However, it omits several meaningful parameters such as raw_read_budget, max_files, return_snippets, and command_allowlist_profile, leaving gaps an agent must infer.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific action verb and resource ('Explore noisy docs, logs, or codebases') and clearly states the output ('return only the evidence a parent agent needs'). It also enumerates the supported objective types, making the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool ('discovery or triage before reading raw files itself'), lists best cases, and gives a concrete exclusion ('Do not use it for final file edits or precise patch generation'). This gives an agent clear routing guidance even without sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedrun_discovery

TDQS

A4.4/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion or overlap between tools. run_discovery has a clearly defined purpose and inputs, making it unambiguous for an agent to select.

Naming Consistency5/5

The single tool name run_discovery follows a clear verb_noun convention. There are no other tools to create naming inconsistencies or mixed conventions.

Tool Count4/5

One tool is minimal, but it is appropriate for a narrowly scoped discovery/triage subagent. The count feels slightly thin compared to typical multi-tool servers, but the server's purpose is focused enough that a single tool can reasonably fulfill it.

Completeness4/5

The tool covers the stated discovery domain across docs research, logs triage, and codebase discovery, with support for scope paths, query hints, and safe terminal commands. Minor gaps exist around iterative refinement or returning raw context, but the core triage workflow is well covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to pre-compute repository structure and access structured intelligence briefs, including dependency graphs, hotspots, and blast radius, reducing token usage and improving code understanding.
    101 npm
    1
    -
  • F
    license
    A
    quality
    C
    maintenance
    Enables main agents to delegate memory retrieval, web research, and multi-step tasks to internal sub-agents, returning concise conclusions while keeping detailed tool calls and raw content out of the main context.
    3
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables coding agents to delegate file reads, command output triage, page fetching, and image inspection to cheap flash models, returning concise answers and verified pointers while keeping raw dumps out of the main model's context.
    MIT