Claude Code CLI MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Claude Code CLI MCP Serverrun a synchronous task to analyze the project structure"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Claude Code CLI MCP Server
Version: 0.2.0 (post-sprint-N+1)
A local STDIO MCP server that exposes tools and reusable prompts for running the Anthropic Claude Code CLI (claude) inside a controlled workspace.
Links
Model Context Protocol (MCP): https://modelcontextprotocol.io/
FastMCP: https://github.com/PrefectHQ/fastmcp
FastMCP docs: https://gofastmcp.com/
uv docs: https://docs.astral.sh/uv/
Pydantic: https://github.com/pydantic/pydantic
Pydantic Settings: https://github.com/pydantic/pydantic-settings
Claude Code CLI: https://claude.ai/
Related MCP server: claudecode-mcp
Why This Project Exists
This project is a multi-provider fork of agy-mcp-server (the sister project wrapping the Google Antigravity CLI). Both projects share the same underlying architecture to provide a thin, secure CLI shim that exposes terminal-based AI tools as Model Context Protocol (MCP) servers.
Using the Claude Code CLI directly inside an editor or orchestrator workspace poses security and configuration challenges. This MCP server encapsulates the claude process, enforcing directory access controls, standardizing outputs, and providing a file-based persistent memory layer so that system prompts and session history survive across project reloads.
Now you can leverage your Claude Pro/Team subscription or ANTHROPIC_API_KEY seamlessly inside any MCP-enabled IDE:
Cursor
Windsurf
Trae
Features
Tools
claude_health: Checks that theclaudebinary is installed, authenticated, and returns its version and authentication status.claude_run_task: Runs a synchronous (blocking) task inside the target workspace directory. Acceptsmodelandfallback_modelto select the Claude model per-run.claude_start_task: Starts an asynchronous background task inside the target workspace. Accepts the samemodelandfallback_modelparameters.claude_poll_task: Polls the status, stdout, stderr, and output changes of an active background task.claude_cancel_task: Cancels a running background task (supporting an optional force signal).claude_list_runs: Lists recent tasks managed by the server and their statuses.claude_init_persistence: Initializes the persistence layer folder and creates template markdown files.claude_read_persistence: Reads the contents of a persistent memory file (AGENTS.md,PROJECTS.md, orMEMORY.md).claude_append_persistence: Appends high-signal updates (e.g., session memories) to a persistence file.claude_update_persistence: Replaces or appends to a specific section of a persistence file by heading anchor.claude_load_persistence_context: Loads the persistent files as truncated excerpts to inject into session context.
Prompts
claude_sync_orchestration: Guidance playbook for executing synchronous tasks safely viaclaude_run_task. Includes parse-error / timeout / list_runs recovery recipes.claude_async_orchestration: Playbook for orchestrating background tasks viaclaude_start_task+claude_poll_task+claude_cancel_task. Includes backoff schedule (1s → 10s) and post-restart recovery viaclaude_list_runs.claude_model_selection_guidance: Lists all 4 model aliases (sonnet / fable / opus / haiku) with tier, cost, and multi-file-safety annotations. Source of truth:MODEL_REGISTRY.claude_security_and_workspace_rules: Summarizes sandbox limits and rules for safe and permissive environments.claude_persistence_protocol: Guides the orchestrator on maintaining the persistent memory files (load context, append session notes, update AGENTS.md with confirm gate).claude_timeout_help: Decision matrix + code snippet for(task_class, files_to_edit, model_alias) → (timeout_s, must_use_async). Matrix reflects the requested model, not always sonnet.claude_quickstart: First-call cheat-sheet for new orchestrators (workspace_path discipline, tool catalog, common gotchas + troubleshoot).claude_troubleshoot: Pattern-matches an error string and returns a canonical fix recipe (NOT_ALLOWED, NOT_LOGGED_IN, MODEL_NOT_ALLOWED, CLAUDE_NOT_FOUND, etc.).
Companion Agent: Femtobot
Recommendation: this server is the canonical
claude_*tool source forfemtobot, the CLI-first AI agent foundation in the percival.OS ecosystem.
Femtobot ships first-class support for the tools exposed here:
/mcpslash command —status,reload,tools claude-code-cli-mcp,restart claude-code-cli-mcpfor runtime inspection and recovery without restarting the agent.mcp-routerbuiltin skill — teaches the LLM when to delegate toclaude_run_taskvs. solving locally withread_file,apply_patch, etc.Capability tags — tool hints show
[long-running, safe-mode:confirm]so the model recognizes theconfirmgate before invokingclaude_run_task.Workspace auto-fill —
claude_run_taskcalls getworkspace_pathfilled in automatically from the active request context.System-prompt block —
## MCP Servers in this workspacelists this server and its tools so the model sees them at planning time.CLAUDE_MCP_PERSISTENCE_LOCATION=workspaceis auto-resolved by femtobot to<cwd_parent>/.open-cli-router/claude-code-cli-mcp/, so persistence files (AGENTS.md,MEMORY.md,PROJECTS.md) end up inside your project when configured that way.
See femtobot/docs/mcp.md
§8 "Femtobot-specific patterns" for the full integration reference, and
the CLI-router-project analysis
for the design rationale.
claude_self_test
Inspect every registered tool's input schema and report robustness. This is a metadata-only check — no tools are invoked, no subprocess is spawned, no quota is consumed. Safe to run in production or CI as a sanity probe.
Output: per-tool schema introspection (allow_extra_keys vs allow type), focused on accidental strictness regressions.
Persistent Memory
This MCP server features a file-based persistence layer stored by default under ~/.open-cli-router/claude-code/. The directory contains three editable markdown files:
AGENTS.md: The system prompt or agent persona instructions.PROJECTS.md: Summaries and structures of active projects.MEMORY.md: Chronological log of high-signal session takeaways and permanent learning.
When persistence is enabled, the server automatically loads these files and prepends their formatted excerpts to the prompt sent to the claude CLI, ensuring cross-session state persistence.
Storage location: global vs workspace
The persistence directory can live in two places, controlled by
CLAUDE_MCP_PERSISTENCE_LOCATION:
Mode | Location | Use case |
|
| User-level, persists across projects, survives |
|
| Project-level, can be committed (use |
<cwd_parent> is the parent of the server's CWD — for a typical setup,
the server's CWD is the server project directory (e.g.
/home/user/CLI-router-project/claude-code-cli-mcp), so the workspace
mode resolves to /home/user/CLI-router-project/.open-cli-router/claude-code/.
Escape hatch: CLAUDE_MCP_PERSISTENCE_BASE_DIR="$cwd_parent/.my-persistence" lets
you pick any custom subdirectory under the workspace root.
⚠️ When using
workspacemode, add.open-cli-router/to.gitignoreto avoid accidentally committing agent memory to source control.
Two-level configuration
Persistence is controlled by both server-level environment variables and runtime MCP tool calls:
Server level (set in the MCP client env block — see MCP Client Configuration):
Variable | Purpose | Default |
| Master switch for the persistence feature |
|
|
|
|
| Base directory; the namespace |
|
| Maximum size per file before writes are rejected |
|
| Create |
|
| Number of |
|
| Seed default markdown content when initializing |
|
| Fraction of |
|
Runtime level (called by the orchestrator via MCP tools):
Initialize once — call
claude_init_persistenceto create the directory and seed the three files.Load context — call
claude_load_persistence_contextat the start of each session to inject excerpts into the next prompt.Append session notes — after meaningful work, call
claude_append_persistenceonMEMORY.md.Update structured sections — when the user changes
AGENTS.mdorPROJECTS.md, callclaude_update_persistenceto persist. Note: updatingAGENTS.mdin safe mode requiresconfirm=true.
Without step 1, persistence is enabled but uninitialized — the server will not inject any context until the directory exists.
How to initialize the persistence directory
The persistence directory is created lazily — it does not exist by default. There are two equivalent ways to create it:
Option A — via MCP (recommended, normal flow):
Once both the MCP client configuration and the server are running, ask the orchestrator agent to call:
Please call claude_init_persistence to create the persistence directory.The tool seeds AGENTS.md, PROJECTS.md, MEMORY.md, and a .initialized marker under ~/.open-cli-router/claude-code/.
Option B — directly via Python (one-shot, useful for verification or first-time setup):
From the project root (claude-code-cli-mcp/):
uv run python -c "from claude_code_mcp.persistence import PersistenceStore; from pathlib import Path; PersistenceStore(base_dir=Path.home()/'.open-cli-router', max_file_bytes=524288, backup_on_write=False, seed_templates=True).init()"This call is idempotent — running it twice does not destroy existing data unless you pass force=True.
The seed templates are written in English so they can be edited by any language-aware agent later.
Quickstart
Prerequisites:
Python 3.11+
uvinstalledclaudeCLI installed (viacurl -fsSL https://claude.ai/install.sh | bash) and authenticated (via a Claude Pro/Team subscription stored at~/.claude.jsonor theANTHROPIC_API_KEYenvironment variable).
Sync dependencies:
uv syncLaunch the server using STDIO transport:
uv run python -m fastmcp.cli run src/claude_code_mcp/server.py --transport stdioMCP Client Configuration
The server is launched by an MCP client (Trae, Cursor, Windsurf, etc.) over STDIO. The recommended setup uses uvx to install the package from a local source path on demand — no global Python install required.
Trae / Cursor / Windsurf (uvx from local source)
Add to your MCP client configuration (~/.trae/mcp.json, .cursor/mcp.json, .windsurf/mcp.json, or the IDE's MCP settings panel):
{
"mcpServers": {
"claude-code-cli-mcp": {
"command": "uvx",
"args": [
"--refresh",
"--from",
"/path/to/claude-code-cli-mcp",
"fastmcp",
"run",
"src/claude_code_mcp/server.py"
],
"cwd": "/path/to/claude-code-cli-mcp",
"env": {
"CLAUDE_MCP_MODE": "safe",
"CLAUDE_MCP_ALLOWED_ROOTS": "[\"/path/to/your/projects\"]",
"CLAUDE_MCP_FORCE_SANDBOX_IN_SAFE_MODE": "true",
"CLAUDE_MCP_ALLOWED_MODELS": "[\"sonnet\", \"opus\"]",
"START_MCP_TIMEOUT_MS": "30000",
"RUN_MCP_TIMEOUT_MS": "600000"
}
}
}
}Tip:
--refreshforcesuvxto re-resolve the local source on every start. Drop it once you stop iterating on the server.Tip:
START_MCP_TIMEOUT_MSandRUN_MCP_TIMEOUT_MSare client-side timeouts consumed by the Trae IDE (not by this server).
Relevant environment variables
The most relevant variables for the env block are listed below. See Configuration for the full reference.
Variable | Purpose |
|
|
| JSON list of workspace roots the server is allowed to access |
|
|
| JSON list of permitted model aliases / full names (e.g. |
| Enables the persistent memory layer ( |
| Base directory for the persistence layer ( |
Persistence is a two-level configuration. The env vars above only enable the feature and pick the base directory. To actually create the files and start injecting context, the orchestrator must call
claude_init_persistenceonce at startup. See Persistent Memory for the full lifecycle.
Step-by-step Trae setup (Portuguese)
For a guided walkthrough in Portuguese, see USO_TRAE.md.
Running Tests
To install dev dependencies and execute the test suite:
uv sync --extra dev
uv run pytestUsing This Server in Trae
For step-by-step instructions in Portuguese on setting up and invoking this server inside the Trae IDE, see USO_TRAE.md.
Configuration
The server is configured using environment variables prefixed with CLAUDE_MCP_ via Pydantic Settings.
Variable | Description | Default |
| Server access mode ( |
|
| JSON list of allowed workspace roots. | Current working directory |
| Path or command to invoke the |
|
| Fallback absolute paths to search for the binary. |
|
| Pass |
|
| Enforce sandbox execution in safe mode. |
|
| Permission mode flag ( |
|
| Maximum allowed execution time for tasks in seconds. |
|
| Default wait time between polling iterations. |
|
| Maximum concurrent background executions allowed. |
|
| History size of completed tasks in the run store. |
|
| Maximum stdout bytes captured from the child process. |
|
| Maximum stderr bytes captured from the child process. |
|
| JSON set of permitted model names or aliases. |
|
| JSON list of allowed env keys to pass in permissive mode. |
|
| JSON list of allowed CLI arguments in permissive mode. |
|
| Enables the persistent markdown memory layer. |
|
| Base directory for the persistence layer. |
|
| Maximum file size for persistence files before rejecting writes. |
|
| Create a |
|
| Seed default markdown files if missing on init. |
|
| Optional Logfire token for application telemetry. |
|
Model Selection
The MCP server exposes two optional parameters on claude_run_task and claude_start_task for selecting the model on a per-request basis:
model: The primary model used for the run. Accepts a short alias ("sonnet","opus","haiku") or a full model name (e.g.,"claude-sonnet-4-6").fallback_model: An optional automatic fallback model used when the primary model is overloaded. Print-mode only in the CLI — since bothclaude_run_task(sync) andclaude_start_task(async) invoke the CLI via-p, this flag is safe to use in both modes.
Examples:
{
"workspace_path": "/abs/path/to/project",
"prompt": "Refactor the auth module",
"model": "sonnet",
"fallback_model": "haiku"
}{
"workspace_path": "/abs/path/to/project",
"prompt": "Plan a migration to event sourcing",
"model": "claude-sonnet-4-6"
}Both fields are validated against the CLAUDE_MCP_ALLOWED_MODELS allowlist. When set, a model outside the allowlist returns MODEL_NOT_ALLOWED; the same check applies to fallback_model. An empty allowlist disables validation and permits any model requested.
For deeper guidance on choosing models and aliases, see the claude_model_selection_guidance prompt.
Model Selection & Timeout Policy
As of sprint N+1 (commits 3a90258..1f07117), this MCP server supports 4 Claude models and a task-class-aware timeout policy. Both features are backward-compatible additions.
Supported models
Alias | Tier | Typical cost | Latency | Multi-file safe |
| standard | ~$0.50/run | ~8 min | yes |
| mid_tier | ~$0.30/run | ~12 min | yes |
| flagship | ~$1.50/run | ~25 min | yes |
| cheap | ~$0.02/run | ~1.5 min | no (>5 files = warning) |
The CLI strings live in Settings.claude_model_aliases and are
overridable via env vars (CLAUDE_MCP_CLAUDE_MODEL_ALIASES__SONNET,
etc). Default values are placeholders (claude-sonnet-5-...) — set
these to your installed claude --list-models output before
relying on a specific alias.
Timeout policy
Each task has a complexity class (TaskClass enum, 10 values:
trivial_edit, smoke_test, single_feature, docs_update,
test_suite, review, multi_file_refactor, architecture,
migration, long_running). The helper
claude_code_mcp.timeout_policy.compute_timeout(task_class, model, files_to_edit, max_budget_usd) returns a (timeout_s, must_use_async)
recommendation.
Sync ceiling: 600s (FastMCP wrapper hard cap). Async ceiling: 3600s (Pydantic validator upper bound).
Task class | Default timeout | Sync/Async |
| 120s | Sync |
| 180s | Sync |
| 240s | Sync |
| 300s | Sync |
| 360s | Sync |
| 600s | Sync |
| 900s | Async |
| 1500s | Async |
| 1800s | Async |
| 1500s | Async |
| 3600s | Async |
The compute_timeout helper bumps these up automatically based on
files_to_edit (≥20 files → ≥900s, ≥50 files → ≥1800s) and warns
when Haiku is used on >5 files.
Using the decision matrix
Orchestrators can call the MCP prompt prompt_timeout_help (name
claude_timeout_help) to get a structured recommendation. The matrix
in the response reflects the requested model (e.g. opus shows the
opus timeouts, not sonnet's), so you can compare apples to apples.
# In your orchestrator
from claude_code_mcp.server import prompt_timeout_help
guide = prompt_timeout_help(
task_class="multi_file_refactor",
files_to_edit=20,
model_alias="opus",
)
# Returns: timeout_s=900+, must_use_async=True, decision matrix
# (rows = task classes, columns = model-specific timeouts), and a
# pre-formatted python snippet using `claude_start_task`.Or directly use the helper:
from claude_code_mcp.models import MODEL_REGISTRY, TaskClass
from claude_code_mcp.timeout_policy import compute_timeout
profile = MODEL_REGISTRY["opus"]
rec = compute_timeout(TaskClass.MULTI_FILE_REFACTOR, profile, files_to_edit=20)
if rec.must_use_async:
run_id = claude_start_task(req={"prompt": "...", "timeout_s": rec.timeout_s})
# Poll with claude_poll_task(drain=true) or via start/poll loop
else:
result = claude_run_task(req={"prompt": "...", "timeout_s": rec.timeout_s})Feature flag
The new types are always available, but the policy helper is gated
behind CLAUDE_MCP_TIMEOUT_POLICY_ENABLED=false (default OFF) for
safe rollout. Set it to true in .env to enable automatic
timeout recommendations in your orchestrator's request layer.
Sync vs async decision (recap)
claude_run_task(sync, ≤600s wrapper cap) —trivial_edit,smoke_test,review,single_feature,docs_update,test_suite, smallmulti_file_refactor(≤20 files).claude_start_task+claude_poll_task(async, ≤3600s) — largemulti_file_refactor(>20 files),architecture,migration,long_running.
If a sync call returns parse_error (response is not valid JSON),
do not retry sync — the subprocess likely returned truncated /
non-JSON output due to a wrapper-level timeout. Switch to
claude_start_task (async) which buffers output incrementally and
survives longer walls.
Security
Safe Mode (Default)
In safe mode (CLAUDE_MCP_MODE=safe):
Sandbox execution is enforced (if
CLAUDE_MCP_FORCE_SANDBOX_IN_SAFE_MODEis enabled).Custom environment overrides are blocked.
Custom extra command-line arguments are blocked.
Bypassing permissions (
bypassPermissionsordontAskpermission modes) is rejected.
Permissive Mode
In permissive mode (CLAUDE_MCP_MODE=permissive):
Environment variable overrides are permitted only if they appear in the
CLAUDE_MCP_ALLOW_ENV_KEYSlist.Extra arguments are allowed only if they appear in
CLAUDE_MCP_ALLOW_EXTRA_ARGS.Skipping permissions (
bypassPermissionsor--dangerously-skip-permissions) requires"--dangerously-skip-permissions"to be explicitly listed inCLAUDE_MCP_ALLOW_EXTRA_ARGS.
Troubleshooting
CLAUDE_NOT_FOUND
The server cannot locate the claude executable. Make sure it is installed (e.g. via the official installer) and added to your PATH. Alternatively, set CLAUDE_MCP_CLAUDE_PATH to the absolute path of the binary.
NOT_ALLOWED: workspace_path is outside allowed roots
Your workspace path is outside of the configured roots. Include the target directory in the CLAUDE_MCP_ALLOWED_ROOTS JSON array, or launch the server from within the target folder.
⚠️
CLAUDE_MCP_ALLOWED_ROOTSand other server env vars are read at MCP server STARTUP. Editing.envafter the server is running has no effect — restart the MCP server (in your client's MCP panel) for changes to apply.
NOT_LOGGED_IN / result.text == "Not logged in · Please run /login"
The claude CLI cannot find its OAuth token. The most common cause is the --bare flag being passed to the subprocess — --bare bypasses ~/.claude/ entirely, so ~/.claude.json (which holds the token) is invisible. Verify this server's force_bare=False (the default). Alternative: run claude login interactively in your shell to seed ~/.claude.json, then restart this MCP server.
MODEL_NOT_ALLOWED
The requested model (or fallback_model) is not in Settings.allowed_models (configurable via CLAUDE_MCP_ALLOWED_MODELS). Empty allowlist disables validation.
PERSISTENCE_FILE_TOO_LARGE
A persistence file has reached the maximum allowed bytes (default 1 MiB). Clean up or truncate obsolete entries in the file (~/.open-cli-router/claude-code/MEMORY.md or PROJECTS.md) to allow new writes.
CONFIRM_REQUIRED
Modifying AGENTS.md (the system prompt) in safe mode requires setting the confirm parameter to true to ensure the override is intentional.
~/.open-cli-router/claude-code/ does not exist
The persistence directory is created lazily. The server will not create it on its own — you must initialize it once via claude_init_persistence (or the Python one-liner under Persistent Memory → How to initialize). Without this, the server is enabled but uninitialized and no context is injected into prompts.
Available Tools
12 toolsclaude_append_persistenceA
Append content to one of the persistence files.
Required: file (agents|projects|memory), content. Optional: section_header (str|None) — if provided, the append is placed under a heading; otherwise content is appended at the end of the file.
Safe-mode constraint: in Settings.mode == "safe", updating AGENTS.md
requires confirm=true. The same applies to update_persistence.
Do not store secrets, credentials, or full file dumps — keep entries small and high-signal.
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| file | Yes | |
| timestamp | Yes | |
| appended_bytes | Yes | |
| new_size_bytes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral disclosure burden. It reveals the non-obvious safe-mode constraint requiring confirm=true for AGENTS.md, explains placement behavior for section_header, and adds content policy constraints. It omits error/idempotency details, but the output schema covers return expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action and requirements, then moves to optional behavior, a conditional constraint, and content policy. Every sentence carries operational value with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, the description covers target files, required vs optional parameters, the conditional confirm flag, and content constraints. Output details are left to the output schema, so nothing critical blocks correct invocation, though usage-vs-alternatives guidance is only implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema description coverage is 0%, so the description must compensate, and it does: it defines file values, marks file and content as required, explains the optional section_header behavior, and explains when confirm matters. This adds substantive meaning beyond the otherwise bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb ('Append'), a clear target resource ('one of the persistence files'), and enumerates the allowed file values ('agents|projects|memory'). The operation is semantically distinct from read/update siblings, and the safe-mode note explicitly connects to update_persistence without blurring the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for adding content to persistence files and gives required/optional fields, but it does not explicitly say when to choose append over update_persistence or read_persistence, nor does it describe conditions for when not to use it. The safe-mode note hints at a related tool but stops short of clear alternative-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_cancel_taskA
Cancel a running task.
Two escalation levels: - force=false (default): sends SIGTERM (graceful). The subprocess has ~5s to clean up before the OS escalates. Try this first. - force=true: sends SIGKILL (immediate). Use only if the subprocess doesn't respond to SIGTERM within ~5s.
Returns: ClaudeCancelTaskResponse with canceled (bool) and status ("cancelled" if actively stopped, "already_done" if run finished naturally, "not_found" if run_id is unknown).
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| status | Yes | |
| canceled | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It clearly explains the signal mechanism (SIGTERM vs SIGKILL), the ~5s cleanup window, OS escalation, and the possible return statuses. This gives the agent a solid mental model of the tool's side effects and outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized with a front-loaded purpose statement, bulleted escalation levels, and a concise return-value explanation. Every sentence adds meaningful information without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers both parameters, the escalation behavior, and the full set of return statuses. The only minor gap is the behavior when req is null, but the schema already indicates this via the nullable type. Overall, an agent has the information needed to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It thoroughly explains the force parameter, including default and recommended sequence. The run_id is only implicitly described via the 'not_found' status, but its role as the task identifier is understandable from context. Some additional detail on run_id sourcing would strengthen it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Cancel a running task.' It clearly identifies the tool's function and differentiates it from siblings like claude_run_task, claude_start_task, and claude_poll_task by being the only cancellation operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on how to use the two escalation levels, including 'Try this first' for force=false and 'Use only if the subprocess doesn't respond' for force=true. It does not explicitly name alternative tools for selection, but the usage context is clear enough for an agent to know when cancellation is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_healthA
Health check for the Claude Code CLI binary.
Verifies the claude binary is on PATH, captures its --version output,
and reports auth status (whether ~/.claude.json has an OAuth token).
Use this as the first call in any orchestration session to validate the
environment.
Optional fields: - expected_version (str): if set, returns ok=false on version mismatch
Returns: - claude_path, claude_version, ok, auth_status, notes
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| notes | No | |
| auth_status | No | |
| claude_path | Yes | |
| claude_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavior, and it delivers: it states it verifies the binary on PATH, captures --version, inspects ~/.claude.json for an OAuth token, and reports auth_status. It also defines the effect of expected_version on ok, exceeding a typical health-check disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and logically ordered: purpose, behavior, usage timing, optional parameter, return fields. No filler sentences; every line adds necessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 optional parameter, no annotations), the description covers purpose, usage context, behavior, parameters, and return values. The output schema is said to exist, so the return field list is a bonus. An agent has everything needed to call it first and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning. It explains the sole parameter expected_version, noting it is optional and that setting it returns ok=false on mismatch — exactly what an agent needs beyond the raw anyOf schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete verb+resource: 'Health check for the Claude Code CLI binary' and lists specific checks (PATH, --version, auth status). This clearly differentiates it from sibling task-execution tools like claude_run_task or claude_self_test.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to call this 'as the first call in any orchestration session to validate the environment.' This provides clear context for when to invoke, though it does not name alternatives or exclusions, preventing a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_init_persistenceA
Initialize the persistence directory and seed the three markdown files.
Idempotent: re-running without force=true is a no-op if files already
exist. Creates the directory at the location resolved by
Settings.resolve_persistence_base_dir() and writes AGENTS.md,
PROJECTS.md, MEMORY.md (unless they already exist).
Optional fields: - force (bool): re-create files even if they exist - seed_templates (bool|None): whether to seed the default templates
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| created | No | |
| base_dir | Yes | |
| seed_version | Yes | |
| already_existed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and it delivers: it discloses idempotency, the no-op behavior when files exist, directory creation, the exact files written, and the effect of force. This goes well beyond a simple 'init' label.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action statement is front-loaded, followed by a tight idempotency note and a compact optional-fields list. No sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity initialization tool with zero required parameters and an output schema, the description covers location, file names, idempotency, force behavior, and template seeding. Nothing essential for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It names force and seed_templates and gives a functional meaning for each, although it leaves the null-vs-false distinction for seed_templates slightly implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Begins with a specific verb and resource: 'Initialize the persistence directory and seed the three markdown files' and even names AGENTS.md, PROJECTS.md, MEMORY.md. This clearly separates it from sibling persistence tools that read, append, update, or load.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context through the idempotency statement: re-running is a no-op unless force=true, and explains what force does. It does not explicitly compare against alternatives or state when not to use it, so it falls just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_list_runsA
List recent runs (active + recently completed).
Returns ClaudeRunSummary entries ordered newest-first, with active
claude_start_task runs at the top followed by completed/cancelled/
timed-out runs from the in-memory store (bounded by Settings.max_runs).
Use this for recovery after orchestrator restart: active async runs
survive across MCP client restarts and can be polled/cancelled via
claude_poll_task and claude_cancel_task using the run_id from this
listing.
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| runs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals ordering (newest-first, active at top), source (in-memory store), bound (Settings.max_runs), and cross-restart persistence. It does not explicitly state read-only behavior, but 'list' strongly implies that, and 'recent' is not precisely defined—minor gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all informative: the scope, the ordering/source/bound detail, and the recovery use case. The text is front-loaded and wastes no words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main behavior, ordering, bound, and recovery workflow, and an output schema exists so return-value detail is unnecessary. It is incomplete only in omitting the optional limit parameter and not explicitly confirming that this is a read-only operation—minor for a tool that works fine with no arguments.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the req.limit parameter, but it never mentions it. The only bound referenced is Settings.max_runs, which is a different limit. An agent reading the description would have no idea that an optional limit parameter with default 50 exists, especially since the schema itself provides no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource ('List recent runs') and immediately scopes it to active plus recently completed. It distinguishes the tool from siblings by referencing claude_start_task, claude_poll_task, and claude_cancel_task for different roles, and specifies the returned ClaudeRunSummary entries with ordering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: 'Use this for recovery after orchestrator restart.' It goes further to explain that active async runs survive restarts and can be polled/cancelled via claude_poll_task and claude_cancel_task, guiding the complete workflow. No alternative list tool exists among siblings, so no exclusion is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_load_persistence_contextA
Load the persistence files as context for the current session.
Returns head+tail excerpts of each file (head_ratio controlled by Settings.persistence_truncation_head_ratio). Use this at the start of each session to hydrate orchestrator memory before dispatching tasks.
Optional fields: - include (list[str]|None): subset of files to load (agents|projects|memory). None = all three. - max_chars_per_file (int): per-file char cap (default from settings).
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| base_dir | No | |
| initialized | No | |
| total_chars | No | |
| agents_excerpt | No | |
| memory_excerpt | No | |
| truncated_flags | No | |
| projects_excerpt | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses useful behavioral traits: returns head+tail excerpts, head_ratio is settings-controlled, and max_chars_per_file caps output. It does not directly state that the operation is non-mutating or describe failure behavior, but for a context-loading operation this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and uses a clear optional-fields section. There is no filler or repetition of schema structure; every sentence adds functional value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple optional-parameter interface, the description covers purpose, usage timing, return truncation behavior, and parameter semantics. It does not mention what happens if persistence files are missing or how this differs from claude_read_persistence, but the presence of an output schema and the startup-context guidance make this sufficiently complete for agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add parameter meaning. It does: 'include' is a subset of files (agents|projects|memory), None means all three, and max_chars_per_file is a per-file character cap. It slightly defers the default to settings while the schema specifies 20000, but overall it compensates well for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Load the persistence files as context for the current session.' It clearly identifies what the tool does and adds return-behavior detail, but it does not explicitly differentiate itself from the sibling claude_read_persistence, so the agent must infer the distinction from context signals.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use guidance: 'Use this at the start of each session to hydrate orchestrator memory before dispatching tasks.' However, it does not provide when-not-to-use guidance or name an alternative tool such as claude_read_persistence, so the exclusion side is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_poll_taskA
Poll an asynchronous task.
Two polling modes: - drain=false (default): returns immediately with current state (new messages, stdout/stderr byte counts, elapsed time). Use for tight control loops with explicit backoff. - drain=true: blocks until status is no longer 'running' (fire-and-wait). Note: this is bounded by your MCP client's request timeout, not the async task's timeout_s. For long blocks, prefer async drain with orchestrator-level polling.
Returns: ClaudePollTaskResponse with status (running|done|error|timeout|cancelled), new_messages, stdout_len/stderr_len, elapsed_seconds, and (once terminal) the full result with total_cost_usd, model_usage, and changes.
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| notes | No | |
| result | No | |
| run_id | Yes | |
| status | Yes | |
| stderr_len | No | |
| stdout_len | No | |
| new_messages | No | |
| elapsed_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries behavioral disclosure. It states that drain=false returns immediately, drain=true blocks and is bounded by the client request timeout, and it lists the return fields and terminal statuses. This is transparent about the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with bullet points and a return section, making it easy to scan. It front-loads the core purpose and then details modes, without unnecessary sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return details are partially covered, but the description still outlines the response structure. It covers the main behavioral differences and timeout bound. Missing details like wait_seconds are minor, making it sufficiently complete for a polling tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It thoroughly explains the 'drain' parameter (the core behavior) but does not mention 'wait_seconds' or the nullable 'req' field. The omission of wait_seconds leaves a partial gap in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Poll an asynchronous task' with a specific verb and resource. It distinguishes itself from siblings like claude_start_task, claude_cancel_task, and claude_list_runs, leaving no ambiguity about its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly describes two polling modes—drain=false for tight control loops with explicit backoff and drain=true for fire-and-wait—and advises orchestrator-level polling for long blocks. While it doesn't name alternatives directly, the mode guidance provides clear when-to-use context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_read_persistenceA
Read one of the three persistence files (agents | projects | memory).
Optional fields: - file (str): which file to read (default: memory) - offset (int): start reading from line N (0-indexed) - limit (int|None): max lines to return (None = no limit)
Large files are automatically truncated at
Settings.persistence_max_file_bytes; the response includes a
truncated flag if this happens.
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| file | Yes | |
| content | Yes | |
| truncated | Yes | |
| size_bytes | Yes | |
| modified_at | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden well by disclosing line-based offset/limit behavior, automatic truncation at Settings.persistence_max_file_bytes, and the truncated flag in the response. It does not cover error handling or file-not-found behavior, but the most operationally surprising behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and well-structured: a one-line purpose, a short bullet list for optional parameters, and a single note about truncation. Nothing is redundant with the schema, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers parameters, defaults, and truncation behavior, and an output schema exists for return-value details. It could be slightly more complete by naming sibling tools or clarifying when load_persistence_context is preferable, but nothing critical is missing for calling the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema. It explains all parameters clearly: file choices and default, offset being 0-indexed, and limit accepting None for no limit. This adds real meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource ('Read one of the three persistence files') and enumerates the allowed files: agents, projects, memory. It is clear, though it does not explicitly distinguish itself from the sibling claude_load_persistence_context, which may also involve reading persistence data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when raw contents of the agents, projects, or memory persistence files are needed. However, it provides no explicit when-not-to-use guidance or mention of alternatives such as load_persistence_context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_run_taskA
Run a single Claude task synchronously (bounded by FastMCP's 600s wrapper cap).
Required: workspace_path, prompt. Common options: model (alias from Settings.claude_model_aliases), fallback_model, permission_mode, options.timeout_s (default 300, max 600), capture_changes.
Estimated cost depends on model and prompt size — check total_cost_usd
in the response. Typical ranges: haiku ~$0.02, sonnet ~$0.50, opus ~$1.50.
For tasks that may exceed 600s (architecture, migration, large refactors),
use claude_start_task (async) instead. For timeout routing, call the
claude_timeout_help prompt first.
Returns: ClaudeRunTaskResponse with result, stdout/stderr, exit_code, timed_out, total_cost_usd, model_usage, and (if capture_changes=true) a unified git diff of workspace changes.
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| notes | No | |
| result | Yes | |
| run_id | Yes | |
| status | Yes | |
| changes | No | |
| exit_code | Yes | |
| num_turns | No | |
| final_text | Yes | |
| session_id | No | |
| duration_ms | No | |
| model_usage | No | |
| total_cost_usd | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses synchronous execution, the 600s cap, cost expectations, and the full response shape including timed_out and optional git diff, which implies workspace mutation. It does not explicitly warn about side effects or permission risks beyond mentioning permission_mode, but the disclosure is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose, then structured into required/common, cost, routing, and return sections. Every sentence adds information, and the line breaks make it scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the rich detail, the description omits the required `req` object wrapper, which is a critical invocation detail. It also contains a factual contradiction with the schema on default and maximum timeout values. Given the tool's complexity, the undocumented advanced parameters further reduce completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description names required and common options, but it does not mention the top-level `req` wrapper that the schema actually requires, so an agent may pass workspace_path/prompt incorrectly. Its timeout_s claim (default 300, max 600) directly contradicts the schema (default 600, max 3600), and with 0% schema coverage, many parameters (effort, max_turns, change_scope, max_budget_usd, etc.) are left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Run a single Claude task synchronously' and explicitly notes the 600s wrapper cap. It clearly differentiates from the async sibling claude_start_task, so an agent can immediately tell this is the sync execution tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly lists required and common options and gives concrete routing rules: use claude_start_task for tasks exceeding 600s, and call claude_timeout_help for timeout routing. This is exactly the when-to-use-vs-alternatives guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_self_testA
Inspect every registered tool's input schema and report robustness.
This is a metadata-only check — no tools are actually invoked.
Args shape:
The MCP client MUST pass arguments wrapped in a req object:
{"req": {"include": ["claude_health"], "only_show_tolerant": true}}
For backwards-compatibility, the server also accepts args={}.
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| tools | No | |
| summary | Yes | |
| server_info | No | |
| total_tools | Yes | |
| tolerant_count | Yes | |
| requires_req_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that no tools are actually invoked, which is a key safety trait. It also discloses the exact argument wrapping requirement (the 'req' object) and the backwards-compatibility fallback ('args={}'), giving the agent precise calling behavior. It does not cover error handling or output format, but those are partially covered by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with a clear purpose statement, a note on metadata-only behavior, and then the argument shape. It includes an example and a note about backwards compatibility, which are useful. It is somewhat verbose but each part contributes to understanding the tool. The key information is front-loaded with the purpose and the no-invocation note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one parameter with a nested structure and no schema descriptions, the description is not fully complete. It explains the wrapper structure but does not define the meaning of 'include' or 'only_show_tolerant', which are essential for correct use. The existence of an output schema covers return values, but the parameter semantics are left underspecified, making it less than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It explains the overall shape: arguments must be wrapped in a 'req' object, and it provides an example with 'include' and 'only_show_tolerant'. However, it does not explain the semantics of these fields (what 'include' filters or what 'tolerant' means), leaving some ambiguity. The example helps, but the meaning of the fields is only implied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: inspecting every registered tool's input schema and reporting robustness. It is specific with a verb ('inspect') and a resource ('registered tool's input schema'), and it explicitly distinguishes itself as a metadata-only check that does not invoke tools. This differentiates it from sibling tools like claude_health or claude_run_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives, nor does it mention any exclusions or conditions. It implies that it is for schema inspection and robustness checking, but it does not provide guidance on when a user would prefer this over other tools. The absence of any alternative comparison leaves the agent to infer usage from the purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_start_taskA
Start a Claude Code CLI task asynchronously (up to 3600s).
Returns immediately with a run_id. The subprocess continues running in
the background; use claude_poll_task to monitor and claude_cancel_task
to stop. Use this for any task that may exceed 600s (architecture,
migration, large multi-file refactors, long-running data agents).
Required: workspace_path, prompt. Optional: model, fallback_model, permission_mode, options.timeout_s (default 300, max 3600 for async), capture_changes.
Concurrent run limit: bounded by Settings.max_concurrent_runs. New calls beyond that limit return MAX_CONCURRENT_RUNS_EXCEEDED.
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| run_id | Yes | |
| session_id | Yes | |
| started_at | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does well: async background execution, immediate return, timeout boundaries, and the MAX_CONCURRENT_RUNS_EXCEEDED error are all disclosed. It could go further on side effects or cleanup expectations, but the core behavioral traits are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly written and well structured, with the most important operational facts front-loaded. Each paragraph earns its place: async behavior, return semantics, when to use it, required parameters, and concurrency limits.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main async workflow and monitoring guidance well, but the tool is complex with many optional parameters and no annotations. It omits the likely synchronous alternative claude_run_task and does not clarify how to pass required fields inside the 'req' object, which is a meaningful invocation gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only names a subset of parameters and does not explain the required 'req' object wrapper. More importantly, it states options.timeout_s default is 300 while the schema says 600, creating a direct contradiction that could lead to incorrect invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Start a Claude Code CLI task asynchronously'. It clearly distinguishes itself from siblings by emphasizing async execution, immediate return of a run_id, and the ability to run up to 3600s.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance for when to use this tool: 'Use this for any task that may exceed 600s' with concrete examples. It also directs the agent to claude_poll_task and claude_cancel_task for lifecycle management. However, it does not explicitly contrast with claude_run_task, which appears to be the likely synchronous alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_update_persistenceA
Replace or append to a section in one of the persistence files.
Required: file, section_anchor, new_content. Optional: mode (replace|append) — replace the entire section vs append inside it (default replace).
Safe-mode constraint: in Settings.mode == "safe", updating AGENTS.md
requires confirm=true. Returns matched=true when the section_anchor
was found; matched=false when the anchor was not found and no edit
happened (so you can detect typos before silent appending).
| Name | Required | Description | Default |
|---|---|---|---|
| req | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| file | Yes | |
| matched | Yes | |
| new_size_bytes | Yes | |
| section_anchor | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it does well: it discloses the safe-mode constraint, the default mode, and the exact matched=true/matched=false return semantics including that no edit happens when the anchor is not found. It does not discuss destructive risk beyond the word 'replace' or atomicity/ordering details, but the key failure behavior is exposed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary action and uses short labeled lines for required/optional/safe-mode/return behavior. Each section adds information, though the safe-mode paragraph is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating persistence tool with no annotations, it covers the essential edge cases: anchor not found, safe mode, mode default, and return signal. An output schema exists for the return shape, so not explaining all return fields is acceptable. The nested req wrapper is the main missing call-construction detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters, and it does: it defines mode's replace vs append behavior, explains when confirm=true is required, and clarifies section_anchor matching. The only gap is that it lists file/section_anchor/new_content as top-level required inputs while the schema nests them under a req object, which could confuse an agent constructing the call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Replace or append to a section in one of the persistence files.' This clearly conveys the tool's core operation. It does not explicitly distinguish itself from the sibling claude_append_persistence, but the mode parameter partially covers that overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It lists required and optional parameters and explains the safe-mode confirm condition, which is actionable guidance. However, it does not say when to prefer this tool over claude_append_persistence or other persistence tools; the alternative routing is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.2.0- First observed
claude_append_persistence - First observed
claude_cancel_task - First observed
claude_health - First observed
claude_init_persistence - First observed
claude_list_runs - First observed
claude_load_persistence_context - First observed
claude_poll_task - First observed
claude_read_persistence - First observed
claude_run_task - First observed
claude_self_test - First observed
claude_start_task - First observed
claude_update_persistence
TDQS
Scored across 12 tools
The task lifecycle tools (run/start/poll/cancel/list) are clearly separated by sync vs async execution and state transitions, and the persistence tools have distinct read/append/update/load operations. The only mild ambiguity is between run_task and start_task, and between update_persistence and append_persistence, but the descriptions explicitly route usage.
All tools share a consistent claude_ prefix and snake_case style, with most following a verb_noun pattern like claude_run_task, claude_poll_task, and claude_read_persistence. claude_health is a noun-only outlier, and claude_load_persistence_context is less consistent with the shorter read/append/update names, but the overall pattern remains predictable.
Twelve tools is well-scoped for a server covering Claude Code CLI task execution and persistent memory files, with no redundant duplicates. Each tool addresses a distinct operational need within those two clear domains.
The surface covers the full task lifecycle: synchronous run, asynchronous start, poll, cancel, and list, plus health verification and self-test. Persistence is also complete with init, read, append, update, and load-context operations, leaving no obvious dead ends for the stated purpose.
Maintenance
Related MCP Connectors
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Augments MCP Server - A comprehensive framework documentation provider for Claude Code
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA server implementation for the Model Context Protocol (MCP) that allows Claude AI to execute commands through a command-line interface, enabling direct system interactions from within Claude.-
- AlicenseAqualityCmaintenanceLocal MCP server that wraps the headless Claude Code CLI as MCP tools, providing stateless access to Claude's coding capabilities through prompt-based interactions. It enables users to execute Claude Code commands with various prompt formats and structured outputs directly from MCP clients.3MIT
- AlicenseAqualityDmaintenanceA Model Context Protocol (MCP) server that provides a bridge to Anthropic's Claude CLI. It allows MCP-compliant clients like Claude Desktop or Gemini to start new chat sessions or continue existing ones.287 npmISC
- FlicenseNot gradedqualityDmaintenanceExposes headless Claude Code as a remote MCP server with a voice client, enabling hands-free task execution and session management via OpenAI's Realtime API.-