Skip to main content
Glama
bill-kopp-ai-dev

Claude Code CLI MCP Server

Claude Code CLI MCP Server

Version: 0.2.0 (post-sprint-N+1)

A local STDIO MCP server that exposes tools and reusable prompts for running the Anthropic Claude Code CLI (claude) inside a controlled workspace.

Related MCP server: claudecode-mcp

Why This Project Exists

This project is a multi-provider fork of agy-mcp-server (the sister project wrapping the Google Antigravity CLI). Both projects share the same underlying architecture to provide a thin, secure CLI shim that exposes terminal-based AI tools as Model Context Protocol (MCP) servers.

Using the Claude Code CLI directly inside an editor or orchestrator workspace poses security and configuration challenges. This MCP server encapsulates the claude process, enforcing directory access controls, standardizing outputs, and providing a file-based persistent memory layer so that system prompts and session history survive across project reloads.

Now you can leverage your Claude Pro/Team subscription or ANTHROPIC_API_KEY seamlessly inside any MCP-enabled IDE:

  • Cursor

  • Windsurf

  • Trae

Features

Tools

  • claude_health: Checks that the claude binary is installed, authenticated, and returns its version and authentication status.

  • claude_run_task: Runs a synchronous (blocking) task inside the target workspace directory. Accepts model and fallback_model to select the Claude model per-run.

  • claude_start_task: Starts an asynchronous background task inside the target workspace. Accepts the same model and fallback_model parameters.

  • claude_poll_task: Polls the status, stdout, stderr, and output changes of an active background task.

  • claude_cancel_task: Cancels a running background task (supporting an optional force signal).

  • claude_list_runs: Lists recent tasks managed by the server and their statuses.

  • claude_init_persistence: Initializes the persistence layer folder and creates template markdown files.

  • claude_read_persistence: Reads the contents of a persistent memory file (AGENTS.md, PROJECTS.md, or MEMORY.md).

  • claude_append_persistence: Appends high-signal updates (e.g., session memories) to a persistence file.

  • claude_update_persistence: Replaces or appends to a specific section of a persistence file by heading anchor.

  • claude_load_persistence_context: Loads the persistent files as truncated excerpts to inject into session context.

Prompts

  • claude_sync_orchestration: Guidance playbook for executing synchronous tasks safely via claude_run_task. Includes parse-error / timeout / list_runs recovery recipes.

  • claude_async_orchestration: Playbook for orchestrating background tasks via claude_start_task + claude_poll_task + claude_cancel_task. Includes backoff schedule (1s → 10s) and post-restart recovery via claude_list_runs.

  • claude_model_selection_guidance: Lists all 4 model aliases (sonnet / fable / opus / haiku) with tier, cost, and multi-file-safety annotations. Source of truth: MODEL_REGISTRY.

  • claude_security_and_workspace_rules: Summarizes sandbox limits and rules for safe and permissive environments.

  • claude_persistence_protocol: Guides the orchestrator on maintaining the persistent memory files (load context, append session notes, update AGENTS.md with confirm gate).

  • claude_timeout_help: Decision matrix + code snippet for (task_class, files_to_edit, model_alias) → (timeout_s, must_use_async). Matrix reflects the requested model, not always sonnet.

  • claude_quickstart: First-call cheat-sheet for new orchestrators (workspace_path discipline, tool catalog, common gotchas + troubleshoot).

  • claude_troubleshoot: Pattern-matches an error string and returns a canonical fix recipe (NOT_ALLOWED, NOT_LOGGED_IN, MODEL_NOT_ALLOWED, CLAUDE_NOT_FOUND, etc.).

Companion Agent: Femtobot

Recommendation: this server is the canonical claude_* tool source for femtobot, the CLI-first AI agent foundation in the percival.OS ecosystem.

Femtobot ships first-class support for the tools exposed here:

  • /mcp slash command — status, reload, tools claude-code-cli-mcp, restart claude-code-cli-mcp for runtime inspection and recovery without restarting the agent.

  • mcp-router builtin skill — teaches the LLM when to delegate to claude_run_task vs. solving locally with read_file, apply_patch, etc.

  • Capability tags — tool hints show [long-running, safe-mode:confirm] so the model recognizes the confirm gate before invoking claude_run_task.

  • Workspace auto-fill — claude_run_task calls get workspace_path filled in automatically from the active request context.

  • System-prompt block — ## MCP Servers in this workspace lists this server and its tools so the model sees them at planning time.

  • CLAUDE_MCP_PERSISTENCE_LOCATION=workspace is auto-resolved by femtobot to <cwd_parent>/.open-cli-router/claude-code-cli-mcp/, so persistence files (AGENTS.md, MEMORY.md, PROJECTS.md) end up inside your project when configured that way.

See femtobot/docs/mcp.md §8 "Femtobot-specific patterns" for the full integration reference, and the CLI-router-project analysis for the design rationale.

claude_self_test

Inspect every registered tool's input schema and report robustness. This is a metadata-only check — no tools are invoked, no subprocess is spawned, no quota is consumed. Safe to run in production or CI as a sanity probe.

Output: per-tool schema introspection (allow_extra_keys vs allow type), focused on accidental strictness regressions.

Persistent Memory

This MCP server features a file-based persistence layer stored by default under ~/.open-cli-router/claude-code/. The directory contains three editable markdown files:

  1. AGENTS.md: The system prompt or agent persona instructions.

  2. PROJECTS.md: Summaries and structures of active projects.

  3. MEMORY.md: Chronological log of high-signal session takeaways and permanent learning.

When persistence is enabled, the server automatically loads these files and prepends their formatted excerpts to the prompt sent to the claude CLI, ensuring cross-session state persistence.

Storage location: global vs workspace

The persistence directory can live in two places, controlled by CLAUDE_MCP_PERSISTENCE_LOCATION:

Mode

Location

Use case

global (default)

~/.open-cli-router/claude-code/

User-level, persists across projects, survives cd

workspace

<cwd_parent>/.open-cli-router/claude-code/

Project-level, can be committed (use .gitignore!), portable with the repo

<cwd_parent> is the parent of the server's CWD — for a typical setup, the server's CWD is the server project directory (e.g. /home/user/CLI-router-project/claude-code-cli-mcp), so the workspace mode resolves to /home/user/CLI-router-project/.open-cli-router/claude-code/.

Escape hatch: CLAUDE_MCP_PERSISTENCE_BASE_DIR="$cwd_parent/.my-persistence" lets you pick any custom subdirectory under the workspace root.

⚠️ When using workspace mode, add .open-cli-router/ to .gitignore to avoid accidentally committing agent memory to source control.

Two-level configuration

Persistence is controlled by both server-level environment variables and runtime MCP tool calls:

Server level (set in the MCP client env block — see MCP Client Configuration):

Variable

Purpose

Default

CLAUDE_MCP_PERSISTENCE_ENABLED

Master switch for the persistence feature

true

CLAUDE_MCP_PERSISTENCE_LOCATION

"global" (in ~) or "workspace" (in <cwd_parent>/.open-cli-router/)

global

CLAUDE_MCP_PERSISTENCE_BASE_DIR

Base directory; the namespace claude-code is appended automatically. Supports $cwd_parent token for custom workspace paths.

~/.open-cli-router

CLAUDE_MCP_PERSISTENCE_MAX_FILE_BYTES

Maximum size per file before writes are rejected

524288 (512 KiB; aligned with agy in Phase 5)

CLAUDE_MCP_PERSISTENCE_BACKUP_ON_WRITE

Create .bak before each modification

false

CLAUDE_MCP_PERSISTENCE_BACKUP_KEEP

Number of .bak files to retain per source file (rotation)

10

CLAUDE_MCP_PERSISTENCE_SEED_TEMPLATES

Seed default markdown content when initializing

true

CLAUDE_MCP_PERSISTENCE_TRUNCATION_HEAD_RATIO

Fraction of max_chars_per_file preserved at head (rest is tail). Lower = more recency.

0.2 (20% head / 80% tail)

Runtime level (called by the orchestrator via MCP tools):

  1. Initialize once — call claude_init_persistence to create the directory and seed the three files.

  2. Load context — call claude_load_persistence_context at the start of each session to inject excerpts into the next prompt.

  3. Append session notes — after meaningful work, call claude_append_persistence on MEMORY.md.

  4. Update structured sections — when the user changes AGENTS.md or PROJECTS.md, call claude_update_persistence to persist. Note: updating AGENTS.md in safe mode requires confirm=true.

Without step 1, persistence is enabled but uninitialized — the server will not inject any context until the directory exists.

How to initialize the persistence directory

The persistence directory is created lazily — it does not exist by default. There are two equivalent ways to create it:

Option A — via MCP (recommended, normal flow):

Once both the MCP client configuration and the server are running, ask the orchestrator agent to call:

Please call claude_init_persistence to create the persistence directory.

The tool seeds AGENTS.md, PROJECTS.md, MEMORY.md, and a .initialized marker under ~/.open-cli-router/claude-code/.

Option B — directly via Python (one-shot, useful for verification or first-time setup):

From the project root (claude-code-cli-mcp/):

uv run python -c "from claude_code_mcp.persistence import PersistenceStore; from pathlib import Path; PersistenceStore(base_dir=Path.home()/'.open-cli-router', max_file_bytes=524288, backup_on_write=False, seed_templates=True).init()"

This call is idempotent — running it twice does not destroy existing data unless you pass force=True.

The seed templates are written in English so they can be edited by any language-aware agent later.

Quickstart

Prerequisites:

  • Python 3.11+

  • uv installed

  • claude CLI installed (via curl -fsSL https://claude.ai/install.sh | bash) and authenticated (via a Claude Pro/Team subscription stored at ~/.claude.json or the ANTHROPIC_API_KEY environment variable).

Sync dependencies:

uv sync

Launch the server using STDIO transport:

uv run python -m fastmcp.cli run src/claude_code_mcp/server.py --transport stdio

MCP Client Configuration

The server is launched by an MCP client (Trae, Cursor, Windsurf, etc.) over STDIO. The recommended setup uses uvx to install the package from a local source path on demand — no global Python install required.

Trae / Cursor / Windsurf (uvx from local source)

Add to your MCP client configuration (~/.trae/mcp.json, .cursor/mcp.json, .windsurf/mcp.json, or the IDE's MCP settings panel):

{
  "mcpServers": {
    "claude-code-cli-mcp": {
      "command": "uvx",
      "args": [
        "--refresh",
        "--from",
        "/path/to/claude-code-cli-mcp",
        "fastmcp",
        "run",
        "src/claude_code_mcp/server.py"
      ],
      "cwd": "/path/to/claude-code-cli-mcp",
      "env": {
        "CLAUDE_MCP_MODE": "safe",
        "CLAUDE_MCP_ALLOWED_ROOTS": "[\"/path/to/your/projects\"]",
        "CLAUDE_MCP_FORCE_SANDBOX_IN_SAFE_MODE": "true",
        "CLAUDE_MCP_ALLOWED_MODELS": "[\"sonnet\", \"opus\"]",
        "START_MCP_TIMEOUT_MS": "30000",
        "RUN_MCP_TIMEOUT_MS": "600000"
      }
    }
  }
}

Tip: --refresh forces uvx to re-resolve the local source on every start. Drop it once you stop iterating on the server.

Tip: START_MCP_TIMEOUT_MS and RUN_MCP_TIMEOUT_MS are client-side timeouts consumed by the Trae IDE (not by this server).

Relevant environment variables

The most relevant variables for the env block are listed below. See Configuration for the full reference.

Variable

Purpose

CLAUDE_MCP_MODE

safe (default) or permissive

CLAUDE_MCP_ALLOWED_ROOTS

JSON list of workspace roots the server is allowed to access

CLAUDE_MCP_FORCE_SANDBOX_IN_SAFE_MODE

true enforces sandboxing under safe mode

CLAUDE_MCP_ALLOWED_MODELS

JSON list of permitted model aliases / full names (e.g. ["sonnet", "opus"])

CLAUDE_MCP_PERSISTENCE_ENABLED

Enables the persistent memory layer (true by default)

CLAUDE_MCP_PERSISTENCE_BASE_DIR

Base directory for the persistence layer (~/.open-cli-router)

Persistence is a two-level configuration. The env vars above only enable the feature and pick the base directory. To actually create the files and start injecting context, the orchestrator must call claude_init_persistence once at startup. See Persistent Memory for the full lifecycle.

Step-by-step Trae setup (Portuguese)

For a guided walkthrough in Portuguese, see USO_TRAE.md.

Running Tests

To install dev dependencies and execute the test suite:

uv sync --extra dev
uv run pytest

Using This Server in Trae

For step-by-step instructions in Portuguese on setting up and invoking this server inside the Trae IDE, see USO_TRAE.md.

Configuration

The server is configured using environment variables prefixed with CLAUDE_MCP_ via Pydantic Settings.

Variable

Description

Default

CLAUDE_MCP_MODE

Server access mode (safe or permissive).

"safe"

CLAUDE_MCP_ALLOWED_ROOTS

JSON list of allowed workspace roots.

Current working directory

CLAUDE_MCP_CLAUDE_PATH

Path or command to invoke the claude CLI.

"claude"

CLAUDE_MCP_CLAUDE_PATH_FALLBACKS

Fallback absolute paths to search for the binary.

["/usr/local/bin/claude", "/opt/homebrew/bin/claude", "~/.local/bin/claude"]

CLAUDE_MCP_FORCE_BARE

Pass --bare to disable local workspace profiles and hooks.

true

CLAUDE_MCP_FORCE_SANDBOX_IN_SAFE_MODE

Enforce sandbox execution in safe mode.

true

CLAUDE_MCP_DEFAULT_PERMISSION_MODE

Permission mode flag (default, acceptEdits, plan, dontAsk, bypassPermissions).

"acceptEdits"

CLAUDE_MCP_DEFAULT_TIMEOUT_S

Maximum allowed execution time for tasks in seconds.

600

CLAUDE_MCP_POLL_DEFAULT_WAIT_SECONDS

Default wait time between polling iterations.

0.5

CLAUDE_MCP_MAX_CONCURRENT_RUNS

Maximum concurrent background executions allowed.

10

CLAUDE_MCP_MAX_RUNS

History size of completed tasks in the run store.

50

CLAUDE_MCP_MAX_STDOUT_BYTES

Maximum stdout bytes captured from the child process.

1000000 (1 MiB)

CLAUDE_MCP_MAX_STDERR_BYTES

Maximum stderr bytes captured from the child process.

200000

CLAUDE_MCP_ALLOWED_MODELS

JSON set of permitted model names or aliases.

["sonnet", "opus"]

CLAUDE_MCP_ALLOW_ENV_KEYS

JSON list of allowed env keys to pass in permissive mode.

[]

CLAUDE_MCP_ALLOW_EXTRA_ARGS

JSON list of allowed CLI arguments in permissive mode.

[]

CLAUDE_MCP_PERSISTENCE_ENABLED

Enables the persistent markdown memory layer.

true

CLAUDE_MCP_PERSISTENCE_BASE_DIR

Base directory for the persistence layer.

~/.open-cli-router

CLAUDE_MCP_PERSISTENCE_MAX_FILE_BYTES

Maximum file size for persistence files before rejecting writes.

1048576 (1 MiB)

CLAUDE_MCP_PERSISTENCE_BACKUP_ON_WRITE

Create a .bak backup copy of files before modification.

false

CLAUDE_MCP_PERSISTENCE_SEED_TEMPLATES

Seed default markdown files if missing on init.

true

CLAUDE_MCP_LOGFIRE_TOKEN

Optional Logfire token for application telemetry.

None

Model Selection

The MCP server exposes two optional parameters on claude_run_task and claude_start_task for selecting the model on a per-request basis:

  • model: The primary model used for the run. Accepts a short alias ("sonnet", "opus", "haiku") or a full model name (e.g., "claude-sonnet-4-6").

  • fallback_model: An optional automatic fallback model used when the primary model is overloaded. Print-mode only in the CLI — since both claude_run_task (sync) and claude_start_task (async) invoke the CLI via -p, this flag is safe to use in both modes.

Examples:

{
  "workspace_path": "/abs/path/to/project",
  "prompt": "Refactor the auth module",
  "model": "sonnet",
  "fallback_model": "haiku"
}
{
  "workspace_path": "/abs/path/to/project",
  "prompt": "Plan a migration to event sourcing",
  "model": "claude-sonnet-4-6"
}

Both fields are validated against the CLAUDE_MCP_ALLOWED_MODELS allowlist. When set, a model outside the allowlist returns MODEL_NOT_ALLOWED; the same check applies to fallback_model. An empty allowlist disables validation and permits any model requested.

For deeper guidance on choosing models and aliases, see the claude_model_selection_guidance prompt.

Model Selection & Timeout Policy

As of sprint N+1 (commits 3a90258..1f07117), this MCP server supports 4 Claude models and a task-class-aware timeout policy. Both features are backward-compatible additions.

Supported models

Alias

Tier

Typical cost

Latency

Multi-file safe

sonnet

standard

~$0.50/run

~8 min

yes

fable

mid_tier

~$0.30/run

~12 min

yes

opus

flagship

~$1.50/run

~25 min

yes

haiku

cheap

~$0.02/run

~1.5 min

no (>5 files = warning)

The CLI strings live in Settings.claude_model_aliases and are overridable via env vars (CLAUDE_MCP_CLAUDE_MODEL_ALIASES__SONNET, etc). Default values are placeholders (claude-sonnet-5-...) — set these to your installed claude --list-models output before relying on a specific alias.

Timeout policy

Each task has a complexity class (TaskClass enum, 10 values: trivial_edit, smoke_test, single_feature, docs_update, test_suite, review, multi_file_refactor, architecture, migration, long_running). The helper claude_code_mcp.timeout_policy.compute_timeout(task_class, model, files_to_edit, max_budget_usd) returns a (timeout_s, must_use_async) recommendation.

Sync ceiling: 600s (FastMCP wrapper hard cap). Async ceiling: 3600s (Pydantic validator upper bound).

Task class

Default timeout

Sync/Async

smoke_test

120s

Sync

trivial_edit, review

180s

Sync

docs_update

240s

Sync

single_feature

300s

Sync

test_suite

360s

Sync

multi_file_refactor (S)

600s

Sync

multi_file_refactor (M)

900s

Async

multi_file_refactor (L)

1500s

Async

architecture

1800s

Async

migration

1500s

Async

long_running

3600s

Async

The compute_timeout helper bumps these up automatically based on files_to_edit (≥20 files → ≥900s, ≥50 files → ≥1800s) and warns when Haiku is used on >5 files.

Using the decision matrix

Orchestrators can call the MCP prompt prompt_timeout_help (name claude_timeout_help) to get a structured recommendation. The matrix in the response reflects the requested model (e.g. opus shows the opus timeouts, not sonnet's), so you can compare apples to apples.

# In your orchestrator
from claude_code_mcp.server import prompt_timeout_help

guide = prompt_timeout_help(
    task_class="multi_file_refactor",
    files_to_edit=20,
    model_alias="opus",
)
# Returns: timeout_s=900+, must_use_async=True, decision matrix
# (rows = task classes, columns = model-specific timeouts), and a
# pre-formatted python snippet using `claude_start_task`.

Or directly use the helper:

from claude_code_mcp.models import MODEL_REGISTRY, TaskClass
from claude_code_mcp.timeout_policy import compute_timeout

profile = MODEL_REGISTRY["opus"]
rec = compute_timeout(TaskClass.MULTI_FILE_REFACTOR, profile, files_to_edit=20)
if rec.must_use_async:
    run_id = claude_start_task(req={"prompt": "...", "timeout_s": rec.timeout_s})
    # Poll with claude_poll_task(drain=true) or via start/poll loop
else:
    result = claude_run_task(req={"prompt": "...", "timeout_s": rec.timeout_s})

Feature flag

The new types are always available, but the policy helper is gated behind CLAUDE_MCP_TIMEOUT_POLICY_ENABLED=false (default OFF) for safe rollout. Set it to true in .env to enable automatic timeout recommendations in your orchestrator's request layer.

Sync vs async decision (recap)

  • claude_run_task (sync, ≤600s wrapper cap) — trivial_edit, smoke_test, review, single_feature, docs_update, test_suite, small multi_file_refactor (≤20 files).

  • claude_start_task + claude_poll_task (async, ≤3600s) — large multi_file_refactor (>20 files), architecture, migration, long_running.

If a sync call returns parse_error (response is not valid JSON), do not retry sync — the subprocess likely returned truncated / non-JSON output due to a wrapper-level timeout. Switch to claude_start_task (async) which buffers output incrementally and survives longer walls.

Security

Safe Mode (Default)

In safe mode (CLAUDE_MCP_MODE=safe):

  • Sandbox execution is enforced (if CLAUDE_MCP_FORCE_SANDBOX_IN_SAFE_MODE is enabled).

  • Custom environment overrides are blocked.

  • Custom extra command-line arguments are blocked.

  • Bypassing permissions (bypassPermissions or dontAsk permission modes) is rejected.

Permissive Mode

In permissive mode (CLAUDE_MCP_MODE=permissive):

  • Environment variable overrides are permitted only if they appear in the CLAUDE_MCP_ALLOW_ENV_KEYS list.

  • Extra arguments are allowed only if they appear in CLAUDE_MCP_ALLOW_EXTRA_ARGS.

  • Skipping permissions (bypassPermissions or --dangerously-skip-permissions) requires "--dangerously-skip-permissions" to be explicitly listed in CLAUDE_MCP_ALLOW_EXTRA_ARGS.

Troubleshooting

CLAUDE_NOT_FOUND

The server cannot locate the claude executable. Make sure it is installed (e.g. via the official installer) and added to your PATH. Alternatively, set CLAUDE_MCP_CLAUDE_PATH to the absolute path of the binary.

NOT_ALLOWED: workspace_path is outside allowed roots

Your workspace path is outside of the configured roots. Include the target directory in the CLAUDE_MCP_ALLOWED_ROOTS JSON array, or launch the server from within the target folder.

⚠️ CLAUDE_MCP_ALLOWED_ROOTS and other server env vars are read at MCP server STARTUP. Editing .env after the server is running has no effect — restart the MCP server (in your client's MCP panel) for changes to apply.

NOT_LOGGED_IN / result.text == "Not logged in · Please run /login"

The claude CLI cannot find its OAuth token. The most common cause is the --bare flag being passed to the subprocess — --bare bypasses ~/.claude/ entirely, so ~/.claude.json (which holds the token) is invisible. Verify this server's force_bare=False (the default). Alternative: run claude login interactively in your shell to seed ~/.claude.json, then restart this MCP server.

MODEL_NOT_ALLOWED

The requested model (or fallback_model) is not in Settings.allowed_models (configurable via CLAUDE_MCP_ALLOWED_MODELS). Empty allowlist disables validation.

PERSISTENCE_FILE_TOO_LARGE

A persistence file has reached the maximum allowed bytes (default 1 MiB). Clean up or truncate obsolete entries in the file (~/.open-cli-router/claude-code/MEMORY.md or PROJECTS.md) to allow new writes.

CONFIRM_REQUIRED

Modifying AGENTS.md (the system prompt) in safe mode requires setting the confirm parameter to true to ensure the override is intentional.

~/.open-cli-router/claude-code/ does not exist

The persistence directory is created lazily. The server will not create it on its own — you must initialize it once via claude_init_persistence (or the Python one-liner under Persistent Memory → How to initialize). Without this, the server is enabled but uninitialized and no context is injected into prompts.

Available Tools

12 tools
claude_append_persistenceA

Append content to one of the persistence files.

Required: file (agents|projects|memory), content. Optional: section_header (str|None) — if provided, the append is placed under a heading; otherwise content is appended at the end of the file.

Safe-mode constraint: in Settings.mode == "safe", updating AGENTS.md requires confirm=true. The same applies to update_persistence.

Do not store secrets, credentials, or full file dumps — keep entries small and high-signal.

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileYes
timestampYes
appended_bytesYes
new_size_bytesYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral disclosure burden. It reveals the non-obvious safe-mode constraint requiring confirm=true for AGENTS.md, explains placement behavior for section_header, and adds content policy constraints. It omits error/idempotency details, but the output schema covers return expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action and requirements, then moves to optional behavior, a conditional constraint, and content policy. Every sentence carries operational value with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, the description covers target files, required vs optional parameters, the conditional confirm flag, and content constraints. Output details are left to the output schema, so nothing critical blocks correct invocation, though usage-vs-alternatives guidance is only implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema description coverage is 0%, so the description must compensate, and it does: it defines file values, marks file and content as required, explains the optional section_header behavior, and explains when confirm matters. This adds substantive meaning beyond the otherwise bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb ('Append'), a clear target resource ('one of the persistence files'), and enumerates the allowed file values ('agents|projects|memory'). The operation is semantically distinct from read/update siblings, and the safe-mode note explicitly connects to update_persistence without blurring the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for adding content to persistence files and gives required/optional fields, but it does not explicitly say when to choose append over update_persistence or read_persistence, nor does it describe conditions for when not to use it. The safe-mode note hints at a related tool but stops short of clear alternative-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_cancel_taskA

Cancel a running task.

Two escalation levels: - force=false (default): sends SIGTERM (graceful). The subprocess has ~5s to clean up before the OS escalates. Try this first. - force=true: sends SIGKILL (immediate). Use only if the subprocess doesn't respond to SIGTERM within ~5s.

Returns: ClaudeCancelTaskResponse with canceled (bool) and status ("cancelled" if actively stopped, "already_done" if run finished naturally, "not_found" if run_id is unknown).

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
canceledYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It clearly explains the signal mechanism (SIGTERM vs SIGKILL), the ~5s cleanup window, OS escalation, and the possible return statuses. This gives the agent a solid mental model of the tool's side effects and outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with a front-loaded purpose statement, bulleted escalation levels, and a concise return-value explanation. Every sentence adds meaningful information without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers both parameters, the escalation behavior, and the full set of return statuses. The only minor gap is the behavior when req is null, but the schema already indicates this via the nullable type. Overall, an agent has the information needed to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It thoroughly explains the force parameter, including default and recommended sequence. The run_id is only implicitly described via the 'not_found' status, but its role as the task identifier is understandable from context. Some additional detail on run_id sourcing would strengthen it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Cancel a running task.' It clearly identifies the tool's function and differentiates it from siblings like claude_run_task, claude_start_task, and claude_poll_task by being the only cancellation operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on how to use the two escalation levels, including 'Try this first' for force=false and 'Use only if the subprocess doesn't respond' for force=true. It does not explicitly name alternative tools for selection, but the usage context is clear enough for an agent to know when cancellation is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_healthA

Health check for the Claude Code CLI binary.

Verifies the claude binary is on PATH, captures its --version output, and reports auth status (whether ~/.claude.json has an OAuth token). Use this as the first call in any orchestration session to validate the environment.

Optional fields: - expected_version (str): if set, returns ok=false on version mismatch

Returns: - claude_path, claude_version, ok, auth_status, notes

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
notesNo
auth_statusNo
claude_pathYes
claude_versionYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavior, and it delivers: it states it verifies the binary on PATH, captures --version, inspects ~/.claude.json for an OAuth token, and reports auth_status. It also defines the effect of expected_version on ok, exceeding a typical health-check disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and logically ordered: purpose, behavior, usage timing, optional parameter, return fields. No filler sentences; every line adds necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (1 optional parameter, no annotations), the description covers purpose, usage context, behavior, parameters, and return values. The output schema is said to exist, so the return field list is a bonus. An agent has everything needed to call it first and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add meaning. It explains the sole parameter expected_version, noting it is optional and that setting it returns ok=false on mismatch — exactly what an agent needs beyond the raw anyOf schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete verb+resource: 'Health check for the Claude Code CLI binary' and lists specific checks (PATH, --version, auth status). This clearly differentiates it from sibling task-execution tools like claude_run_task or claude_self_test.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call this 'as the first call in any orchestration session to validate the environment.' This provides clear context for when to invoke, though it does not name alternatives or exclusions, preventing a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_init_persistenceA

Initialize the persistence directory and seed the three markdown files.

Idempotent: re-running without force=true is a no-op if files already exist. Creates the directory at the location resolved by Settings.resolve_persistence_base_dir() and writes AGENTS.md, PROJECTS.md, MEMORY.md (unless they already exist).

Optional fields: - force (bool): re-create files even if they exist - seed_templates (bool|None): whether to seed the default templates

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
createdNo
base_dirYes
seed_versionYes
already_existedNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and it delivers: it discloses idempotency, the no-op behavior when files exist, directory creation, the exact files written, and the effect of force. This goes well beyond a simple 'init' label.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action statement is front-loaded, followed by a tight idempotency note and a compact optional-fields list. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity initialization tool with zero required parameters and an output schema, the description covers location, file names, idempotency, force behavior, and template seeding. Nothing essential for a correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It names force and seed_templates and gives a functional meaning for each, although it leaves the null-vs-false distinction for seed_templates slightly implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Begins with a specific verb and resource: 'Initialize the persistence directory and seed the three markdown files' and even names AGENTS.md, PROJECTS.md, MEMORY.md. This clearly separates it from sibling persistence tools that read, append, update, or load.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context through the idempotency statement: re-running is a no-op unless force=true, and explains what force does. It does not explicitly compare against alternatives or state when not to use it, so it falls just short of full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_list_runsA

List recent runs (active + recently completed).

Returns ClaudeRunSummary entries ordered newest-first, with active claude_start_task runs at the top followed by completed/cancelled/ timed-out runs from the in-memory store (bounded by Settings.max_runs).

Use this for recovery after orchestrator restart: active async runs survive across MCP client restarts and can be polled/cancelled via claude_poll_task and claude_cancel_task using the run_id from this listing.

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
runsNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It reveals ordering (newest-first, active at top), source (in-memory store), bound (Settings.max_runs), and cross-restart persistence. It does not explicitly state read-only behavior, but 'list' strongly implies that, and 'recent' is not precisely defined—minor gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all informative: the scope, the ordering/source/bound detail, and the recovery use case. The text is front-loaded and wastes no words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main behavior, ordering, bound, and recovery workflow, and an output schema exists so return-value detail is unnecessary. It is incomplete only in omitting the optional limit parameter and not explicitly confirming that this is a read-only operation—minor for a tool that works fine with no arguments.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the req.limit parameter, but it never mentions it. The only bound referenced is Settings.max_runs, which is a different limit. An agent reading the description would have no idea that an optional limit parameter with default 50 exists, especially since the schema itself provides no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and resource ('List recent runs') and immediately scopes it to active plus recently completed. It distinguishes the tool from siblings by referencing claude_start_task, claude_poll_task, and claude_cancel_task for different roles, and specifies the returned ClaudeRunSummary entries with ordering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: 'Use this for recovery after orchestrator restart.' It goes further to explain that active async runs survive restarts and can be polled/cancelled via claude_poll_task and claude_cancel_task, guiding the complete workflow. No alternative list tool exists among siblings, so no exclusion is required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_load_persistence_contextA

Load the persistence files as context for the current session.

Returns head+tail excerpts of each file (head_ratio controlled by Settings.persistence_truncation_head_ratio). Use this at the start of each session to hydrate orchestrator memory before dispatching tasks.

Optional fields: - include (list[str]|None): subset of files to load (agents|projects|memory). None = all three. - max_chars_per_file (int): per-file char cap (default from settings).

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
base_dirNo
initializedNo
total_charsNo
agents_excerptNo
memory_excerptNo
truncated_flagsNo
projects_excerptNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses useful behavioral traits: returns head+tail excerpts, head_ratio is settings-controlled, and max_chars_per_file caps output. It does not directly state that the operation is non-mutating or describe failure behavior, but for a context-loading operation this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the core purpose, and uses a clear optional-fields section. There is no filler or repetition of schema structure; every sentence adds functional value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple optional-parameter interface, the description covers purpose, usage timing, return truncation behavior, and parameter semantics. It does not mention what happens if persistence files are missing or how this differs from claude_read_persistence, but the presence of an output schema and the startup-context guidance make this sufficiently complete for agent invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add parameter meaning. It does: 'include' is a subset of files (agents|projects|memory), None means all three, and max_chars_per_file is a per-file character cap. It slightly defers the default to settings while the schema specifies 20000, but overall it compensates well for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Load the persistence files as context for the current session.' It clearly identifies what the tool does and adds return-behavior detail, but it does not explicitly differentiate itself from the sibling claude_read_persistence, so the agent must infer the distinction from context signals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear when-to-use guidance: 'Use this at the start of each session to hydrate orchestrator memory before dispatching tasks.' However, it does not provide when-not-to-use guidance or name an alternative tool such as claude_read_persistence, so the exclusion side is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_poll_taskA

Poll an asynchronous task.

Two polling modes: - drain=false (default): returns immediately with current state (new messages, stdout/stderr byte counts, elapsed time). Use for tight control loops with explicit backoff. - drain=true: blocks until status is no longer 'running' (fire-and-wait). Note: this is bounded by your MCP client's request timeout, not the async task's timeout_s. For long blocks, prefer async drain with orchestrator-level polling.

Returns: ClaudePollTaskResponse with status (running|done|error|timeout|cancelled), new_messages, stdout_len/stderr_len, elapsed_seconds, and (once terminal) the full result with total_cost_usd, model_usage, and changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesNo
resultNo
run_idYes
statusYes
stderr_lenNo
stdout_lenNo
new_messagesNo
elapsed_secondsNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries behavioral disclosure. It states that drain=false returns immediately, drain=true blocks and is bounded by the client request timeout, and it lists the return fields and terminal statuses. This is transparent about the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with bullet points and a return section, making it easy to scan. It front-loads the core purpose and then details modes, without unnecessary sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return details are partially covered, but the description still outlines the response structure. It covers the main behavioral differences and timeout bound. Missing details like wait_seconds are minor, making it sufficiently complete for a polling tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It thoroughly explains the 'drain' parameter (the core behavior) but does not mention 'wait_seconds' or the nullable 'req' field. The omission of wait_seconds leaves a partial gap in parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Poll an asynchronous task' with a specific verb and resource. It distinguishes itself from siblings like claude_start_task, claude_cancel_task, and claude_list_runs, leaving no ambiguity about its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly describes two polling modes—drain=false for tight control loops with explicit backoff and drain=true for fire-and-wait—and advises orchestrator-level polling for long blocks. While it doesn't name alternatives directly, the mode guidance provides clear when-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_read_persistenceA

Read one of the three persistence files (agents | projects | memory).

Optional fields: - file (str): which file to read (default: memory) - offset (int): start reading from line N (0-indexed) - limit (int|None): max lines to return (None = no limit)

Large files are automatically truncated at Settings.persistence_max_file_bytes; the response includes a truncated flag if this happens.

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileYes
contentYes
truncatedYes
size_bytesYes
modified_atYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden well by disclosing line-based offset/limit behavior, automatic truncation at Settings.persistence_max_file_bytes, and the truncated flag in the response. It does not cover error handling or file-not-found behavior, but the most operationally surprising behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and well-structured: a one-line purpose, a short bullet list for optional parameters, and a single note about truncation. Nothing is redundant with the schema, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers parameters, defaults, and truncation behavior, and an output schema exists for return-value details. It could be slightly more complete by naming sibling tools or clarifying when load_persistence_context is preferable, but nothing critical is missing for calling the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the schema. It explains all parameters clearly: file choices and default, offset being 0-indexed, and limit accepting None for no limit. This adds real meaning beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource ('Read one of the three persistence files') and enumerates the allowed files: agents, projects, memory. It is clear, though it does not explicitly distinguish itself from the sibling claude_load_persistence_context, which may also involve reading persistence data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: when raw contents of the agents, projects, or memory persistence files are needed. However, it provides no explicit when-not-to-use guidance or mention of alternatives such as load_persistence_context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_run_taskA

Run a single Claude task synchronously (bounded by FastMCP's 600s wrapper cap).

Required: workspace_path, prompt. Common options: model (alias from Settings.claude_model_aliases), fallback_model, permission_mode, options.timeout_s (default 300, max 600), capture_changes.

Estimated cost depends on model and prompt size — check total_cost_usd in the response. Typical ranges: haiku ~$0.02, sonnet ~$0.50, opus ~$1.50.

For tasks that may exceed 600s (architecture, migration, large refactors), use claude_start_task (async) instead. For timeout routing, call the claude_timeout_help prompt first.

Returns: ClaudeRunTaskResponse with result, stdout/stderr, exit_code, timed_out, total_cost_usd, model_usage, and (if capture_changes=true) a unified git diff of workspace changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesNo
resultYes
run_idYes
statusYes
changesNo
exit_codeYes
num_turnsNo
final_textYes
session_idNo
duration_msNo
model_usageNo
total_cost_usdNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses synchronous execution, the 600s cap, cost expectations, and the full response shape including timed_out and optional git diff, which implies workspace mutation. It does not explicitly warn about side effects or permission risks beyond mentioning permission_mode, but the disclosure is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then structured into required/common, cost, routing, and return sections. Every sentence adds information, and the line breaks make it scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the rich detail, the description omits the required `req` object wrapper, which is a critical invocation detail. It also contains a factual contradiction with the schema on default and maximum timeout values. Given the tool's complexity, the undocumented advanced parameters further reduce completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description names required and common options, but it does not mention the top-level `req` wrapper that the schema actually requires, so an agent may pass workspace_path/prompt incorrectly. Its timeout_s claim (default 300, max 600) directly contradicts the schema (default 600, max 3600), and with 0% schema coverage, many parameters (effort, max_turns, change_scope, max_budget_usd, etc.) are left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: 'Run a single Claude task synchronously' and explicitly notes the 600s wrapper cap. It clearly differentiates from the async sibling claude_start_task, so an agent can immediately tell this is the sync execution tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly lists required and common options and gives concrete routing rules: use claude_start_task for tasks exceeding 600s, and call claude_timeout_help for timeout routing. This is exactly the when-to-use-vs-alternatives guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_self_testA

Inspect every registered tool's input schema and report robustness.

This is a metadata-only check — no tools are actually invoked.

Args shape: The MCP client MUST pass arguments wrapped in a req object: {"req": {"include": ["claude_health"], "only_show_tolerant": true}} For backwards-compatibility, the server also accepts args={}.

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolsNo
summaryYes
server_infoNo
total_toolsYes
tolerant_countYes
requires_req_countYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that no tools are actually invoked, which is a key safety trait. It also discloses the exact argument wrapping requirement (the 'req' object) and the backwards-compatibility fallback ('args={}'), giving the agent precise calling behavior. It does not cover error handling or output format, but those are partially covered by the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with a clear purpose statement, a note on metadata-only behavior, and then the argument shape. It includes an example and a note about backwards compatibility, which are useful. It is somewhat verbose but each part contributes to understanding the tool. The key information is front-loaded with the purpose and the no-invocation note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one parameter with a nested structure and no schema descriptions, the description is not fully complete. It explains the wrapper structure but does not define the meaning of 'include' or 'only_show_tolerant', which are essential for correct use. The existence of an output schema covers return values, but the parameter semantics are left underspecified, making it less than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate. It explains the overall shape: arguments must be wrapped in a 'req' object, and it provides an example with 'include' and 'only_show_tolerant'. However, it does not explain the semantics of these fields (what 'include' filters or what 'tolerant' means), leaving some ambiguity. The example helps, but the meaning of the fields is only implied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: inspecting every registered tool's input schema and reporting robustness. It is specific with a verb ('inspect') and a resource ('registered tool's input schema'), and it explicitly distinguishes itself as a metadata-only check that does not invoke tools. This differentiates it from sibling tools like claude_health or claude_run_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives, nor does it mention any exclusions or conditions. It implies that it is for schema inspection and robustness checking, but it does not provide guidance on when a user would prefer this over other tools. The absence of any alternative comparison leaves the agent to infer usage from the purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_start_taskA

Start a Claude Code CLI task asynchronously (up to 3600s).

Returns immediately with a run_id. The subprocess continues running in the background; use claude_poll_task to monitor and claude_cancel_task to stop. Use this for any task that may exceed 600s (architecture, migration, large multi-file refactors, long-running data agents).

Required: workspace_path, prompt. Optional: model, fallback_model, permission_mode, options.timeout_s (default 300, max 3600 for async), capture_changes.

Concurrent run limit: bounded by Settings.max_concurrent_runs. New calls beyond that limit return MAX_CONCURRENT_RUNS_EXCEEDED.

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
session_idYes
started_atYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does well: async background execution, immediate return, timeout boundaries, and the MAX_CONCURRENT_RUNS_EXCEEDED error are all disclosed. It could go further on side effects or cleanup expectations, but the core behavioral traits are clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tightly written and well structured, with the most important operational facts front-loaded. Each paragraph earns its place: async behavior, return semantics, when to use it, required parameters, and concurrency limits.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main async workflow and monitoring guidance well, but the tool is complex with many optional parameters and no annotations. It omits the likely synchronous alternative claude_run_task and does not clarify how to pass required fields inside the 'req' object, which is a meaningful invocation gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only names a subset of parameters and does not explain the required 'req' object wrapper. More importantly, it states options.timeout_s default is 300 while the schema says 600, creating a direct contradiction that could lead to incorrect invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Start a Claude Code CLI task asynchronously'. It clearly distinguishes itself from siblings by emphasizing async execution, immediate return of a run_id, and the ability to run up to 3600s.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance for when to use this tool: 'Use this for any task that may exceed 600s' with concrete examples. It also directs the agent to claude_poll_task and claude_cancel_task for lifecycle management. However, it does not explicitly contrast with claude_run_task, which appears to be the likely synchronous alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claude_update_persistenceA

Replace or append to a section in one of the persistence files.

Required: file, section_anchor, new_content. Optional: mode (replace|append) — replace the entire section vs append inside it (default replace).

Safe-mode constraint: in Settings.mode == "safe", updating AGENTS.md requires confirm=true. Returns matched=true when the section_anchor was found; matched=false when the anchor was not found and no edit happened (so you can detect typos before silent appending).

ParametersJSON Schema
NameRequiredDescriptionDefault
reqNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileYes
matchedYes
new_size_bytesYes
section_anchorYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden, and it does well: it discloses the safe-mode constraint, the default mode, and the exact matched=true/matched=false return semantics including that no edit happens when the anchor is not found. It does not discuss destructive risk beyond the word 'replace' or atomicity/ordering details, but the key failure behavior is exposed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the primary action and uses short labeled lines for required/optional/safe-mode/return behavior. Each section adds information, though the safe-mode paragraph is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating persistence tool with no annotations, it covers the essential edge cases: anchor not found, safe mode, mode default, and return signal. An output schema exists for the return shape, so not explaining all return fields is acceptable. The nested req wrapper is the main missing call-construction detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters, and it does: it defines mode's replace vs append behavior, explains when confirm=true is required, and clarifies section_anchor matching. The only gap is that it lists file/section_anchor/new_content as top-level required inputs while the schema nests them under a req object, which could confuse an agent constructing the call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Replace or append to a section in one of the persistence files.' This clearly conveys the tool's core operation. It does not explicitly distinguish itself from the sibling claude_append_persistence, but the mode parameter partially covers that overlap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It lists required and optional parameters and explains the safe-mode confirm condition, which is actionable guidance. However, it does not say when to prefer this tool over claude_append_persistence or other persistence tools; the alternative routing is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv0.2.0
    • First observedclaude_append_persistence
    • First observedclaude_cancel_task
    • First observedclaude_health
    • First observedclaude_init_persistence
    • First observedclaude_list_runs
    • First observedclaude_load_persistence_context
    • First observedclaude_poll_task
    • First observedclaude_read_persistence
    • First observedclaude_run_task
    • First observedclaude_self_test
    • First observedclaude_start_task
    • First observedclaude_update_persistence

TDQS

A4.2/5.0

Scored across 12 tools

Disambiguation4/5

The task lifecycle tools (run/start/poll/cancel/list) are clearly separated by sync vs async execution and state transitions, and the persistence tools have distinct read/append/update/load operations. The only mild ambiguity is between run_task and start_task, and between update_persistence and append_persistence, but the descriptions explicitly route usage.

Naming Consistency4/5

All tools share a consistent claude_ prefix and snake_case style, with most following a verb_noun pattern like claude_run_task, claude_poll_task, and claude_read_persistence. claude_health is a noun-only outlier, and claude_load_persistence_context is less consistent with the shorter read/append/update names, but the overall pattern remains predictable.

Tool Count5/5

Twelve tools is well-scoped for a server covering Claude Code CLI task execution and persistent memory files, with no redundant duplicates. Each tool addresses a distinct operational need within those two clear domains.

Completeness5/5

The surface covers the full task lifecycle: synchronous run, asynchronous start, poll, cancel, and list, plus health verification and self-test. Persistence is also complete with init, read, append, update, and load-context operations, leaving no obvious dead ends for the stated purpose.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    A server implementation for the Model Context Protocol (MCP) that allows Claude AI to execute commands through a command-line interface, enabling direct system interactions from within Claude.
    -
  • A
    license
    A
    quality
    C
    maintenance
    Local MCP server that wraps the headless Claude Code CLI as MCP tools, providing stateless access to Claude's coding capabilities through prompt-based interactions. It enables users to execute Claude Code commands with various prompt formats and structured outputs directly from MCP clients.
    3
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Exposes headless Claude Code as a remote MCP server with a voice client, enabling hands-free task execution and session management via OpenAI's Realtime API.
    -