Skip to main content
Glama

Coding Tools MCP

English | 简体中文

Give any AI chat or agent a safe pair of hands on your codebase.

PyPI npm Python compliance release License

Coding Tools MCP is a model-neutral coding runtime served over the Model Context Protocol: file reading and search, structured multi-file patches, command execution, interactive sessions, and git — one server that any MCP client can drive. Claude Desktop, Claude Code, Cursor, Cline, or an agent you build yourself all get the same 20 battle-tested tools, confined to one workspace, gated by permission modes.

Watch the demo

Why people use it

  • It turns a chat app into a coding agent. Claude Desktop — or any MCP chat client — gets real repo access with the subscription you already have. No extra product required.

  • Safety is the product, not an afterthought. One workspace root per server. Absolute paths, .. traversal, and symlink escapes are rejected. Permission modes gate network access, shell expansion, inline scripts, and destructive commands. On Linux, Landlock adds kernel-level filesystem confinement.

  • It is model- and vendor-neutral. A fixed, truthfully annotated catalog — no profile switching, no annotation games. Swap models or clients freely; the runtime and its behavior stay put.

  • It is engineered for context windows. Results are summarized, paginated, and capped by design; serialized tool-result bytes dropped 37% release-over-release on the deterministic dogfood workload with unchanged task completion.

Related MCP server: Cloud Harness MCP

Quickstart

Run it with whichever toolchain you already have (the server is Python ≥ 3.11 from PyPI; the npm package is a thin launcher that starts it via uv or pipx):

uvx coding-tools-mcp --stdio --workspace /path/to/repo   # Python toolchain
npx coding-tools-mcp --stdio --workspace /path/to/repo   # Node toolchain

Wire it into Claude Desktop, Claude Code, Cursor, or Cline — the JSON is the same everywhere (swap uvx for npx if you prefer Node):

{
  "mcpServers": {
    "coding-tools": {
      "command": "uvx",
      "args": ["coding-tools-mcp", "--stdio", "--workspace", "/path/to/repo"]
    }
  }
}

Then ask your client: "run the test suite and fix the first failure."

Prefer HTTP? Drop --stdio and the server speaks Streamable HTTP on http://127.0.0.1:8765/mcp (MCP 2025-11-25, with 2025-06-18 compatibility). A one-line installer, per-client walkthroughs, and troubleshooting live in docs/quickstart.md and docs/mcp-client-config.md.

Seven things to try

1. Make Claude Desktop your coding agent. The config above is all it takes — the chat window you already pay for can now read, patch, test, and commit-review a real repository.

2. Code on your own machine from anywhere.

CODING_TOOLS_MCP_AUTH_MODE=bearer ./scripts/tunnel.sh cloudflared /path/to/repo

Loopback bind + authenticated HTTPS tunnel (cloudflared, ngrok, or Microsoft Dev Tunnel). Point claude.ai on your phone at https://<tunnel-host>/mcp and drive your home workstation from anywhere. Bearer tokens and OAuth 2.1 + PKCE (with RFC 7591 dynamic registration) are built in. → docs/remote-mcp.md

3. Let an agent loose on untrusted code — inside a disposable sandbox.

docker build -t coding-tools-mcp-sandbox:local .
docker run --rm --init -it -p 8765:8765 -v "$PWD:/workspace" coding-tools-mcp-sandbox:local

A containerized server with toolchains and caches preconfigured, safe to point at a sketchy PR and destroy afterwards. → docs/docker.md

4. Spin up a cloud sandbox with one MCP call. The bundled Cloudflare Worker control plane exposes start_coding_tools_sandbox as an MCP tool: one call dispatches a GitHub Actions runner that boots the Docker sandbox and publishes it behind an authenticated Cloudflare Tunnel. Ephemeral compute, no server of your own.

5. Drive it from a GUI.

python -m pip install "coding-tools-mcp[desktop]"
coding-tools-mcp-desktop

Per-workspace profiles, server and tunnel start/stop, credential setup with clipboard helpers, live health checks. English and 简体中文.

6. Keep an interactive session alive. exec_command starts a REPL or debugger under a real PTY; write_stdin feeds it across turns; read_output pages long output; kill_session cleans up. Long-running processes are first-class, with deadline watchdogs and bounded buffers.

7. Give your own agent production-grade hands. Building an agent loop with the Anthropic SDK or anything else? Don't hand-roll file and exec tools — speak MCP to this server and inherit the whole safety boundary. → docs/embedding.md

The tool catalog

One stable, truthfully annotated set — permission modes change command policy, never which tools the model sees. apply_patch is the sole file-mutation primitive: staged, baseline-checked, atomic across files, with rollback.

Group

Tools

Files & search

read_file · list_dir · list_files · search_text · apply_patch · view_image

Execution

exec_command · write_stdin · read_output · kill_session · request_permissions

Git

git_status · git_diff · git_log · git_show · git_blame

Runtime

server_info · check_exec_environment · get_default_cwd · set_default_cwd

Root AGENTS.md/CLAUDE.md files load into the initialize context automatically. Tool content is concise agent-facing text; structuredContent carries the complete machine result. Schemas and result envelopes: docs/tools-and-schemas.md · docs/runtime-contract-v0.2.md.

Safety Boundary

Mode

Meant for

What it allows

safe (default)

day-to-day agent work

file tools and vetted commands; network-looking commands, shell expansion, inline scripts, and destructive commands all require explicit permission

trusted

local development

opens network, shell expansion, and inline scripts; keeps secret filtering and destructive-command checks

dangerous

isolated containers/VMs only

disables exec_command permission gates; workspace path boundaries still apply

Recursive listing and search exclude .git, node_modules, build outputs, virtualenvs, and caches. Commands run with workspace-bound cwd, scrubbed environment, timeouts, and output caps. Linux hosts with Landlock get kernel-enforced filesystem confinement; other platforms get an explicit warning — this is still not a complete OS sandbox, so use the Docker image or a VM for genuinely untrusted work. Details: SECURITY.md · docs/security-boundary.md · docs/permission-modes.md

Telemetry

The server sends anonymous usage telemetry (per-tool success/latency counters and version/platform dimensions — never paths, arguments, commands, or file contents) to help prioritize fixes. Disable it with CODING_TOOLS_MCP_TELEMETRY=off or DO_NOT_TRACK=1; it is automatically off in CI. CODING_TOOLS_MCP_TELEMETRY=debug prints every event to stderr instead of sending. The full event list and guarantees are in docs/telemetry.md.

Evidence, Dogfood and SWE-bench

Every release ships through a tag-triggered pipeline in which the compliance suite, real-workload benchmark, and SWE-bench harness run from the same commit that publishes to PyPI and npm — both via trusted publishing, npm with provenance. Dogfood efficiency metrics are reproducible (make dogfood-smoke) and checked in under reports/. This repository does not claim a model-generated SWE-bench leaderboard result — see docs/swe-bench.md for exactly what is and is not measured. More: COMPLIANCE.md · BENCHMARK.md · docs/dogfood.md

Documentation

Development

python -m pip install -e ".[dev]"
make ci        # lint, typecheck, tests, protocol/integration suites, gates

The full gate matrix is in docs/ci-and-tests.md.

License

This project is licensed under the Apache License 2.0.

If you use code, documentation, substantial implementation details, or derivative work from this project, preserve the copyright notice, license notice, and NOTICE file, and clearly attribute the original project.

Project: Coding Tools MCP
Author: Coding Tools MCP Contributors
Source: https://github.com/xyTom/coding-tools-mcp

Citation metadata is available in CITATION.cff.

Available Tools

18 tools
apply_patchApply patchA
Destructive

Stage, validate, and atomically apply a patch envelope. Example: *** Begin Patch *** Update File: app.py @@ -old +new *** End Patch

ParametersJSON Schema
NameRequiredDescriptionDefault
patchYes
dry_runNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as destructive and non-idempotent. The description adds a useful layer: it validates before applying and applies atomically, which tells the agent that a failed validation should not leave partial edits. It does not spell out failure output, but it enriches the behavioral contract beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise lead sentence states the key behavior, followed immediately by a compact example that demonstrates the exact envelope syntax. Every line earns its place; there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema covers return-value expectations, and the example covers patch syntax. Still, the dry_run semantics and what happens when validation fails are left implicit, which are meaningful gaps for an agent deciding whether to perform a destructive write or a trial run.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must carry parameter meaning. It does teach the required patch envelope format through a concrete example, which is valuable. However, the dry_run parameter is never mentioned or explained; an agent must infer its behavior from the name and default value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Stage, validate, and atomically apply a patch envelope.' This is far more informative than the title and clearly identifies a distinct file-modification capability among siblings like read_file and exec_command. The inline example pins down the exact input format, so an agent cannot confuse it with other tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The patch-envelope example implies that this is the intended mechanism for structured file edits, but the description never states when to prefer it over alternatives (e.g., exec_command or git apply) or when not to use it. It provides context but no explicit exclusions or alternative tool names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_exec_environmentCheck exec environmentA
Read-onlyIdempotent

Return lightweight exec_command sandbox and environment status known to the server.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds 'lightweight' and 'known to the server,' suggesting a cheap cached status view, which is useful context. It does not cover authentication or staleness, but annotations carry the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the action 'Return' and the key subject immediately. Every word contributes value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool is a parameterless, read-only status check with an output schema available, the description sufficiently covers the call contract. The high-level wording leaves specifics of the returned status to the output schema, which is appropriate. It could mention typical usage relative to exec_command, but that falls under usage guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters and an empty input schema, schema coverage is complete by definition. The description does not need to explain parameters, and its statement is consistent with a parameterless call, earning the baseline 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as returning exec_command sandbox and environment status, using the specific verb 'Return' and naming the resource. It is distinguishable from sibling tools like exec_command and server_info, though it does not enumerate what status fields are included.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to call this tool versus alternatives such as exec_command or server_info. It does not state preconditions, use cases, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exec_commandExecute commandA
Destructive

Run a bounded command under runtime policy. Pass workdir explicitly for reconnect-safe paths. A still-running command returns command_id. Example: {"cmd":"pytest -q","workdir":".","yield_time_ms":30000}. Retained output is bounded per stream; for very large output redirect to a file (cmd > out.log 2>&1) and page it with read_file or search_text.

ParametersJSON Schema
NameRequiredDescriptionDefault
cmdYes
cwdNo
envNo
ttyNo
stdinNo
workdirNo.
verbosityNo
timeout_msNo
preview_bytesNo
yield_time_msNo
max_output_bytesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the description doesn't need to repeat those. It adds useful context: commands are bounded, still-running commands return a command_id, and retained output is bounded per stream. This goes beyond annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences plus an example, front-loaded with the core purpose. Every sentence adds value; no filler. The example is compact and clarifies usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While an output schema exists (so return values aren't the description's job), the description does not explain several important parameters and their effects (e.g., verbosity, max_output_bytes). It covers typical usage but not advanced cases, so it's incomplete for a tool with this parameter count.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions workdir (explicitly for reconnect-safety), yield_time_ms (via example), and cmd (in redirect usage), but leaves env, tty, stdin, verbosity, timeout_ms, preview_bytes, and max_output_bytes unexplained. This is partial compensation, insufficient for a tool with 11 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Run') and resource ('bounded command under runtime policy'), and gives a concrete example. It clearly differentiates from file/sibling tools like read_file, list_dir, and search_text by describing command execution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when to use: 'Pass workdir explicitly for reconnect-safe paths' and recommends redirecting large output to a file and paging with read_file or search_text. Does not explicitly state when to avoid this tool, but gives clear alternatives for output handling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_blameGit blameA
Read-onlyIdempotent

Return bounded git blame metadata for a workspace file.

ParametersJSON Schema
NameRequiredDescriptionDefault
revNo
pathYes
end_lineNo
max_linesNo
start_lineNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds the useful behavioral clue that results are 'bounded' and scoped to a workspace file, which meaningfully narrows expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One succinct, front-loaded sentence with no filler or repetition. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a standard read-only git tool with an output schema and strong safety annotations, the description is nearly sufficient. Minor gaps like line-range inclusivity and rev interpretation remain, but they are partially inferable and do not block correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate for parameter meaning. It only hints at 'bounded' behavior and never explains path, rev, start_line, end_line, or max_lines semantics; the agent must infer meaning from names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says 'Return bounded git blame metadata for a workspace file' – a specific verb, a clear resource (git blame), and a narrow scope. It distinguishes the tool from siblings like git_log, git_diff, and git_show without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use git_blame versus the sibling git_* tools. It does not state exclusions, prerequisites, or conditions that would route an agent to an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_diffGit diffC
Read-onlyIdempotent

Return unified git diff for workspace changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
pathsNo
stagedNo
unstagedNo
max_bytesNo
context_linesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds 'unified' and 'workspace changes' as useful context, but it does not disclose defaults, truncation behavior via max_bytes, or how staged/unstaged changes are handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one concise sentence with no filler and front-loads the core action and output format. It is efficient, though the brevity is partly a product of omitting parameter and usage detail, so it is not a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six parameters and zero schema descriptions, this is too thin. It states the core behavior and annotations cover safety, but it omits how to restrict paths, choose staged vs unstaged diffs, control context lines, or handle large outputs. Better than a pure tautology, but not adequate for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description mentions none of the six parameters: path, paths, staged, unstaged, max_bytes, or context_lines. The description therefore adds no semantic meaning beyond the bare parameter names and defaults already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and clearly identifies the resource: a unified git diff scoped to workspace changes. The 'workspace changes' scope separates it broadly from commit-history siblings like git_show and git_log, though it does not explicitly name any alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance about when to use git_diff versus sibling tools such as git_status, git_show, or git_blame. It also does not mention exclusions, prerequisites, or typical workflows, so the agent must infer the usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_logGit logB
Read-onlyIdempotent

Return recent git commits with bounded structured metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoHEAD
pathNo.
skipNo
max_countNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds 'bounded structured metadata', which hints at limited, schema-shaped output, but does not detail ordering, pagination, or other behavioral traits. Given the strong annotations, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with the key action and scope front-loaded. Every word earns its place and there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool with no required parameters and an output schema present, the description is partially complete. It lacks sibling-usage guidance and parameter semantics, but the annotations and schema defaults cover much of the operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the four parameters. It does not explain ref, path, skip, or max_count; 'recent' and 'bounded' only loosely hint at defaults and limits. The property names are self-explanatory, but the description itself adds minimal parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('recent git commits'), and adds the useful qualifier 'bounded structured metadata'. It is distinguishable from siblings like git_diff, git_show, and git_blame, though it does not explicitly name them or the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use git_log versus sibling tools such as git_show or git_diff, nor are there exclusions or alternative references. The word 'recent' implies history inspection, but the decision context is left entirely to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_showGit showB
Read-onlyIdempotent

Return bounded git show output for a revision.

ParametersJSON Schema
NameRequiredDescriptionDefault
revNoHEAD
pathNo
pathsNo
max_bytesNo
include_diffNo
context_linesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the 'bounded' qualifier, which hints at output size constraints, but does not elaborate on behavior when limits are exceeded or how the output is structured. Given annotation coverage, this is a minimal but acceptable addition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence with zero redundancy. It communicates the core purpose efficiently and is appropriately front-loaded with the action and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has six parameters (all optional) and an output schema (though not shown), but the description lacks explanations for parameter behavior, edge cases, or how the 'bounded' aspect is enforced. It is adequate for a simple read-only tool but does not fully cover the contextual nuances an agent might need for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning no parameter documentation exists. The description does not compensate by explaining any of the six parameters (rev, path, paths, max_bytes, include_diff, context_lines). While parameter names are somewhat self-explanatory, the description adds no semantic value beyond what the schema already provides, failing to fill the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and names the resource ('git show output') with a qualifier ('bounded') that hints at output limits. It is clear about the core action but does not explicitly distinguish it from siblings like git_log or git_diff, though the 'bounded' qualifier provides a slight differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as git_log or git_diff. The description only states what it does without any contextual cues for selection, making it purely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_statusGit statusB
Read-onlyIdempotent

Return git working tree status for the workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
max_entriesNo
include_untrackedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnlyHint, idempotentHint, and destructiveHint, so the description is not required to restate safety. It adds the workspace scope and confirms the return value is status, but does not disclose behavior like result truncation via max_entries or untracked-file handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every word contributes to identifying the resource and operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with annotations and an output schema, the description is mostly complete: it states the workspace scope and the core function. It is slightly incomplete because it does not clarify how path or max_entries affect the result, though the schema partially covers these.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the three parameters. It does not mention path, max_entries, or include_untracked, leaving the agent to rely on parameter names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair, 'Return git working tree status,' which clearly identifies the operation. It does not explicitly differentiate from sibling git tools like git_diff or git_log, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the sibling git tools (git_diff, git_log, git_show, git_blame). The usage is only implied by the tool's name and purpose, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kill_commandKill commandB
Destructive

Terminate a server-managed command by command_id. Example: {"command_id":"abc","signal":"KILL"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
signalNoTERM
wait_msNo
verbosityNo
command_idYes
kill_wait_msNo
preview_bytesNo
max_output_bytesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, and the description's 'Terminate' is consistent with that rather than adding much new behavioral detail. It adds a useful scoping trait ('server-managed') and hints at signal selection via the example, but it does not discuss consequences like force-kill escalation, output loss, or irreversibility beyond what the destructive hint implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence plus a compact JSON example: it is front-loaded and has no filler. The example is useful, though it showcases the non-default 'KILL' signal while the schema default is 'TERM', making the example slightly less representative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema and destructive annotations, the tool has seven parameters and the description only covers two. It does not explain the wait/kill-wait process, the output-cap limits, or what happens after termination, leaving meaningful gaps for an agent deciding how to call the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate, but it only mentions command_id and signal in the example. The six other parameters (notably wait_ms vs. kill_wait_ms, verbosity, preview_bytes, max_output_bytes) are left unexplained, so an agent cannot reason about whether to adjust defaults or how they affect the call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action verb ('Terminate'), names the exact resource ('server-managed command'), and names the identifying parameter ('command_id'). This makes it clearly distinct from sibling tools like exec_command or read_output, and the short JSON example reinforces intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives, when not to use it, or what prerequisites exist (e.g., command must be running and managed by the server). The phrase 'server-managed command' conveys only a broad context; no conditions, exclusions, or signal-selection advice are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dirList directoryC
Read-onlyIdempotent

List directory entries inside the configured workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
sortNoname
max_depthNo
recursiveNo
max_entriesNo
include_hiddenNo
include_ignoredNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the workspace boundary ('inside the configured workspace'), which is useful, but it does not disclose behaviors like symlink handling, recursion defaults, or entry filtering beyond what the schema already exposes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler or repetition. Every word contributes to stating the tool's purpose and workspace constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, no parameter descriptions, and a sibling tool named list_files that could easily be confused with this one, the description is too sparse for an agent to select and invoke the tool correctly in all cases. The output schema and annotations cover return values and safety, but the description still leaves usage and parameter semantics under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no meaning for any of the 7 parameters. The agent gets no guidance on 'path' interpretation, sorting semantics, recursion/depth limits, or hidden/ignored entry behavior from the description itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'List directory entries inside the configured workspace.' It is not a tautology and conveys scope, but it does not differentiate from the sibling list_files or other listing tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives such as list_files or read_file. There are no exclusions, prerequisites, or routing cues beyond the general implication that it lists directory contents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_filesList filesB
Read-onlyIdempotent

List workspace files using glob filters.

ParametersJSON Schema
NameRequiredDescriptionDefault
globNo
pathNo.
sortNopath
patternsNo
max_resultsNo
include_hiddenNo
include_ignoredNo
exclude_patternsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds little behavioral context beyond 'workspace files' and 'glob filters'; it does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no redundant filler. Every word contributes to stating the tool's scope and mechanism.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 8 parameters, no parameter descriptions, and a sibling list_dir that could overlap, this terse description is not enough for an agent to confidently choose and invoke the tool correctly. The output schema covers return structure, but invocation semantics remain under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description mentions only 'glob filters,' which loosely maps to glob/patterns but leaves path, sort, max_results, include_hidden, include_ignored, and exclude_patterns unexplained. The description does not compensate for the undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('List'), a resource ('workspace files'), and a mechanism ('glob filters'), making the tool's core purpose clear. It does not explicitly contrast itself with siblings like list_dir, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'using glob filters' implies that this tool is for pattern-based listing, which gives some usage context. However, it provides no explicit guidance on when to choose this over list_dir, read_file, or search_text, and no exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_fileRead fileB
Read-onlyIdempotent

Read a UTF-8 text file slice inside the configured workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
encodingNoutf-8
end_lineNo
max_bytesNo
max_linesNo
start_lineNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey read-only, idempotent, and non-destructive behavior. The description adds useful context about UTF-8 text and workspace confinement, but does not disclose behavior like truncation, line/byte limits, or handling of binary files.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word earns its place by conveying the resource type, scope, and slicing capability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite output schema and annotations covering some context, the tool has six parameters with zero schema descriptions and no usage guidance. The description alone does not sufficiently equip an agent to form correct calls or interpret parameter interactions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for six undocumented parameters. It only hints at 'slice' and 'UTF-8', providing no meaningful semantics for path, start_line, end_line, max_lines, max_bytes, or encoding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and resource ('UTF-8 text file slice') and scopes it to the configured workspace. It clearly distinguishes this from siblings like read_output and view_image, which operate on command output or images.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives such as list_files, search_text, or read_output. It does not state exclusions or prerequisites beyond the implicit workspace boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_outputRead outputA
Read-onlyIdempotent

Read retained command output using an output_ref returned by exec_command/write_stdin. Each stream retains the earliest output (head) plus the most recent output (rolling tail); bytes between them may be evicted and are reported via evicted_gap_bytes. Example: {"output_ref":"command:abc:stdout","offset":0,"limit":4096}.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
streamNo
output_refYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses non-obvious retention behavior: earliest output plus rolling tail, with evicted bytes reported via evicted_gap_bytes. This goes beyond the annotations (readOnlyHint, idempotentHint) and gives the agent an accurate mental model of what data may be missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences plus a compact JSON example. Every sentence adds value: source of output_ref, retention behavior, and a realistic invocation. No filler or repetition of schema/annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the essential behavior, the provenance of output_ref, and an example invocation. Since an output schema exists, return-value details are not required. The stream parameter and offset pagination behavior are minor gaps that could be clarified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It provides a concrete example showing output_ref format ('command:abc:stdout') and uses of offset/limit, but it does not explain the stream parameter or offset semantics in prose. Partial compensation leaves some parameters under-documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Read retained command output using an output_ref returned by exec_command/write_stdin'. This clearly distinguishes read_output from read_file and other siblings by tying it to command-execution output rather than file contents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly ties usage to output_refs returned by exec_command/write_stdin, establishing a clear when-to-use context. It does not list exclusions or name alternatives, but the precondition is specific enough for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_permissionsRequest permissionsA
Read-only

Report scoped permission-request status without silently granting operations.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoonce
reasonYes
argumentsYes
tool_nameYes
permissionYes
ttl_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and destructiveHint=false. The description adds meaningful behavioral context by clarifying that the tool reports permission-request status and does not silently grant operations, which is especially valuable given the potentially misleading tool name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler or repetition. The core scoping qualifier is front-loaded, and the behavioral caveat is placed immediately after, making the definition easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and reduces the need to describe return values, the description omits critical invocation context: how arguments should be structured, when permission requests apply, and what the status report covers. For a tool with 6 parameters and no parameter descriptions, this is not complete enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description explains none of the six parameters. It gives tool-level context only, leaving agents to infer the meaning of tool_name, permission, scope, ttl_seconds, reason, and arguments from names and enums alone. The description does not compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is specific: it names a verb ('Report'), a resource ('scoped permission-request status'), and an important exclusion ('without silently granting operations'). This clearly distinguishes it from sibling tools like exec_command and apply_patch that actually execute operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no explicit guidance on when to call this tool versus alternatives, nor does it mention any prerequisites or workflow context. The relationship to exec_command and apply_patch is only implied by the schema's tool_name enum, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_textSearch textB
Read-onlyIdempotent

Search UTF-8 workspace files for text or regex matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
globNo
pathNo.
queryYes
regexNo
max_resultsNo
context_linesNo
exclude_globsNo
include_globsNo
case_sensitiveNo
max_preview_bytesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, so the safety profile is covered. The description adds useful context that the search is confined to UTF-8 workspace files, but it does not disclose result truncation, preview limits, or other search behavior; this is acceptable given the annotations but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single active-voice sentence with no filler. It front-loads the action, scope, and supported modes, making it easy to scan while avoiding repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter tool with zero parameter descriptions, the description provides only a high-level overview. It is sufficient for a simple query-only invocation because defaults are in the schema and the output schema exists, but it leaves filtering, pathing, and result-control semantics largely unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only hints at query and regex via 'text or regex matches'. It does not explain glob, path, include/exclude_globs, max_results, context_lines, case_sensitive, or max_preview_bytes, leaving most of the 10 parameters semantically undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Search'), a clear resource ('UTF-8 workspace files'), and the supported modes ('text or regex matches'). This is distinct from siblings like read_file or list_files, so an agent can tell this is a content-search tool rather than a file-reading or listing tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as read_file or grep-like behavior, nor any exclusions or conditions that would route an agent here. The intended usage must be inferred entirely from the tool name and the generic search verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

server_infoServer infoA
Read-onlyIdempotent

Return server, workspace, project-context, auth, policy, and fixed-tool metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe read operation. The description adds useful content detail (the metadata categories) but discloses no additional behavioral traits such as output size, permission requirements, or latency. This is acceptable for a zero-parameter info tool but does not go beyond the annotation baseline significantly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, directly stated sentence with no filler. The main verb and object are front-loaded, and the category list is compact and useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only metadata tool with an output schema present, the description is complete. It tells the agent what categories of information are available capital and the annotations cover safety, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters SB, so parameter semantics are largely moot; the baseline is 4. The description adds value by enumerating what the returned metadata covers, which helps the agent understand the semantic scope even though no parameters exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') with a clear resource (server metadata) and enumerates the exact categories included. This clearly distinguishes it from sibling tools that focus on files, commands, git, or permissions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The use case is implied: an agent would call this when it needs server/workspace/auth/policy context. However, the description gives no explicit guidance on when to prefer this over alternatives such as check_exec_environment or when not to use it. The 'when to use' context is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

view_imageView imageC
Read-onlyIdempotent

Return a workspace image as MCP image content.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
max_bytesNo
max_widthNo
max_heightNo
auto_resizeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=true, idempotentHint=true, and destructiveHint=false, the description adds the output format ('MCP image content') which is useful but not substantive behavioral context. It does not disclose error behavior, side effects, or resizing logic, but the annotation coverage lowers the bar.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the core action and output type. It earns its place but provides little beyond the barest statement of purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five parameters and zero schema descriptions, the description should compensate by explaining parameter behavior or usage context. It does neither. While an output schema exists, the tool remains significantly under-described for an agent to invoke it correctly, especially regarding resizing and byte limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no information about the five parameters (path, max_bytes, max_width, max_height, auto_resize). An agent has no way to know what these parameters control or how they affect the output.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and resource ('workspace image'), and specifies the output format ('MCP image content'). This distinguishes it from sibling tools like read_file (text) and list_dir (directory listing), though it does not explicitly name those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. An agent is not told to use it for image files or that read_file handles text. The context is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_stdinWrite stdinB

Poll or interact with a running command by command_id. Empty chars wait for output; non-empty chars writes to stdin. Example: {"command_id":"abc","chars":"","yield_time_ms":10000}.

ParametersJSON Schema
NameRequiredDescriptionDefault
charsNo
verbosityNo
command_idYes
preview_bytesNo
yield_time_msNo
max_output_bytesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
okYes
errorNo

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide no behavioral hints (all false), so the description carries the burden. It discloses the key non-obvious behavior: empty chars poll for output while non-empty chars write to stdin. This is critical for correct invocation and is not evident from the schema. It also provides a concrete example. It does not disclose side effects or limitations like whether writing is blocking, but the main behavioral trait is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the purpose is stated first, followed by the critical chars rule and a concrete example. No wasted words. The JSON example is a useful structural aid. Could be considered slightly terse, but every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need no explanation. The core interaction modes are covered, but the description lacks guidance on how command_id relates to exec_command, what verbosity levels mean, and how output limits interplay with polling. For a 6-parameter tool with 0% schema coverage, this leaves gaps an agent must resolve elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so all parameter semantics must come from the description. The description explains 'chars' (empty vs non-empty) and implicitly command_id and yield_time_ms via the example, but verbosity, preview_bytes, and max_output_bytes are left unexplained. Names and defaults give hints, but the description does not add sufficient meaning for these parameters, making correct selection harder.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb-resource pair: 'poll or interact with a running command by command_id.' It distinguishes from siblings like exec_command (starts) and kill_command (terminates) by focusing on stdin interaction and polling. The empty/non-empty chars rule adds specificity. However, it doesn't explicitly name a sibling alternative, so not a perfect 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus read_output or exec_command. The description implies usage via behavior but never states conditions or alternatives. An agent must infer that this is for sending input, not retrieving output. Lacks when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 18 tool updatesv0.2.2
    • First observedapply_patch
    • First observedcheck_exec_environment
    • First observedexec_command
    • First observedgit_blame
    • First observedgit_diff
    • First observedgit_log
    • First observedgit_show
    • First observedgit_status
    • First observedkill_command
    • First observedlist_dir
    • First observedlist_files
    • First observedread_file
    • First observedread_output
    • First observedrequest_permissions
    • First observedsearch_text
    • First observedserver_info
    • First observedview_image
    • First observedwrite_stdin

TDQS

B3.4/5.0

Scored across 18 tools

Disambiguation4/5

Most tools target a distinct action or resource, and the descriptions clarify their roles. Minor overlap exists between list_dir and list_files for workspace browsing, and server_info and check_exec_environment both report environment metadata, but an agent can generally disambiguate them.

Naming Consistency4/5

The naming is largely consistent with snake_case verb_noun patterns like read_file, apply_patch, and kill_command, plus a coherent git_* prefix group. A few names like server_info and check_exec_environment are more noun-like or verbose, but the overall pattern is predictable.

Tool Count4/5

Eighteen tools is slightly above the typical well-scoped range, but each tool serves a plausible purpose for a coding workspace server: file inspection, patching, command execution, git inspection, and permissions. It is not bloated enough to feel overwhelming or redundant.

Completeness4/5

The server covers core coding workflows well: reading and searching files, applying patches, running and interacting with commands, and inspecting git history. Some operations like explicit file creation/deletion or git writing operations are not first-class tools, but exec_command and apply_patch provide workarounds, so the gaps are minor.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Exposes a secure, path-confined bridge to a local workspace and git remotes, enabling MCP clients to search, read, write, reset files, and perform git operations.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Gives any MCP-compatible AI chat or agent a safe, model-neutral coding runtime with file read/search, structured multi-file patches, command execution, interactive sessions, and git operations, all confined to a single workspace and gated by permission modes.
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Gives MCP-compatible AI clients safe, hands-on access to local codebases: file reading/search, multi-file patches, command execution, interactive sessions, Git inspection, and coordination of local agent providers such as Antigravity and OpenCode.
    18
    Apache 2.0