Coding Tools MCP
Provides Git repository tools for inspecting status, diffs, commit history, file details, and blame information.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Coding Tools MCPrun the test suite and fix the first failure"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Coding Tools MCP
English | 简体中文
Give any AI chat or agent a safe pair of hands on your codebase.
Coding Tools MCP is a model-neutral coding runtime served over the Model Context Protocol: file reading and search, structured multi-file patches, command execution, interactive sessions, and git — one server that any MCP client can drive. Claude Desktop, Claude Code, Cursor, Cline, or an agent you build yourself all get the same 20 battle-tested tools, confined to one workspace, gated by permission modes.

Why people use it
It turns a chat app into a coding agent. Claude Desktop — or any MCP chat client — gets real repo access with the subscription you already have. No extra product required.
Safety is the product, not an afterthought. One workspace root per server. Absolute paths,
..traversal, and symlink escapes are rejected. Permission modes gate network access, shell expansion, inline scripts, and destructive commands. On Linux, Landlock adds kernel-level filesystem confinement.It is model- and vendor-neutral. A fixed, truthfully annotated catalog — no profile switching, no annotation games. Swap models or clients freely; the runtime and its behavior stay put.
It is engineered for context windows. Results are summarized, paginated, and capped by design; serialized tool-result bytes dropped 37% release-over-release on the deterministic dogfood workload with unchanged task completion.
Related MCP server: Cloud Harness MCP
Quickstart
Run it with whichever toolchain you already have (the server is Python ≥ 3.11
from PyPI; the npm package is a thin launcher that starts it via uv or
pipx):
uvx coding-tools-mcp --stdio --workspace /path/to/repo # Python toolchain
npx coding-tools-mcp --stdio --workspace /path/to/repo # Node toolchainWire it into Claude Desktop, Claude Code, Cursor, or Cline — the JSON is the
same everywhere (swap uvx for npx if you prefer Node):
{
"mcpServers": {
"coding-tools": {
"command": "uvx",
"args": ["coding-tools-mcp", "--stdio", "--workspace", "/path/to/repo"]
}
}
}Then ask your client: "run the test suite and fix the first failure."
Prefer HTTP? Drop --stdio and the server speaks Streamable HTTP on
http://127.0.0.1:8765/mcp (MCP 2025-11-25, with 2025-06-18
compatibility). A one-line installer, per-client walkthroughs, and
troubleshooting live in docs/quickstart.md and
docs/mcp-client-config.md.
Seven things to try
1. Make Claude Desktop your coding agent. The config above is all it takes — the chat window you already pay for can now read, patch, test, and commit-review a real repository.
2. Code on your own machine from anywhere.
CODING_TOOLS_MCP_AUTH_MODE=bearer ./scripts/tunnel.sh cloudflared /path/to/repoLoopback bind + authenticated HTTPS tunnel (cloudflared, ngrok, or
Microsoft Dev Tunnel). Point claude.ai on your phone at
https://<tunnel-host>/mcp and drive your home workstation from anywhere.
Bearer tokens and OAuth 2.1 + PKCE (with RFC 7591 dynamic registration) are
built in. → docs/remote-mcp.md
3. Let an agent loose on untrusted code — inside a disposable sandbox.
docker build -t coding-tools-mcp-sandbox:local .
docker run --rm --init -it -p 8765:8765 -v "$PWD:/workspace" coding-tools-mcp-sandbox:localA containerized server with toolchains and caches preconfigured, safe to point at a sketchy PR and destroy afterwards. → docs/docker.md
4. Spin up a cloud sandbox with one MCP call. The bundled
Cloudflare Worker control plane exposes
start_coding_tools_sandbox as an MCP tool: one call dispatches a GitHub
Actions runner that boots the Docker sandbox and publishes it behind an
authenticated Cloudflare Tunnel. Ephemeral compute, no server of your own.
5. Drive it from a GUI.
python -m pip install "coding-tools-mcp[desktop]"
coding-tools-mcp-desktopPer-workspace profiles, server and tunnel start/stop, credential setup with clipboard helpers, live health checks. English and 简体中文.
6. Keep an interactive session alive. exec_command starts a REPL or
debugger under a real PTY; write_stdin feeds it across turns; read_output
pages long output; kill_session cleans up. Long-running processes are
first-class, with deadline watchdogs and bounded buffers.
7. Give your own agent production-grade hands. Building an agent loop with the Anthropic SDK or anything else? Don't hand-roll file and exec tools — speak MCP to this server and inherit the whole safety boundary. → docs/embedding.md
The tool catalog
One stable, truthfully annotated set — permission modes change command
policy, never which tools the model sees. apply_patch is the sole
file-mutation primitive: staged, baseline-checked, atomic across files, with
rollback.
Group | Tools |
Files & search |
|
Execution |
|
Git |
|
Runtime |
|
Root AGENTS.md/CLAUDE.md files load into the initialize context
automatically. Tool content is concise agent-facing text;
structuredContent carries the complete machine result. Schemas and result
envelopes: docs/tools-and-schemas.md ·
docs/runtime-contract-v0.2.md.
Safety Boundary
Mode | Meant for | What it allows |
| day-to-day agent work | file tools and vetted commands; network-looking commands, shell expansion, inline scripts, and destructive commands all require explicit permission |
| local development | opens network, shell expansion, and inline scripts; keeps secret filtering and destructive-command checks |
| isolated containers/VMs only | disables |
Recursive listing and search exclude .git, node_modules, build outputs,
virtualenvs, and caches. Commands run with workspace-bound cwd, scrubbed
environment, timeouts, and output caps. Linux hosts with Landlock get
kernel-enforced filesystem confinement; other platforms get an explicit
warning — this is still not a complete OS sandbox, so use the Docker image or
a VM for genuinely untrusted work. Details:
SECURITY.md · docs/security-boundary.md ·
docs/permission-modes.md
Telemetry
The server sends anonymous usage telemetry (per-tool success/latency counters
and version/platform dimensions — never paths, arguments, commands, or file
contents) to help prioritize fixes. Disable it with
CODING_TOOLS_MCP_TELEMETRY=off or DO_NOT_TRACK=1; it is automatically off
in CI. CODING_TOOLS_MCP_TELEMETRY=debug prints every event to stderr instead
of sending. The full event list and guarantees are in
docs/telemetry.md.
Evidence, Dogfood and SWE-bench
Every release ships through a tag-triggered pipeline in which the compliance
suite, real-workload benchmark, and SWE-bench harness run from the same commit
that publishes to PyPI and npm — both via trusted publishing, npm with
provenance. Dogfood efficiency metrics are reproducible (make dogfood-smoke)
and checked in under reports/. This repository does not claim a
model-generated SWE-bench leaderboard result — see
docs/swe-bench.md for exactly what is and is not
measured. More: COMPLIANCE.md · BENCHMARK.md ·
docs/dogfood.md
Documentation
Getting started | |
Remote & sandboxed | |
Tools & contract | |
Execution | |
Integration | |
Security & quality | Security policy · Security boundary · CI and tests · Limitations · Competitive analysis |
Development
python -m pip install -e ".[dev]"
make ci # lint, typecheck, tests, protocol/integration suites, gatesThe full gate matrix is in docs/ci-and-tests.md.
License
This project is licensed under the Apache License 2.0.
If you use code, documentation, substantial implementation details, or derivative work from this project, preserve the copyright notice, license notice, and NOTICE file, and clearly attribute the original project.
Project: Coding Tools MCP
Author: Coding Tools MCP Contributors
Source: https://github.com/xyTom/coding-tools-mcp
Citation metadata is available in CITATION.cff.
Available Tools
18 toolsapply_patchApply patchADestructive
Stage, validate, and atomically apply a patch envelope. Example: *** Begin Patch *** Update File: app.py @@ -old +new *** End Patch
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive and non-idempotent. The description adds a useful layer: it validates before applying and applies atomically, which tells the agent that a failed validation should not leave partial edits. It does not spell out failure output, but it enriches the behavioral contract beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise lead sentence states the key behavior, followed immediately by a compact example that demonstrates the exact envelope syntax. Every line earns its place; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema covers return-value expectations, and the example covers patch syntax. Still, the dry_run semantics and what happens when validation fails are left implicit, which are meaningful gaps for an agent deciding whether to perform a destructive write or a trial run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It does teach the required patch envelope format through a concrete example, which is valuable. However, the dry_run parameter is never mentioned or explained; an agent must infer its behavior from the name and default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Stage, validate, and atomically apply a patch envelope.' This is far more informative than the title and clearly identifies a distinct file-modification capability among siblings like read_file and exec_command. The inline example pins down the exact input format, so an agent cannot confuse it with other tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The patch-envelope example implies that this is the intended mechanism for structured file edits, but the description never states when to prefer it over alternatives (e.g., exec_command or git apply) or when not to use it. It provides context but no explicit exclusions or alternative tool names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_exec_environmentCheck exec environmentARead-onlyIdempotent
Return lightweight exec_command sandbox and environment status known to the server.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds 'lightweight' and 'known to the server,' suggesting a cheap cached status view, which is useful context. It does not cover authentication or staleness, but annotations carry the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loading the action 'Return' and the key subject immediately. Every word contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is a parameterless, read-only status check with an output schema available, the description sufficiently covers the call contract. The high-level wording leaves specifics of the returned status to the output schema, which is appropriate. It could mention typical usage relative to exec_command, but that falls under usage guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters and an empty input schema, schema coverage is complete by definition. The description does not need to explain parameters, and its statement is consistent with a parameterless call, earning the baseline 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as returning exec_command sandbox and environment status, using the specific verb 'Return' and naming the resource. It is distinguishable from sibling tools like exec_command and server_info, though it does not enumerate what status fields are included.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to call this tool versus alternatives such as exec_command or server_info. It does not state preconditions, use cases, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exec_commandExecute commandADestructive
Run a bounded command under runtime policy. Pass workdir explicitly for reconnect-safe paths. A still-running command returns command_id. Example: {"cmd":"pytest -q","workdir":".","yield_time_ms":30000}. Retained output is bounded per stream; for very large output redirect to a file (cmd > out.log 2>&1) and page it with read_file or search_text.
| Name | Required | Description | Default |
|---|---|---|---|
| cmd | Yes | ||
| cwd | No | ||
| env | No | ||
| tty | No | ||
| stdin | No | ||
| workdir | No | . | |
| verbosity | No | ||
| timeout_ms | No | ||
| preview_bytes | No | ||
| yield_time_ms | No | ||
| max_output_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the description doesn't need to repeat those. It adds useful context: commands are bounded, still-running commands return a command_id, and retained output is bounded per stream. This goes beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus an example, front-loaded with the core purpose. Every sentence adds value; no filler. The example is compact and clarifies usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While an output schema exists (so return values aren't the description's job), the description does not explain several important parameters and their effects (e.g., verbosity, max_output_bytes). It covers typical usage but not advanced cases, so it's incomplete for a tool with this parameter count.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions workdir (explicitly for reconnect-safety), yield_time_ms (via example), and cmd (in redirect usage), but leaves env, tty, stdin, verbosity, timeout_ms, preview_bytes, and max_output_bytes unexplained. This is partial compensation, insufficient for a tool with 11 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Run') and resource ('bounded command under runtime policy'), and gives a concrete example. It clearly differentiates from file/sibling tools like read_file, list_dir, and search_text by describing command execution.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use: 'Pass workdir explicitly for reconnect-safe paths' and recommends redirecting large output to a file and paging with read_file or search_text. Does not explicitly state when to avoid this tool, but gives clear alternatives for output handling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_blameGit blameARead-onlyIdempotent
Return bounded git blame metadata for a workspace file.
| Name | Required | Description | Default |
|---|---|---|---|
| rev | No | ||
| path | Yes | ||
| end_line | No | ||
| max_lines | No | ||
| start_line | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds the useful behavioral clue that results are 'bounded' and scoped to a workspace file, which meaningfully narrows expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One succinct, front-loaded sentence with no filler or repetition. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a standard read-only git tool with an output schema and strong safety annotations, the description is nearly sufficient. Minor gaps like line-range inclusivity and rev interpretation remain, but they are partially inferable and do not block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate for parameter meaning. It only hints at 'bounded' behavior and never explains path, rev, start_line, end_line, or max_lines semantics; the agent must infer meaning from names and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says 'Return bounded git blame metadata for a workspace file' – a specific verb, a clear resource (git blame), and a narrow scope. It distinguishes the tool from siblings like git_log, git_diff, and git_show without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use git_blame versus the sibling git_* tools. It does not state exclusions, prerequisites, or conditions that would route an agent to an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_diffGit diffCRead-onlyIdempotent
Return unified git diff for workspace changes.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| paths | No | ||
| staged | No | ||
| unstaged | No | ||
| max_bytes | No | ||
| context_lines | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds 'unified' and 'workspace changes' as useful context, but it does not disclose defaults, truncation behavior via max_bytes, or how staged/unstaged changes are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise sentence with no filler and front-loads the core action and output format. It is efficient, though the brevity is partly a product of omitting parameter and usage detail, so it is not a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six parameters and zero schema descriptions, this is too thin. It states the core behavior and annotations cover safety, but it omits how to restrict paths, choose staged vs unstaged diffs, control context lines, or handle large outputs. Better than a pure tautology, but not adequate for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description mentions none of the six parameters: path, paths, staged, unstaged, max_bytes, or context_lines. The description therefore adds no semantic meaning beyond the bare parameter names and defaults already present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and clearly identifies the resource: a unified git diff scoped to workspace changes. The 'workspace changes' scope separates it broadly from commit-history siblings like git_show and git_log, though it does not explicitly name any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to use git_diff versus sibling tools such as git_status, git_show, or git_blame. It also does not mention exclusions, prerequisites, or typical workflows, so the agent must infer the usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_logGit logBRead-onlyIdempotent
Return recent git commits with bounded structured metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | HEAD | |
| path | No | . | |
| skip | No | ||
| max_count | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds 'bounded structured metadata', which hints at limited, schema-shaped output, but does not detail ordering, pagination, or other behavioral traits. Given the strong annotations, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with the key action and scope front-loaded. Every word earns its place and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with no required parameters and an output schema present, the description is partially complete. It lacks sibling-usage guidance and parameter semantics, but the annotations and schema defaults cover much of the operational context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the four parameters. It does not explain ref, path, skip, or max_count; 'recent' and 'bounded' only loosely hint at defaults and limits. The property names are self-explanatory, but the description itself adds minimal parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('recent git commits'), and adds the useful qualifier 'bounded structured metadata'. It is distinguishable from siblings like git_diff, git_show, and git_blame, though it does not explicitly name them or the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use git_log versus sibling tools such as git_show or git_diff, nor are there exclusions or alternative references. The word 'recent' implies history inspection, but the decision context is left entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_showGit showBRead-onlyIdempotent
Return bounded git show output for a revision.
| Name | Required | Description | Default |
|---|---|---|---|
| rev | No | HEAD | |
| path | No | ||
| paths | No | ||
| max_bytes | No | ||
| include_diff | No | ||
| context_lines | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the 'bounded' qualifier, which hints at output size constraints, but does not elaborate on behavior when limits are exceeded or how the output is structured. Given annotation coverage, this is a minimal but acceptable addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with zero redundancy. It communicates the core purpose efficiently and is appropriately front-loaded with the action and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has six parameters (all optional) and an output schema (though not shown), but the description lacks explanations for parameter behavior, edge cases, or how the 'bounded' aspect is enforced. It is adequate for a simple read-only tool but does not fully cover the contextual nuances an agent might need for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning no parameter documentation exists. The description does not compensate by explaining any of the six parameters (rev, path, paths, max_bytes, include_diff, context_lines). While parameter names are somewhat self-explanatory, the description adds no semantic value beyond what the schema already provides, failing to fill the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and names the resource ('git show output') with a qualifier ('bounded') that hints at output limits. It is clear about the core action but does not explicitly distinguish it from siblings like git_log or git_diff, though the 'bounded' qualifier provides a slight differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as git_log or git_diff. The description only states what it does without any contextual cues for selection, making it purely implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_statusGit statusBRead-onlyIdempotent
Return git working tree status for the workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | . | |
| max_entries | No | ||
| include_untracked | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnlyHint, idempotentHint, and destructiveHint, so the description is not required to restate safety. It adds the workspace scope and confirms the return value is status, but does not disclose behavior like result truncation via max_entries or untracked-file handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word contributes to identifying the resource and operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with annotations and an output schema, the description is mostly complete: it states the workspace scope and the core function. It is slightly incomplete because it does not clarify how path or max_entries affect the result, though the schema partially covers these.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three parameters. It does not mention path, max_entries, or include_untracked, leaving the agent to rely on parameter names and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair, 'Return git working tree status,' which clearly identifies the operation. It does not explicitly differentiate from sibling git tools like git_diff or git_log, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the sibling git tools (git_diff, git_log, git_show, git_blame). The usage is only implied by the tool's name and purpose, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kill_commandKill commandBDestructive
Terminate a server-managed command by command_id. Example: {"command_id":"abc","signal":"KILL"}.
| Name | Required | Description | Default |
|---|---|---|---|
| signal | No | TERM | |
| wait_ms | No | ||
| verbosity | No | ||
| command_id | Yes | ||
| kill_wait_ms | No | ||
| preview_bytes | No | ||
| max_output_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description's 'Terminate' is consistent with that rather than adding much new behavioral detail. It adds a useful scoping trait ('server-managed') and hints at signal selection via the example, but it does not discuss consequences like force-kill escalation, output loss, or irreversibility beyond what the destructive hint implies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence plus a compact JSON example: it is front-loaded and has no filler. The example is useful, though it showcases the non-default 'KILL' signal while the schema default is 'TERM', making the example slightly less representative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema and destructive annotations, the tool has seven parameters and the description only covers two. It does not explain the wait/kill-wait process, the output-cap limits, or what happens after termination, leaving meaningful gaps for an agent deciding how to call the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, but it only mentions command_id and signal in the example. The six other parameters (notably wait_ms vs. kill_wait_ms, verbosity, preview_bytes, max_output_bytes) are left unexplained, so an agent cannot reason about whether to adjust defaults or how they affect the call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action verb ('Terminate'), names the exact resource ('server-managed command'), and names the identifying parameter ('command_id'). This makes it clearly distinct from sibling tools like exec_command or read_output, and the short JSON example reinforces intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives, when not to use it, or what prerequisites exist (e.g., command must be running and managed by the server). The phrase 'server-managed command' conveys only a broad context; no conditions, exclusions, or signal-selection advice are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dirList directoryCRead-onlyIdempotent
List directory entries inside the configured workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | . | |
| sort | No | name | |
| max_depth | No | ||
| recursive | No | ||
| max_entries | No | ||
| include_hidden | No | ||
| include_ignored | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the workspace boundary ('inside the configured workspace'), which is useful, but it does not disclose behaviors like symlink handling, recursion defaults, or entry filtering beyond what the schema already exposes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler or repetition. Every word contributes to stating the tool's purpose and workspace constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters, no parameter descriptions, and a sibling tool named list_files that could easily be confused with this one, the description is too sparse for an agent to select and invoke the tool correctly in all cases. The output schema and annotations cover return values and safety, but the description still leaves usage and parameter semantics under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no meaning for any of the 7 parameters. The agent gets no guidance on 'path' interpretation, sorting semantics, recursion/depth limits, or hidden/ignored entry behavior from the description itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'List directory entries inside the configured workspace.' It is not a tautology and conveys scope, but it does not differentiate from the sibling list_files or other listing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as list_files or read_file. There are no exclusions, prerequisites, or routing cues beyond the general implication that it lists directory contents.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_filesList filesBRead-onlyIdempotent
List workspace files using glob filters.
| Name | Required | Description | Default |
|---|---|---|---|
| glob | No | ||
| path | No | . | |
| sort | No | path | |
| patterns | No | ||
| max_results | No | ||
| include_hidden | No | ||
| include_ignored | No | ||
| exclude_patterns | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds little behavioral context beyond 'workspace files' and 'glob filters'; it does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no redundant filler. Every word contributes to stating the tool's scope and mechanism.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, no parameter descriptions, and a sibling list_dir that could overlap, this terse description is not enough for an agent to confidently choose and invoke the tool correctly. The output schema covers return structure, but invocation semantics remain under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description mentions only 'glob filters,' which loosely maps to glob/patterns but leaves path, sort, max_results, include_hidden, include_ignored, and exclude_patterns unexplained. The description does not compensate for the undocumented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List'), a resource ('workspace files'), and a mechanism ('glob filters'), making the tool's core purpose clear. It does not explicitly contrast itself with siblings like list_dir, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'using glob filters' implies that this tool is for pattern-based listing, which gives some usage context. However, it provides no explicit guidance on when to choose this over list_dir, read_file, or search_text, and no exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileRead fileBRead-onlyIdempotent
Read a UTF-8 text file slice inside the configured workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| encoding | No | utf-8 | |
| end_line | No | ||
| max_bytes | No | ||
| max_lines | No | ||
| start_line | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only, idempotent, and non-destructive behavior. The description adds useful context about UTF-8 text and workspace confinement, but does not disclose behavior like truncation, line/byte limits, or handling of binary files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word earns its place by conveying the resource type, scope, and slicing capability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite output schema and annotations covering some context, the tool has six parameters with zero schema descriptions and no usage guidance. The description alone does not sufficiently equip an agent to form correct calls or interpret parameter interactions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for six undocumented parameters. It only hints at 'slice' and 'UTF-8', providing no meaningful semantics for path, start_line, end_line, max_lines, max_bytes, or encoding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and resource ('UTF-8 text file slice') and scopes it to the configured workspace. It clearly distinguishes this from siblings like read_output and view_image, which operate on command output or images.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives such as list_files, search_text, or read_output. It does not state exclusions or prerequisites beyond the implicit workspace boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_outputRead outputARead-onlyIdempotent
Read retained command output using an output_ref returned by exec_command/write_stdin. Each stream retains the earliest output (head) plus the most recent output (rolling tail); bytes between them may be evicted and are reported via evicted_gap_bytes. Example: {"output_ref":"command:abc:stdout","offset":0,"limit":4096}.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| stream | No | ||
| output_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses non-obvious retention behavior: earliest output plus rolling tail, with evicted bytes reported via evicted_gap_bytes. This goes beyond the annotations (readOnlyHint, idempotentHint) and gives the agent an accurate mental model of what data may be missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences plus a compact JSON example. Every sentence adds value: source of output_ref, retention behavior, and a realistic invocation. No filler or repetition of schema/annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the essential behavior, the provenance of output_ref, and an example invocation. Since an output schema exists, return-value details are not required. The stream parameter and offset pagination behavior are minor gaps that could be clarified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides a concrete example showing output_ref format ('command:abc:stdout') and uses of offset/limit, but it does not explain the stream parameter or offset semantics in prose. Partial compensation leaves some parameters under-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read retained command output using an output_ref returned by exec_command/write_stdin'. This clearly distinguishes read_output from read_file and other siblings by tying it to command-execution output rather than file contents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly ties usage to output_refs returned by exec_command/write_stdin, establishing a clear when-to-use context. It does not list exclusions or name alternatives, but the precondition is specific enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_permissionsRequest permissionsARead-only
Report scoped permission-request status without silently granting operations.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | once | |
| reason | Yes | ||
| arguments | Yes | ||
| tool_name | Yes | ||
| permission | Yes | ||
| ttl_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and destructiveHint=false. The description adds meaningful behavioral context by clarifying that the tool reports permission-request status and does not silently grant operations, which is especially valuable given the potentially misleading tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler or repetition. The core scoping qualifier is front-loaded, and the behavioral caveat is placed immediately after, making the definition easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and reduces the need to describe return values, the description omits critical invocation context: how arguments should be structured, when permission requests apply, and what the status report covers. For a tool with 6 parameters and no parameter descriptions, this is not complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description explains none of the six parameters. It gives tool-level context only, leaving agents to infer the meaning of tool_name, permission, scope, ttl_seconds, reason, and arguments from names and enums alone. The description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is specific: it names a verb ('Report'), a resource ('scoped permission-request status'), and an important exclusion ('without silently granting operations'). This clearly distinguishes it from sibling tools like exec_command and apply_patch that actually execute operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no explicit guidance on when to call this tool versus alternatives, nor does it mention any prerequisites or workflow context. The relationship to exec_command and apply_patch is only implied by the schema's tool_name enum, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_textSearch textBRead-onlyIdempotent
Search UTF-8 workspace files for text or regex matches.
| Name | Required | Description | Default |
|---|---|---|---|
| glob | No | ||
| path | No | . | |
| query | Yes | ||
| regex | No | ||
| max_results | No | ||
| context_lines | No | ||
| exclude_globs | No | ||
| include_globs | No | ||
| case_sensitive | No | ||
| max_preview_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, so the safety profile is covered. The description adds useful context that the search is confined to UTF-8 workspace files, but it does not disclose result truncation, preview limits, or other search behavior; this is acceptable given the annotations but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single active-voice sentence with no filler. It front-loads the action, scope, and supported modes, making it easy to scan while avoiding repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with zero parameter descriptions, the description provides only a high-level overview. It is sufficient for a simple query-only invocation because defaults are in the schema and the output schema exists, but it leaves filtering, pathing, and result-control semantics largely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only hints at query and regex via 'text or regex matches'. It does not explain glob, path, include/exclude_globs, max_results, context_lines, case_sensitive, or max_preview_bytes, leaving most of the 10 parameters semantically undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Search'), a clear resource ('UTF-8 workspace files'), and the supported modes ('text or regex matches'). This is distinct from siblings like read_file or list_files, so an agent can tell this is a content-search tool rather than a file-reading or listing tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as read_file or grep-like behavior, nor any exclusions or conditions that would route an agent here. The intended usage must be inferred entirely from the tool name and the generic search verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
server_infoServer infoARead-onlyIdempotent
Return server, workspace, project-context, auth, policy, and fixed-tool metadata.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe read operation. The description adds useful content detail (the metadata categories) but discloses no additional behavioral traits such as output size, permission requirements, or latency. This is acceptable for a zero-parameter info tool but does not go beyond the annotation baseline significantly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, directly stated sentence with no filler. The main verb and object are front-loaded, and the category list is compact and useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only metadata tool with an output schema present, the description is complete. It tells the agent what categories of information are available capital and the annotations cover safety, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters SB, so parameter semantics are largely moot; the baseline is 4. The description adds value by enumerating what the returned metadata covers, which helps the agent understand the semantic scope even though no parameters exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') with a clear resource (server metadata) and enumerates the exact categories included. This clearly distinguishes it from sibling tools that focus on files, commands, git, or permissions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied: an agent would call this when it needs server/workspace/auth/policy context. However, the description gives no explicit guidance on when to prefer this over alternatives such as check_exec_environment or when not to use it. The 'when to use' context is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
view_imageView imageCRead-onlyIdempotent
Return a workspace image as MCP image content.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| max_bytes | No | ||
| max_width | No | ||
| max_height | No | ||
| auto_resize | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring readOnlyHint=true, idempotentHint=true, and destructiveHint=false, the description adds the output format ('MCP image content') which is useful but not substantive behavioral context. It does not disclose error behavior, side effects, or resizing logic, but the annotation coverage lowers the bar.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the core action and output type. It earns its place but provides little beyond the barest statement of purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With five parameters and zero schema descriptions, the description should compensate by explaining parameter behavior or usage context. It does neither. While an output schema exists, the tool remains significantly under-described for an agent to invoke it correctly, especially regarding resizing and byte limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about the five parameters (path, max_bytes, max_width, max_height, auto_resize). An agent has no way to know what these parameters control or how they affect the output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('workspace image'), and specifies the output format ('MCP image content'). This distinguishes it from sibling tools like read_file (text) and list_dir (directory listing), though it does not explicitly name those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives. An agent is not told to use it for image files or that read_file handles text. The context is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_stdinWrite stdinB
Poll or interact with a running command by command_id. Empty chars wait for output; non-empty chars writes to stdin. Example: {"command_id":"abc","chars":"","yield_time_ms":10000}.
| Name | Required | Description | Default |
|---|---|---|---|
| chars | No | ||
| verbosity | No | ||
| command_id | Yes | ||
| preview_bytes | No | ||
| yield_time_ms | No | ||
| max_output_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | |
| error | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide no behavioral hints (all false), so the description carries the burden. It discloses the key non-obvious behavior: empty chars poll for output while non-empty chars write to stdin. This is critical for correct invocation and is not evident from the schema. It also provides a concrete example. It does not disclose side effects or limitations like whether writing is blocking, but the main behavioral trait is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the purpose is stated first, followed by the critical chars rule and a concrete example. No wasted words. The JSON example is a useful structural aid. Could be considered slightly terse, but every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need no explanation. The core interaction modes are covered, but the description lacks guidance on how command_id relates to exec_command, what verbosity levels mean, and how output limits interplay with polling. For a 6-parameter tool with 0% schema coverage, this leaves gaps an agent must resolve elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so all parameter semantics must come from the description. The description explains 'chars' (empty vs non-empty) and implicitly command_id and yield_time_ms via the example, but verbosity, preview_bytes, and max_output_bytes are left unexplained. Names and defaults give hints, but the description does not add sufficient meaning for these parameters, making correct selection harder.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pair: 'poll or interact with a running command by command_id.' It distinguishes from siblings like exec_command (starts) and kill_command (terminates) by focusing on stdin interaction and polling. The empty/non-empty chars rule adds specificity. However, it doesn't explicitly name a sibling alternative, so not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus read_output or exec_command. The description implies usage via behavior but never states conditions or alternatives. An agent must infer that this is for sending input, not retrieving output. Lacks when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v0.2.2- First observed
apply_patch - First observed
check_exec_environment - First observed
exec_command - First observed
git_blame - First observed
git_diff - First observed
git_log - First observed
git_show - First observed
git_status - First observed
kill_command - First observed
list_dir - First observed
list_files - First observed
read_file - First observed
read_output - First observed
request_permissions - First observed
search_text - First observed
server_info - First observed
view_image - First observed
write_stdin
TDQS
Scored across 18 tools
Most tools target a distinct action or resource, and the descriptions clarify their roles. Minor overlap exists between list_dir and list_files for workspace browsing, and server_info and check_exec_environment both report environment metadata, but an agent can generally disambiguate them.
The naming is largely consistent with snake_case verb_noun patterns like read_file, apply_patch, and kill_command, plus a coherent git_* prefix group. A few names like server_info and check_exec_environment are more noun-like or verbose, but the overall pattern is predictable.
Eighteen tools is slightly above the typical well-scoped range, but each tool serves a plausible purpose for a coding workspace server: file inspection, patching, command execution, git inspection, and permissions. It is not bloated enough to feel overwhelming or redundant.
The server covers core coding workflows well: reading and searching files, applying patches, running and interacting with commands, and inspecting git history. Some operations like explicit file creation/deletion or git writing operations are not first-class tools, but exec_command and apply_patch provide workarounds, so the gaps are minor.
Maintenance
Related MCP Connectors
A MCP server built for developers enabling Git based project management with project and personal…
An MCP server that gives your AI access to the source code and docs of all public github repos
Permission-aware onboarding MCP server: answers about a codebase, filtered by the caller's role.
The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceExposes a secure, path-confined bridge to a local workspace and git remotes, enabling MCP clients to search, read, write, reset files, and perform git operations.-
- AlicenseNot gradedqualityAmaintenanceEnables AI clients to securely operate isolated coding workspaces with file, command, Git, and deployment tools via authenticated remote MCP.11MIT
- AlicenseNot gradedqualityCmaintenanceGives any MCP-compatible AI chat or agent a safe, model-neutral coding runtime with file read/search, structured multi-file patches, command execution, interactive sessions, and git operations, all confined to a single workspace and gated by permission modes.Apache 2.0
- AlicenseNot gradedqualityAmaintenanceGives MCP-compatible AI clients safe, hands-on access to local codebases: file reading/search, multi-file patches, command execution, interactive sessions, Git inspection, and coordination of local agent providers such as Antigravity and OpenCode.18Apache 2.0