Free Codin Agent
Enables video rendering and media helper capabilities through FFmpeg.
Provides Git integration for version control and project inspection, enabling agents to manage repositories and inspect history.
Provides headless language-server operations for Python projects.
Provides headless language-server operations for TypeScript projects.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Free Codin Agentrun the test suite and fix the first failing test"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Free Codin Agent 4 Ever
A client-agnostic Model Context Protocol (MCP) server that gives an AI coding host controlled access to local developer capabilities: files, shell, Git, LSP, browser automation, Windows UI automation, VS Code state, long-running jobs, checkpoints, and resource-aware execution.
This repository is intentionally provider-neutral. It does not require or dispatch work to any hosted model vendor.
What this public repository contains
Standard MCP over stdio for local MCP clients.
Standard MCP over Streamable HTTP at
/mcp.Legacy MCP HTTP+SSE compatibility at
/sse+/messages.Transactional file edits with checkpoints and optimistic SHA-256 preconditions.
Shell, PTY, background processes, durable jobs, Git, and project inspection.
Headless language-server operations.
Isolated Playwright browser automation.
Optional companion Chrome extension for the user's real browser.
Optional companion VS Code extension for editor diagnostics/references/debug state.
Windows UI automation and media helpers.
Filesystem root allowlisting and resource admission controls.
Related MCP server: AI Knowledge Center MCP
Arena / LM Arena
The server itself is plain MCP and is not tied to a specific chat product. Use it with an MCP-capable host or with an Arena deployment/product surface that accepts standard MCP servers. No private Arena API, browser cookie, session token, or reverse-engineered endpoint is included here.
Fastest installation on Windows
Option A — download and double-click
Download the repository as a ZIP and extract it.
Double-click
install.cmd.Enter the folder that the coding agent is allowed to access.
Copy the MCP configuration printed by the installer into your MCP client.
If Node.js 20+ is missing and Windows Package Manager is available, the installer offers to install Node.js LTS automatically. The agent itself is installed globally, so the extracted ZIP can be deleted afterwards.
After installation the MCP client does not need a repository path. It launches:
free-codin-agent stdioThe generated MCP configuration is:
{
"mcpServers": {
"free-codin-agent": {
"command": "free-codin-agent",
"args": ["stdio"]
}
}
}Option B — npm package
The package is structured for global npm installation:
npm install -g free-codin-agent-4-ever
free-codin-agent setupOnce the package is published to the npm registry, these two commands are enough.
Useful commands:
free-codin-agent start MCP over stdio
free-codin-agent setup choose/configure the allowed workspace
free-codin-agent doctor verify the local installation
free-codin-agent print-config
free-codin-agent http start the local HTTP transportConfiguration and runtime state are stored in the current user's standard application-data directories. They are never stored in this public repository.
Do I need a domain?
No for local MCP. A stdio-capable MCP client starts free-codin-agent
directly on the same machine, so there is no website, public IP, DNS record,
domain, TLS certificate, or tunnel involved.
Yes, an HTTPS-reachable endpoint is needed when the MCP host runs somewhere else and must connect back to the user's computer. That endpoint can be a per-user tunnel URL, a domain owned by that user, or a secure relay service. Every installation should have its own authenticated endpoint; users should not share one unrestricted public bridge.
A marketing/documentation website is optional and is unrelated to the MCP transport itself.
Requirements
Node.js 20 or newer.
Windows for the full Windows automation feature set.
Optional: Chrome for browser automation.
Optional: FFmpeg for video rendering.
Optional: Go / TypeScript / Python language servers depending on the projects you inspect.
Manual developer install
npm install
Copy-Item .env.example .envFor normal end users, prefer install.cmd / free-codin-agent setup.
The repository-local .env workflow is retained only for development and advanced deployments.
Start over stdio
This is the preferred local MCP transport:
npm startExample generic MCP client configuration for a source checkout:
{
"mcpServers": {
"free-codin-agent-4-ever": {
"command": "node",
"args": ["<ABSOLUTE_PATH_TO_REPOSITORY>/stdio.js"]
}
}
}Replace <ABSOLUTE_PATH_TO_REPOSITORY> with the location where you cloned
the repository. No machine path, username, drive letter, workspace, or account
is built into the project.
Start over HTTP
npm run start:httpDefault endpoint:
http://127.0.0.1:3000/mcpWhen MCP_API_KEY is set, clients must send:
Authorization: Bearer <MCP_API_KEY>Without a key, HTTP access is accepted only from a direct loopback connection. Requests arriving through a proxy/tunnel are rejected.
Security model
This server can edit files and execute commands. Treat it like local developer access.
Recommended public defaults:
Keep
MCP_HOST=127.0.0.1unless you intentionally expose the HTTP server.Set
MCP_ALLOWED_ROOTSto the repositories the agent may touch.Set a strong
MCP_API_KEYbefore any non-loopback exposure.Do not commit
.env, generated logs, browser profiles, screenshots, clips, checkpoints, memory, or job state.Review destructive tool calls in the MCP client.
Use the Chrome pairing token only with the unpacked extension instance you control.
Configuration
See .env.example. Important settings:
Variable | Purpose |
| Known repositories, separated by comma or semicolon |
| Default directory for relative paths |
| Optional filesystem sandbox |
| Bearer token for HTTP transport |
| Browser origins allowed to call HTTP endpoints |
| Enable/disable the companion Chrome bridge |
| Chrome extension pairing secret |
| Optional extension identity pin |
| Local VS Code companion endpoint |
Chrome companion extension
Open Chrome's extension management page.
Enable Developer mode.
Load
chrome-extensionas an unpacked extension.Set
CHROME_BRIDGE_TOKENin.env.Open the extension popup and save the same token and port.
Optionally set
CHROME_EXTENSION_IDafter Chrome assigns the unpacked extension an ID.
The extension has no hardcoded identity key in this repository.
VS Code companion extension
The source is under vscode-extension. It exposes only a loopback HTTP helper for diagnostics, references, definitions, open files, and debug state.
Its configuration key is:
freeCodinAgent.portDefault port: 3005.
Verification
npm run check
npm run privacy
npm run preflight
npm testPrivacy / publication hygiene
The public source tree intentionally excludes:
local
.envfiles and API keys;tunnel URLs and connector URLs;
browser profiles, screenshots, clips, logs, checkpoints, job/task state, and local memory;
prior conversation archives;
fixed Chrome extension identity keys;
machine-specific repository paths;
model-provider orchestration code and account-specific integrations.
Run your own secret scanner before publishing any later local changes.
License
ISC. See LICENSE.
Available Tools
114 toolsapply_patchADestructive
Apply a multi-file patch atomically (nothing is written unless every hunk applies). Preferred way to change code. Format: *** Begin Patch *** Update File: src/app.ts @@ function main() { (optional anchor line to jump near the change) context line (space prefix) -removed line +added line *** Add File: src/new.ts +file content line *** Delete File: src/old.ts *** Update File: a.ts *** Move to: b.ts *** End Patch Include ~3 unchanged context lines around each change. The result reports compiler/linter problems in the changed files and a checkpoint id for undo.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Base directory for relative paths (default: default workspace) | |
| check | No | Run compilers/linters on the changed files afterwards (default true) | |
| patch | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=false, and the description adds behavior they cannot express: the operation is atomic and nothing is written unless every hunk applies, the result reports compiler/linter problems, and a checkpoint id is returned for undo. Atomicity and undo are exactly the non-obvious traits an agent needs before mutating many files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the atomicity guarantee and preference statement in the first sentence, then spends the remainder on the format example, which earns its space for a grammar-driven parameter. It is long, but nearly every line carries required syntax or behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly describes what comes back (diagnostics plus checkpoint id) and covers the destructive/atomic semantics. It omits how failures are surfaced and how this relates to edit_file/write_file, which are the remaining gaps for a multi-file mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and documents cwd and check, but leaves the required 'patch' parameter undescribed. The description compensates by specifying the exact patch grammar (Begin/Update/Add/Delete/Move/End, +/-/space prefixes, optional anchor) and the context-line convention, which is what makes the required parameter usable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (apply) and resource (multi-file patch) plus a distinguishing trait: atomic, all-or-nothing. Combined with 'Preferred way to change code', an agent can separate it from nearby file-mutation siblings like edit_file and write_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions itself as the preferred mechanism for changing code, which is real usage guidance. It stops short of naming the alternatives (edit_file, write_file) or stating when a single-file edit is preferable, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batchBDestructive
Run several tool calls in one request. Parallel mode is allowed only when every called tool is read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| calls | Yes | ||
| parallel | No | ||
| stopOnError | No | ||
| failOnAnyError | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the risk profile is known. The description adds the read-only precondition for parallel execution, which is useful. But it omits error-handling behavior (what stopOnError vs failOnAnyError actually do, especially with both defaulting to false) and the 20-call cap from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with what the tool does followed by the key constraint. No filler. It is terse to the point of under-specification, but that is a completeness problem rather than a verbosity one.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a meta-tool with 4 parameters, 0% schema description coverage, and no output schema, the description is thin. It should define batch entry structure, the 20-call ceiling, and the distinct roles of stopOnError/failOnAnyError to make the tool safely invocable. Annotations cover safety but not orchestration semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate — and largely does not. Only the conditional meaning of 'parallel' is addressed; 'calls', 'stopOnError', and 'failOnAnyError' are never explained, and the maxItems:20 limit is not mentioned. This leaves the most consequential semantics undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Run several tool calls in one request.' An agent immediately knows this is a multiplexer/batch executor, distinct from any single-call sibling. However, it does not clarify what a 'call' entry must contain (tool name + args) or how results are returned, leaving some ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one explicit usage rule: parallel mode is permitted only when every called tool is read-only. That is genuine guidance. But it says nothing about when to prefer batch over issuing calls individually, nor about how stopOnError/failOnAnyError affect execution, so the guidance is partial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clickBDestructive
Click an element by ref (preferred), selector or text. Returns the updated snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from browser_snapshot, e.g. "e12" | |
| text | No | Visible text to match (alternative to ref) | |
| button | No | ||
| selector | No | CSS or Playwright selector (alternative to ref) | |
| snapshot | No | ||
| modifiers | No | ||
| doubleClick | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and non-idempotent, so the safety profile is largely covered. The description adds the return behavior ('Returns the updated snapshot'), but it omits meaningful behavioral context for a click tool — e.g. that clicking may navigate, submit forms, or trigger irreversible actions, which matters given destructiveHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the targeting preference front-loaded and the return value second; every word earns its place and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, zero-required, destructive mutation tool with no output schema and only 43% parameter documentation, the description is thin. It does not explain the button, modifiers, doubleClick, or snapshot options, nor the side effects an agent should anticipate before clicking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, and four parameters (button, snapshot, modifiers, doubleClick) have no schema description at all. The description only hints at the ref/selector/text triage, which the schema already labels as alternatives, so it does not compensate for the undocumented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (click) and resource (an element), plus the three supported targeting modes (ref, selector, text). It is clearly distinguishable from sibling read-only tools like browser_get_text or browser_snapshot, though it does not distinguish itself from the near-identical chrome_click or adjacent interaction tools such as browser_hover.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(preferred)' on ref implies a targeting precedence, which is light usage guidance. However, there is no explicit when-to-use vs. when-not, and no mention of prerequisites (e.g. that a ref must come from a prior browser_snapshot) or which siblings to prefer for hovering, typing, or pressing keys.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_closeADestructive
Close the Playwright browser (the persistent profile stays on disk).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is covered. The description adds a valuable non-obvious behavioral detail: the persistent profile is not destroyed, which meaningfully clarifies the scope of the destructive action. It still does not say what live session state or tabs are lost when the browser closes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. The essential action is stated first, and the important caveat about the persistent profile is appropriately appended.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter close operation with no output schema, the description plus annotations provide sufficient context to invoke the tool correctly. Minor omissions remain, such as behavior when no browser is open or confirmation of what session state is discarded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero input parameters, so there are no parameter semantics for the description to explain. The schema itself has 100% description coverage, and the baseline for a zero-parameter tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Close the Playwright browser.' The parenthetical narrows the scope by clarifying that the persistent profile remains on disk, which distinguishes this from destructive cleanup that would delete profile state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use or when-not-to-use guidance. It implies end-of-session cleanup, but it does not name alternatives such as browser_tabs, chrome_close_tab, or other browser lifecycle tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_consoleCRead-onlyIdempotent
Console messages, page errors, failed requests and dialogs captured from the browser.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | ||
| limit | No | ||
| onlyErrors | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, open-world behavior. The description usefully adds the content types captured (errors, failed requests, dialogs), but says nothing about the mutating effect of the 'clear' parameter, buffering, or whether logs persist across navigations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, and the most important content types are front-loaded. It reads as a fragment rather than a verb-led instruction, but no words are wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With three undocumented parameters, no output schema, and a 'clear' flag that appears to wipe captured data, the definition is too thin to invoke safely. An agent cannot know what clear does or what limit/onlyErrors control.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three parameters, so the description is the only place semantics could be added — and it adds none. 'clear', 'limit', and 'onlyErrors' are never mentioned; 'onlyErrors' is only obliquely implied by the phrase 'page errors'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resources retrieved — console messages, page errors, failed requests, dialogs — so an agent knows exactly what it gets. It omits an action verb and gives no differentiation from the sibling chrome_console, which appears to expose the same data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as chrome_console or browser_dialog despite heavy sibling overlap. The agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_dialogBDestructive
Inspect or answer a JavaScript dialog (alert/confirm/prompt) in the Playwright browser. Dialogs are never accepted automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | status | |
| promptText | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the write-capable nature is covered structurally. The description usefully adds that dialogs are never auto-accepted (so the agent must handle them explicitly), but it does not explain what accept vs dismiss does to the underlying page or how promptText applies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the resource and scope. No filler, though the second sentence is a caveat rather than a usage instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with an undocumented enum and no output schema, the description leaves real gaps: it never maps the action values to behavior or explains promptText, so an agent must infer the accept/dismiss/prompt semantics from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with 2 parameters, and the description explains neither 'action' (whose enum values status/accept/dismiss map directly onto the stated capabilities) nor 'promptText' (presumably the text submitted on accept of a prompt). Only the verbs 'inspect'/'answer' loosely hint at the action values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: inspect or answer a JavaScript dialog, scoped to the Playwright browser. It reasonably distinguishes it from the sibling chrome_dialog, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied — the tool is obviously for when a dialog is blocking the browser, but the description gives no guidance on when to use status vs accept vs dismiss, nor any prerequisite or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_evaluateBDestructive
Run JavaScript in the page. Pass an expression ("document.title") or a function body using return.
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=true, so the risk profile is covered by structured data. The description adds nothing behavioral beyond that: it says nothing about execution context/privileges, sandboxing, timeouts, or error behavior for arbitrary code execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and then the argument format. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one parameter and annotations covering safety, the description is nearly sufficient. But there is no output schema, so the return value (serialized result, or what happens on a thrown error / non-serializable value) should have been mentioned to let an agent call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden and does so usefully: it clarifies that 'script' can be a bare expression ("document.title") or a function body requiring an explicit return. That is real semantic guidance beyond the bare string type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Run JavaScript in the page.' This is clearly distinguishable from read-oriented siblings like browser_get_text or browser_snapshot, though it never explicitly names them or notes its overlap with browser_console/chrome_evaluate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the two argument forms but gives no guidance on when to reach for this tool versus browser_get_text, browser_snapshot, or browser_console. No prerequisites, no when-not, no alternatives are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_get_textCRead-onlyIdempotent
Visible text of the page or of an element (CSS selector).
| Name | Required | Description | Default |
|---|---|---|---|
| maxChars | No | ||
| selector | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive) and open-world behavior, so the bar is lower. The description adds the useful nuance that only 'visible text' is returned, but it omits truncation behavior for maxChars, handling of missing selectors, and default selector behavior, leaving meaningful gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no wasted words, and the core resource is front-loaded. It is arguably too thin, but as a pure conciseness measure it is efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-output-schema tool, the description leaves key behaviors unspecified: what happens when selector is omitted (whole page?), how maxChars truncation works, and failure behavior when a selector matches nothing. Annotations cover safety, but the description is insufficient for the agent to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that the selector is a CSS selector, but says nothing about the maxChars parameter, leaving one of two parameters undocumented in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('visible text') and its two scopes ('page or of an element'), making clear what kind of content is returned. It lacks an explicit verb and does not distinguish itself from siblings like browser_snapshot or browser_evaluate, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as browser_snapshot or browser_evaluate. The phrase 'page or of an element (CSS selector)' hints at two modes but provides no conditions or preferences for selecting between them or among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_hoverCDestructive
Hover over an element.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from browser_snapshot, e.g. "e12" | |
| text | No | Visible text to match (alternative to ref) | |
| selector | No | CSS or Playwright selector (alternative to ref) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (destructiveHint=true, openWorldHint=true, idempotentHint=false), so the description is not burdened with that. However, it adds nothing beyond them: no mention of whether hovering triggers hover-only UI, side effects, or failures on missing elements. No contradiction with the annotations, just zero added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four words, front-loaded, no filler. It is as concise as possible, though the brevity borders on under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-optional-param interaction tool with full schema coverage and no output schema, the minimal description is borderline adequate. It omits target-resolution precedence (ref vs text vs selector) and what hovering actually produces, which an agent would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each of ref, text, and selector documented as alternatives, so the schema carries semantics. The description adds nothing about precedence among the three targeting modes, which is the one thing missing. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: hover over an element. It is distinguishable from the other browser_* tools (click, type, scroll) by its unique action, though it offers no explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to hover versus clicking, typing, or waiting, and no note about prerequisites such as having a snapshot ref available. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_press_keyBDestructive
Press a key or chord on the page, e.g. "Enter", "Escape", "Control+A", "ArrowDown".
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| snapshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, readOnlyHint=false, openWorldHint=true, and non-idempotency, so the safety profile is covered. The description adds the chord syntax convention but says nothing about focus requirements, that keystrokes can mutate page state, or any rate/confirmation concerns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the action front-loaded and examples inline. No wasted wording, though it is sparse enough that a clause about focus or the snapshot flag could have earned its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutating tool with no output schema, the description covers the primary input but omits the snapshot parameter entirely and gives no behavioral context beyond annotations. Adequate but with a visible gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load. It does clarify the 'key' parameter with chord examples like 'Control+A' and 'ArrowDown', which is genuinely useful syntax information absent from the schema. However, the second parameter 'snapshot' is never mentioned, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Press a key or chord on the page') and supplies concrete examples, which separates it from browser_type and browser_click. It does not explicitly name the sibling it differs from, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to press a key versus using browser_type, browser_click, or a chord-based alternative, and no stated prerequisites such as an element needing focus. Usage is only implied by the examples.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_record_startADestructive
Start recording a video of a browser session. Opens a dedicated recording tab (with your profile logins); every browser_* action afterwards is filmed until browser_record_stop.
| Name | Required | Description | Default |
|---|---|---|---|
| fps | No | ||
| url | No | ||
| name | No | ||
| width | No | ||
| format | No | mp4 | |
| height | No | ||
| headless | No | ||
| resourceWaitMs | No | ||
| useProfileLogins | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructive/openWorld/non-idempotent, and the description adds real behavioral detail beyond them: a dedicated tab is opened, it uses the caller's profile logins, and the recording is scoped to subsequent browser_* actions until stop. It does not explain why the operation is flagged destructive or what happens to the recording output, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and followed by the lifecycle constraint. No filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the description covers the lifecycle adequately but says nothing about how the resulting video is identified or retrieved, nor about the parameter space (resolution, format, fps, headless behavior). It is minimally viable rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage across 9 parameters (fps, format, width/height, headless, resourceWaitMs, url, name, useProfileLogins), and the description documents none of them. It only obliquely hints at useProfileLogins via 'with your profile logins', so the agent must infer the meaning and units of nearly every parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start recording a video of a browser session') and immediately clarifies the mechanism ('Opens a dedicated recording tab'). An agent can distinguish this from browser_record_stop and from browser_screenshot without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context — recording begins here and captures 'every browser_* action afterwards ... until browser_record_stop' — which effectively names the terminating sibling. It lacks an explicit when-not-to-use (e.g., versus record_clip or one-off screenshots), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_record_stopBDestructive
Stop the recording and save the clip (mp4 via ffmpeg). Returns the file path.
| Name | Required | Description | Default |
|---|---|---|---|
| discard | No | ||
| trimStart | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds useful behavior beyond annotations by stating the output is an mp4 via ffmpeg and that it returns a file path, but it does not explain how the destructive discard option affects saved output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the action, output format, and return value without filler. Every clause contributes to understanding the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with two completely undocumented parameters and no output schema, the description is too sparse. It covers stop/save/return format, but omits any explanation of discard or trimStart, leaving significant gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter. It provides no meaning for 'discard' or 'trimStart', so an agent cannot infer how these options alter the stop/save behavior from the description alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Stop') and resource ('the recording'), and specifies the save behavior and output format. It distinguishes itself from browser_record_start by describing the stop/save action rather than starting a recording.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Stop the recording and save the clip' implies this should be used after a recording has been started, which is clear context. However, it does not name alternatives such as browser_record_start or record_clip, nor does it say when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_resizeCDestructive
Set the viewport size.
| Name | Required | Description | Default |
|---|---|---|---|
| width | Yes | ||
| height | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=false, destructiveHint=true, openWorldHint=true, and idempotentHint=false. The description adds no behavioral context beyond the basic action, such as whether the resize affects subsequent screenshots, evaluations, or page layout. With annotations covering the safety profile, the description's lack of additional detail is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. However, it is so terse that it borders on under-specification for a tool with two required parameters and no parameter documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter mutation tool, annotations cover the safety profile, but the description provides no parameter semantics, units, or usage context. With no output schema and 0% schema description coverage, an agent lacks enough guidance to invoke this correctly without guessing units or intent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required parameters (width and height). The description does not mention units (pixels?), valid ranges, or any meaning beyond what the bare property names already imply, so it fails to compensate for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Set the viewport size.' This clearly distinguishes it from siblings like browser_screenshot or browser_evaluate, which capture or inspect browser state. It lacks any further scope or differentiation, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no prerequisites, and no context about how resizing interacts with other browser operations. The description only states what it does, not when or why to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotARead-onlyIdempotent
Screenshot of the page (or one element). The image is returned so you can see it, and saved to disk.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from browser_snapshot, e.g. "e12" | |
| text | No | Visible text to match (alternative to ref) | |
| fullPage | No | ||
| selector | No | CSS or Playwright selector (alternative to ref) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds a genuine behavioral fact not in annotations: the image is both returned to the model and persisted to disk. It omits where the file lands and what format, which caps it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that loses no words and immediately states scope and return behavior. Nothing redundant or padding-like.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so partially, explaining that an image comes back and is saved to disk. It leaves unspecified the save location, image format, and whether the returned image is inline, which an agent might need for follow-up steps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so ref/text/selector are already documented in the schema; the description adds only the page-vs-element framing. The boolean fullPage has no description in either place, leaving one parameter's semantics unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Screenshot of the page (or one element)') and clarifies scope (whole page vs. a single element), which lets an agent distinguish it from browser_snapshot or browser_get_text. It does not name those siblings explicitly, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by 'The image is returned so you can see it,' which suggests reaching for this when visual output is needed rather than text. No when-to-use/when-not or comparison against browser_get_text or browser_snapshot is given, so the routing decision is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_scrollCDestructive
Scroll the page by deltaY/deltaX pixels, or scroll an element (ref/selector/text) into view.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from browser_snapshot, e.g. "e12" | |
| text | No | Visible text to match (alternative to ref) | |
| deltaX | No | ||
| deltaY | No | ||
| selector | No | CSS or Playwright selector (alternative to ref) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation/safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false), and the description adds nothing behavioral on top — no note on prerequisites (page must be loaded), what happens if the ref/text/selector is not found, or whether the scroll delta is bounded. For a state-changing tool with no output schema this is a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with both use modes front-loaded and no filler. It could have been split into a target-selection clause and a delta clause for faster scanning, but it wastes nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Five optional parameters, no output schema, and the description only broadly names both modes without covering default deltas, overloaded-parameter precedence, or failure behavior when a target is not found. It is minimally adequate but not complete for a tool with this many knobs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, and the description contributes the useful unit detail that deltaY/deltaX are in pixels, which the schema omits. However it does not explain precedence when ref, text, and selector are all supplied, nor the deltaY=600 default behavior, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scroll) plus both targeting modes (pixel delta vs element into view), so the agent knows exactly what the tool does. It does not differentiate itself from the near-identical sibling chrome_scroll or from other viewport-affecting tools, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance at all: nothing about when to scroll versus calling browser_snapshot, browser_wait_for, or a click that auto-scrolls. No prerequisites or sequencing hints are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_select_optionCDestructive
Select option(s) in a .
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from browser_snapshot, e.g. "e12" | |
| text | No | Visible text to match (alternative to ref) | |
| value | No | ||
| values | No | ||
| selector | No | CSS or Playwright selector (alternative to ref) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=false, and openWorldHint=true, covering the safety profile. The description adds no further behavioral context: it does not mention event triggering, multi-select behavior, side effects, or what 'destructive' might mean in this context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words, but it is under-specified for a tool with five parameters and browser-interaction complexity. Brevity here reflects missing information rather than efficient communication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only partial schema parameter coverage, the description should explain how to choose between ref, text, and selector, and what value/values do. It provides none of this, leaving the agent with insufficient context beyond the annotations and schema for a five-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 60% (ref, text, and selector have descriptions; value and values do not). The description does not explain any parameter semantics or clarify how ref, text, value, values, and selector interact, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: select option(s) in a <select>. It is clear enough to know the operation, but it does not distinguish this tool from siblings like browser_click or browser_type, nor does it mention the browser-automation context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as browser_click, browser_type, or browser_evaluate. The description only restates the operation with no context about preconditions or selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_snapshotARead-onlyIdempotent
Accessibility snapshot of the current page with [ref=eN] handles. Use it before clicking/typing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds meaningful context by disclosing the return format (ref handles) and the workflow dependency for click/type, but says nothing about snapshot freshness, staleness after navigation, or size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the output format stated first and the usage trigger second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only snapshot tool with no output schema, the description covers what the tool returns (ref handles) and when to reach for it. It could mention that refs become invalid after navigation, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing for the description to clarify beyond what the empty schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource (accessibility snapshot of the current page) and specifies the output shape ([ref=eN] handles), which distinguishes it from browser_screenshot and browser_get_text by implication. It stops short of explicitly contrasting itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use it before clicking/typing" gives a concrete trigger condition and positions the tool in the interaction workflow. It does not state when-not to use it or name the visual/text alternatives (browser_screenshot, browser_get_text) explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tabsCDestructive
List, open (new), select or close tabs by index.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| index | No | ||
| action | No | list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=true, so the agent knows some actions mutate and destroy state. The description adds the useful fact that one tool multiplexes four distinct operations addressed by index, but it never says that 'close' is irreversible or that index is mandatory for select/close. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core purpose front-loaded and zero filler. It is efficient, though the terseness is partly under-specification rather than true compression.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-action multiplexer with zero required parameters, no output schema, and no action/parameter mapping rules, the description is too thin. An agent still cannot tell which parameters are valid per action or how this tool relates to the chrome_* tab siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and no parameter is required, so the description carries the full burden — yet it only obliquely implies 'index' is the addressing key and 'open (new)' hints that url pairs with the new action. It does not state that url is used only with action='new', nor that index is required for select/close, leaving invalid combinations undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names four concrete verbs (list, open, select, close) and the resource (tabs), plus the addressing mechanism (by index), so an agent immediately understands it is a tab-management multiplexer. It falls short of a 5 because it never distinguishes itself from the closely-overlapping siblings chrome_list_tabs, chrome_new_tab, chrome_switch_tab, and chrome_close_tab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'by index' is the only usage signal; there is no statement of when to pick this tool over the four chrome_* tab siblings, no prerequisites, and no guidance on which action to use in which situation. The four actions are simply enumerated without routing logic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_typeADestructive
Type into an input/textarea/contenteditable. Replaces its value unless slowly: true (types key by key). submit: true presses Enter.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from browser_snapshot, e.g. "e12" | |
| text | Yes | ||
| slowly | No | ||
| submit | No | ||
| selector | No | CSS or Playwright selector (alternative to ref) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, but the description is what explains why: the default behavior replaces the field's existing value, and only slowly:true preserves key-by-key entry. It also discloses that submit:true fires Enter, a side effect not captured by annotations. Focus/permission requirements remain unmentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses with no filler, and the destination plus the default overwrite behavior are front-loaded. The telegraphic 'slowly: true' shorthand is compact but slightly cryptic without a comma-joined explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with annotations and no output schema, the description covers the main non-obvious behaviors (overwrite default, slow typing, Enter submission). It omits failure handling and ref-vs-selector precedence, which are minor given the schema already defines both targeting options.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% and slowly/submit have no schema descriptions, yet the description supplies their exact semantics (slow = per-key typing, submit = Enter press) plus the default replace behavior of text. ref and selector are already documented in the schema, so the remaining gap is small.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (type) and the exact target set (input/textarea/contenteditable), which is more precise than the tool name alone. It does not name a sibling or draw an explicit boundary against browser_press_key or browser_click, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the target-element list, and the slowly/submit flags hint at when each mode applies (character-by-character vs instant, Enter submission). There is no explicit when-to-use/when-not guidance and no routing to browser_press_key for plain key presses.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_upload_fileCDestructive
Upload local file(s) through a file input or an upload button (ref/selector/text).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from browser_snapshot, e.g. "e12" | |
| path | No | ||
| text | No | Visible text to match (alternative to ref) | |
| paths | No | ||
| selector | No | CSS or Playwright selector (alternative to ref) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and non-idempotent behavior, so safety context exists. The description adds virtually nothing beyond them — no note that the named local path must exist, that it interacts with the host filesystem, or what happens if the element is not a valid file input.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded. Effectively sized, though the parenthetical overloads it with locator terminology that arguably belongs in usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, open-world mutation tool with five parameters, no output schema, and no required params, the description omits the prerequisite (ref acquisition), the path/paths semantics, and any failure or side-effect behavior. What an agent needs to call it correctly is largely missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description names ref/selector/text, mapping to three of the five parameters, but says nothing about the single-vs-multiple file distinction between 'path' and 'paths', which is the most consequential semantic in the schema. With 60% schema coverage, some compensation occurs but the critical path/paths ambiguity is left unresolved.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Upload local file(s)') and mentions the target forms (file input or upload button), which distinguishes it from sibling tools like browser_click or browser_type. The trailing '(ref/selector/text)' is confusingly framed as if those were upload methods rather than element locators, slightly muddying the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, and no mention of prerequisites such as obtaining a ref from browser_snapshot before calling this. The agent is left to infer the workflow entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_wait_forBDestructive
Wait for a selector or text to appear (state visible/hidden/attached), or just wait timeoutMs.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | ||
| state | No | ||
| selector | No | ||
| timeoutMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=true, openWorldHint=true, idempotentHint=false), so the description's marginal job is to explain wait semantics — and it does not. The single most important behavior for a wait tool, what happens when the timeout elapses (throw vs. return a negative result), is never stated, nor is whether it blocks other operations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence, front-loaded with the primary mode and trailing the fallback mode. Minor inaccuracy in the parenthetical state list (says 'appear ... visible/hidden/attached' while the enum also includes 'detached') is the only blemish.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should at minimum establish return/timeout behavior and the full state vocabulary; it does neither. For a four-parameter, zero-required tool with no schema documentation, this leaves a real gap an agent must guess around.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden and only partially does. It maps 'text'/'selector' to the two wait modes and names usable states, but it calls out only three of the four enum values (omitting 'detached') and never defines what 'attached' vs 'visible' means or that text and selector should not both be supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (wait) and the two wait-able resources (selector or text), plus the alternate timeout-only mode, so the agent knows this is a blocking-until-condition tool rather than a poll/read tool. It does not explicitly contrast itself with siblings like browser_get_text or browser_snapshot, which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: pass selector/text to block on a condition, or pass only timeoutMs to pause. There are no explicit when-to-use/when-not rules, no note that text and selector are alternatives, and no mention of a sibling for non-blocking inspection. Useful but inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capability_callCDestructive
Fallback dispatcher for a capability present on the live server but missing from a cached client catalog.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | ||
| tool | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose the risky profile (destructive, openWorld, non-idempotent, not read-only), so the safety burden is largely covered. The description adds that this is a fallback path, but omits critical behavior for a dispatcher: what happens when 'tool' is unknown, whether args are forwarded verbatim, and whether the call is proxied elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler or redundancy. It is efficient, though brevity here borders on under-specification rather than tight editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a generic dispatcher that can invoke arbitrary capabilities with arbitrary arguments, the description is thin: no guidance on valid tool identifiers, argument-passing contract, error behavior for missing capabilities, or reference to capability_list for discovery. With no output schema and 0% param coverage, the description should carry far more.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are two parameters, including a nested 'args' object with no documented shape. The description says nothing about what 'tool' should contain (a capability name) or how 'args' must be structured to match the target capability's schema, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'fallback dispatcher' conveys a generic act of routing, and the qualifier 'capability present on the live server but missing from a cached client catalog' is really a usage condition rather than a statement of what the tool does. An agent can infer it invokes a named capability, but the verb is vague and the resource is abstract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the trigger condition (a capability is on the live server but absent from the cached catalog), which is a real usage cue. However, it never names the natural partner tool (capability_list) for discovering capabilities, nor does it state when this fallback should NOT be used versus a normal direct call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capability_listBRead-onlyIdempotent
List the live MCP capability catalog. Use prefix to filter tool names.
| Name | Required | Description | Default |
|---|---|---|---|
| prefix | No | ||
| includeDescriptions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description only adds the word 'live', hinting the catalog is dynamically generated, without saying whether it reflects runtime-registered capabilities or anything about result shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the core purpose leads and the parameter hint follows. Slightly terse relative to what is left unexplained, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-required-input read tool with no output schema, the definition is minimally workable. It still leaves includeDescriptions behavior and the return shape (names only vs. full metadata) to inference, which matters for an agent choosing between this and capability_call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and only partly does: it explains prefix ('filter tool names') but leaves includeDescriptions completely undocumented, and it does not say whether prefix matches tool names, capability IDs, or both, nor whether it is case-sensitive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'List the live MCP capability catalog' — which clearly distinguishes it as the read/enumeration counterpart to capability_call. It stops short of naming that sibling explicitly, but the pairing is unambiguous from the verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use prefix to filter tool names' addresses how to use a parameter rather than when to choose this tool over capability_call or when a full unfiltered listing is appropriate. Usage is implied by the name/description but never stated as a condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_filesBRead-onlyIdempotent
Run the relevant compilers/linters (tsc, eslint, node --check, go vet, gofmt, py_compile, ruff, JSON) on specific files and report problems.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | File | |
| paths | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description usefully adds that problems are reported rather than auto-fixed, but omits output format, whether a nonzero exit is surfaced, and whether all listed linters always run or only those matching the file type.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the enumerated linters are the substance of the tool and earn their length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and two loosely documented parameters, the description carries real burden. It explains what runs and that problems are reported, but leaves the path/paths semantics and the result shape unexplained, so it is only marginally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: 'paths' carries no description at all. The description says 'on specific files' but never clarifies the relationship between the singular 'path' and plural 'paths' parameters, nor whether both are required or mutually exclusive, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Run') and resource ('compilers/linters') and enumerates the actual tools invoked (tsc, eslint, node --check, go vet, etc.). This is far more concrete than a bare 'check' name, though it never distinguishes itself from lsp_diagnostics, vscode_diagnostics, or run_command, which also surface problems.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites, and no routing to alternatives despite several siblings (lsp_diagnostics, vscode_diagnostics, run_command) that overlap with static analysis. The agent must infer the selection criterion entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkpoint_listARead-onlyIdempotent
List recent checkpoints (one per file-changing tool call) with the files each one touched.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this as a read-only, idempotent, non-destructive operation, so safety is covered. The description usefully adds the semantics of what a checkpoint represents (one per file-changing tool call) and that it reports touched files, but stays silent on ordering, retention window, or pagination behavior for older checkpoints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tightly written sentence with no filler; the resource is front-loaded and the parenthetical definition earns its place by clarifying a project-specific term.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only one optional parameter, the description conveys enough to call it correctly and previews the return shape (checkpoints with their touched files). It falls just short on the meaning of the limit parameter and return ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'limit' parameter has no description in the schema (0% coverage), and the description only gestures at it via the word 'recent'. It does not state whether limit counts checkpoints or files, or what happens when there are fewer than the limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (checkpoints), and even defines what a checkpoint is — one per file-changing tool call — plus the payload (files touched). It does not, however, distinguish itself from the sibling checkpoint_rewind, which an agent might confuse for the read side of this feature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the discovery step before calling checkpoint_rewind, but the description never says so and offers no when-to-use or when-not-to-use guidance. No alternative is named despite checkpoint_rewind being an obvious adjacent tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkpoint_rewindADestructive
Undo: restore files to their state before the given checkpoint (default: the latest), undoing it and every later checkpoint.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuine value beyond that by disclosing exactly what is destroyed: the target checkpoint 'and every later checkpoint', plus the default-target behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the destructive verb 'Undo', with the scope and default packed in without filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent single-param tool with no output schema, the description covers the action, the default target, and the collateral scope. It could still note that the undo is itself not trivially reversible, but the annotations carry the risk signal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter 'id' is undocumented in the schema. The description compensates by explaining that id refers to a checkpoint and that omitting it defaults to the latest checkpoint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action (restore/undo) and resource (files at a named checkpoint), and states the scope includes the named checkpoint plus every later one. It is clearly distinguishable from checkpoint_list, the only nearby sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the default (latest checkpoint when no id is given) and the blast radius of the operation, but never states when to prefer this over alternatives (e.g. checkpoint_list to inspect first) or any precondition. Usage is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_call_site_toolADestructive
Invoke a tool the web page itself declares (window.mcp / window.mcp_tools / meta[name=webmcp-tool]). Take a chrome_snapshot first to see the catalog.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | ||
| tabId | No | Chrome tab id (default: active tab) | |
| frameId | No | Frame id, default 0 | |
| toolName | Yes | ||
| snapshotId | No | snapshotId returned by chrome_snapshot; rejects stale actions |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, so the agent knows this is an unsafe, side-effecting call into page-supplied code. The description adds the mechanism (page-declared tools) but says nothing about trust boundaries, failure modes, or that the invoked behavior is entirely page-controlled. Adds some context; not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the purpose and followed by the prerequisite action. Nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and for an open-world, destructive tool that proxies arbitrary page-declared APIs the description never indicates what is returned or how errors from the page tool surface. The snapshot prerequisite is covered, but return/error behavior is left entirely implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%; tabId, frameId and snapshotId are documented in the schema, but the required toolName and the free-form nested args object are undocumented anywhere. The description only indirectly supports snapshotId via the snapshot-first instruction and adds no syntax for args, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (invoke) and resource (a tool the web page itself declares), and names the concrete declaration surfaces (window.mcp / window.__mcp_tools__ / meta[name=webmcp-tool]). This clearly separates it from browser_evaluate or chrome_evaluate, which run agent-supplied code rather than page-declared tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite in 'Take a chrome_snapshot first to see the catalog,' routing the agent to the sibling that produces the toolName/snapshotId it needs. No when-not guidance or statement about what to do if the page declares no tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_clickADestructive
Click an element by its [#N] ref using a real (trusted) mouse click via the Chrome debugger; falls back to a DOM click if the debugger cannot attach.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| tabId | No | Chrome tab id (default: active tab) | |
| frameId | No | Frame id, default 0 | |
| snapshotId | No | snapshotId returned by chrome_snapshot; rejects stale actions | |
| doubleClick | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds real value beyond them: it explains the trusted/debugger click path and the DOM-click fallback when the debugger cannot attach, which tells the agent about reliability trade-offs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler: mechanism, parameter semantics, and fallback all in one clause chain. Nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an action tool with annotations covering the safety profile and no output schema, the description covers mechanism, ref semantics, and failure fallback adequately. Minor gaps: no mention of whether it scrolls into view or what a failed click reports, but these are small for a click primitive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, and the description clarifies the ref semantics ([#N], element reference) which the schema alone does not. It says nothing about doubleClick, tabId, frameId, or snapshotId beyond what the schema already documents, so it lands at the baseline for partial coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (click an element) and specifies the mechanism (trusted mouse click via Chrome debugger) with a fallback. The [#N] ref format is called out, which is useful given the ref parameter is only typed as integer. Does not explicitly distinguish itself from the sibling browser_click, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage (you need a [#N] ref, presumably from chrome_snapshot) and discloses the fallback path, but never states when to prefer this over browser_click or other click siblings. The condition selecting the alternative is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_close_tabCDestructive
Close a tab.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | Yes | Chrome tab id (default: active tab) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true. The description adds no behavioral context beyond that, such as irreversibility, what happens to the active tab, or permission requirements. It merely restates the action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
'Close a tab.' is extremely concise with no wasted words, but it is under-specified for a destructive tool. It front-loads the action but omits any caveat or scoping detail that would help an agent use it safely.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with rich annotations, the definition is minimally functional. However, it lacks any mention of the default active-tab behavior or how it relates to sibling close tools, leaving the agent to rely entirely on the schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema documents tabId with a default value. The description adds no parameter meaning beyond what the schema provides, so the baseline of 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Close' and resource 'tab', making the action clear. However, it provides no differentiation from sibling tools such as browser_close or chrome_switch_tab, so an agent cannot easily distinguish this tool's scope without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives, and no prerequisites are given. The agent must infer appropriate usage entirely from the tool name and context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_consoleCRead-onlyIdempotent
Console messages and exceptions captured from a tab since the debugger attached to it.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | ||
| tabId | No | Chrome tab id (default: active tab) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful scope context ('since the debugger attached' implies a bounded, non-persistent buffer), but says nothing about retention limits, when the buffer resets, or what clearing does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the scoping constraint front-loaded and no filler. It is a sentence fragment without a verb, but that costs little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should characterize the returned message structure and the effect of 'clear'; it does neither. It never states the prerequisite that the debugger must already be attached for any messages to exist.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: tabId is documented in the schema (default: active tab) but 'clear' is bare in both schema and description. The description never mentions that clear discards captured messages, which is the one parameter most likely to surprise a caller.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource precisely — console messages and exceptions from a tab — with the scoping qualifier 'since the debugger attached'. It is clear what this tool returns, though it never states an explicit verb (retrieve/list) and gives no differentiation from the near-identical sibling browser_console.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites (e.g., debugger must be attached), and no routing advice. The sibling list contains both browser_console and chrome_console, and the description does nothing to help an agent choose between them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_dialogADestructive
Inspect or answer a JavaScript dialog (alert/confirm/prompt/beforeunload) that is blocking a tab. Dialogs are never accepted automatically: read the message, then accept or dismiss (promptText for prompt()).
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | Chrome tab id (default: active tab) | |
| action | No | status | |
| promptText | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, readOnlyHint=false, and non-idempotence. The description adds genuine context: dialogs are never auto-accepted, and the required read-then-answer sequence. It doesn't detail consequences of accept vs dismiss, which would push it higher.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the constraint ('never accepted automatically') front-loaded, followed by the procedure. Zero wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a dialog handler with no output schema, the description covers the essential behavior and sequence. It omits what 'status' returns for a non-blocking tab, a minor gap given the small tool surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 33%; tabId is documented in the schema itself, and the enum action values are self-explanatory. The description usefully clarifies that promptText applies to prompt() dialogs, but adds nothing for the status vs accept vs dismiss distinction beyond the schema enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('inspect or answer') and resource ('JavaScript dialog'), and enumerates the dialog types (alert/confirm/prompt/beforeunload). No sibling tool handles dialogs, so the agent can identify it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies clear context ('that is blocking a tab') and a workflow ('read the message, then accept or dismiss'). It stops short of naming alternatives or explicit when-not-to-use conditions, but the trigger condition is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_downloadsBRead-onlyIdempotent
Recent Chrome downloads (file path, state, size, source URL, referrer).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| since | No | ISO time; only downloads started after it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so safety and repeatability are covered structurally. The description adds little behavioral context beyond enumerating return fields; it says nothing about ordering, truncation at the default limit, or whether the list is scoped to the active browser profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero filler; the resource leads and the fields follow. It is perhaps too terse to earn a 5 relative to the gaps above, but there is no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the parenthetical field list is genuinely useful for anticipating the response shape, which partially compensates. However, an agent still cannot tell ordering, default result count, or whether 'since' is required for meaningful output, so the definition is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: 'since' is documented in the schema as an ISO time filter, but 'limit' has no description at all in either place. With low coverage, the description should compensate, yet it never mentions pagination, the default of 10, or ordering of the returned items.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource ('Recent Chrome downloads') and enumerates the returned fields (file path, state, size, source URL, referrer), so an agent knows exactly what data this yields. It is not a verb+resource phrasing, but the retrieval intent is unambiguous and the tool is clearly distinct from browser_* and file_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives, no prerequisites, and no exclusions. The word 'Recent' implies a recency filter but the agent is not told how recency interacts with the 'since' parameter or when to prefer this over chrome_status, chrome_snapshot, or resource_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_evaluateBDestructive
Run JavaScript in a tab's page context (via the debugger). Pass an expression; the JSON-serializable result is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | Chrome tab id (default: active tab) | |
| expression | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is non-read-only, destructive, open-world, and non-idempotent. The description adds the useful 'via the debugger' context and the fact that the result must be JSON-serializable, but says nothing about what can be mutated, error/timing behavior, or the consequences of running arbitrary code.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and scope, zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description reasonably covers the return ('the JSON-serializable result is returned'), and annotation coverage handles the safety profile. However, for a tool that executes arbitrary code, the lack of any note on error handling, async/promise handling, or scope limits leaves real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: tabId is documented in the schema (default active tab) while expression has no schema description. The description compensates partially by stating an expression must be passed and that its result must be JSON-serializable, but adds no format or syntax detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — run JavaScript in a tab's page context — and notes the mechanism (via the debugger). It is clearly distinguishable from retrieval siblings like chrome_get_text or chrome_console, though it does not explicitly contrast itself with browser_evaluate or chrome_console.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no naming of alternatives such as chrome_console or browser_evaluate. The only hint is the procedural 'Pass an expression', which is invocation mechanics rather than selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_extension_reloadADestructive
Reload the Free Codin Agent Chrome extension from disk (after it was updated) and report the new version.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag this as non-read-only, destructive, non-idempotent and open-world. The description usefully adds that the extension is loaded 'from disk' and that a version is reported back, but it does not warn that reloading can interrupt an active browser session or require the extension to already be installed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every clause carries information (what is reloaded, from where, the precondition, and the returned datum).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description explains that the new version is reported. With zero parameters and annotations covering the safety profile, the only minor gap is what happens if the extension is absent or a session is active.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to clarify, and it does not introduce confusion about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (reload) and resource (the Free Codin Agent Chrome extension), plus the side effect (report the new version). This is clearly distinguishable from siblings like chrome_status or chrome_pairing_setup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(after it was updated)' gives a clear triggering condition, so the agent knows when this is appropriate. No explicit alternatives or when-not-to-use guidance, but no sibling overlaps this operation either.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_get_textBRead-onlyIdempotent
Visible text of a tab (document.body.innerText, truncated).
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | Chrome tab id (default: active tab) | |
| maxChars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description's one added behavioral fact is that output is truncated, which is real value, but it omits what truncation does to the result (cut-off point, indicator) and whether auth/session state matters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded parenthetical fragment with zero filler. It is efficient, though the brevity borders on under-specification rather than deliberate concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should say what the return value looks like (plain string? object with metadata?), and it only says the text is truncated. Annotations do carry the behavior/safety load, so the tool is callable, but key return-shape information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: tabId is documented in the schema but maxChars has just a default (40000) with no description, and the description does not compensate for either parameter. The mention of 'truncated' loosely hints at maxChars but gives no units, limits, or behavior when exceeded.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names the resource and the exact extraction method ('Visible text of a tab (document.body.innerText, truncated)'), so an agent knows this returns rendered visible text rather than HTML or a DOM dump. It does not, however, distinguish itself from close siblings like chrome_snapshot, chrome_evaluate, or browser_get_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus chrome_snapshot (structure), chrome_screenshot (pixels), or chrome_evaluate (arbitrary JS). The agent must infer the choice from the terse purpose line alone; no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_list_tabsARead-onlyIdempotent
List all tabs in the user's real Chrome.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds only the "real Chrome" scope, which is mildly useful context consistent with the open-world annotation, but it does not disclose auth needs, return shape, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no waste. Every word contributes to stating what is listed and where.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter read operation with rich annotations, the description is nearly complete enough to call the tool correctly. The main omission is any indication of what tab information is returned, though that is a minor gap for a list-tabs tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to explain. Per the rubric, a zero-parameter tool has a baseline of 4, and the description does not need to compensate for undocumented inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ("List") and resource ("tabs") and scopes it to the user's real Chrome, making the purpose immediately clear. It does not explicitly differentiate itself from the similarly named sibling browser_tabs, which prevents a top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as browser_tabs or chrome_switch_tab. The description also omits prerequisites or context for invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_new_tabCDestructive
Open a new tab in the real Chrome.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | about:blank | |
| active | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose the safety profile (readOnly=false, destructive=true, openWorld=true, idempotent=false). The description adds only 'real Chrome' as context and does not explain what gets destroyed, how 'active' affects focus, or any side effects beyond what the annotations state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. For a simple open-a-tab action, that is appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, two undocumented parameters, and annotations indicating destructive and open-world behavior. The description does not cover parameter meanings, side effects, or return behavior, so it is not complete enough for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention either parameter. It does not explain the 'url' default of about:blank or the 'active' default of true, leaving both parameters undocumented. With two undocumented parameters, the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Open a new tab' in 'real Chrome.' It clearly states what the tool does, but does not differentiate from siblings like chrome_navigate, chrome_switch_tab, or browser_tabs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of alternatives. An agent must infer that this is for creating a new tab rather than navigating or switching, but the description gives no explicit routing information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_pairing_setupADestructive
Securely configure or rotate the Chrome bridge pairing token in both the connected extension and .env. The secret is never returned. Restart the MCP worker afterwards to enforce it.
| Name | Required | Description | Default |
|---|---|---|---|
| rotate | No | Rotate even if token authentication is already configured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation destructive, non-idempotent, and open-world. The description adds valuable context beyond that: the secret is never returned, both the extension and .env are updated, and a worker restart is required to enforce the change. It does not cover permissions or failure modes, so it is not fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, followed by the security guarantee and the enforcement step. Every sentence carries practical information for invoking the tool correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter setup tool with no output schema, the definition covers the action, dual update target, non-return of the secret, and required restart. It could state prerequisites such as an active connected extension, but annotations already cover the safety profile and the description is otherwise sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single rotate parameter is already documented in the schema as rotating even if token authentication is configured. The description only implies configure-vs-rotate behavior and adds no syntax or additional parameter meaning beyond the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (configure/rotate), resource (Chrome bridge pairing token), and dual target (connected extension and .env). It is clearly distinct from sibling tools like chrome_status or chrome_extension_reload, which handle page or extension state rather than pairing credentials.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says to restart the MCP worker afterward and the schema notes rotate behavior, so some usage context is implied. However, it does not explicitly say when to use this instead of chrome_status, chrome_extension_reload, or other chrome tools, and it names no when-not conditions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_press_keyADestructive
Press a key or chord (e.g. "Enter", "Tab", "Escape", "Control+a") as a trusted key event in the tab.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| ref | No | Optional element to focus first | |
| tabId | No | Chrome tab id (default: active tab) | |
| frameId | No | Frame id, default 0 | |
| snapshotId | No | snapshotId returned by chrome_snapshot; rejects stale actions |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered. The phrase 'as a trusted key event' is genuine added context (events are dispatched as trusted rather than synthetic), but it is not elaborated on and no focus/consequence behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that gives the action and the input format with zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter action tool with annotations present and no output schema, the definition plus schema descriptions cover what an agent needs to invoke it (target, focus, staleness guard). Only minor behavioral detail is missing, which is acceptable given schema richness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80% and the required 'key' parameter has no schema description, so the examples ('Control+a' chord syntax) materially compensate by showing the accepted input format. The remaining focus/tab/frame semantics are already documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb plus resource ('Press a key or chord') and grounds it with concrete examples ('Enter', 'Tab', 'Escape', 'Control+a') that clearly distinguish it from text-typing siblings like chrome_type. Strong and specific, though it never explicitly names or routes away from a sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus chrome_type or browser_click, no prerequisites, and no exclusions. Usage is only implied by the examples, leaving an agent to infer the right context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_screenshotARead-onlyIdempotent
Screenshot of a tab in the real Chrome (visible viewport). The image is returned so you can see it.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | Chrome tab id (default: active tab) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already cover safety traits (readOnly, idempotent, non-destructive). The description adds valuable context beyond annotations by disclosing that the output is an image and that the capture is limited to the visible viewport rather than the full page.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences and is front-loaded with the core action and scope. The second sentence is helpful for output expectations but slightly redundant, keeping it just short of perfectly economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple screenshot tool with rich annotations and a fully documented schema, the description covers purpose, viewport scope, and return type adequately. It does not address relative usage versus other screenshot tools, but that is a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single optional parameter tabId is fully documented in the schema. The description adds no further meaning about the parameter, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action and resource: taking a screenshot of a tab in the real Chrome, scoped to the visible viewport. It implicitly distinguishes itself from browser_screenshot by specifying 'real Chrome', but does not explicitly name or differentiate against sibling tools such as browser_screenshot or chrome_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives like browser_screenshot, windows_screenshot, or chrome_snapshot. The description only states what the tool does, leaving usage context and exclusions to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_scrollBDestructive
Scroll the inspected tab by y pixels (negative scrolls up).
| Name | Required | Description | Default |
|---|---|---|---|
| y | No | ||
| tabId | No | Chrome tab id (default: active tab) | |
| frameId | No | Frame id, default 0 | |
| snapshotId | No | snapshotId returned by chrome_snapshot; rejects stale actions |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety and mutability profile is known without the description. The description adds the meaning of negative y values, which is useful behavioral detail, but does not explain what makes scrolling destructive, whether a snapshot is required, or how stale actions are handled beyond what the schema already says.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that contains no redundant or filler language. It is sized appropriately for a simple scroll action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple four-parameter scroll tool with no output schema, the description is minimally adequate: it covers the core action and y semantics. It still omits usage routing against browser_scroll and does not discuss the snapshotId stale-action behavior that the schema mentions, which would be helpful context for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, and the undocumented parameter is y. The description compensates by defining y as pixel distance and explaining that negative values scroll upward, which is exactly the missing semantics. The other three parameters are covered by their schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Scroll) and resource (the inspected tab), and specifies the unit and direction semantics ('by y pixels', 'negative scrolls up'). However, it does not distinguish this Chrome-specific tool from the sibling browser_scroll, leaving an agent to infer the platform split.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as browser_scroll or when scrolling may be unnecessary. The direction note implies basic operation but not selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_snapshotARead-onlyIdempotent
Semantic snapshot of a tab: interactive elements with [#N] refs, plus page-declared site tools. Take one before clicking or typing.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | Chrome tab id (default: active tab) | |
| frameId | No | Frame id, default 0 | |
| snapshotId | No | snapshotId returned by chrome_snapshot; rejects stale actions | |
| maxElements | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it as read-only, idempotent, non-destructive, and open-world. The description adds valuable behavioral context about the snapshot's content format and the required ordering before interaction actions, though it does not mention rate limits or pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero waste. The output format is front-loaded, followed immediately by the action-oriented usage instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully summarizes the return shape (interactive elements with refs, site tools) and pairs it with the required workflow. It omits details about maxElements, frameId behavior, and what 'site tools' entail, but is sufficient for correct tool selection and basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so most parameters are documented in the schema. The description adds no parameter-specific semantics beyond the general output format, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('snapshot of a tab') and specifies the output content ('interactive elements with [#N] refs, plus page-declared site tools'). It distinguishes a semantic snapshot from a visual screenshot implicitly, but does not explicitly contrast it with siblings like chrome_screenshot or browser_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: 'Take one before clicking or typing.' This tells the agent when to call it, but does not state when to avoid it or name an alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_statusBRead-onlyIdempotent
Connection status of the Chrome extension bridge (the user's real, logged-in Chrome).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered structurally. The description adds the useful clarification that the status refers to the real logged-in Chrome bridge rather than a headless browser, but says nothing about what the reported status looks like or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the essential identifying information comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument status probe with no output schema, the description leaves the return values (e.g., connected/disconnected, extension version) entirely unspecified, which is the one remaining gap an agent would care about.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters and 100% schema coverage, so there is nothing for the description to disambiguate. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific resource (the Chrome extension bridge connection) and clarifies that the bridge targets the user's real, logged-in Chrome, so an agent can tell it apart from browser_* or chrome_* action tools. It stops short of explicitly contrasting itself with the neighboring chrome_pairing_setup / chrome_extension_reload tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to call this versus alternatives, whether it must be checked before other chrome_* calls, or what a negative result implies. The read-only diagnostic nature is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_switch_tabBDestructive
Activate a tab and focus its window.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | Yes | Chrome tab id (default: active tab) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description usefully adds the window-focus side effect (behavior beyond the annotations), but says nothing about what happens with an invalid or stale tabId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with the action and its window side effect. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema tool whose annotations already carry the safety profile, the description covers the essential behavior. Minor gap: no indication of error behavior for an invalid tabId or whether focus is mandatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single tabId parameter is documented there (including the 'default: active tab' note). The description adds no format, range, or sourcing detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Activate a tab and focus its window' clearly distinguishes it from siblings like chrome_list_tabs, chrome_new_tab, and chrome_close_tab. It stops short of naming those siblings explicitly, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to switch tabs versus creating (chrome_new_tab) or listing (chrome_list_tabs) them, and no note on prerequisites such as the tab already existing or being reachable. The agent must infer the usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chrome_typeADestructive
Focus an element by ref and type text with trusted input (replaces the current value unless append: true). submit: true presses Enter afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| text | Yes | ||
| tabId | No | Chrome tab id (default: active tab) | |
| append | No | ||
| submit | No | ||
| frameId | No | Frame id, default 0 | |
| snapshotId | No | snapshotId returned by chrome_snapshot; rejects stale actions |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds useful behavioral context: it replaces the current value unless append is true, and submit: true presses Enter afterward. It does not mention auth, rate limits, or failure behavior, but this is solid additional detail beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. Core action comes first, followed by the important append and submit modifiers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the destructive annotation, no output schema, and a mutation-oriented tool, the description covers the essential behavioral modifiers. It leaves out explicit ref provenance from chrome_snapshot, but the schema's snapshotId field hints at that dependency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 43% schema description coverage, the description compensates by explaining append and submit semantics, plus the core ref and text roles. However, it does not fully document all seven parameters, and some context such as where ref originates from is left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: focus an element by ref and type text. It distinguishes the action itself, but does not explicitly differentiate it from sibling tools like browser_type or chrome_press_key.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the tool is used to type text into a referenced element, but gives no explicit guidance on when to choose it over browser_type, chrome_click, chrome_press_key, or other alternatives. No when-not-to-use conditions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
copy_pathCDestructive
Copy a file or directory.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| overwrite | No | ||
| destination | Yes | ||
| allowPartialUndo | No | Proceed even if the change is too large to snapshot completely (undo would be partial) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered. The description adds nothing beyond the name: it doesn't say whether directory copies are recursive, whether existing destinations are overwritten, or what happens on partial failure (the allowPartialUndo param hints at snapshot limits). With annotations doing the disclosure work, this minimal description is weak but not contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with zero waste and the verb front-loaded, which is structurally fine. But the brevity comes at the cost of substance rather than being earned conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with two undocumented required parameters (source, destination) and no output schema, the definition is far too thin. Annotations cover the safety signal, but the description leaves overwrite semantics, recursive directory behavior, and the partial-undo limitation entirely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%: source, destination, and overwrite are undocumented in the schema, and the one-sentence description adds no meaning for any of the four parameters. It fails to compensate for the coverage gap, e.g. it never mentions that copying can overwrite an existing destination.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Copy a file or directory'), so the operation is unambiguous. However, it offers no differentiation from the closely related sibling move_path or delete_path, which an agent selecting among filesystem tools would benefit from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative guidance is provided. The purpose implies copying rather than moving or deleting, but nothing routes the agent between copy_path, move_path, and delete_path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_pathADestructive
Delete a file, or a directory when recursive is true. Undoable with checkpoint_rewind.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path | |
| recursive | No | ||
| allowPartialUndo | No | Proceed even if the change is too large to snapshot completely (undo would be partial) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and non-idempotent, so the safety profile is covered. The description adds genuinely new behavioral context beyond that: deletions are recoverable via checkpoint_rewind, and the recursive flag governs whether directories are affected at all. It stops short of describing failure modes or snapshot limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the primary action front-loaded and the recovery mechanism stated second. Nothing is wasted and no clause is self-evident filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers undo and recursive scope, which is the essentials. It omits edge cases an agent needs: behavior on non-empty directories, permission requirements, and whether path must exist or can be relative.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the description adds real meaning for 'recursive' (controls directory deletion) and echoes partial-undo semantics for allowPartialUndo. But 'path' remains an uninformative 'Path' string in the schema with no format guidance (relative vs absolute), so the description does not fully compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a file') and explicitly handles the directory case via the recursive flag, so the scope is unambiguous. There is no sibling delete tool to differentiate from, so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the branch condition for recursive (directories) and points at checkpoint_rewind for recovery, which is useful implicit guidance. However, it never says when to prefer this tool over alternatives such as moving to trash, nor what happens when a non-empty directory is deleted without recursive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_fileADestructive
Replace exact text in one file. Either oldText/newText, or edits: [{oldText,newText,replaceAll}] applied atomically in order. oldText must match exactly once unless replaceAll. Reports diagnostics and a checkpoint id.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path | |
| check | No | Run compilers/linters on the changed files afterwards (default true) | |
| edits | No | ||
| newText | No | ||
| oldText | No | ||
| replaceAll | No | ||
| expectedSha256 | No | Optional optimistic precondition from file_info; refuse the edit if the file changed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: unique-match semantics for oldText, replaceAll override, atomic in-order application of edits, and reporting of diagnostics plus a checkpoint id. Annotations cover destructive/idempotent flags; the description enriches with matching and atomicity rules the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action, then input modes, then matching/atomicity rules, then return info. Every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description gives matching rules, atomicity, diagnostics, and checkpoint id – the essentials an agent needs. It omits how to recover via checkpoint_rewind and leaves expectedSha256 unexplained, so it stops short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43%, and the description compensates by explaining the exact-match requirement (once unless replaceAll), the dual mode (oldText/newText or edits array), and atomic sequencing. It leaves path, check, and expectedSha256 without added meaning, but covers the highest-risk parameters well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Replace) and resource (exact text in one file), and distinguishes itself from write_file and apply_patch by restricting to exact-text replacement. An agent can tell this apart from its sibling editing tools immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the two mutually exclusive input modes (oldText/newText vs edits array), which is usage-relevant, but never states when to prefer this over write_file or apply_patch. No explicit when-to-use reasoning is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_urlARead-onlyIdempotent
Fetch a URL (follows redirects) and return readable text. raw: true returns the unprocessed body (useful for JSON/APIs).
| Name | Required | Description | Default |
|---|---|---|---|
| raw | No | ||
| url | Yes | ||
| maxChars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavior: redirects are followed automatically and output is converted to readable text versus a raw body. It omits truncation behavior from maxChars, error/timeout handling, and auth requirements, which matters for open-world fetching.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the core behavior is front-loaded and the mode flag is explained immediately after. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description is the only source for return-value expectations; it hints at 'readable text' vs raw body but says nothing about truncation limits or failure modes. Adequate for a simple fetch, but leaves an important gap around maxChars given zero schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the semantic load. It explains 'raw' well, but leaves maxChars (default 40000) undocumented — an agent cannot tell whether exceeding it truncates, errors, or paginates — and says nothing about url beyond the obvious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (URL), plus a behavioral detail (follows redirects) and the output shape (readable text). It is clear what the tool does, though it does not explicitly differentiate itself from siblings like browser_navigate or browser_get_text that also retrieve page content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'raw: true' note gives a concrete use case (JSON/APIs) for one mode, which is real usage guidance. However, there is no guidance on when to prefer this over browser_get_text, browser_navigate, or web_search in a large sibling set, leaving the primary routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
file_infoBRead-onlyIdempotent
Check whether a path exists and get its type, size and modification time.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is fully covered by structured data. The description usefully discloses the concrete return fields (type, size, modification time), which is the main value it adds and which matters since there is no output schema. It does not state behavior on missing paths (error vs. false) or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that conveys purpose and return values with zero waste. Size is well matched to a simple one-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only probe with full annotation coverage and no output schema, the description adequately conveys both the operation and the returned fields. It only falls short on edge-case behavior for non-existent paths and routing away from check_files.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter with 100% schema coverage, so the schema already documents 'path' (albeit minimally as 'Path'). The description adds no format or meaning beyond what the schema provides, making baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (check path existence) plus the returned attributes (type, size, mtime), so the purpose is unambiguous. It does not distinguish itself from the overlapping sibling check_files, leaving some ambiguity about which to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, and no mention of the natural alternative check_files or the broader list_directory. The agent must infer usage context on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_filesBRead-onlyIdempotent
Find files by glob pattern, e.g. "**/*.test.js" or "package.json".
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Root directory | |
| pattern | Yes | ||
| maxResults | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the glob semantics that make the operation predictable, but it omits the maxResults default cap (500) that bounds results, which is relevant behavior for a finder.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the purpose front-loaded and examples doing useful work. No filler, though the brevity contributes to the coverage gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only file finder with annotations covering safety and no output schema, this is nearly complete, but gaps remain: the maxResults default/cap, the meaning of path, and any routing versus list_directory or search_code.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate, and it partly does by illustrating valid pattern syntax ('**/*.test.js'). However it says nothing about path (root directory) or the maxResults cap, leaving two of three parameters to the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find files') and adds glob-pattern examples that clarify the mechanism. It does not explicitly distinguish itself from siblings like list_directory or search_code, but the glob framing gives it a recognizable identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance and names no alternatives. It never says when to prefer globbing here versus search_code or list_directory, so the agent must infer selection from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_symbolBRead-onlyIdempotent
Find where a function/class/type/method is defined (name or Container.name) and where it is used (whole-word references).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter: function, method, class, type, interface, struct, … | |
| name | Yes | ||
| path | No | Project directory (default: default workspace) | |
| exact | No | false = substring match on names | |
| references | No | ||
| maxReferences | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, so safety is covered. The description adds that references are matched as whole words, which is useful behavioral detail, but it says nothing about reference truncation despite the maxReferences default of 80.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the primary action (finding definitions) before the secondary one (finding usages). No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema, the description omits what results look like and how maxReferences truncates them. It covers the core intent but leaves the agent guessing about result shape and truncation behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, so half the parameters (name, references, maxReferences) are undocumented in the schema. The description partly compensates by clarifying the name format and whole-word reference semantics, but leaves the references toggle and the 80-reference cap unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find) and resource (symbol definitions and usages), and names the accepted input formats (name or Container.name). It is distinguishable from generic siblings like search_code, though it never distinguishes itself from the very similar lsp_definition/lsp_references/vscode_find_references.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this versus the closely related lsp_definition, lsp_references, vscode_find_references, or search_code siblings. Usage is only implied by the description of what it searches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_system_infoARead-onlyIdempotent
Host info, configured workspaces, default workspace and available integrations. Call this first in a new conversation.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds the 'call first' behavioral cue, which is useful context. It doesn't disclose return format or any other behavioral traits, but with annotations doing the heavy lifting, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences: the first lists the returned information, the second gives the usage directive. It is front-loaded with purpose and wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only discovery tool with no output schema, the description covers what the tool returns and when to call it. It is nearly complete, though it could briefly mention if the result is cached or if it reflects live state, but that's minor given the simplicity of the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameters for the description to explain. The description's enumeration of returned data gives some sense of scope, but parameter semantics is not applicable. Baseline 4 is correct for zero parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the resource: host info, workspaces, default workspace, and available integrations. It's not a tautology like the name 'get_system_info', and it enumerates what is returned. However, it doesn't distinguish this from sibling tools like list_workspaces or capability_list, which may overlap in scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it: 'Call this first in a new conversation.' This is a clear, actionable directive that tells the agent exactly when to invoke the tool, with no ambiguity or exclusions needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_commitBDestructive
Stage files (all changes by default) and create a commit.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Repository directory (default: default workspace) | |
| files | No | Specific files to stage; default all | |
| message | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so the agent knows this mutates repository state. The description usefully adds the default staging behavior ('all changes by default'), but says nothing about failure modes, empty-commit handling, or what happens to unstaged/untracked files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with the staging behavior front-loaded before the commit action. Nothing wasted, though it is arguably too terse given the mutation's risk profile.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter destructive mutation with no output schema and no annotations covering commit semantics, the description covers the core operation but omits commit-failure behavior, partial-staging implications, and the interaction with the required message parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: path and files are documented in the schema, message is not (though it is self-evident and required). The description adds the meaningful default for the files parameter ('all changes by default'), which reinforces rather than merely repeats the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific compound verb+resource: stage files and create a commit, with the default staging scope ('all changes by default') made explicit. It is clearly distinguishable from read-oriented siblings like git_status, git_diff, and git_log, though it never names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or prerequisite guidance. It doesn't say a commit requires staged changes, that a message is mandatory, or how it relates to git_diff/git_status as the recommended pre-step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_diffARead-onlyIdempotent
Full unified diff of the working tree (or staged changes, or against a ref), optionally for one file.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| file | No | ||
| path | No | Repository directory (default: default workspace) | |
| staged | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that the output is a full unified diff, but does not disclose output size, error behavior, or pagination details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the core purpose and parenthetically listing the supported modes. Every phrase carries information and no words are wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still communicates that the return value is a full unified diff, and annotations cover the read-only safety profile. The low schema coverage and lack of ref/file syntax details leave minor gaps, but the definition is nearly complete for a simple read-only git inspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, so the description must compensate for three undocumented parameters. It maps to 'staged changes' (staged), 'against a ref' (ref), and 'optionally for one file' (file), adding meaningful semantics, but does not explain ref syntax or how staged and ref interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Full unified diff of the working tree (or staged changes, or against a ref), optionally for one file.' It clearly distinguishes the tool from siblings like git_status and git_log by naming the exact output type and scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates modes of use: working tree, staged changes, against a ref, optionally one file. However, it does not explicitly compare this tool to alternatives such as git_status or git_log, nor does it state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_logCRead-onlyIdempotent
Recent commit history.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Repository directory (default: default workspace) | |
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond that — no ordering (newest-first?), no pagination semantics for the limit, no indication of what a truncated history means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single four-word fragment with zero filler, so it is maximally brief, but brevity here comes from under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool this could be adequate, but with no output schema and one undocumented parameter, the description leaves the return shape, ordering, and limit behavior entirely unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema description coverage, the description is expected to compensate, but it mentions neither parameter. The 'path' parameter is documented in the schema, while 'limit' has no description anywhere, so an agent cannot tell what it caps or how it interacts with the default of 15.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a bare noun phrase, 'Recent commit history,' that essentially restates the tool name git_log. It gives no verb, no scope (which repo, which branch, whole history vs. HEAD), and does nothing to distinguish it from siblings like git_status, git_diff, or git_commit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus git_diff, git_status, or git_commit. The agent must infer that this returns a listing of past commits rather than diffs or working-tree state, with no stated conditions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_statusBRead-onlyIdempotent
git status (short) plus diff stat against HEAD.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Repository directory (default: default workspace) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered by structured data. The description adds useful behavioral content in that it reveals output composition (short status format plus a diffstat against HEAD), but says nothing about untracked files, staging areas, or what happens in a non-repo directory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with zero waste and the key output qualifiers ('short', 'against HEAD') front-loaded. It is arguably under-specified rather than over-long, but it is efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey what is returned; it does partially by naming status-short plus diffstat. However, it omits whether untracked/ignored files appear, staged vs unstaged separation, and failure behavior outside a repository, leaving meaningful gaps for a tool an agent will use to decide next steps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one optional parameter and its schema description ('Repository directory (default: default workspace)') already carries 100% coverage of its meaning. The description adds no syntax or default-behavior detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific git operation and its scope ('status (short) plus diff stat against HEAD'), so an agent knows it reports working-tree state rather than commits or diffs alone. It does not explicitly distinguish itself from siblings like git_diff or git_log, which keeps it out of 5 territory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of the obvious alternative git_diff even though the description itself includes diff-stat output. The agent must infer selection context on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_cancelBDestructive
Cancel a running job (kills its process tree).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true. The description adds useful behavioral context by specifying that cancellation kills the process tree. It does not cover authorization needs, return format, or whether the operation waits for termination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. The key action and its destructive effect are communicated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple destructive tool with rich annotations and no output schema, the description states the effect but leaves the required 'id' parameter unexplained and gives no usage guidance versus alternatives. It is minimally adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter 'id' is undocumented in both schema and description. The description does not explain the format, source, or meaning of the job identifier beyond the obvious implication from 'job'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Cancel a running job'. The parenthetical clarifies the scope by noting it kills the process tree. It is clear but does not explicitly differentiate from sibling tools like kill_process or process_stop, which could cause confusion for an agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use job_cancel versus alternatives such as kill_process, process_stop, or job_status. Usage is only implied by the tool name and the phrase 'running job'. No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_listBRead-onlyIdempotent
Recent background jobs with status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds two small facts beyond that: results are time-scoped ('recent') and include status. It says nothing about ordering, result size, or how the job set is bounded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded fragment with no filler is efficient, though the brevity shades into under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param list tool with no output schema, the description should at least hint at what is returned (fields, ordering, how 'recent' is bounded) and how it differs from job_status. One of those gaps is closed ('with status'), the rest are not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document. Schema coverage is trivially 100% and the baseline for an argument-less tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Recent background jobs with status' names the resource (background jobs) and a scope (recent) but uses no verb, so it reads as a noun caption rather than a statement of what the tool does. It is distinguishable from job_output/job_cancel only by inference, and the relationship to job_status (whose status? how many jobs?) is left unclear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of the obvious alternatives job_status, job_output, or process_list. The agent must guess whether this lists many jobs or reports one job's status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_outputBRead-onlyIdempotent
Read a job log: tail (default 200 lines), a line range (fromLine/lines), or grep (regex with context lines).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| grep | No | ||
| tail | No | ||
| lines | No | ||
| context | No | ||
| fromLine | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds the default tail size (200 lines), which is useful behavioral context. It does not disclose pagination behavior or log growth considerations, so a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and packs three modes efficiently. It is not overly verbose, though the parenthetical grammar slightly obscures the mode-to-parameter mapping.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six undocumented parameters and no output schema, the description covers the main read modes but leaves id, context, and mode-interaction semantics unexplained. Adequate but with clear gaps for an agent invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema documents none of the six parameters. The description partially compensates by explaining tail default, fromLine/lines range semantics, and grep with context lines, but the meaning of id, the interaction between modes, and the context parameter remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'Read a job log.' It then enumerates the three access modes (tail, line range, grep), which lets an agent distinguish it from siblings like job_status or process_output. It does not explicitly name those siblings, keeping it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The three modes imply when each is useful (default tail, explicit line range, regex grep), but there is no explicit when-to-use guidance, no mention of alternatives such as process_output or read_file, and no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_startADestructive
Run a long command (test suite, build, install, benchmark) as a background job. Waits up to waitMs (default 25s): if it finishes you get a smart summary (pass/fail counts, failing tests with context, output tail) instead of the raw log; otherwise poll job_status.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| title | No | ||
| waitMs | No | ||
| command | Yes | ||
| timeoutMs | No | Kill after this long (default 1h) | |
| resourceClass | No | Admission class; defaults to automatic command classification | |
| resourceWaitMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavior: the wait window, the smart-summary return shape, and the polling fallback when the job outlives the wait. It doesn't disclose that the job keeps running after waitMs elapses, which is implied but not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences: the capability first, then the wait/poll contract and what you get back. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Return behavior is well explained despite no output schema, and the wait/poll loop is clear. But for a 7-parameter destructive mutation tool, five parameters are undocumented in both the description and the schema, leaving an agent guessing about timeout, resource class, and working directory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 29%, so the description needs to compensate, but it only restates waitMs and its default (already in the schema). cwd, title, timeoutMs, resourceClass, and resourceWaitMs receive no explanation, and the 50s cap on waitMs is left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (run a long command as a background job) and enumerates concrete use cases (test suite, build, install, benchmark). It doesn't explicitly contrast with the similarly named process_start/run_command/pty_start siblings, so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: use for long-running commands, then wait up to waitMs, otherwise poll job_status. The fallback path to a named sibling is explicit. It lacks guidance on when to prefer run_command or process_start over a background job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_statusARead-onlyIdempotent
Status of a job; waitMs (≤50s) waits for it to finish. Finished jobs return the failure summary.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| waitMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and open-world behavior. The description adds useful context beyond that: the wait is bounded to ≤50s, and finished jobs return a failure summary. This extra behavioral detail is helpful, though it omits other traits like retry behavior or status value semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with a semicolon; purpose is front-loaded, followed by the wait behavior and the return note. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a simple two-parameter status tool with rich annotations and no output schema, the description covers what it does, the optional wait limit, and the failure summary for finished jobs. The only notable gap is the lack of any explanation for the required id parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains waitMs clearly (waits for finish, ≤50s), but the required id parameter is left completely undocumented, leaving the agent to infer that it refers to a job identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource: retrieving the status of a job. The scope is specific enough to differentiate from siblings like job_output and job_cancel, but the description never explicitly names or contrasts with those alternatives, so it misses the top mark.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage by describing the optional waitMs parameter as a way to wait for completion, but gives no explicit when-to-use guidance or alternatives (e.g., use job_output for results, job_list for all jobs). The agent must infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kill_processBDestructive
Kill an OS process tree by PID.
| Name | Required | Description | Default |
|---|---|---|---|
| pid | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds useful scope beyond annotations by specifying that it kills the process tree, not just a single process, but does not mention permissions, forcefulness, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single, front-loaded sentence with no wasted words. The core action and targeting mechanism are stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter mutation tool with annotations covering the destructive nature and no output schema, the description is minimally adequate. It omits routing guidance against similar stop/cancel siblings and does not describe behavior when the PID is invalid or already terminated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single pid parameter. The description partially compensates by identifying the parameter as a PID, but adds no format, validation, or lookup details beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb, resource, and scope: kill an OS process tree by PID. It is clear what the tool does, but it does not explicitly distinguish itself from nearby siblings like process_stop or pty_stop.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is provided. The agent must infer that this is for terminating an entire process tree rather than using process_stop, pty_stop, or job_cancel.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_directoryBRead-onlyIdempotent
Tree listing of a directory (skips node_modules, .git, build outputs unless includeIgnored).
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Directory (default: default workspace) | |
| depth | No | ||
| includeIgnored | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false. The description adds useful context beyond that: it skips node_modules, .git, and build outputs by default unless includeIgnored is set. It still does not describe depth behavior or return format, so it remains adequate rather than rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. The parenthetical adds the key filtering behavior without disrupting readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only directory listing tool whose annotations already cover safety and idempotence, the description is mostly sufficient. However, because there is no output schema and depth is undocumented, an agent still lacks complete context on recursion behavior and result format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate for undocumented parameters. It explains includeIgnored implicitly through the ignore-skipping behavior, but depth is not mentioned at all, and its default and maximum of 10 are left for the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specifies the verb 'Tree listing' and the resource 'directory', so the core action is clear. It does not distinguish itself from nearby filesystem siblings such as find_files or search_code, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states what the tool does but gives no when-to-use guidance, no alternatives, and no exclusions beyond the default ignore behavior. Usage is only weakly implied by the purpose itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_processesARead-onlyIdempotent
List OS processes (optionally filtered by name substring).
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds only the filter's substring behavior and does not disclose return format, ordering, or scope beyond OS processes, making it adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with zero wasted words. The core purpose and the optional filter behavior are communicated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with one optional parameter, full annotation coverage, and no output schema, the description is nearly complete. A minor gap remains around distinguishing it from process_list and describing ordering or scope, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for the single optional parameter. It successfully defines 'filter' as a name substring match, which is meaningful semantic detail not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'List OS processes'. It is clear what the tool does, but it does not explicitly distinguish itself from the sibling tool process_list, leaving some ambiguity for the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that the optional filter is used for a name substring, which implies usage. However, it does not say when to use this tool versus alternatives such as process_list or get_system_info, so guidance beyond parameter usage is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspacesARead-onlyIdempotent
List configured project workspaces with git branch and number of changed files.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint, idempotentHint and non-destructive behavior, so the safety profile is covered. The description earns credit by disclosing what the listing returns (git branch + changed-file count), which goes beyond the structured fields and is valuable given there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The resource comes before the payload detail, so the agent can stop reading as soon as it has the gist.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only lister with full annotation coverage, the description supplies the essential missing piece: what the result contains. Only the absence of any route to a sibling/alternative tool keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4; there are no parameter semantics to document or omit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb ('List') plus a specific resource ('configured project workspaces') and even the payload ('git branch and number of changed files'). It is distinguishable from the browser/pty/git siblings, though it does not name a specific alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by naming the resource, but gives no explicit when-to-use or when-not-to-use guidance relative to nearby tools such as git_status, project_context, or list_directory. Nothing tells the agent why it would pick this over those.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_callsBRead-onlyIdempotent
Call hierarchy of the function at file:line: incoming (who calls it) or outgoing (what it calls).
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Source file | |
| line | Yes | 1-based line | |
| symbol | No | Identifier on that line; its column is found automatically | |
| character | No | 1-based column (or give symbol instead) | |
| direction | No | incoming |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, openWorld=false, and destructive=false, so the safety profile is covered. The description adds semantic meaning for incoming/outgoing, but does not mention language-server prerequisites, failure modes, or return format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler, and the core operation and direction options are presented efficiently. Every part of the sentence contributes to understanding the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does not explain what the returned call hierarchy looks like or how results are structured. It also omits when-to-use guidance relative to sibling LSP tools, leaving some gaps despite clear annotations and mostly descriptive parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so parameters are largely documented. The description adds useful meaning beyond the enum by defining incoming as 'who calls it' and outgoing as 'what it calls', which clarifies the key choice the caller must make.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation (call hierarchy) on a specific resource (function at file:line) and explains the two direction modes. It is clear but does not explicitly differentiate itself from siblings such as lsp_references or lsp_definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what incoming and outgoing mean but gives no guidance on when to use this tool versus alternatives like lsp_references or lsp_definition. It describes output rather than usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_code_actionsARead-onlyIdempotent
List language-server code actions/quick fixes/refactors for a source range without applying them.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Source file | |
| line | Yes | 1-based line | |
| only | No | Optional CodeActionKind filter, e.g. quickfix or source.organizeImports | |
| limit | No | ||
| symbol | No | Identifier on that line; its column is found automatically | |
| endLine | No | ||
| character | No | 1-based column (or give symbol instead) | |
| endCharacter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world, so the safety profile is fully covered; 'without applying them' reinforces this but adds little. The description does not disclose return shape, result limit behavior, or what happens when no actions exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the key constraint (non-applying, range-scoped, language-server) front-loaded and zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, no-output-schema tool with 63% coverage, the description is adequate but leaves gaps: it doesn't explain the default limit of 50, the endLine/endCharacter range semantics, or the shape of the returned actions list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 63% schema coverage across 8 params, the schema already documents file, line, only, symbol and character. The description only implies a 'source range' but adds no meaning for limit, endLine, endCharacter, or the CodeActionKind filter, so it neither duplicates nor compensates meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('language-server code actions/quick fixes/refactors') scoped to a source range, and the qualifier 'without applying them' sharply separates it from any apply/rename/format counterpart among the lsp_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Without applying them' implies the usage context (discovery/preview rather than mutation), but no alternative tool or explicit when-not condition is named, so the agent must infer routing from sibling names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_definitionBRead-onlyIdempotent
Go to definition (kind: definition | typeDefinition | implementation) of the symbol at file:line (give character or symbol name).
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Source file | |
| kind | No | ||
| line | Yes | 1-based line | |
| symbol | No | Identifier on that line; its column is found automatically | |
| character | No | 1-based column (or give symbol instead) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a read-only, idempotent, non-destructive operation, so the safety profile is covered. The description adds the positional requirement (file:line plus a symbol/character marker) but says nothing about the response (a location, possibly multiple at kind=implementation) or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, front-loaded with the action and immediately followed by the kind options. No wasted words, though the parenthetical phrasing is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, no output schema, and moderate complexity, the description covers the inputs adequately but omits what a caller gets back and how to disambiguate from the sibling vscode_go_to_definition. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema largely documents the parameters including the kind enum and the 1-based semantics of line/character. The description adds only the mutual-exclusion hint '(give character or symbol name)', which is useful but marginal beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Go to definition') and resource (the symbol at file:line), and enumerates the three supported kinds. It is distinguishable from siblings like lsp_references and lsp_hover, though it does not explicitly differentiate from the near-duplicate vscode_go_to_definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the 'kind' enumeration and the note to supply character or symbol name, but there is no explicit when-to-use vs when-not, and no guidance on choosing between this and vscode_go_to_definition or lsp_references.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_diagnosticsARead-onlyIdempotent
Compiler/type-checker diagnostics from the real language server (gopls, typescript-language-server, pyright) for files or a directory — no VS Code needed. all: true also returns diagnostics the server reported for other files of the project.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | ||
| path | No | ||
| paths | No | ||
| waitMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that diagnostics come from a live language server and that all:true widens the scope, but omits any note on the 30s waitMs latency, server-not-ready behavior, or result format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with purpose and backends, then the one parameter caveat that matters. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so return shape is unaddressed, and the ambiguous path/paths pair plus waitMs timing semantics are unexplained. Adequate for a read-only tool but short of what an agent needs to call it precisely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must carry the burden. It only explains 'all' (returns diagnostics for other project files) and leaves path vs paths selection and the waitMs timeout default undocumented — the latter being important since a 30s wait implies asynchronous server startup.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (compiler/type-checker diagnostics) and names the concrete backends (gopls, typescript-language-server, pyright). The phrase 'no VS Code needed' implicitly distinguishes it from the sibling vscode_diagnostics, so an agent can route correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (real language-server diagnostics without VS Code) but gives no explicit when/when-not against the near-identical sibling vscode_diagnostics, nor guidance on file vs directory scope. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_format_previewBRead-onlyIdempotent
Preview language-server formatting edits for a file without applying them.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | ||
| limit | No | ||
| tabSize | No | ||
| insertSpaces | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description's only added behavior is that edits are shown and not applied, which reinforces but does not go beyond the read-only annotation; it says nothing about output shape or size limits on the preview.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence, front-loaded with the verb and the non-applying constraint. Nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should convey the shape of the preview, and it only hints at 'formatting edits'. Combined with four undocumented parameters and no clarity on where the preview result goes, the definition leaves real gaps for a formatting tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four parameters, and the description only implicitly addresses 'file'. The format-controlling parameters (limit, tabSize, insertSpaces) are never mentioned, so an agent gets no guidance on what they mean or which to set, and with low coverage the description is expected to compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (preview) and resource (language-server formatting edits) and bounds it with 'for a file without applying them', which cleanly distinguishes it from write_file/edit_file/apply_patch. It stops short of naming the sibling that actually applies the edits, so it is clear but not sibling-complete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without applying them' implies the intended workflow (preview before committing a format), but there is no explicit when-to-use/when-not guidance nor any pointer to the apply step. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_hoverBRead-onlyIdempotent
Type signature and documentation of the symbol at file:line.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Source file | |
| line | Yes | 1-based line | |
| symbol | No | Identifier on that line; its column is found automatically | |
| character | No | 1-based column (or give symbol instead) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and a closed world, so the safety profile is fully covered by structured data. The description adds some value by disclosing what is returned (type signature plus docs), but says nothing about behavior when no symbol exists at the position or about server availability errors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the location and the returned content are both stated immediately. Nothing could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the important work of naming the return payload (type signature and documentation), which is genuinely necessary. However, it omits edge-case behavior such as empty results when no symbol is under the cursor, leaving a gap for a tool an agent will call speculatively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema itself documents both the file/line anchors and the symbol-vs-character alternative mechanism. The description only echoes 'file:line' and adds no syntax, format, or precedence detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource returned (type signature and documentation) at a precise location (file:line), which is far more informative than a bare 'hover'. It does not explicitly contrast itself with close siblings like lsp_definition or lsp_references, but the returned artifact is distinct enough that an agent can differentiate them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives (lsp_definition, lsp_references, find_symbol), nor any prerequisites such as needing a running LSP server. The agent must infer usage purely from the name and the LSP convention.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_referencesARead-onlyIdempotent
All references to the symbol at file:line, resolved by the language server (exact, not text search).
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Source file | |
| line | Yes | 1-based line | |
| limit | No | ||
| symbol | No | Identifier on that line; its column is found automatically | |
| character | No | 1-based column (or give symbol instead) | |
| includeDeclaration | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered structurally. The description adds genuine context that results are LSP-resolved (semantic) rather than textual, but it omits return shape, result volume, and the effect of includeDeclaration.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core scope front-loaded and the disambiguating qualifier immediately after. No filler, nothing to cut.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description says nothing about what is returned (list of locations, ordering, effect of limit). Combined with silent includeDeclaration semantics, the definition is adequate for identification but leaves behavioral gaps for a 6-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, above the 50% baseline, so the schema already documents file, line, symbol, and character. The description adds nothing about the two undocumented/under-documented parameters (limit's default, includeDeclaration's true-default meaning), so it fails to compensate for the remaining gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('all references to the symbol') anchored at file:line, and adds a sharp qualifier ('resolved by the language server (exact, not text search)'). This is clear and self-contained, but it does not distinguish itself from close siblings like lsp_definition or vscode_find_references, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'exact, not text search' clause implicitly tells the agent to prefer this over a text-based search, which is a form of usage context. However, it never names an alternative tool or states when-not-to-use, and the presence of lsp_definition and vscode_find_references as siblings leaves the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_rename_previewARead-onlyIdempotent
Preview the exact workspace edits a semantic rename would make. Does not write files; apply reviewed changes through apply_patch/edit_file.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Source file | |
| line | Yes | 1-based line | |
| limit | No | ||
| symbol | No | Identifier on that line; its column is found automatically | |
| newName | Yes | ||
| character | No | 1-based column (or give symbol instead) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so safety is covered structurally; the description's 'does not write files' reinforces but does not extend that. It adds the useful workflow fact that reviewed changes go through apply_patch/edit_file, but says nothing about truncation of the preview (relevant given the limit parameter) or how edits are ordered/grouped.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the read-only constraint and the handoff to apply_patch/edit_file are both front-loaded and immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description adequately signals the return value type (a set of workspace edits) and the next action. The remaining gap is the limit/truncation behavior, which matters for an agent reading a large preview, but overall the definition is usable as-is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so file/line/symbol/character are documented by the schema itself, but the description adds no meaning to any parameter. The 'limit' parameter (default 200) is undocumented in both places, leaving its truncation semantics unexplained despite being the main behavioral risk of a preview.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Preview the exact workspace edits a semantic rename would make'), which is unambiguous and clearly distinct from the browser/pty/file siblings. An agent knows it produces a dry-run of a rename, not the rename itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Does not write files' and names the follow-up tools (apply_patch/edit_file), which tells the agent both what this tool is not for and how to complete the workflow. It stops short of stating when a preview is required versus optional, so it is not fully prescriptive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_statusARead-onlyIdempotent
Which language servers are installed and running (language, project root, pid).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world behavior. The description adds useful disclosure beyond annotations by naming the returned fields (language, project root, pid), which is important because no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It efficiently communicates the resource and the returned fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only status tool with rich annotations and no output schema, the description is nearly complete. It covers the returned data but could state more about scope or limitations, such as whether results are project-specific.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so parameter semantics baseline is 4. The description does not need to document parameters and correctly focuses on output instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and result: language servers that are installed and running, with returned fields language, project root, and pid. It is clear enough to distinguish from diagnostic or navigation LSP tools, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit when-to-use guidance, prerequisites, or alternatives among the many sibling tools. Usage is only inferable from the tool name and the returned information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lsp_symbolsBRead-onlyIdempotent
Symbols: of one file (file), or matching a query across the project (query + path).
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | ||
| path | No | ||
| limit | No | ||
| query | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the scope distinction (one file vs. whole project), which is real behavioral context about cost/breadth, but it says nothing about result volume, truncation, or how the limit interacts with project-wide scans.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the two modes front-loaded and no filler. It is efficient, though the terse 'Symbols:' fragment reads more like a schema note than a description, which slightly hurts first-glance readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a 4-parameter tool with 0% schema coverage and no output schema, so the description bears heavy weight. It maps three of four parameters loosely but omits `limit` entirely, leaves `path` ambiguous, and gives no guidance on output shape or choice among the numerous sibling symbol tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning — and it does clarify that `file` selects single-file mode while `query`+`path` drive project search. It still leaves `path` semantics (project root? relative? filter scope?) and `limit` behavior unstated, so the compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the resource (symbols) and distinguishes two retrieval modes — single file vs. project-wide query — which is more than a restatement of the name. However, it does not differentiate this tool from close siblings like find_symbol, outline, or repo_map, which appear to overlap in purpose, so an agent still has to guess which symbol tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical param pairings (file) vs (query + path) implicitly tell the agent which argument combination triggers which mode, which is useful. But there is no explicit when-to-use guidance relative to the many sibling symbol/search tools, and no statement of exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_pathCDestructive
Move or rename a file or directory.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| overwrite | No | ||
| destination | Yes | ||
| allowPartialUndo | No | Proceed even if the change is too large to snapshot completely (undo would be partial) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and non-idempotency, so the safety profile is partly covered. The description adds only the rename clarification and says nothing about overwriting an existing destination, snapshot/undo limits, or permissions for a destructive filesystem operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single six-word sentence, front-loaded with the verb and free of filler. Nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive 4-parameter filesystem mutation with no output schema and only 25% parameter documentation, the description is far too thin. It omits overwrite behavior, undo semantics, and the meaning of the undocumented parameters, so an agent cannot call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (just allowPartialUndo is documented), yet the description says nothing about source, destination, or overwrite semantics. It does not compensate at all for the undocumented parameters, leaving overwrite behavior and undo implications opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear specific verb+resource: 'Move or rename a file or directory' tells the agent exactly what operation it performs, and the 'or rename' clause usefully notes that a move doubles as a rename. However, it offers no differentiation from near siblings like copy_path or delete_path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites, and no mention of the copy_path/delete_path alternatives that accomplish related outcomes. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outlineARead-onlyIdempotent
All symbols (functions, classes, methods, types) of a file with line numbers, so you can read_file just the part you need. A directory gives a repo map of it.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is covered. The description adds valuable behavioral detail the annotations can't: it returns symbols with line numbers and switches behavior between file and directory inputs. It leaves pagination/limits unstated, but for a simple read tool that is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste. The primary behavior (file symbols with line numbers) is front-loaded, followed by the directory variant and the motivating use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool with no output schema, the description conveys what is returned, how it varies by input type, and why it's useful. Only minor gaps remain, such as output size or ordering guarantees.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single 'path' parameter, so the description carries the burden and does so meaningfully: it clarifies that path may be a file (returns symbols) or a directory (returns a repo map). This is a real semantic clarification beyond the bare schema type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: enumerating all symbols (functions, classes, methods, types) of a file with line numbers. It also describes the directory mode. It is clear, but it doesn't explicitly differentiate itself from siblings like repo_map, find_symbol, or lsp_symbols, which overlap in purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'so you can read_file just the part you need' implies a use case and a follow-up workflow, which is helpful context. However, there is no explicit guidance on when to choose this over find_symbol, repo_map, or lsp_symbols, and no when-not conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_inputADestructive
Write text to stdin of a process started with process_start (REPLs, prompts, interactive CLIs). close: true writes the data and closes stdin.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| data | Yes | ||
| close | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the agent knows this mutates a running process irreversibly. The description adds one genuinely useful behavioral detail beyond the annotations: close:true writes the data and closes stdin. It does not say what happens if the id is stale or the process has already exited.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler. The core action is front-loaded and the close:true caveat follows immediately after.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter, no-output-schema tool, this covers the essential action and the one non-obvious parameter. The remaining gap is minimal: id and data are only implicit, and there is no note on error behavior for a dead process.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the schema gives no help. The description explains the optional close parameter's effect well, and 'text to stdin' implies data while the process_start reference implies id, but neither id nor data is described explicitly, leaving the agent to infer their meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Write text to stdin of a process') and ties it to the sibling that creates the process (process_start), so the agent knows these two go together. It falls short of 5 because it never distinguishes itself from pty_write, which reads and writes to an interactive terminal and could easily be confused with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete usage context ('REPLs, prompts, interactive CLIs') and a prerequisite ('a process started with process_start'). It does not state any when-not condition or name an alternative tool, so it stops short of explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_listARead-onlyIdempotent
List background processes started by this bridge.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint false, so the safety profile is covered. The description adds useful scoping context ('started by this bridge'), but it does not describe return format, pagination, or any behavioral details beyond that scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words, and the scope ('started by this bridge') is front-loaded. It is appropriately sized for a zero-parameter listing tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-parameter read-only tool with rich annotations and no output schema, the description covers the essential purpose and scope. The only gap is lack of differentiation from sibling list_processes, but the annotations already handle safety and idempotency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the schema is empty with 100% description coverage. The description appropriately adds no parameter details because none are needed; baseline for zero-parameter tools is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('background processes') and adds the scope 'started by this bridge.' This is clear, but it does not distinguish the tool from sibling list_processes or job_list, leaving some ambiguity about what makes this list distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this tool versus alternatives like list_processes, process_output, or job_list. The description only states what it does, not the context or conditions that would select it over a sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_outputARead-onlyIdempotent
Read new output from a background process since the last read (or all buffered output with all: true).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| all | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds useful context beyond those annotations by disclosing the stateful cursor behavior ('since the last read') and the buffering scope ('all buffered output with all: true'). It does not mention whether the call blocks when no new output is available, but the added cursor semantics are a meaningful contribution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the core action and then qualifies it with the parameter override. No wasted words; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 parameters, no output schema), and the description tells the agent what is returned and how to change the read scope. It omits whether the call blocks when no new output exists and does not describe the output format, which is a minor gap given no output schema is present and annotations cover safety.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the full burden. It explains the optional 'all' boolean's effect (return all buffered output) and implies that 'id' is the background process identifier, but it does not specify the format or type of the id beyond context. Partial compensation for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (Read) and resource (new output from a background process) with a scoping qualifier (since the last read or all buffered output). It implicitly distinguishes itself from pty/job siblings by specifying 'background process', but it does not name an alternative sibling tool by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by explaining the 'since the last read' default and the 'all: true' override, which guides the agent on how to call it. However, it never states when to prefer this tool over siblings like job_output or pty_read, leaving alternatives to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_startADestructive
Start a long-running background command (dev server, watcher, long build). Returns an id; read logs with process_output, stop with process_stop.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| name | No | ||
| waitMs | No | How long to wait before returning initial output | |
| command | Yes | ||
| resourceWaitMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive, non-idempotent, and open-world behavior. The description adds lifecycle context: it returns an id and points to process_output for logs and process_stop for termination. It omits waitMs/resourceWaitMs effects, but annotations carry the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, with the primary purpose front-loaded and no filler. Every clause provides actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no output schema and poor parameter descriptions, the description covers the return id and follow-up tools but leaves parameter semantics and wait behavior unexplained. Adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (waitMs only). The description only implies a 'command' parameter and does not explain cwd, name, or resourceWaitMs. With low coverage, it fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb 'Start' and resource 'long-running background command', with examples (dev server, watcher, long build). It does not explicitly differentiate from run_command, job_start, or pty_start among the many siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides when to use (long-running background commands) and names the follow-up tools process_output and process_stop. It does not give explicit exclusions or compare against run_command, job_start, or pty_start.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_stopADestructive
Stop a background process started with process_start (kills its whole process tree).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavioral detail: stopping kills the entire process tree, not just the parent, which an agent could not infer from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, 14 words, front-loaded with the action and immediately qualified by the process_start linkage and the tree-kill behavior. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation tool with no output schema and annotations covering the destructive/idempotency profile, the description supplies the essential extras: the originating tool and the tree-kill scope. Only minor gaps remain, such as behavior when the id is stale or already terminated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter 'id' with 0% schema description coverage, so the description is the only source. It implies the id refers to a process previously launched via process_start but does not state whether it is a PID, opaque handle, or whether the id remains valid after the process exits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (stop) plus resource (background process) and it names the originating sibling tool process_start, which anchors the resource type. It does not, however, distinguish itself from the other termination-flavored siblings such as kill_process, pty_stop, or job_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'started with process_start' implies the intended lifecycle pairing, but it never states when to prefer this over kill_process or how it differs from pty_stop/job_cancel. Usage is implied rather than explicitly routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_contextARead-onlyIdempotent
Start here for any project: detected stack and build/test/lint commands, git state, top-level layout, project instruction files, README start and saved memory notes.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project directory (default: default workspace) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world, so the safety profile is covered. The description then adds what is actually returned (stack, commands, git state, layout, instruction files, memory notes), which is genuine content-level information beyond the annotations. It does not discuss truncation or freshness of the gathered data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the 'start here' positioning and then a tight enumeration of contents. Every clause maps to a distinct piece of returned information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only orientation tool with complete annotations and no output schema, the description enumerates the returned sections well enough for an agent to know what it will get and when to call it first. It could say more about scope limits or ordering, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional parameter ('path', with default workspace behavior) fully documented in the schema. The description adds no path semantics of its own, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists exactly what the tool aggregates: detected stack, build/test/lint commands, git state, top-level layout, instruction files, README start, and memory notes. That distinguishes it from narrower siblings like repo_map, git_status, and project_memory. It lacks an explicit verb ('returns'/'aggregates'), but the resource and scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Start here for any project' gives clear, explicit when-to-use guidance and positions it as the entry point before the more specialized siblings. It stops short of naming when not to use it or which sibling to reach for instead once a specific need appears, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_memoryADestructive
Durable per-project notes shown by project_context in future conversations. action: list | add (note) | remove (index). Save conventions, gotchas and verified commands — not temporary state.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| path | No | Project directory (default: default workspace) | |
| index | No | ||
| action | No | list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered. The description adds genuinely useful context beyond them: the notes are 'durable' and resurface via project_context 'in future conversations', and remove operates by index, signaling irreversible per-entry deletion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the persistence purpose, then the action syntax, then the content rule. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, no-output-schema mutation tool, the description covers purpose, action modes, parameter mapping, and the persistence behavior an agent needs. The only gap is the shape of the 'list' return (e.g. whether indexes are shown), which matters since remove relies on index.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only path is documented), so the description must compensate and largely does: it binds 'add' to note and 'remove' to index, revealing the action-to-parameter mapping the schema does not express. path's directory semantics are left to the schema, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (durable per-project notes) and enumerates the three action modes (list/add/remove), which lets an agent tell it apart from the many browser/process siblings. It does not name the closest alternative by comparison (project_context is mentioned only as the consumer), so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit content guidance — 'Save conventions, gotchas and verified commands — not temporary state' — which is a clear when-to-use plus a when-not-to-use. It stops short of naming a competing tool for ephemeral state, so no full routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pty_listBRead-onlyIdempotent
List interactive PTY sessions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds nothing beyond that — no note on what a session entry contains, whether stopped sessions are included, or how results are ordered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words, appropriate for a no-argument list operation. It is arguably too terse, but nothing in it is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, non-destructive list tool with no output schema, the description is minimally viable. It omits any indication of what a returned session looks like or how to use the result (e.g., feeding an ID into pty_read or pty_stop), which would help an agent act on the output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. The description correctly implies an unfiltered full listing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (interactive PTY sessions), which is clear enough to distinguish from pty_start/pty_read/pty_write/pty_stop in the sibling set. It does not, however, explicitly differentiate itself from other listing tools like process_list or job_list, which could return overlapping or confusable results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to call this versus alternatives such as process_list or job_list, nor any preconditions. An agent must infer usage purely from the tool name and sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pty_readARead-onlyIdempotent
Read new terminal output from a PTY session (or all buffered output with all: true). ANSI escapes are stripped by default.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| all | No | ||
| raw | No | ||
| maxChars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false. The description adds useful behavioral context beyond annotations: ANSI escapes are stripped by default and all: true returns buffered output. It stops short of explaining truncation, blocking behavior, or error cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence states the core action, and the parenthetical efficiently covers the all option. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should convey return semantics. It indicates that terminal text is returned with ANSI stripped by default, but omits maxChars truncation/default behavior and does not fully explain raw or id. Annotations cover safety, but parameter and return gaps remain for a four-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for four parameters, so the description must carry the burden. It explicitly explains all: true and indirectly hints at raw behavior via 'ANSI escapes are stripped by default', but id and maxChars are undocumented in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Read' and resource 'terminal output from a PTY session', and distinguishes scope via 'new' versus 'all buffered output'. This clearly separates it from sibling operations like pty_write and pty_list, even without naming process_output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage after starting or writing to a PTY and gives parameter-level guidance for all: true, but it does not state when to choose pty_read over process_output, job_output, or other read tools, nor does it provide any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pty_resizeCDestructive
Resize a PTY terminal.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| cols | No | ||
| rows | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate this is a destructive, non-idempotent, non-read-only operation, implying it modifies state. The description doesn't add any behavioral context beyond what annotations already say — it doesn't explain what gets destroyed, whether resizing affects running processes, or any side effects. With annotations covering the safety profile, the description could add value but doesn't. It also doesn't contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the action. It's efficient and not verbose, but it may be too terse given the complexity and missing details. However, it avoids waste, so it scores well on structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters, 0% schema coverage, no output schema, and destructive annotations, the description is completely inadequate. It doesn't explain parameter usage, return values, or behavioral implications. An agent would struggle to invoke this correctly without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. It fails to mention the required 'id' parameter or the optional 'cols' and 'rows' parameters, their meanings, formats, or constraints. The agent gets no additional semantic guidance beyond the schema, which itself has no descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Resize) and resource (PTY terminal), so the agent understands the general action. However, it lacks any detail about what resizing entails (e.g., terminal dimensions, cursor behavior) and doesn't differentiate this from other terminal control tools like pty_write or pty_read. It's minimally viable but generic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as browser_resize or other terminal tools. The description doesn't specify prerequisites, timing, or scenarios where resizing is needed. It simply states the action without context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pty_startBDestructive
Start a real interactive terminal backed by Windows ConPTY. Use for REPLs, debuggers, SSH/TUI programs and CLIs that require a terminal. By default starts the configured interactive shell; command runs inside it, or file+args starts an executable directly.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| args | No | ||
| cols | No | ||
| file | No | ||
| name | No | ||
| rows | No | ||
| waitMs | No | ||
| command | No | ||
| resourceWaitMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=false, so the agent knows this is a local, non-read-only, non-idempotent operation. The description adds that a command runs inside the configured shell or that file+args starts an executable directly, but it does not describe persistence, interaction lifecycle, or other side effects beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and terminal backing, then moves to usage and invocation details. Every sentence adds distinct information with no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no schema descriptions and no output schema, the description is materially incomplete. It explains the main command/file/args invocation pattern but omits behavioral meaning for most parameters, making it insufficient for reliable parameter selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must compensate. It clarifies the relationship between command, file, and args, but leaves cwd, cols, rows, name, waitMs, and resourceWaitMs completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Start a real interactive terminal backed by Windows ConPTY.' It also identifies the use cases (REPLs, debuggers, SSH/TUI programs, terminal-requiring CLIs), which implicitly distinguishes it from non-terminal process and command tools. However, it does not explicitly name or contrast with any sibling tool, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for REPLs, debuggers, SSH/TUI programs and CLIs that require a terminal' gives clear positive usage guidance. It does not provide when-not conditions or name alternatives such as run_command or process_start, so it is not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pty_stopCDestructive
Stop and remove an interactive PTY session.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description adds the 'remove' nuance (the session is deleted, not just paused), which is meaningful beyond 'stop', but it omits what happens to running processes, whether an invalid id errors, or whether it blocks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the action front-loaded and no padding. Brevity is appropriate, though it contributes nothing to filling the semantic gaps.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema and an entirely undocumented parameter, the description should explain the identifier source and the consequences of removal. It leaves both to inference, which is insufficient even with annotations carrying the safety signal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required 'id' parameter. The description never states that this is a PTY session identifier obtained from pty_list/pty_start or what format it takes, so nothing compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair and resource: 'Stop and remove an interactive PTY session.' An agent can distinguish it from pty_start/pty_read/pty_write by the verb, though it doesn't explicitly contrast with close siblings like process_stop or kill_process.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives named. The agent must infer that this is the cleanup counterpart to pty_start and is distinct from process_stop on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pty_writeBDestructive
Write to an interactive PTY. enter: true appends the Enter key.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| data | Yes | ||
| enter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral detail — that enter:true appends the Enter key — but says nothing about whether the write blocks, whether it can fail if the PTY is dead, or how it interacts with pty_read buffering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences, front-loaded with the action and followed by the one non-obvious parameter behavior. Nothing is wasted, though the brevity is partly a symptom of under-specification rather than discipline.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent 3-parameter tool with no output schema and zero schema coverage, the description is thin. It omits the lifecycle dependency on pty_start, the meaning of 'id', and how the caller observes the effect — the agent is left to guess at the basic usage loop.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains only the optional 'enter' flag; the required 'id' (presumably a PTY session identifier) and 'data' (payload semantics, encoding, newline handling) are entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Write to an interactive PTY'), which clearly separates it from pty_read, pty_start, and pty_stop. It does not, however, clarify what 'id' refers to or how this differs from process_input, so differentiation from the broader process family is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus pty_read, process_input, or browser_type. The description never states that a PTY must already exist (via pty_start) or that output must be retrieved with pty_read, both of which an agent needs to invoke this correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileARead-onlyIdempotent
Read a text file with line numbers. Relative paths resolve against the default workspace. Use startLine/endLine for large files.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path | |
| endLine | No | ||
| startLine | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real value beyond that: relative paths resolve against the default workspace, and output includes line numbers. It doesn't disclose error behavior for missing files, but for a read tool this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, and the core action is front-loaded. The path-resolution and large-file hints follow in priority order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description conveys the key return characteristic (line numbers) and the path-resolution rule. For a simple 3-param read tool that is close to complete, with only edge-case semantics (out-of-range lines, missing file) left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% — only 'path' is documented as 'File path'. The description compensates partially by naming startLine/endLine and tying them to large files, but it doesn't clarify indexing base, inclusivity of endLine, or what happens when omitted. Baseline 3 fits given the partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Read a text file') plus an output trait ('with line numbers'), which cleanly separates it from write_file, edit_file, and read_many_files. It stops short of naming the sibling it differs from (read_many_files for multi-file reads), so it isn't a full 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It offers one usage hint — 'Use startLine/endLine for large files' — which is genuinely actionable, but there is no guidance on when to choose this over read_many_files or file_info, and no stated preconditions. Usage is implied rather than articulated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_many_filesARead-onlyIdempotent
Read up to 50 files in one call (strings, or {path,startLine,endLine}). Much faster than several read_file calls.
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | ||
| maxTotalChars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds one useful behavioral fact, the 50-file batch cap, but says nothing about truncation behavior or what happens when the output exceeds the default maxTotalChars of 150000.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence of capability plus one sentence of motivation, zero filler, with the batch limit front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with no output schema, the essential call shape is conveyed, but an agent is left guessing about maxTotalChars and about truncation when many large files are requested. Adequate but with a clear documentation gap on the second parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description has to carry the load: it does document the polymorphic file entry forms (string vs {path, startLine, endLine}), which mirrors the anyOf but gives an agent a usable summary. However, the second parameter maxTotalChars (default 150000) is never mentioned, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) plus resource (files) with an explicit scope qualifier (up to 50 in one call). The contrast with the sibling read_file is implicit in the name and made explicit in the description, so an agent can distinguish the batch reader from the single-file reader immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly positions itself against several read_file calls, which is a genuine routing signal (use this when you'd otherwise call read_file repeatedly). It stops short of stating when-not-to-use it (e.g. for a single file, or when output would exceed maxTotalChars), so it's clear context but no explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_clipADestructive
Render a video clip from code: pass a complete HTML document (CSS/JS/canvas/SVG animation) or a URL; it is played headlessly at the given size for durationMs and saved as mp4/webm/gif. Default 1080x1920 (vertical, Shorts/TikTok/Reels). If the page defines window.startClip(), it is called when recording begins.
| Name | Required | Description | Default |
|---|---|---|---|
| fps | No | ||
| url | No | ||
| html | No | ||
| name | No | File name inside the clips folder | |
| width | No | ||
| format | No | mp4 | |
| height | No | ||
| output | No | Explicit output path (overrides name) | |
| durationMs | No | ||
| readySelector | No | Wait for this selector before starting the clock | |
| resourceWaitMs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag this as a destructive, open-world, non-idempotent write, so the safety bar is partly covered. The description adds real behavioral detail: playback is headless, the clip is recorded for durationMs, the pages defines window.startClip() which is invoked at recording start, and readySelector waits before the clock starts. It doesn't disclose errors, rate limits, or the returned handle.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses, front-loaded with the action and input, then output defaults, then the optional startClip hook. Every sentence earns its place, though the input-vs-output sentence is dense and packs several unrelated facts together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter destructive tool with no output schema and low schema coverage, the description still leaves gaps: fps and resourceWaitMs are never explained, the name vs output resolution is left to the schema, and there is no indication of what is returned (file path/handle) after recording. It is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 27%, so the description must carry weight, and it does: it defines html vs url input, explains width/height defaults as vertical 1080x1920, clarifies format options and durationMs semantics. It omits fps, resourceWaitMs, and the readySelector/name/output relationships, but covers most of the functionally critical parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Render a video clip from code') plus the input forms (HTML document or URL) and the output artifacts (mp4/webm/gif). This clearly separates it from browser_screenshot, browser_record_start/stop, and windows_snapshot in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains what to pass (complete HTML document or a URL) and mentions the window.startClip() hook, which is genuine usage context. However, it never says when to choose this over browser_record_start/browser_record_stop, and it does not note that at least one of html/url is functionally required despite required=0.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_mapBRead-onlyIdempotent
Compact map of a codebase: the most-referenced files with their key functions/classes/types and line numbers. Understand a large project in one call. focus: file names or symbols to prioritise.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project directory (default: default workspace) | |
| focus | No | ||
| maxChars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds useful behavioral context about the output shape (compact, most-referenced-first, includes line numbers), but says nothing about cost, truncation behavior, or what happens when the map is too large.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the output description before the usage note and the parameter hint. No filler, though the trailing 'focus:' fragment reads as a schema note tacked onto prose rather than integrated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of describing returns – it does reasonably well (files, symbols, line numbers). However, with a 12k default character budget and truncation likely, an agent gets no guidance on what is dropped or how to react, and one of three params is unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only `path` is documented in the schema), so the description must compensate. It does explain `focus` as 'file names or symbols to prioritise', which adds real meaning, but `maxChars` and its default-truncation implications are left entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: producing a compact map of a codebase showing most-referenced files with key functions/classes/types and line numbers. That is concrete enough to distinguish it from outline or search_code at a glance, though it never names a sibling to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Understand a large project in one call" implies the usage context (orientation in a big codebase), but there is no explicit when-to-use vs. when-not, and no mention of alternatives such as outline, project_context, or find_symbol that an agent might otherwise pick for similar goals.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resource_statusARead-onlyIdempotent
Current workstation resource pressure and admission-controller slots for commands, builds/tests, LSP, diagnostics, background processes and video work.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false. The description adds useful scope by naming the resource categories it covers, but it does not disclose freshness guarantees, latency, or whether the values are cached or live beyond the word "Current".
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tightly written sentence that front-loads the resource being reported and then enumerates the categories. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter status tool with rich annotations, the description is nearly complete. It identifies what is reported but does not describe the shape or fields of the return value, which would be useful since no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to clarify. The baseline for a 0-parameter tool is 4, and the description adds no confusion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool returns: current workstation resource pressure and admission-controller slots across specific work categories. It lacks any explicit differentiation from siblings such as get_system_info or lsp_status, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It does not say when to use this tool, when not to use it, or what alternatives exist. The reader must infer that it is for checking resource status before starting heavy work, but no guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_commandADestructive
Run a shell command (/bin/bash) and wait for it to finish. Returns exitCode, stdout, stderr. Use for builds, tests, git, npm, python, ffmpeg, etc. For servers/watchers that never exit use process_start.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory (default: default workspace) | |
| stdin | No | Optional text piped to stdin | |
| command | Yes | ||
| timeoutMs | No | Default 120000 | |
| resourceClass | No | Admission class; defaults to automatic command classification | |
| resourceWaitMs | No | How long to wait for CPU/RAM capacity before refusing the command |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the destructive, non-idempotent, open-world profile, so credit goes to what the description adds: blocking-until-completion semantics and the return contract (exitCode, stdout, stderr). It does not explain what happens on timeout (whether the process is killed, partially run, or left alive), which matters for a destructive tool with a 120s default timeout.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and completion behavior, then usage examples, then the routing exclusion. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully names the return fields and the blocking behavior. The remaining gap is timeout semantics and the resource-admission refusal path (resourceWaitMs), which an agent running heavy/destructive commands would benefit from knowing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so the schema already documents cwd, stdin, timeoutMs, resourceClass, and resourceWaitMs. The description adds no parameter-level detail beyond the schema baseline, so the standard 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run a shell command (/bin/bash)') plus a key execution trait ('wait for it to finish'), and names the sibling it is not (process_start) for the never-exiting case. An agent can distinguish it from job_start/pty_start/process_start without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete positive examples (builds, tests, git, npm, python, ffmpeg) and an explicit exclusion with the alternative to use instead ('For servers/watchers that never exit use process_start'). This is exactly the when/when-not/alternative pattern.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_codeARead-onlyIdempotent
Search file contents with a regex (ripgrep when installed, Node fallback otherwise). Returns path:line: text; context adds surrounding lines.
| Name | Required | Description | Default |
|---|---|---|---|
| glob | No | File filter, e.g. "*.ts" | |
| path | No | Directory to search (default: default workspace) | |
| query | Yes | Regex (or literal text with literal: true) | |
| context | No | ||
| literal | No | ||
| maxResults | No | ||
| caseSensitive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuinely new context: the ripgrep-vs-Node backend fallback and the return shape 'path:line: text'. It omits any mention of maxResults truncation or rate/performance behavior, keeping it short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action, then the return format. Zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description supplies the return format and context behavior. For a 7-param read tool this is nearly complete, with the minor gap being maxResults truncation semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 43%, so the description must compensate. It clarifies the 'query' semantics (regex, or literal with literal: true) and explains 'context' (surrounding lines), which the schema leaves bare. However, maxResults and caseSensitive remain undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Search file contents with a regex.' The phrase 'file contents' implicitly separates it from find_files (filename search), but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no when-not, and no routing to alternatives like find_files or find_symbol. The agent must infer from the 'file contents' scope that this is content search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_add_breakpointCDestructive
Add a breakpoint (optionally conditional).
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Absolute file path | |
| line | Yes | 1-based line | |
| condition | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and non-idempotent, but the description adds nothing: it doesn't say whether breakpoints persist across sessions, whether re-adding the same line is harmless, or that a debug session must be running. With annotations present the bar is lower, yet the description contributes no behavioral context at all.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with no filler and the optional aspect front-loaded. It is efficient, though borderline under-specified rather than maximally informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating debug tool with no output schema, the description omits the debug-session prerequisite, where the breakpoint attaches, and any indication of success/failure. An agent could not confidently invoke this without inferring context from outside the definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: 'file' and 'line' are documented in the schema, while 'condition' has no schema description. The description's 'optionally conditional' partially compensates by telling the agent what the parameter is for, but it gives no expression syntax or scope for the condition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a breakpoint') plus a scope modifier ('optionally conditional'). No sibling performs this action, so differentiation is unnecessary, but the description stops short of saying where the breakpoint lands or its lifecycle.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites. Critically, it never states that an active debug session is required, or how this relates to vscode_debug_state — the one sibling an agent would plausibly pair with it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_debug_stateCRead-onlyIdempotent
Active debug session and breakpoints.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. Beyond that the description discloses nothing - no indication of what 'debug state' includes, whether it errors with no active session, or what the response contains (and there is no output schema to fall back on).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is certainly short, but this reads as under-specification rather than tight writing: a single verbless fragment that omits the operation entirely. Brevity here costs clarity rather than buying it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations describing return values, the description is the only place the agent could learn what comes back, and it does not. 'Debug state' is left undefined for a tool that presumably returns session status and the breakpoint list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. There is no parameter behavior for the description to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Active debug session and breakpoints' names the resource but contains no verb, so the agent must infer that this is a read operation. It never distinguishes itself from the sibling vscode_add_breakpoint, which it is clearly the read counterpart to. Purpose is guessable but vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives such as vscode_add_breakpoint or vscode_diagnostics. No prerequisites, no context, no exclusions are given. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_diagnosticsARead-onlyIdempotent
Live compiler/linter errors and warnings from VS Code language servers (optionally filtered to one file).
| Name | Required | Description | Default |
|---|---|---|---|
| file | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description contributes the useful 'Live' qualifier, implying diagnostics reflect current editor state rather than a persisted snapshot, but adds nothing about freshness, staleness, or what happens when no server is attached.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the resource first and the scoping qualifier in parentheses. No filler, no restatement of annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter, read-only diagnostic query with full annotation coverage and no output schema, this is nearly sufficient. What is missing is the relationship to lsp_diagnostics and the expected return shape (severity/files grouping), which matters more here than for a trivial read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden, and it does: 'optionally filtered to one file' tells the agent the single 'file' parameter is optional and scoping, not required. It does not specify path format (absolute vs workspace-relative), which is the remaining gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb-resource pair: compiler/linter errors and warnings sourced from VS Code language servers, with an optional file scope. It is distinguishable from file-reading siblings, but it never addresses the near-identical lsp_diagnostics sibling, leaving overlap ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisite (e.g. whether the editor/language server must be running), and no direction to alternatives such as lsp_diagnostics for the same information. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_find_referencesBRead-onlyIdempotent
All references to the symbol at a position (via the language server).
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Absolute file path | |
| line | Yes | 1-based line | |
| character | Yes | 1-based column |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is well covered. The description adds that the operation uses the language server, which is useful behavioral context beyond the annotations. It does not describe language-server availability requirements, latency, or result completeness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It states the core operation and mechanism efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, read-only language-server tool with 100% schema coverage and rich annotations, the description supplies the essential operation. It implies a list of references without describing output format, but the absence of an output schema makes that a minor omission rather than a critical gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so file, line, and character are already fully documented. The description adds only the concept of a 'symbol at a position,' which maps naturally to the line/character parameters but does not add syntax or format details beyond the schema. The baseline of 3 for high schema coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: find all references to the symbol at a given position. It also specifies the mechanism (via the language server), which clarifies the semantic nature of the search. However, it does not differentiate this tool from close siblings such as lsp_references or vscode_go_to_definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use guidance, prerequisites, or alternatives. It only implies the use case by describing the operation. With a sibling like lsp_references appearing to perform a similar function, the lack of routing guidance is a meaningful gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_go_to_definitionCRead-onlyIdempotent
Definition location of the symbol at a position.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Absolute file path | |
| line | Yes | 1-based line | |
| character | Yes | 1-based column |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is covered structurally. The description adds nothing beyond that: no note on what is returned, whether the definition may be in another file, or behavior when no definition exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no waste, and the operative concept (definition location of the symbol) is front-loaded. It is terse to the point of being slightly under-specified rather than verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, fully-schema-documented, read-only navigation tool, the description is minimally sufficient. It omits any handling of the no-definition case and does not clarify its relationship to the overlapping lsp_definition sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear notes that file is absolute and line/character are 1-based. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (definition location of the symbol) and the positioning input, so an agent knows this resolves a symbol to its definition. However, it is stated as a noun phrase rather than a verb, and it does not differentiate itself from the sibling lsp_definition, which appears to do the same thing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no alternatives named, despite the presence of lsp_definition among siblings that likely overlaps in function. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_open_fileBDestructive
Open a file in VS Code at a line so the user can see it.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Absolute file path | |
| line | No | 1-based line |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false), so the description does not need to restate that. It adds modest context by framing the action as making the file visible to the user, but it never explains the surprising destructive flag, what happens if the path does not exist, or any side effects on editor state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the action front-loaded and zero filler; nothing repeats the schema or annotations. The phrasing 'at a line' is slightly loose, which keeps it short of ideal precision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema, the description covers the core contract adequately. It omits failure behavior (missing file, no VS Code instance) and any statement about editor state after the call, which an agent would benefit from knowing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – both 'file' (absolute path) and 'line' (1-based) are documented in the schema itself. The description only gestures at the line parameter ('at a line') and adds nothing about path format or line defaults beyond what the schema already says, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Open) and resource (a file in VS Code) and adds the intent 'so the user can see it', which implicitly separates it from read_file, which returns content to the agent. It does not explicitly name or contrast any of the ~90 sibling tools, so the differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this instead of read_file, vscode_open_files, or file_info, nor any prerequisites (e.g. VS Code must be running, workspace must be open). The only hint is the trailing purpose clause, which implies user-facing display but never says so as guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vscode_open_filesCRead-onlyIdempotent
Open editor tabs and the active file in VS Code.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds almost nothing beyond that — it does not explain what tabs actually get opened, whether the active file is derived from context, or what side effects (focus stealing, view changes) occur.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding. It is concise, though its brevity stems partly from under-specification rather than tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters, no output schema, and a near-identical singular sibling, the description should explain what state or context drives the file selection. Instead it leaves the central question — which files and why — unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters there is nothing for the description to disambiguate; the baseline for a no-param tool is 4. The description does not need to compensate for any schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb (open) and resource (editor tabs / active file in VS Code), but the plural name with a zero-parameter schema leaves it unclear which files it opens or how they are chosen. It also fails to distinguish itself from the sibling vscode_open_file, which an agent would naturally confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus vscode_open_file or vscode_debug_state, and no prerequisites or context are given. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchBRead-onlyIdempotent
Search the web (DuckDuckGo). Returns title, url, snippet for the top results.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds some useful context by naming the engine and return fields, but it does not disclose result limits, pagination, or other operational behavior beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded, and contains no wasted words. It states the core action and return shape in one efficient sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple search tool with rich annotations, no output schema, and one required parameter, the description is mostly complete: it names the engine and return fields. It could still mention result count or limitations, but the essential context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single query parameter has 0% schema description coverage, so the description must carry parameter meaning. It only implies that the query is a web search string, adding no syntax, operator, or formatting guidance beyond the bare schema type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Search the web (DuckDuckGo)', and states the return fields. It clearly identifies a web-search tool, but it does not explicitly differentiate itself from sibling tools like fetch_url or browser_navigate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a web-search use case through its name and text, but it gives no explicit guidance on when to choose web_search over fetch_url, browser_navigate, or other retrieval siblings. There are no exclusions or alternative-routing instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
windows_actionADestructive
Act on a native Windows UI Automation element from windows_snapshot: invoke, setValue, toggle, select, expand, collapse, focus or scrollIntoView. Re-snapshot after meaningful UI changes.
| Name | Required | Description | Default |
|---|---|---|---|
| pid | No | Process ID from windows_list | |
| ref | Yes | ||
| title | No | Top-level window title substring (alternative to pid) | |
| value | No | ||
| action | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=false and openWorldHint=false, so the safety profile is largely covered. The description adds the re-snapshot workflow constraint, which is genuine behavioral value, but it does not indicate which of the eight actions are state-mutating (invoke/setValue) versus benign (focus/scrollIntoView), nor how failures surface when a ref goes stale.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the core capability plus its action set is front-loaded before the operational follow-up. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, non-idempotent, destructive UI-automation tool with no output schema, the description covers the action surface and one workflow rule but omits what happens on success/failure, ref staleness behavior, and any permission or focus prerequisites. It is workable but leaves an agent to infer too much before a destructive call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%: pid and title are documented in the schema, but ref, value, and action have no field descriptions. The description partially compensates by listing the full action vocabulary and identifying windows_snapshot as where ref comes from, but it says nothing about the ref format or which actions consume value, leaving real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear resource (a native Windows UI Automation element from windows_snapshot) and enumerates the eight actions it can perform, which concretely defines the tool's scope. It implicitly differentiates itself from the snapshot tool by naming it as the source of the ref, though 'Act on' is a dispatcher-style verb that relies on the action list for precision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction to re-snapshot after meaningful UI changes is a useful workflow cue, and referencing windows_snapshot as the ref source implies the correct sequence. However, it never states when to choose windows_action versus alternatives (browser_*, chrome_*, windows_screenshot) or any preconditions such as needing a valid ref from a fresh snapshot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
windows_listBRead-onlyIdempotent
List visible top-level native Windows application windows with process IDs and UI Automation refs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/closed-world, so the safety profile is covered. The description adds real behavioral detail beyond that: only visible, top-level, native windows are enumerated, and the response carries process IDs and UI Automation refs. It does not state pagination or ordering behavior, but with no output schema the return-shape hint is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the scope front-loaded and zero filler. Every clause (visible, top-level, native, returned IDs/refs) carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only enumeration tool with one optional parameter and no output schema, the description covers what is listed and roughly what comes back. The undocumented 'limit' parameter is the only meaningful omission, and safety semantics are already in the annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'limit' parameter has 0% schema description coverage and is not mentioned anywhere in the description. The agent must guess that it caps the number of returned windows rather than compensating from the text; the description does nothing to fill this gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (visible top-level native Windows application windows) plus the returned artifacts (process IDs, UI Automation refs). The scoping to 'top-level native' implicitly separates it from chrome_list_tabs/browser_tabs and from windows_snapshot, though it never names those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as windows_snapshot or windows_action. The agent must infer that this is the enumeration step preceding an action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
windows_screenshotARead-onlyIdempotent
Capture the visible bounds of a native Windows application window selected semantically by pid/title. Returns the PNG to the client and saves it under screenshots/.
| Name | Required | Description | Default |
|---|---|---|---|
| pid | No | Process ID from windows_list | |
| title | No | Top-level window title substring (alternative to pid) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive). Beyond that, the description discloses two useful traits: the PNG is returned to the client, and a copy is persisted under screenshots/ — a side effect not visible in the schema. It stops short of noting behaviors like occlusion/foreground requirements or failure when a window is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the core capability is front-loaded before the return/persistence note. Nothing is repeated from the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and zero required parameters, the description supplies the missing return-channel information (PNG returned and saved to screenshots/). It is nearly complete; the only unaddressed items are error/edge conditions such as a nonexistent pid or a minimized window.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters carry their own descriptions, so the schema does the heavy lifting. The description only echoes the pid/title alternative selection ('selected semantically by pid/title') without adding format or precedence detail beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Capture the visible bounds of a native Windows application window') and names the selection mechanism (pid/title). The qualifier 'native Windows application window' cleanly separates it from browser_screenshot, chrome_screenshot, and windows_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'native' implicitly signals when this tool applies instead of the browser/chrome screenshot siblings, and the pid reference to windows_list hints at a prerequisite. However, it never explicitly says when to choose this over browser_screenshot or record_clip, nor states any exclusion or ordering requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
windows_snapshotBRead-onlyIdempotent
Semantic Windows UI Automation snapshot of a native app window. Returns stable runtime refs, control names/types, supported patterns and bounds. Prefer this over coordinate clicking.
| Name | Required | Description | Default |
|---|---|---|---|
| pid | No | Process ID from windows_list | |
| title | No | Top-level window title substring (alternative to pid) | |
| maxElements | No | ||
| includeOffscreen | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuine value by disclosing the return contents (stable runtime refs, control names/types, supported patterns, bounds), which no output schema exists to convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the mechanism and purpose, then return values, then the routing hint. Nothing is padded, but the routing sentence is a fragment that could be folded in.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, so listing return contents is the right move. Remaining gaps: no guidance on pid vs title selection strategy, and no explanation of the undocumented maxElements/includeOffscreen parameters. Adequate for invocation basics, incomplete on tuning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: pid and title are documented (including 'alternative to pid'), but maxElements and includeOffscreen have no schema description and the tool description says nothing about parameters at all. The description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Semantic Windows UI Automation snapshot of a native app window') and clarifies the underlying mechanism (UI Automation), which separates it from the pixel-based windows_screenshot. It stops short of naming the sibling it replaces explicitly, only gesturing at 'coordinate clicking'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Prefer this over coordinate clicking' gives a directional preference, but no when-not condition and no named alternative tool (windows_screenshot or windows_action are never mentioned). Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
windows_waitBRead-onlyIdempotent
Wait for a semantic native Windows control to appear, become enabled, or disappear. Matches by accessible name substring, AutomationId and/or control type.
| Name | Required | Description | Default |
|---|---|---|---|
| pid | No | Process ID from windows_list | |
| name | No | ||
| state | No | present | |
| title | No | Top-level window title substring (alternative to pid) | |
| enabled | No | ||
| timeoutMs | No | ||
| controlType | No | ||
| maxElements | No | ||
| automationId | No | ||
| includeOffscreen | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds the semantic state model (present/enabled/gone), which is useful, but says nothing about polling behavior, what timeoutMs does on expiry, or what the call returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and immediately followed by the matching criteria. No filler, no repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A 10-parameter blocking wait tool with no output schema and 20% schema coverage demands more than two sentences. Timeout semantics, default 15s behavior, match-combination logic, and return shape are all missing, leaving the agent unable to predict behavior on failure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (just pid and title), leaving eight parameters undocumented. The description partially compensates by naming the match keys (accessible name substring, AutomationId, control type), but it never explains state (present/gone), enabled, timeoutMs, maxElements, or includeOffscreen, and the 'and/or' matching logic between the three keys is ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Wait for a semantic native Windows control') and enumerates the three wait conditions (appear, become enabled, disappear) plus the matching keys. It is clearly distinguishable from windows_list/windows_snapshot, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (call this before acting on a control that may not yet exist or be ready), but there is no explicit when-to-use or when-not guidance and no alternatives mentioned, e.g. using windows_snapshot to check current state instead of blocking.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_fileADestructive
Create or fully overwrite a file (parent folders are created). Prefer apply_patch/edit_file for changes to existing files.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File path | |
| check | No | Run compilers/linters on the changed files afterwards (default true) | |
| content | Yes | Complete new file content | |
| expectedSha256 | No | Optional optimistic precondition from file_info; refuse overwrite if the file changed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavioral context – parent folders are auto-created and the operation is a full overwrite rather than a merge – though it doesn't say what happens on failure or how the SHA precondition rejection surfaces.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste; the destructive overwrite semantics are front-loaded and the alternative-tool guidance follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with destructive annotations, a fully-covered schema, and no output schema, the description covers what it must: overwrite semantics, side effects, and routing to safer alternatives. Only failure behavior is unstated, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all four parameters including check and expectedSha256 are documented in the schema itself. The description adds no parameter-level detail beyond the overwrite scope, which is the expected baseline when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create or fully overwrite a file') and clarifies scope with 'fully overwrite' plus the parent-folder side effect. It also names sibling alternatives, so an agent can distinguish it from apply_patch/edit_file without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Prefer apply_patch/edit_file for changes to existing files,' giving both the preferred alternative and the condition that selects it. This is an explicit when-not-to-use statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
114 tool updates
v1.1.0- First observed
apply_patch - First observed
batch - First observed
browser_click - First observed
browser_close - First observed
browser_console - First observed
browser_dialog - First observed
browser_evaluate - First observed
browser_get_text - First observed
browser_hover - First observed
browser_navigate - First observed
browser_press_key - First observed
browser_record_start - First observed
browser_record_stop - First observed
browser_resize - First observed
browser_screenshot - First observed
browser_scroll - First observed
browser_select_option - First observed
browser_snapshot - First observed
browser_tabs - First observed
browser_type - First observed
browser_upload_file - First observed
browser_wait_for - First observed
capability_call - First observed
capability_list - First observed
check_files - First observed
checkpoint_list - First observed
checkpoint_rewind - First observed
chrome_call_site_tool - First observed
chrome_click - First observed
chrome_close_tab - First observed
chrome_console - First observed
chrome_dialog - First observed
chrome_downloads - First observed
chrome_evaluate - First observed
chrome_extension_reload - First observed
chrome_get_text - First observed
chrome_list_tabs - First observed
chrome_navigate - First observed
chrome_new_tab - First observed
chrome_pairing_setup - First observed
chrome_press_key - First observed
chrome_screenshot - First observed
chrome_scroll - First observed
chrome_snapshot - First observed
chrome_status - First observed
chrome_switch_tab - First observed
chrome_type - First observed
copy_path - First observed
delete_path - First observed
edit_file - First observed
fetch_url - First observed
file_info - First observed
find_files - First observed
find_symbol - First observed
get_system_info - First observed
git_commit - First observed
git_diff - First observed
git_log - First observed
git_status - First observed
job_cancel - First observed
job_list - First observed
job_output - First observed
job_start - First observed
job_status - First observed
kill_process - First observed
list_directory - First observed
list_processes - First observed
list_workspaces - First observed
lsp_calls - First observed
lsp_code_actions - First observed
lsp_definition - First observed
lsp_diagnostics - First observed
lsp_format_preview - First observed
lsp_hover - First observed
lsp_references - First observed
lsp_rename_preview - First observed
lsp_status - First observed
lsp_symbols - First observed
move_path - First observed
outline - First observed
process_input - First observed
process_list - First observed
process_output - First observed
process_start - First observed
process_stop - First observed
project_context - First observed
project_memory - First observed
pty_list - First observed
pty_read - First observed
pty_resize - First observed
pty_start - First observed
pty_stop - First observed
pty_write - First observed
read_file - First observed
read_many_files - First observed
record_clip - First observed
repo_map - First observed
resource_status - First observed
run_command - First observed
search_code - First observed
vscode_add_breakpoint - First observed
vscode_debug_state - First observed
vscode_diagnostics - First observed
vscode_find_references - First observed
vscode_go_to_definition - First observed
vscode_open_file - First observed
vscode_open_files - First observed
web_search - First observed
windows_action - First observed
windows_list - First observed
windows_screenshot - First observed
windows_snapshot - First observed
windows_wait - First observed
write_file
TDQS
Scored across 114 tools
Distinct domain prefixes (browser_, chrome_, lsp_, vscode_, pty_, job_) and detailed descriptions help, but there are multiple overlapping suites for similar actions, such as browser_* vs chrome_* and run_command vs process_start vs job_start. An agent can usually pick correctly from context, yet the parallel tool families create real misselection risk.
Most tools use snake_case with a predictable domain prefix and action-oriented names, especially browser_, chrome_, windows_, lsp_, vscode_, pty_, job_, and process_. There are minor deviations like outline, batch, repo_map, and project_context, but the overall naming pattern is readable and mostly consistent.
The server exposes 114 tools, far beyond the 3–15 range for a well-scoped set. Even with broad coding-agent capabilities, this volume makes discovery, selection, and maintenance unusually heavy.
The surface is very broad: file editing, code search, LSP, terminals, background jobs, browser and Chrome automation, Windows UI automation, video recording, git, web search, and project memory. Some lifecycle gaps remain, such as git branch/push/pull/merge operations and dedicated test/package management, but most agent workflows are covered.
Maintenance
Related MCP Connectors
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Personal assistant MCP server with search, execute, packages, jobs, secrets, and integrations.
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Related MCP Servers
- AlicenseBqualityDmaintenanceProduction-grade MCP server that gives AI agents safe access to your local dev environment: filesystem, databases, processes, and OpenAPI specs.1543 npm3MIT
- AlicenseNot gradedqualityAmaintenanceLocal-first MCP server that provides project context, verification gates, and structured tools for coding agents to discover knowledge, run diagnostics, and execute allowlisted commands within a repository.16 npmMIT
- AlicenseNot gradedqualityCmaintenanceA self-hosted MCP server that gives AI agents controlled access to a machine: filesystem, shell, background processes, git, web fetching and persistent key-value memory.GPL 3.0
- AlicenseAqualityBmaintenanceA lightweight MCP server that enables AI assistants to execute local development tools and retrieve system status with low latency over stdio or HTTP.193 npm3MIT