Skip to main content
Glama
office233
by office233

Free Coding Agent

Give frontier AI models real hands on a real developer workstation.

Free Coding Agent is an open-source, provider-neutral MCP execution layer that turns a capable model from a chat box into a hands-on software engineer. The model supplies the intelligence; Free Coding Agent supplies controlled access to files, terminals, Git, language servers, browsers, Windows UI, jobs, checkpoints, tests, and diagnostics.

LM Arena describes Code Arena as a place to build and compare with frontier AI models. Free Coding Agent is designed to be a compatible execution layer for any host that supports MCP or equivalent remote tools. It is independent software and is not an LM Arena product.

Core principle: no shared project infrastructure

Free Coding Agent does not require a relay, domain, account, API key, tunnel, or backend operated by this project.

Every installation is independent:

  • its own local configuration;

  • its own workspace allowlist;

  • its own MCP bearer token;

  • its own HTTP server;

  • its own HTTPS hostname/tunnel when remote access is enabled;

  • its own third-party account, if the user chooses a tunneling provider.

There is no common Free Coding Agent relay and no common remote credential.

Related MCP server: Workspace MCP

What it can do

  • transactional filesystem reads/writes;

  • atomic multi-file patches and rollback checkpoints;

  • optimistic SHA-256 concurrency protection;

  • shell commands and long-running processes;

  • real PTY / ConPTY for REPLs, debuggers, SSH and TUIs;

  • Git status, diff, history and repository operations;

  • headless LSP diagnostics, definitions, references, hover and symbols;

  • project inspection and project-local memory;

  • Playwright browser automation;

  • optional real-Chrome companion extension;

  • optional VS Code companion extension;

  • Windows UI automation;

  • FFmpeg/media workflows;

  • resource admission to avoid CPU/RAM stampedes;

  • MCP batching and capability discovery;

  • local stdio MCP;

  • authenticated local HTTP MCP;

  • optional personal public HTTPS MCP.

Durable Agent Mode

Free Coding Agent includes a provider-neutral task kernel for autonomous and multi-agent coding. It does not call or hardcode a model vendor: any MCP-capable model can act as a worker.

The durable loop is:

task_submit
    ↓
task_next / task_claim
    ↓
worker edits + targeted checks
    ↓
task_ready
    ↓
task_verify
   ↙       ↘
repairing  succeeded
   ↓
task_claim → repair → task_ready → task_verify

Important guarantees:

  • append-only task journal with fsync;

  • idempotency keys prevent duplicate submissions;

  • dependencies form explicit task DAGs;

  • exclusive leaseKey ownership prevents two workers from editing the same worktree at once;

  • cross-process journal locking prevents two MCP processes from racing the same task;

  • lease expiry and dead-owner recovery fail to interrupted, never fake success;

  • task_next atomically assigns the highest-priority runnable work;

  • verification commands are executed as real processes in the task workspace;

  • tasks can require an independent logical verifier identity distinct from every implementation worker;

  • failed gates create an evidence-backed repair loop;

  • configurable maximum repair rounds prevent infinite retry loops;

  • only task_verify can transition verified work to succeeded;

  • task events provide a durable audit trail.

Core tools:

task_submit
task_next
task_claim
task_heartbeat
task_ready
task_verify
task_update
task_reconcile
task_contract
task_status
task_list
task_events
task_stats
task_cancel

This is the portable agent-control layer: the model may come from OpenCode, an MCP desktop host, an IDE, a self-hosted model, or another compatible client. The task kernel only coordinates ownership, state, evidence and verification.

Worker/verifier IDs are logical identities supplied by the MCP host for coordination and audit; they are not authentication credentials.

Install on Windows

The normal installation flow is:

Download ZIP
→ extract
→ double-click install.cmd
→ choose the folder the agent may access
→ optionally configure YOUR personal HTTPS endpoint
→ Ready

The installer:

  1. verifies Node.js 20+;

  2. can install Node.js LTS when Windows Package Manager is available;

  3. installs free-coding-agent globally;

  4. configures the user's own workspace;

  5. runs free-coding-agent doctor;

  6. optionally guides the user through personal HTTPS setup.

The extracted installer folder can be deleted afterwards.

Local mode — no domain, no internet tunnel

For an MCP client running on the same machine, nothing public is needed:

{
  "mcpServers": {
    "free-coding-agent": {
      "command": "free-coding-agent",
      "args": ["stdio"]
    }
  }
}

No domain. No TLS. No exposed port. No external service.

For HTTP clients, Free Coding Agent uses the current MCP Streamable HTTP transport at POST /mcp. Deprecated legacy HTTP+SSE endpoints are intentionally not exposed in v3.

Personal HTTPS mode

When the model host is remote/cloud-based, it needs a public HTTPS URL that reaches the user's machine.

Free Coding Agent's official remote model is personal self-configuration, not a shared relay.

Run:

free-coding-agent remote setup

The setup uses the current user's own Tailscale account.

Tailscale Funnel gives that machine its own stable HTTPS hostname under the user's tailnet:

https://<device>.<tailnet>.ts.net/mcp

Free Coding Agent then:

  • generates a strong MCP bearer token locally;

  • keeps the HTTP origin bound to 127.0.0.1;

  • exposes it only through that user's Funnel;

  • stores the URL/token only in that user's application-data directory;

  • adds a current-user startup entry for the local HTTP server;

  • prints the exact MCP configuration.

Example output shape:

{
  "mcpServers": {
    "free-coding-agent": {
      "url": "https://YOUR-PERSONAL-HOST.ts.net/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_LOCAL_TOKEN"
      }
    }
  }
}

Tailscale documents Funnel hostnames as predictable/stable, automatically TLS-protected, and publicly reachable while the user's machine is online.

Alternative: personal Cloudflare Tunnel

Users who own a domain on Cloudflare can use a locally managed named Cloudflare Tunnel instead.

The ownership model remains the same:

their Cloudflare account
+ their domain
+ their tunnel
+ their DNS record
+ their local MCP token

Cloudflare's documented flow is:

cloudflared tunnel login
cloudflared tunnel create <name>
cloudflared tunnel route dns <name-or-id> <hostname>
cloudflared tunnel run <name-or-id>

A named tunnel can also be installed as an OS service. Free Coding Agent does not own or operate it.

Why personal HTTPS instead of a project relay?

Because the open-source project should keep working independently of its author.

With personal endpoints:

  • no project-operated server can go offline and break everybody;

  • no project operator sees MCP traffic;

  • no shared credential exists;

  • no shared customer/device database exists;

  • no central billing requirement exists;

  • users can replace Tailscale/Cloudflare with another provider;

  • users can self-host everything indefinitely.

CLI

free-coding-agent
free-coding-agent stdio
free-coding-agent setup
free-coding-agent remote setup
free-coding-agent remote start
free-coding-agent remote status
free-coding-agent remote off
free-coding-agent doctor
free-coding-agent print-config

npm

The package is structured for public npm installation:

npm install -g free-coding-agent
free-coding-agent setup

The npm package must exist in the registry before those commands can be used from npm directly.

Security model

The agent can modify files and execute code, so remote exposure must fail closed.

The personal HTTPS setup enforces:

  • a random local MCP_API_KEY;

  • bearer authentication for every remote MCP request;

  • loopback-only local HTTP binding;

  • filesystem allowlisting via FCA_ALLOWED_ROOTS;

  • no anonymous proxy access;

  • no secret committed to Git;

  • no project-wide tunnel or account;

  • no remote credentials shared between users.

A remote URL without the user's bearer token cannot use the agent.

User configuration

Normal users should use the CLI instead of editing config manually.

Important variables:

Variable

Purpose

FCA_WORKSPACES

User's known workspaces

FCA_DEFAULT_WORKSPACE

Default workspace

FCA_ALLOWED_ROOTS

Filesystem sandbox

FCA_REMOTE_PROVIDER

Personal remote provider

FCA_REMOTE_URL

User's personal public MCP URL

FCA_TAILSCALE_DNS_NAME

User's personal Tailscale DNS name

MCP_API_KEY

User's own HTTP bearer token

MCP_HOST

Local HTTP bind address

PORT

Local HTTP port

TASKS_DIR

Durable task journal directory

TASK_LEASE_TTL_MS

Worker lease lifetime

TASK_MAX_ACTIVE

Maximum simultaneously running task workers

TASK_MAX_REPAIR_ROUNDS

Maximum evidence-backed repair rounds

TASK_VERIFY_TIMEOUT_MS

Per verification-command timeout

FCA_CHROME_BRIDGE_ENABLED

Enable real-Chrome integration

CHROME_BRIDGE_TOKEN

Local Chrome pairing secret

CHROME_EXTENSION_ID

Optional extension identity pin

VSCODE_BRIDGE_URL

Local VS Code companion endpoint

Configuration and runtime state live in standard per-user application-data directories, outside the repository.

OpenCode

OpenCode supports both local and remote MCP servers. A personal Free Coding Agent HTTPS endpoint can be configured as a remote MCP server with bearer authentication.

Conceptually:

{
  "type": "remote",
  "url": "https://YOUR-PERSONAL-HOST/mcp",
  "oauth": false,
  "headers": {
    "Authorization": "Bearer YOUR_LOCAL_TOKEN"
  }
}

Use OpenCode's current configuration schema/documentation for the surrounding config structure.

LM Arena / frontier models

Arena's public documentation describes Agent Mode and Code Arena as environments for agentic coding and frontier-model comparison.

Free Coding Agent does not scrape or automate Arena's private UI/API. If Arena exposes a supported MCP or compatible remote-tool integration, the user's personal HTTPS MCP endpoint is ready for it.

This distinction is intentional: the open-source agent remains standards-based and does not depend on undocumented consumer-site behavior.

References:

Chrome companion extension

The optional extension gives the local MCP server controlled access to the user's real Chrome.

  1. Load chrome-extension as an unpacked extension.

  2. Configure a local CHROME_BRIDGE_TOKEN.

  3. Enter the same token in the extension popup.

  4. Optionally pin the assigned extension ID.

No fixed extension identity key is stored in this repository.

VS Code companion extension

The optional extension exposes loopback-only editor state:

  • diagnostics;

  • references;

  • definitions;

  • open files;

  • debugger state.

Configuration key:

freeCodingAgent.port

Default port: 3005.

Development

npm install
npm run check
npm run privacy
npm run preflight
npm test

Privacy / publication hygiene

The public source tree intentionally excludes:

  • local .env files;

  • API keys;

  • tunnel credentials;

  • machine-specific paths;

  • user names and email addresses;

  • conversation archives;

  • browser profiles;

  • screenshots and clips;

  • logs, checkpoints and job state;

  • local memory;

  • fixed Chrome extension identity keys;

  • provider-account-specific orchestration.

npm run privacy fails if common secrets, email addresses, absolute Windows paths, or user-home paths appear in publishable source.

License

ISC. See LICENSE.

Available Tools

128 tools
apply_patchA
Destructive

Apply a multi-file patch atomically (nothing is written unless every hunk applies). Preferred way to change code. Format: *** Begin Patch *** Update File: src/app.ts @@ function main() { (optional anchor line to jump near the change) context line (space prefix) -removed line +added line *** Add File: src/new.ts +file content line *** Delete File: src/old.ts *** Update File: a.ts *** Move to: b.ts *** End Patch Include ~3 unchanged context lines around each change. The result reports compiler/linter problems in the changed files and a checkpoint id for undo.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoBase directory for relative paths (default: default workspace)
checkNoRun compilers/linters on the changed files afterwards (default true)
patchYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=false, and the description adds behavior they cannot express: the operation is atomic and nothing is written unless every hunk applies, the result reports compiler/linter problems, and a checkpoint id is returned for undo. Atomicity and undo are exactly the non-obvious traits an agent needs before mutating many files.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the atomicity guarantee and preference statement in the first sentence, then spends the remainder on the format example, which earns its space for a grammar-driven parameter. It is long, but nearly every line carries required syntax or behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly describes what comes back (diagnostics plus checkpoint id) and covers the destructive/atomic semantics. It omits how failures are surfaced and how this relates to edit_file/write_file, which are the remaining gaps for a multi-file mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and documents cwd and check, but leaves the required 'patch' parameter undescribed. The description compensates by specifying the exact patch grammar (Begin/Update/Add/Delete/Move/End, +/-/space prefixes, optional anchor) and the context-line convention, which is what makes the required parameter usable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (apply) and resource (multi-file patch) plus a distinguishing trait: atomic, all-or-nothing. Combined with 'Preferred way to change code', an agent can separate it from nearby file-mutation siblings like edit_file and write_file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly positions itself as the preferred mechanism for changing code, which is real usage guidance. It stops short of naming the alternatives (edit_file, write_file) or stating when a single-file edit is preferable, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batchB
Destructive

Run several tool calls in one request. Parallel mode is allowed only when every called tool is read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
callsYes
parallelNo
stopOnErrorNo
failOnAnyErrorNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the risk profile is known. The description adds the read-only precondition for parallel execution, which is useful. But it omits error-handling behavior (what stopOnError vs failOnAnyError actually do, especially with both defaulting to false) and the 20-call cap from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what the tool does followed by the key constraint. No filler. It is terse to the point of under-specification, but that is a completeness problem rather than a verbosity one.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a meta-tool with 4 parameters, 0% schema description coverage, and no output schema, the description is thin. It should define batch entry structure, the 20-call ceiling, and the distinct roles of stopOnError/failOnAnyError to make the tool safely invocable. Annotations cover safety but not orchestration semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description must compensate — and largely does not. Only the conditional meaning of 'parallel' is addressed; 'calls', 'stopOnError', and 'failOnAnyError' are never explained, and the maxItems:20 limit is not mentioned. This leaves the most consequential semantics undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Run several tool calls in one request.' An agent immediately knows this is a multiplexer/batch executor, distinct from any single-call sibling. However, it does not clarify what a 'call' entry must contain (tool name + args) or how results are returned, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one explicit usage rule: parallel mode is permitted only when every called tool is read-only. That is genuine guidance. But it says nothing about when to prefer batch over issuing calls individually, nor about how stopOnError/failOnAnyError affect execution, so the guidance is partial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_clickB
Destructive

Click an element by ref (preferred), selector or text. Returns the updated snapshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from browser_snapshot, e.g. "e12"
textNoVisible text to match (alternative to ref)
buttonNo
selectorNoCSS or Playwright selector (alternative to ref)
snapshotNo
modifiersNo
doubleClickNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and non-idempotent, so the safety profile is largely covered. The description adds the return behavior ('Returns the updated snapshot'), but it omits meaningful behavioral context for a click tool — e.g. that clicking may navigate, submit forms, or trigger irreversible actions, which matters given destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the targeting preference front-loaded and the return value second; every word earns its place and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, zero-required, destructive mutation tool with no output schema and only 43% parameter documentation, the description is thin. It does not explain the button, modifiers, doubleClick, or snapshot options, nor the side effects an agent should anticipate before clicking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, and four parameters (button, snapshot, modifiers, doubleClick) have no schema description at all. The description only hints at the ref/selector/text triage, which the schema already labels as alternatives, so it does not compensate for the undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) and resource (an element), plus the three supported targeting modes (ref, selector, text). It is clearly distinguishable from sibling read-only tools like browser_get_text or browser_snapshot, though it does not distinguish itself from the near-identical chrome_click or adjacent interaction tools such as browser_hover.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(preferred)' on ref implies a targeting precedence, which is light usage guidance. However, there is no explicit when-to-use vs. when-not, and no mention of prerequisites (e.g. that a ref must come from a prior browser_snapshot) or which siblings to prefer for hovering, typing, or pressing keys.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_closeA
Destructive

Close the Playwright browser (the persistent profile stays on disk).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is covered. The description adds a valuable non-obvious behavioral detail: the persistent profile is not destroyed, which meaningfully clarifies the scope of the destructive action. It still does not say what live session state or tabs are lost when the browser closes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. The essential action is stated first, and the important caveat about the persistent profile is appropriately appended.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter close operation with no output schema, the description plus annotations provide sufficient context to invoke the tool correctly. Minor omissions remain, such as behavior when no browser is open or confirmation of what session state is discarded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero input parameters, so there are no parameter semantics for the description to explain. The schema itself has 100% description coverage, and the baseline for a zero-parameter tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Close the Playwright browser.' The parenthetical narrows the scope by clarifying that the persistent profile remains on disk, which distinguishes this from destructive cleanup that would delete profile state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use or when-not-to-use guidance. It implies end-of-session cleanup, but it does not name alternatives such as browser_tabs, chrome_close_tab, or other browser lifecycle tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_consoleC
Read-onlyIdempotent

Console messages, page errors, failed requests and dialogs captured from the browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNo
limitNo
onlyErrorsNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, open-world behavior. The description usefully adds the content types captured (errors, failed requests, dialogs), but says nothing about the mutating effect of the 'clear' parameter, buffering, or whether logs persist across navigations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, and the most important content types are front-loaded. It reads as a fragment rather than a verb-led instruction, but no words are wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With three undocumented parameters, no output schema, and a 'clear' flag that appears to wipe captured data, the definition is too thin to invoke safely. An agent cannot know what clear does or what limit/onlyErrors control.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters, so the description is the only place semantics could be added — and it adds none. 'clear', 'limit', and 'onlyErrors' are never mentioned; 'onlyErrors' is only obliquely implied by the phrase 'page errors'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resources retrieved — console messages, page errors, failed requests, dialogs — so an agent knows exactly what it gets. It omits an action verb and gives no differentiation from the sibling chrome_console, which appears to expose the same data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as chrome_console or browser_dialog despite heavy sibling overlap. The agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_dialogB
Destructive

Inspect or answer a JavaScript dialog (alert/confirm/prompt) in the Playwright browser. Dialogs are never accepted automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNostatus
promptTextNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the write-capable nature is covered structurally. The description usefully adds that dialogs are never auto-accepted (so the agent must handle them explicitly), but it does not explain what accept vs dismiss does to the underlying page or how promptText applies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the resource and scope. No filler, though the second sentence is a caveat rather than a usage instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with an undocumented enum and no output schema, the description leaves real gaps: it never maps the action values to behavior or explains promptText, so an agent must infer the accept/dismiss/prompt semantics from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with 2 parameters, and the description explains neither 'action' (whose enum values status/accept/dismiss map directly onto the stated capabilities) nor 'promptText' (presumably the text submitted on accept of a prompt). Only the verbs 'inspect'/'answer' loosely hint at the action values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: inspect or answer a JavaScript dialog, scoped to the Playwright browser. It reasonably distinguishes it from the sibling chrome_dialog, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied — the tool is obviously for when a dialog is blocking the browser, but the description gives no guidance on when to use status vs accept vs dismiss, nor any prerequisite or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_evaluateB
Destructive

Run JavaScript in the page. Pass an expression ("document.title") or a function body using return.

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=true, so the risk profile is covered by structured data. The description adds nothing behavioral beyond that: it says nothing about execution context/privileges, sandboxing, timeouts, or error behavior for arbitrary code execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and then the argument format. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one parameter and annotations covering safety, the description is nearly sufficient. But there is no output schema, so the return value (serialized result, or what happens on a thrown error / non-serializable value) should have been mentioned to let an agent call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden and does so usefully: it clarifies that 'script' can be a bare expression ("document.title") or a function body requiring an explicit return. That is real semantic guidance beyond the bare string type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Run JavaScript in the page.' This is clearly distinguishable from read-oriented siblings like browser_get_text or browser_snapshot, though it never explicitly names them or notes its overlap with browser_console/chrome_evaluate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two argument forms but gives no guidance on when to reach for this tool versus browser_get_text, browser_snapshot, or browser_console. No prerequisites, no when-not, no alternatives are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_get_textC
Read-onlyIdempotent

Visible text of the page or of an element (CSS selector).

ParametersJSON Schema
NameRequiredDescriptionDefault
maxCharsNo
selectorNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive) and open-world behavior, so the bar is lower. The description adds the useful nuance that only 'visible text' is returned, but it omits truncation behavior for maxChars, handling of missing selectors, and default selector behavior, leaving meaningful gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no wasted words, and the core resource is front-loaded. It is arguably too thin, but as a pure conciseness measure it is efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, no-output-schema tool, the description leaves key behaviors unspecified: what happens when selector is omitted (whole page?), how maxChars truncation works, and failure behavior when a selector matches nothing. Annotations cover safety, but the description is insufficient for the agent to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that the selector is a CSS selector, but says nothing about the maxChars parameter, leaving one of two parameters undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource ('visible text') and its two scopes ('page or of an element'), making clear what kind of content is returned. It lacks an explicit verb and does not distinguish itself from siblings like browser_snapshot or browser_evaluate, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as browser_snapshot or browser_evaluate. The phrase 'page or of an element (CSS selector)' hints at two modes but provides no conditions or preferences for selecting between them or among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_hoverC
Destructive

Hover over an element.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from browser_snapshot, e.g. "e12"
textNoVisible text to match (alternative to ref)
selectorNoCSS or Playwright selector (alternative to ref)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (destructiveHint=true, openWorldHint=true, idempotentHint=false), so the description is not burdened with that. However, it adds nothing beyond them: no mention of whether hovering triggers hover-only UI, side effects, or failures on missing elements. No contradiction with the annotations, just zero added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four words, front-loaded, no filler. It is as concise as possible, though the brevity borders on under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-optional-param interaction tool with full schema coverage and no output schema, the minimal description is borderline adequate. It omits target-resolution precedence (ref vs text vs selector) and what hovering actually produces, which an agent would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each of ref, text, and selector documented as alternatives, so the schema carries semantics. The description adds nothing about precedence among the three targeting modes, which is the one thing missing. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: hover over an element. It is distinguishable from the other browser_* tools (click, type, scroll) by its unique action, though it offers no explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to hover versus clicking, typing, or waiting, and no note about prerequisites such as having a snapshot ref available. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_navigateB
Destructive

Open a URL in the Playwright-controlled Chrome (persistent profile: logins are remembered). Returns the page snapshot with element refs.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
newTabNo
snapshotNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true, idempotentHint=false and readOnlyHint=false. The description adds genuinely useful context not in the annotations — persistent profile with remembered logins, and that it returns a page snapshot with element refs. It does not, however, explain why navigation is flagged destructive (loss of current page state).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and runtime, then the return value. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully states the return value, and annotations cover the safety profile. But it omits newTab semantics and any routing guidance among the many browser/chrome siblings, so it is adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, but it only implies 'url' and 'snapshot' (via 'Returns the page snapshot'). The newTab parameter is undocumented in both schema and description, leaving a meaningful gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Open a URL in the Playwright-controlled Chrome') and distinguishes the runtime from the chrome_* extension siblings. An agent can tell it navigates the Playwright browser, though it does not name which sibling to prefer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of when to choose this over chrome_navigate or fetch_url. The agent must infer usage entirely from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_press_keyB
Destructive

Press a key or chord on the page, e.g. "Enter", "Escape", "Control+A", "ArrowDown".

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
snapshotNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, readOnlyHint=false, openWorldHint=true, and non-idempotency, so the safety profile is covered. The description adds the chord syntax convention but says nothing about focus requirements, that keystrokes can mutate page state, or any rate/confirmation concerns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the action front-loaded and examples inline. No wasted wording, though it is sparse enough that a clause about focus or the snapshot flag could have earned its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutating tool with no output schema, the description covers the primary input but omits the snapshot parameter entirely and gives no behavioral context beyond annotations. Adequate but with a visible gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load. It does clarify the 'key' parameter with chord examples like 'Control+A' and 'ArrowDown', which is genuinely useful syntax information absent from the schema. However, the second parameter 'snapshot' is never mentioned, leaving half the parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Press a key or chord on the page') and supplies concrete examples, which separates it from browser_type and browser_click. It does not explicitly name the sibling it differs from, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to press a key versus using browser_type, browser_click, or a chord-based alternative, and no stated prerequisites such as an element needing focus. Usage is only implied by the examples.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_record_startA
Destructive

Start recording a video of a browser session. Opens a dedicated recording tab (with your profile logins); every browser_* action afterwards is filmed until browser_record_stop.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNo
urlNo
nameNo
widthNo
formatNomp4
heightNo
headlessNo
resourceWaitMsNo
useProfileLoginsNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/openWorld/non-idempotent, and the description adds real behavioral detail beyond them: a dedicated tab is opened, it uses the caller's profile logins, and the recording is scoped to subsequent browser_* actions until stop. It does not explain why the operation is flagged destructive or what happens to the recording output, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and followed by the lifecycle constraint. No filler or restated title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description covers the lifecycle adequately but says nothing about how the resulting video is identified or retrieved, nor about the parameter space (resolution, format, fps, headless behavior). It is minimally viable rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage across 9 parameters (fps, format, width/height, headless, resourceWaitMs, url, name, useProfileLogins), and the description documents none of them. It only obliquely hints at useProfileLogins via 'with your profile logins', so the agent must infer the meaning and units of nearly every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start recording a video of a browser session') and immediately clarifies the mechanism ('Opens a dedicated recording tab'). An agent can distinguish this from browser_record_stop and from browser_screenshot without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context — recording begins here and captures 'every browser_* action afterwards ... until browser_record_stop' — which effectively names the terminating sibling. It lacks an explicit when-not-to-use (e.g., versus record_clip or one-off screenshots), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_record_stopB
Destructive

Stop the recording and save the clip (mp4 via ffmpeg). Returns the file path.

ParametersJSON Schema
NameRequiredDescriptionDefault
discardNo
trimStartNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds useful behavior beyond annotations by stating the output is an mp4 via ffmpeg and that it returns a file path, but it does not explain how the destructive discard option affects saved output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the action, output format, and return value without filler. Every clause contributes to understanding the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with two completely undocumented parameters and no output schema, the description is too sparse. It covers stop/save/return format, but omits any explanation of discard or trimStart, leaving significant gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention either parameter. It provides no meaning for 'discard' or 'trimStart', so an agent cannot infer how these options alter the stop/save behavior from the description alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Stop') and resource ('the recording'), and specifies the save behavior and output format. It distinguishes itself from browser_record_start by describing the stop/save action rather than starting a recording.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Stop the recording and save the clip' implies this should be used after a recording has been started, which is clear context. However, it does not name alternatives such as browser_record_start or record_clip, nor does it say when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_resizeC
Destructive

Set the viewport size.

ParametersJSON Schema
NameRequiredDescriptionDefault
widthYes
heightYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=false, destructiveHint=true, openWorldHint=true, and idempotentHint=false. The description adds no behavioral context beyond the basic action, such as whether the resize affects subsequent screenshots, evaluations, or page layout. With annotations covering the safety profile, the description's lack of additional detail is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. However, it is so terse that it borders on under-specification for a tool with two required parameters and no parameter documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation tool, annotations cover the safety profile, but the description provides no parameter semantics, units, or usage context. With no output schema and 0% schema description coverage, an agent lacks enough guidance to invoke this correctly without guessing units or intent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required parameters (width and height). The description does not mention units (pixels?), valid ranges, or any meaning beyond what the bare property names already imply, so it fails to compensate for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Set the viewport size.' This clearly distinguishes it from siblings like browser_screenshot or browser_evaluate, which capture or inspect browser state. It lacks any further scope or differentiation, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no prerequisites, and no context about how resizing interacts with other browser operations. The description only states what it does, not when or why to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_screenshotA
Read-onlyIdempotent

Screenshot of the page (or one element). The image is returned so you can see it, and saved to disk.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from browser_snapshot, e.g. "e12"
textNoVisible text to match (alternative to ref)
fullPageNo
selectorNoCSS or Playwright selector (alternative to ref)

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds a genuine behavioral fact not in annotations: the image is both returned to the model and persisted to disk. It omits where the file lands and what format, which caps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that loses no words and immediately states scope and return behavior. Nothing redundant or padding-like.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so partially, explaining that an image comes back and is saved to disk. It leaves unspecified the save location, image format, and whether the returned image is inline, which an agent might need for follow-up steps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so ref/text/selector are already documented in the schema; the description adds only the page-vs-element framing. The boolean fullPage has no description in either place, leaving one parameter's semantics unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Screenshot of the page (or one element)') and clarifies scope (whole page vs. a single element), which lets an agent distinguish it from browser_snapshot or browser_get_text. It does not name those siblings explicitly, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by 'The image is returned so you can see it,' which suggests reaching for this when visual output is needed rather than text. No when-to-use/when-not or comparison against browser_get_text or browser_snapshot is given, so the routing decision is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_scrollC
Destructive

Scroll the page by deltaY/deltaX pixels, or scroll an element (ref/selector/text) into view.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from browser_snapshot, e.g. "e12"
textNoVisible text to match (alternative to ref)
deltaXNo
deltaYNo
selectorNoCSS or Playwright selector (alternative to ref)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation/safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false), and the description adds nothing behavioral on top — no note on prerequisites (page must be loaded), what happens if the ref/text/selector is not found, or whether the scroll delta is bounded. For a state-changing tool with no output schema this is a real gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with both use modes front-loaded and no filler. It could have been split into a target-selection clause and a delta clause for faster scanning, but it wastes nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Five optional parameters, no output schema, and the description only broadly names both modes without covering default deltas, overloaded-parameter precedence, or failure behavior when a target is not found. It is minimally adequate but not complete for a tool with this many knobs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description contributes the useful unit detail that deltaY/deltaX are in pixels, which the schema omits. However it does not explain precedence when ref, text, and selector are all supplied, nor the deltaY=600 default behavior, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scroll) plus both targeting modes (pixel delta vs element into view), so the agent knows exactly what the tool does. It does not differentiate itself from the near-identical sibling chrome_scroll or from other viewport-affecting tools, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all: nothing about when to scroll versus calling browser_snapshot, browser_wait_for, or a click that auto-scrolls. No prerequisites or sequencing hints are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_select_optionC
Destructive

Select option(s) in a .

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from browser_snapshot, e.g. "e12"
textNoVisible text to match (alternative to ref)
valueNo
valuesNo
selectorNoCSS or Playwright selector (alternative to ref)

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=false, and openWorldHint=true, covering the safety profile. The description adds no further behavioral context: it does not mention event triggering, multi-select behavior, side effects, or what 'destructive' might mean in this context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words, but it is under-specified for a tool with five parameters and browser-interaction complexity. Brevity here reflects missing information rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and only partial schema parameter coverage, the description should explain how to choose between ref, text, and selector, and what value/values do. It provides none of this, leaving the agent with insufficient context beyond the annotations and schema for a five-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60% (ref, text, and selector have descriptions; value and values do not). The description does not explain any parameter semantics or clarify how ref, text, value, values, and selector interact, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: select option(s) in a <select>. It is clear enough to know the operation, but it does not distinguish this tool from siblings like browser_click or browser_type, nor does it mention the browser-automation context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as browser_click, browser_type, or browser_evaluate. The description only restates the operation with no context about preconditions or selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_snapshotA
Read-onlyIdempotent

Accessibility snapshot of the current page with [ref=eN] handles. Use it before clicking/typing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds meaningful context by disclosing the return format (ref handles) and the workflow dependency for click/type, but says nothing about snapshot freshness, staleness after navigation, or size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the output format stated first and the usage trigger second. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter, read-only snapshot tool with no output schema, the description covers what the tool returns (ref handles) and when to reach for it. It could mention that refs become invalid after navigation, but nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is nothing for the description to clarify beyond what the empty schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource (accessibility snapshot of the current page) and specifies the output shape ([ref=eN] handles), which distinguishes it from browser_screenshot and browser_get_text by implication. It stops short of explicitly contrasting itself with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use it before clicking/typing" gives a concrete trigger condition and positions the tool in the interaction workflow. It does not state when-not to use it or name the visual/text alternatives (browser_screenshot, browser_get_text) explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_tabsC
Destructive

List, open (new), select or close tabs by index.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
indexNo
actionNolist

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=true, so the agent knows some actions mutate and destroy state. The description adds the useful fact that one tool multiplexes four distinct operations addressed by index, but it never says that 'close' is irreversible or that index is mandatory for select/close. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core purpose front-loaded and zero filler. It is efficient, though the terseness is partly under-specification rather than true compression.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-action multiplexer with zero required parameters, no output schema, and no action/parameter mapping rules, the description is too thin. An agent still cannot tell which parameters are valid per action or how this tool relates to the chrome_* tab siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and no parameter is required, so the description carries the full burden — yet it only obliquely implies 'index' is the addressing key and 'open (new)' hints that url pairs with the new action. It does not state that url is used only with action='new', nor that index is required for select/close, leaving invalid combinations undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names four concrete verbs (list, open, select, close) and the resource (tabs), plus the addressing mechanism (by index), so an agent immediately understands it is a tab-management multiplexer. It falls short of a 5 because it never distinguishes itself from the closely-overlapping siblings chrome_list_tabs, chrome_new_tab, chrome_switch_tab, and chrome_close_tab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'by index' is the only usage signal; there is no statement of when to pick this tool over the four chrome_* tab siblings, no prerequisites, and no guidance on which action to use in which situation. The four actions are simply enumerated without routing logic.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_typeA
Destructive

Type into an input/textarea/contenteditable. Replaces its value unless slowly: true (types key by key). submit: true presses Enter.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from browser_snapshot, e.g. "e12"
textYes
slowlyNo
submitNo
selectorNoCSS or Playwright selector (alternative to ref)

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, but the description is what explains why: the default behavior replaces the field's existing value, and only slowly:true preserves key-by-key entry. It also discloses that submit:true fires Enter, a side effect not captured by annotations. Focus/permission requirements remain unmentioned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight clauses with no filler, and the destination plus the default overwrite behavior are front-loaded. The telegraphic 'slowly: true' shorthand is compact but slightly cryptic without a comma-joined explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with annotations and no output schema, the description covers the main non-obvious behaviors (overwrite default, slow typing, Enter submission). It omits failure handling and ref-vs-selector precedence, which are minor given the schema already defines both targeting options.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40% and slowly/submit have no schema descriptions, yet the description supplies their exact semantics (slow = per-key typing, submit = Enter press) plus the default replace behavior of text. ref and selector are already documented in the schema, so the remaining gap is small.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (type) and the exact target set (input/textarea/contenteditable), which is more precise than the tool name alone. It does not name a sibling or draw an explicit boundary against browser_press_key or browser_click, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the target-element list, and the slowly/submit flags hint at when each mode applies (character-by-character vs instant, Enter submission). There is no explicit when-to-use/when-not guidance and no routing to browser_press_key for plain key presses.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_upload_fileC
Destructive

Upload local file(s) through a file input or an upload button (ref/selector/text).

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from browser_snapshot, e.g. "e12"
pathNo
textNoVisible text to match (alternative to ref)
pathsNo
selectorNoCSS or Playwright selector (alternative to ref)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and non-idempotent behavior, so safety context exists. The description adds virtually nothing beyond them — no note that the named local path must exist, that it interacts with the host filesystem, or what happens if the element is not a valid file input.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded. Effectively sized, though the parenthetical overloads it with locator terminology that arguably belongs in usage guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world mutation tool with five parameters, no output schema, and no required params, the description omits the prerequisite (ref acquisition), the path/paths semantics, and any failure or side-effect behavior. What an agent needs to call it correctly is largely missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description names ref/selector/text, mapping to three of the five parameters, but says nothing about the single-vs-multiple file distinction between 'path' and 'paths', which is the most consequential semantic in the schema. With 60% schema coverage, some compensation occurs but the critical path/paths ambiguity is left unresolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Upload local file(s)') and mentions the target forms (file input or upload button), which distinguishes it from sibling tools like browser_click or browser_type. The trailing '(ref/selector/text)' is confusingly framed as if those were upload methods rather than element locators, slightly muddying the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no exclusions, and no mention of prerequisites such as obtaining a ref from browser_snapshot before calling this. The agent is left to infer the workflow entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_wait_forB
Destructive

Wait for a selector or text to appear (state visible/hidden/attached), or just wait timeoutMs.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
stateNo
selectorNo
timeoutMsNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=true, openWorldHint=true, idempotentHint=false), so the description's marginal job is to explain wait semantics — and it does not. The single most important behavior for a wait tool, what happens when the timeout elapses (throw vs. return a negative result), is never stated, nor is whether it blocks other operations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence, front-loaded with the primary mode and trailing the fallback mode. Minor inaccuracy in the parenthetical state list (says 'appear ... visible/hidden/attached' while the enum also includes 'detached') is the only blemish.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should at minimum establish return/timeout behavior and the full state vocabulary; it does neither. For a four-parameter, zero-required tool with no schema documentation, this leaves a real gap an agent must guess around.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden and only partially does. It maps 'text'/'selector' to the two wait modes and names usable states, but it calls out only three of the four enum values (omitting 'detached') and never defines what 'attached' vs 'visible' means or that text and selector should not both be supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (wait) and the two wait-able resources (selector or text), plus the alternate timeout-only mode, so the agent knows this is a blocking-until-condition tool rather than a poll/read tool. It does not explicitly contrast itself with siblings like browser_get_text or browser_snapshot, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: pass selector/text to block on a condition, or pass only timeoutMs to pause. There are no explicit when-to-use/when-not rules, no note that text and selector are alternatives, and no mention of a sibling for non-blocking inspection. Useful but inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capability_callC
Destructive

Fallback dispatcher for a capability present on the live server but missing from a cached client catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
toolYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose the risky profile (destructive, openWorld, non-idempotent, not read-only), so the safety burden is largely covered. The description adds that this is a fallback path, but omits critical behavior for a dispatcher: what happens when 'tool' is unknown, whether args are forwarded verbatim, and whether the call is proxied elsewhere.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler or redundancy. It is efficient, though brevity here borders on under-specification rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a generic dispatcher that can invoke arbitrary capabilities with arbitrary arguments, the description is thin: no guidance on valid tool identifiers, argument-passing contract, error behavior for missing capabilities, or reference to capability_list for discovery. With no output schema and 0% param coverage, the description should carry far more.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are two parameters, including a nested 'args' object with no documented shape. The description says nothing about what 'tool' should contain (a capability name) or how 'args' must be structured to match the target capability's schema, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'fallback dispatcher' conveys a generic act of routing, and the qualifier 'capability present on the live server but missing from a cached client catalog' is really a usage condition rather than a statement of what the tool does. An agent can infer it invokes a named capability, but the verb is vague and the resource is abstract.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the trigger condition (a capability is on the live server but absent from the cached catalog), which is a real usage cue. However, it never names the natural partner tool (capability_list) for discovering capabilities, nor does it state when this fallback should NOT be used versus a normal direct call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capability_listB
Read-onlyIdempotent

List the live MCP capability catalog. Use prefix to filter tool names.

ParametersJSON Schema
NameRequiredDescriptionDefault
prefixNo
includeDescriptionsNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description only adds the word 'live', hinting the catalog is dynamically generated, without saying whether it reflects runtime-registered capabilities or anything about result shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler; the core purpose leads and the parameter hint follows. Slightly terse relative to what is left unexplained, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, no-required-input read tool with no output schema, the definition is minimally workable. It still leaves includeDescriptions behavior and the return shape (names only vs. full metadata) to inference, which matters for an agent choosing between this and capability_call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and only partly does: it explains prefix ('filter tool names') but leaves includeDescriptions completely undocumented, and it does not say whether prefix matches tool names, capability IDs, or both, nor whether it is case-sensitive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'List the live MCP capability catalog' — which clearly distinguishes it as the read/enumeration counterpart to capability_call. It stops short of naming that sibling explicitly, but the pairing is unambiguous from the verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use prefix to filter tool names' addresses how to use a parameter rather than when to choose this tool over capability_call or when a full unfiltered listing is appropriate. Usage is implied by the name/description but never stated as a condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_filesB
Read-onlyIdempotent

Run the relevant compilers/linters (tsc, eslint, node --check, go vet, gofmt, py_compile, ruff, JSON) on specific files and report problems.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoFile
pathsNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description usefully adds that problems are reported rather than auto-fixed, but omits output format, whether a nonzero exit is surfaced, and whether all listed linters always run or only those matching the file type.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the enumerated linters are the substance of the tool and earn their length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and two loosely documented parameters, the description carries real burden. It explains what runs and that problems are reported, but leaves the path/paths semantics and the result shape unexplained, so it is only marginally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: 'paths' carries no description at all. The description says 'on specific files' but never clarifies the relationship between the singular 'path' and plural 'paths' parameters, nor whether both are required or mutually exclusive, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Run') and resource ('compilers/linters') and enumerates the actual tools invoked (tsc, eslint, node --check, go vet, etc.). This is far more concrete than a bare 'check' name, though it never distinguishes itself from lsp_diagnostics, vscode_diagnostics, or run_command, which also surface problems.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement of prerequisites, and no routing to alternatives despite several siblings (lsp_diagnostics, vscode_diagnostics, run_command) that overlap with static analysis. The agent must infer the selection criterion entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkpoint_listA
Read-onlyIdempotent

List recent checkpoints (one per file-changing tool call) with the files each one touched.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish this as a read-only, idempotent, non-destructive operation, so safety is covered. The description usefully adds the semantics of what a checkpoint represents (one per file-changing tool call) and that it reports touched files, but stays silent on ordering, retention window, or pagination behavior for older checkpoints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tightly written sentence with no filler; the resource is front-loaded and the parenthetical definition earns its place by clarifying a project-specific term.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and only one optional parameter, the description conveys enough to call it correctly and previews the return shape (checkpoints with their touched files). It falls just short on the meaning of the limit parameter and return ordering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'limit' parameter has no description in the schema (0% coverage), and the description only gestures at it via the word 'recent'. It does not state whether limit counts checkpoints or files, or what happens when there are fewer than the limit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (checkpoints), and even defines what a checkpoint is — one per file-changing tool call — plus the payload (files touched). It does not, however, distinguish itself from the sibling checkpoint_rewind, which an agent might confuse for the read side of this feature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the discovery step before calling checkpoint_rewind, but the description never says so and offers no when-to-use or when-not-to-use guidance. No alternative is named despite checkpoint_rewind being an obvious adjacent tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkpoint_rewindA
Destructive

Undo: restore files to their state before the given checkpoint (default: the latest), undoing it and every later checkpoint.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuine value beyond that by disclosing exactly what is destroyed: the target checkpoint 'and every later checkpoint', plus the default-target behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the destructive verb 'Undo', with the scope and default packed in without filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent single-param tool with no output schema, the description covers the action, the default target, and the collateral scope. It could still note that the undo is itself not trivially reversible, but the annotations carry the risk signal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter 'id' is undocumented in the schema. The description compensates by explaining that id refers to a checkpoint and that omitting it defaults to the latest checkpoint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific action (restore/undo) and resource (files at a named checkpoint), and states the scope includes the named checkpoint plus every later one. It is clearly distinguishable from checkpoint_list, the only nearby sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the default (latest checkpoint when no id is given) and the blast radius of the operation, but never states when to prefer this over alternatives (e.g. checkpoint_list to inspect first) or any precondition. Usage is implied rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_call_site_toolA
Destructive

Invoke a tool the web page itself declares (window.mcp / window.mcp_tools / meta[name=webmcp-tool]). Take a chrome_snapshot first to see the catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
tabIdNoChrome tab id (default: active tab)
frameIdNoFrame id, default 0
toolNameYes
snapshotIdNosnapshotId returned by chrome_snapshot; rejects stale actions

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, so the agent knows this is an unsafe, side-effecting call into page-supplied code. The description adds the mechanism (page-declared tools) but says nothing about trust boundaries, failure modes, or that the invoked behavior is entirely page-controlled. Adds some context; not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the purpose and followed by the prerequisite action. Nothing redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and for an open-world, destructive tool that proxies arbitrary page-declared APIs the description never indicates what is returned or how errors from the page tool surface. The snapshot prerequisite is covered, but return/error behavior is left entirely implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%; tabId, frameId and snapshotId are documented in the schema, but the required toolName and the free-form nested args object are undocumented anywhere. The description only indirectly supports snapshotId via the snapshot-first instruction and adds no syntax for args, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (invoke) and resource (a tool the web page itself declares), and names the concrete declaration surfaces (window.mcp / window.__mcp_tools__ / meta[name=webmcp-tool]). This clearly separates it from browser_evaluate or chrome_evaluate, which run agent-supplied code rather than page-declared tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit prerequisite in 'Take a chrome_snapshot first to see the catalog,' routing the agent to the sibling that produces the toolName/snapshotId it needs. No when-not guidance or statement about what to do if the page declares no tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_clickA
Destructive

Click an element by its [#N] ref using a real (trusted) mouse click via the Chrome debugger; falls back to a DOM click if the debugger cannot attach.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYes
tabIdNoChrome tab id (default: active tab)
frameIdNoFrame id, default 0
snapshotIdNosnapshotId returned by chrome_snapshot; rejects stale actions
doubleClickNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is partly covered. The description adds real value beyond them: it explains the trusted/debugger click path and the DOM-click fallback when the debugger cannot attach, which tells the agent about reliability trade-offs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler: mechanism, parameter semantics, and fallback all in one clause chain. Nothing wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an action tool with annotations covering the safety profile and no output schema, the description covers mechanism, ref semantics, and failure fallback adequately. Minor gaps: no mention of whether it scrolls into view or what a failed click reports, but these are small for a click primitive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description clarifies the ref semantics ([#N], element reference) which the schema alone does not. It says nothing about doubleClick, tabId, frameId, or snapshotId beyond what the schema already documents, so it lands at the baseline for partial coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (click an element) and specifies the mechanism (trusted mouse click via Chrome debugger) with a fallback. The [#N] ref format is called out, which is useful given the ref parameter is only typed as integer. Does not explicitly distinguish itself from the sibling browser_click, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage (you need a [#N] ref, presumably from chrome_snapshot) and discloses the fallback path, but never states when to prefer this over browser_click or other click siblings. The condition selecting the alternative is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_close_tabC
Destructive

Close a tab.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesChrome tab id (default: active tab)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true. The description adds no behavioral context beyond that, such as irreversibility, what happens to the active tab, or permission requirements. It merely restates the action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

'Close a tab.' is extremely concise with no wasted words, but it is under-specified for a destructive tool. It front-loads the action but omits any caveat or scoping detail that would help an agent use it safely.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with rich annotations, the definition is minimally functional. However, it lacks any mention of the default active-tab behavior or how it relates to sibling close tools, leaving the agent to rely entirely on the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema documents tabId with a default value. The description adds no parameter meaning beyond what the schema provides, so the baseline of 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Close' and resource 'tab', making the action clear. However, it provides no differentiation from sibling tools such as browser_close or chrome_switch_tab, so an agent cannot easily distinguish this tool's scope without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives, and no prerequisites are given. The agent must infer appropriate usage entirely from the tool name and context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_consoleC
Read-onlyIdempotent

Console messages and exceptions captured from a tab since the debugger attached to it.

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNo
tabIdNoChrome tab id (default: active tab)

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful scope context ('since the debugger attached' implies a bounded, non-persistent buffer), but says nothing about retention limits, when the buffer resets, or what clearing does.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the scoping constraint front-loaded and no filler. It is a sentence fragment without a verb, but that costs little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should characterize the returned message structure and the effect of 'clear'; it does neither. It never states the prerequisite that the debugger must already be attached for any messages to exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: tabId is documented in the schema (default: active tab) but 'clear' is bare in both schema and description. The description never mentions that clear discards captured messages, which is the one parameter most likely to surprise a caller.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource precisely — console messages and exceptions from a tab — with the scoping qualifier 'since the debugger attached'. It is clear what this tool returns, though it never states an explicit verb (retrieve/list) and gives no differentiation from the near-identical sibling browser_console.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., debugger must be attached), and no routing advice. The sibling list contains both browser_console and chrome_console, and the description does nothing to help an agent choose between them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_dialogA
Destructive

Inspect or answer a JavaScript dialog (alert/confirm/prompt/beforeunload) that is blocking a tab. Dialogs are never accepted automatically: read the message, then accept or dismiss (promptText for prompt()).

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdNoChrome tab id (default: active tab)
actionNostatus
promptTextNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, readOnlyHint=false, and non-idempotence. The description adds genuine context: dialogs are never auto-accepted, and the required read-then-answer sequence. It doesn't detail consequences of accept vs dismiss, which would push it higher.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the constraint ('never accepted automatically') front-loaded, followed by the procedure. Zero wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a dialog handler with no output schema, the description covers the essential behavior and sequence. It omits what 'status' returns for a non-blocking tab, a minor gap given the small tool surface.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 33%; tabId is documented in the schema itself, and the enum action values are self-explanatory. The description usefully clarifies that promptText applies to prompt() dialogs, but adds nothing for the status vs accept vs dismiss distinction beyond the schema enum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('inspect or answer') and resource ('JavaScript dialog'), and enumerates the dialog types (alert/confirm/prompt/beforeunload). No sibling tool handles dialogs, so the agent can identify it unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies clear context ('that is blocking a tab') and a workflow ('read the message, then accept or dismiss'). It stops short of naming alternatives or explicit when-not-to-use conditions, but the trigger condition is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_downloadsB
Read-onlyIdempotent

Recent Chrome downloads (file path, state, size, source URL, referrer).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
sinceNoISO time; only downloads started after it

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so safety and repeatability are covered structurally. The description adds little behavioral context beyond enumerating return fields; it says nothing about ordering, truncation at the default limit, or whether the list is scoped to the active browser profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero filler; the resource leads and the fields follow. It is perhaps too terse to earn a 5 relative to the gaps above, but there is no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the parenthetical field list is genuinely useful for anticipating the response shape, which partially compensates. However, an agent still cannot tell ordering, default result count, or whether 'since' is required for meaningful output, so the definition is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: 'since' is documented in the schema as an ISO time filter, but 'limit' has no description at all in either place. With low coverage, the description should compensate, yet it never mentions pagination, the default of 10, or ordering of the returned items.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('Recent Chrome downloads') and enumerates the returned fields (file path, state, size, source URL, referrer), so an agent knows exactly what data this yields. It is not a verb+resource phrasing, but the retrieval intent is unambiguous and the tool is clearly distinct from browser_* and file_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, no prerequisites, and no exclusions. The word 'Recent' implies a recency filter but the agent is not told how recency interacts with the 'since' parameter or when to prefer this over chrome_status, chrome_snapshot, or resource_status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_evaluateB
Destructive

Run JavaScript in a tab's page context (via the debugger). Pass an expression; the JSON-serializable result is returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdNoChrome tab id (default: active tab)
expressionYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is non-read-only, destructive, open-world, and non-idempotent. The description adds the useful 'via the debugger' context and the fact that the result must be JSON-serializable, but says nothing about what can be mutated, error/timing behavior, or the consequences of running arbitrary code.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and scope, zero filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description reasonably covers the return ('the JSON-serializable result is returned'), and annotation coverage handles the safety profile. However, for a tool that executes arbitrary code, the lack of any note on error handling, async/promise handling, or scope limits leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: tabId is documented in the schema (default active tab) while expression has no schema description. The description compensates partially by stating an expression must be passed and that its result must be JSON-serializable, but adds no format or syntax detail beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource — run JavaScript in a tab's page context — and notes the mechanism (via the debugger). It is clearly distinguishable from retrieval siblings like chrome_get_text or chrome_console, though it does not explicitly contrast itself with browser_evaluate or chrome_console.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no naming of alternatives such as chrome_console or browser_evaluate. The only hint is the procedural 'Pass an expression', which is invocation mechanics rather than selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_extension_reloadA
Destructive

Reload the Free Coding Agent Chrome extension from disk (after it was updated) and report the new version.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is provided structurally. The description adds what happens operationally (reload from disk, returns new version), but says nothing about the disruption implied by destructiveHint — e.g., whether the extension restarts, loses state, or drops connection — which is the key behavioral risk here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and packs the trigger condition and return value without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema and full annotation coverage, the description covers action, trigger, and return value. The only gap is the nature of the destructive side effect, which annotations flag but do not explain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (reload), a specific resource (the Free Coding Agent Chrome extension), the trigger condition (after it was updated), and the outcome (report the new version). This distinguishes it clearly from chrome_status, which presumably just reports state without reloading from disk.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical 'after it was updated' gives a clear precondition for when to call this. However, no alternative is named — e.g., chrome_status is not referenced as the way to check version without reloading, so the agent must infer that.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_get_textB
Read-onlyIdempotent

Visible text of a tab (document.body.innerText, truncated).

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdNoChrome tab id (default: active tab)
maxCharsNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description's one added behavioral fact is that output is truncated, which is real value, but it omits what truncation does to the result (cut-off point, indicator) and whether auth/session state matters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded parenthetical fragment with zero filler. It is efficient, though the brevity borders on under-specification rather than deliberate concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should say what the return value looks like (plain string? object with metadata?), and it only says the text is truncated. Annotations do carry the behavior/safety load, so the tool is callable, but key return-shape information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: tabId is documented in the schema but maxChars has just a default (40000) with no description, and the description does not compensate for either parameter. The mention of 'truncated' loosely hints at maxChars but gives no units, limits, or behavior when exceeded.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names the resource and the exact extraction method ('Visible text of a tab (document.body.innerText, truncated)'), so an agent knows this returns rendered visible text rather than HTML or a DOM dump. It does not, however, distinguish itself from close siblings like chrome_snapshot, chrome_evaluate, or browser_get_text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus chrome_snapshot (structure), chrome_screenshot (pixels), or chrome_evaluate (arbitrary JS). The agent must infer the choice from the terse purpose line alone; no prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_list_tabsA
Read-onlyIdempotent

List all tabs in the user's real Chrome.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds only the "real Chrome" scope, which is mildly useful context consistent with the open-world annotation, but it does not disclose auth needs, return shape, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no waste. Every word contributes to stating what is listed and where.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter read operation with rich annotations, the description is nearly complete enough to call the tool correctly. The main omission is any indication of what tab information is returned, though that is a minor gap for a list-tabs tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to explain. Per the rubric, a zero-parameter tool has a baseline of 4, and the description does not need to compensate for undocumented inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ("List") and resource ("tabs") and scopes it to the user's real Chrome, making the purpose immediately clear. It does not explicitly differentiate itself from the similarly named sibling browser_tabs, which prevents a top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as browser_tabs or chrome_switch_tab. The description also omits prerequisites or context for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_navigateB
Destructive

Navigate a tab to an http(s) URL and wait for it to load.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
tabIdNoChrome tab id (default: active tab)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already disclose that this is a non-readonly, open-world, non-idempotent, destructive operation. The description adds that the call waits for the page to load, which is useful beyond the annotations, but it does not explain what navigation may disrupt, error behavior, or timeout handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It immediately states the action, target, and loading behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple navigation tool with annotation coverage and no output schema, the description provides the essential action and wait condition. It is mostly complete, though it could better distinguish itself from browser_navigate and mention any timeout or failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and while tabId is documented in the schema, the url parameter has no schema description. The description partially compensates by constraining url to http(s) URLs, but it does not otherwise clarify expected format, redirect behavior, or tabId defaulting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Navigate') and resource ('a tab to an http(s) URL'), and adds the loading condition. It does not distinguish this tool from the similarly named browser_navigate sibling, which is a noticeable gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no indication of when to use chrome_navigate instead of alternatives such as browser_navigate, chrome_switch_tab, chrome_new_tab, or chrome_close_tab. Usage is only implied by the verb 'Navigate'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_new_tabC
Destructive

Open a new tab in the real Chrome.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoabout:blank
activeNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose the safety profile (readOnly=false, destructive=true, openWorld=true, idempotent=false). The description adds only 'real Chrome' as context and does not explain what gets destroyed, how 'active' affects focus, or any side effects beyond what the annotations state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. For a simple open-a-tab action, that is appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, two undocumented parameters, and annotations indicating destructive and open-world behavior. The description does not cover parameter meanings, side effects, or return behavior, so it is not complete enough for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention either parameter. It does not explain the 'url' default of about:blank or the 'active' default of true, leaving both parameters undocumented. With two undocumented parameters, the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'Open a new tab' in 'real Chrome.' It clearly states what the tool does, but does not differentiate from siblings like chrome_navigate, chrome_switch_tab, or browser_tabs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives. An agent must infer that this is for creating a new tab rather than navigating or switching, but the description gives no explicit routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_pairing_setupA
Destructive

Securely configure or rotate the Chrome bridge pairing token in both the connected extension and .env. The secret is never returned. Restart the MCP worker afterwards to enforce it.

ParametersJSON Schema
NameRequiredDescriptionDefault
rotateNoRotate even if token authentication is already configured

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation destructive, non-idempotent, and open-world. The description adds valuable context beyond that: the secret is never returned, both the extension and .env are updated, and a worker restart is required to enforce the change. It does not cover permissions or failure modes, so it is not fully exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action, followed by the security guarantee and the enforcement step. Every sentence carries practical information for invoking the tool correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter setup tool with no output schema, the definition covers the action, dual update target, non-return of the secret, and required restart. It could state prerequisites such as an active connected extension, but annotations already cover the safety profile and the description is otherwise sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single rotate parameter is already documented in the schema as rotating even if token authentication is configured. The description only implies configure-vs-rotate behavior and adds no syntax or additional parameter meaning beyond the structured schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (configure/rotate), resource (Chrome bridge pairing token), and dual target (connected extension and .env). It is clearly distinct from sibling tools like chrome_status or chrome_extension_reload, which handle page or extension state rather than pairing credentials.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says to restart the MCP worker afterward and the schema notes rotate behavior, so some usage context is implied. However, it does not explicitly say when to use this instead of chrome_status, chrome_extension_reload, or other chrome tools, and it names no when-not conditions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_press_keyA
Destructive

Press a key or chord (e.g. "Enter", "Tab", "Escape", "Control+a") as a trusted key event in the tab.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
refNoOptional element to focus first
tabIdNoChrome tab id (default: active tab)
frameIdNoFrame id, default 0
snapshotIdNosnapshotId returned by chrome_snapshot; rejects stale actions

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered. The phrase 'as a trusted key event' is genuine added context (events are dispatched as trusted rather than synthetic), but it is not elaborated on and no focus/consequence behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that gives the action and the input format with zero filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter action tool with annotations present and no output schema, the definition plus schema descriptions cover what an agent needs to invoke it (target, focus, staleness guard). Only minor behavioral detail is missing, which is acceptable given schema richness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80% and the required 'key' parameter has no schema description, so the examples ('Control+a' chord syntax) materially compensate by showing the accepted input format. The remaining focus/tab/frame semantics are already documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb plus resource ('Press a key or chord') and grounds it with concrete examples ('Enter', 'Tab', 'Escape', 'Control+a') that clearly distinguish it from text-typing siblings like chrome_type. Strong and specific, though it never explicitly names or routes away from a sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus chrome_type or browser_click, no prerequisites, and no exclusions. Usage is only implied by the examples, leaving an agent to infer the right context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_screenshotA
Read-onlyIdempotent

Screenshot of a tab in the real Chrome (visible viewport). The image is returned so you can see it.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdNoChrome tab id (default: active tab)

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already cover safety traits (readOnly, idempotent, non-destructive). The description adds valuable context beyond annotations by disclosing that the output is an image and that the capture is limited to the visible viewport rather than the full page.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two sentences and is front-loaded with the core action and scope. The second sentence is helpful for output expectations but slightly redundant, keeping it just short of perfectly economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple screenshot tool with rich annotations and a fully documented schema, the description covers purpose, viewport scope, and return type adequately. It does not address relative usage versus other screenshot tools, but that is a minor gap given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single optional parameter tabId is fully documented in the schema. The description adds no further meaning about the parameter, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action and resource: taking a screenshot of a tab in the real Chrome, scoped to the visible viewport. It implicitly distinguishes itself from browser_screenshot by specifying 'real Chrome', but does not explicitly name or differentiate against sibling tools such as browser_screenshot or chrome_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives like browser_screenshot, windows_screenshot, or chrome_snapshot. The description only states what the tool does, leaving usage context and exclusions to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_scrollB
Destructive

Scroll the inspected tab by y pixels (negative scrolls up).

ParametersJSON Schema
NameRequiredDescriptionDefault
yNo
tabIdNoChrome tab id (default: active tab)
frameIdNoFrame id, default 0
snapshotIdNosnapshotId returned by chrome_snapshot; rejects stale actions

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety and mutability profile is known without the description. The description adds the meaning of negative y values, which is useful behavioral detail, but does not explain what makes scrolling destructive, whether a snapshot is required, or how stale actions are handled beyond what the schema already says.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that contains no redundant or filler language. It is sized appropriately for a simple scroll action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple four-parameter scroll tool with no output schema, the description is minimally adequate: it covers the core action and y semantics. It still omits usage routing against browser_scroll and does not discuss the snapshotId stale-action behavior that the schema mentions, which would be helpful context for safe invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, and the undocumented parameter is y. The description compensates by defining y as pixel distance and explaining that negative values scroll upward, which is exactly the missing semantics. The other three parameters are covered by their schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Scroll) and resource (the inspected tab), and specifies the unit and direction semantics ('by y pixels', 'negative scrolls up'). However, it does not distinguish this Chrome-specific tool from the sibling browser_scroll, leaving an agent to infer the platform split.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as browser_scroll or when scrolling may be unnecessary. The direction note implies basic operation but not selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_snapshotA
Read-onlyIdempotent

Semantic snapshot of a tab: interactive elements with [#N] refs, plus page-declared site tools. Take one before clicking or typing.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdNoChrome tab id (default: active tab)
frameIdNoFrame id, default 0
snapshotIdNosnapshotId returned by chrome_snapshot; rejects stale actions
maxElementsNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it as read-only, idempotent, non-destructive, and open-world. The description adds valuable behavioral context about the snapshot's content format and the required ordering before interaction actions, though it does not mention rate limits or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with zero waste. The output format is front-loaded, followed immediately by the action-oriented usage instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully summarizes the return shape (interactive elements with refs, site tools) and pairs it with the required workflow. It omits details about maxElements, frameId behavior, and what 'site tools' entail, but is sufficient for correct tool selection and basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so most parameters are documented in the schema. The description adds no parameter-specific semantics beyond the general output format, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('snapshot of a tab') and specifies the output content ('interactive elements with [#N] refs, plus page-declared site tools'). It distinguishes a semantic snapshot from a visual screenshot implicitly, but does not explicitly contrast it with siblings like chrome_screenshot or browser_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: 'Take one before clicking or typing.' This tells the agent when to call it, but does not state when to avoid it or name an alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_statusB
Read-onlyIdempotent

Connection status of the Chrome extension bridge (the user's real, logged-in Chrome).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered structurally. The description adds the useful clarification that the status refers to the real logged-in Chrome bridge rather than a headless browser, but says nothing about what the reported status looks like or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the essential identifying information comes first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument status probe with no output schema, the description leaves the return values (e.g., connected/disconnected, extension version) entirely unspecified, which is the one remaining gap an agent would care about.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters and 100% schema coverage, so there is nothing for the description to disambiguate. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a specific resource (the Chrome extension bridge connection) and clarifies that the bridge targets the user's real, logged-in Chrome, so an agent can tell it apart from browser_* or chrome_* action tools. It stops short of explicitly contrasting itself with the neighboring chrome_pairing_setup / chrome_extension_reload tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to call this versus alternatives, whether it must be checked before other chrome_* calls, or what a negative result implies. The read-only diagnostic nature is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_switch_tabB
Destructive

Activate a tab and focus its window.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesChrome tab id (default: active tab)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description usefully adds the window-focus side effect (behavior beyond the annotations), but says nothing about what happens with an invalid or stale tabId.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with the action and its window side effect. Zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, no-output-schema tool whose annotations already carry the safety profile, the description covers the essential behavior. Minor gap: no indication of error behavior for an invalid tabId or whether focus is mandatory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single tabId parameter is documented there (including the 'default: active tab' note). The description adds no format, range, or sourcing detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Activate a tab and focus its window' clearly distinguishes it from siblings like chrome_list_tabs, chrome_new_tab, and chrome_close_tab. It stops short of naming those siblings explicitly, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to switch tabs versus creating (chrome_new_tab) or listing (chrome_list_tabs) them, and no note on prerequisites such as the tab already existing or being reachable. The agent must infer the usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chrome_typeA
Destructive

Focus an element by ref and type text with trusted input (replaces the current value unless append: true). submit: true presses Enter afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYes
textYes
tabIdNoChrome tab id (default: active tab)
appendNo
submitNo
frameIdNoFrame id, default 0
snapshotIdNosnapshotId returned by chrome_snapshot; rejects stale actions

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds useful behavioral context: it replaces the current value unless append is true, and submit: true presses Enter afterward. It does not mention auth, rate limits, or failure behavior, but this is solid additional detail beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. Core action comes first, followed by the important append and submit modifiers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the destructive annotation, no output schema, and a mutation-oriented tool, the description covers the essential behavioral modifiers. It leaves out explicit ref provenance from chrome_snapshot, but the schema's snapshotId field hints at that dependency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 43% schema description coverage, the description compensates by explaining append and submit semantics, plus the core ref and text roles. However, it does not fully document all seven parameters, and some context such as where ref originates from is left implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: focus an element by ref and type text. It distinguishes the action itself, but does not explicitly differentiate it from sibling tools like browser_type or chrome_press_key.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the tool is used to type text into a referenced element, but gives no explicit guidance on when to choose it over browser_type, chrome_click, chrome_press_key, or other alternatives. No when-not-to-use conditions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

copy_pathC
Destructive

Copy a file or directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
overwriteNo
destinationYes
allowPartialUndoNoProceed even if the change is too large to snapshot completely (undo would be partial)

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered. The description adds nothing beyond the name: it doesn't say whether directory copies are recursive, whether existing destinations are overwritten, or what happens on partial failure (the allowPartialUndo param hints at snapshot limits). With annotations doing the disclosure work, this minimal description is weak but not contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero waste and the verb front-loaded, which is structurally fine. But the brevity comes at the cost of substance rather than being earned conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with two undocumented required parameters (source, destination) and no output schema, the definition is far too thin. Annotations cover the safety signal, but the description leaves overwrite semantics, recursive directory behavior, and the partial-undo limitation entirely unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%: source, destination, and overwrite are undocumented in the schema, and the one-sentence description adds no meaning for any of the four parameters. It fails to compensate for the coverage gap, e.g. it never mentions that copying can overwrite an existing destination.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Copy a file or directory'), so the operation is unambiguous. However, it offers no differentiation from the closely related sibling move_path or delete_path, which an agent selecting among filesystem tools would benefit from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is provided. The purpose implies copying rather than moving or deleting, but nothing routes the agent between copy_path, move_path, and delete_path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_pathA
Destructive

Delete a file, or a directory when recursive is true. Undoable with checkpoint_rewind.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath
recursiveNo
allowPartialUndoNoProceed even if the change is too large to snapshot completely (undo would be partial)

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and non-idempotent, so the safety profile is covered. The description adds genuinely new behavioral context beyond that: deletions are recoverable via checkpoint_rewind, and the recursive flag governs whether directories are affected at all. It stops short of describing failure modes or snapshot limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the primary action front-loaded and the recovery mechanism stated second. Nothing is wasted and no clause is self-evident filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers undo and recursive scope, which is the essentials. It omits edge cases an agent needs: behavior on non-empty directories, permission requirements, and whether path must exist or can be relative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% and the description adds real meaning for 'recursive' (controls directory deletion) and echoes partial-undo semantics for allowPartialUndo. But 'path' remains an uninformative 'Path' string in the schema with no format guidance (relative vs absolute), so the description does not fully compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete a file') and explicitly handles the directory case via the recursive flag, so the scope is unambiguous. There is no sibling delete tool to differentiate from, so it earns a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives the branch condition for recursive (directories) and points at checkpoint_rewind for recovery, which is useful implicit guidance. However, it never says when to prefer this tool over alternatives such as moving to trash, nor what happens when a non-empty directory is deleted without recursive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_fileA
Destructive

Replace exact text in one file. Either oldText/newText, or edits: [{oldText,newText,replaceAll}] applied atomically in order. oldText must match exactly once unless replaceAll. Reports diagnostics and a checkpoint id.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path
checkNoRun compilers/linters on the changed files afterwards (default true)
editsNo
newTextNo
oldTextNo
replaceAllNo
expectedSha256NoOptional optimistic precondition from file_info; refuse the edit if the file changed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: unique-match semantics for oldText, replaceAll override, atomic in-order application of edits, and reporting of diagnostics plus a checkpoint id. Annotations cover destructive/idempotent flags; the description enriches with matching and atomicity rules the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the core action, then input modes, then matching/atomicity rules, then return info. Every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description gives matching rules, atomicity, diagnostics, and checkpoint id – the essentials an agent needs. It omits how to recover via checkpoint_rewind and leaves expectedSha256 unexplained, so it stops short of fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43%, and the description compensates by explaining the exact-match requirement (once unless replaceAll), the dual mode (oldText/newText or edits array), and atomic sequencing. It leaves path, check, and expectedSha256 without added meaning, but covers the highest-risk parameters well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Replace) and resource (exact text in one file), and distinguishes itself from write_file and apply_patch by restricting to exact-text replacement. An agent can tell this apart from its sibling editing tools immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two mutually exclusive input modes (oldText/newText vs edits array), which is usage-relevant, but never states when to prefer this over write_file or apply_patch. No explicit when-to-use reasoning is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_urlA
Read-onlyIdempotent

Fetch a URL (follows redirects) and return readable text. raw: true returns the unprocessed body (useful for JSON/APIs).

ParametersJSON Schema
NameRequiredDescriptionDefault
rawNo
urlYes
maxCharsNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavior: redirects are followed automatically and output is converted to readable text versus a raw body. It omits truncation behavior from maxChars, error/timeout handling, and auth requirements, which matters for open-world fetching.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler; the core behavior is front-loaded and the mode flag is explained immediately after. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description is the only source for return-value expectations; it hints at 'readable text' vs raw body but says nothing about truncation limits or failure modes. Adequate for a simple fetch, but leaves an important gap around maxChars given zero schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the semantic load. It explains 'raw' well, but leaves maxChars (default 40000) undocumented — an agent cannot tell whether exceeding it truncates, errors, or paginates — and says nothing about url beyond the obvious.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch) and resource (URL), plus a behavioral detail (follows redirects) and the output shape (readable text). It is clear what the tool does, though it does not explicitly differentiate itself from siblings like browser_navigate or browser_get_text that also retrieve page content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'raw: true' note gives a concrete use case (JSON/APIs) for one mode, which is real usage guidance. However, there is no guidance on when to prefer this over browser_get_text, browser_navigate, or web_search in a large sibling set, leaving the primary routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_infoB
Read-onlyIdempotent

Check whether a path exists and get its type, size and modification time.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is fully covered by structured data. The description usefully discloses the concrete return fields (type, size, modification time), which is the main value it adds and which matters since there is no output schema. It does not state behavior on missing paths (error vs. false) or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that conveys purpose and return values with zero waste. Size is well matched to a simple one-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only probe with full annotation coverage and no output schema, the description adequately conveys both the operation and the returned fields. It only falls short on edge-case behavior for non-existent paths and routing away from check_files.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Single parameter with 100% schema coverage, so the schema already documents 'path' (albeit minimally as 'Path'). The description adds no format or meaning beyond what the schema provides, making baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (check path existence) plus the returned attributes (type, size, mtime), so the purpose is unambiguous. It does not distinguish itself from the overlapping sibling check_files, leaving some ambiguity about which to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, and no mention of the natural alternative check_files or the broader list_directory. The agent must infer usage context on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_filesB
Read-onlyIdempotent

Find files by glob pattern, e.g. "**/*.test.js" or "package.json".

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRoot directory
patternYes
maxResultsNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the glob semantics that make the operation predictable, but it omits the maxResults default cap (500) that bounds results, which is relevant behavior for a finder.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the purpose front-loaded and examples doing useful work. No filler, though the brevity contributes to the coverage gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only file finder with annotations covering safety and no output schema, this is nearly complete, but gaps remain: the maxResults default/cap, the meaning of path, and any routing versus list_directory or search_code.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate, and it partly does by illustrating valid pattern syntax ('**/*.test.js'). However it says nothing about path (root directory) or the maxResults cap, leaving two of three parameters to the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find files') and adds glob-pattern examples that clarify the mechanism. It does not explicitly distinguish itself from siblings like list_directory or search_code, but the glob framing gives it a recognizable identity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance and names no alternatives. It never says when to prefer globbing here versus search_code or list_directory, so the agent must infer selection from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_symbolB
Read-onlyIdempotent

Find where a function/class/type/method is defined (name or Container.name) and where it is used (whole-word references).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoFilter: function, method, class, type, interface, struct, …
nameYes
pathNoProject directory (default: default workspace)
exactNofalse = substring match on names
referencesNo
maxReferencesNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive, so safety is covered. The description adds that references are matched as whole words, which is useful behavioral detail, but it says nothing about reference truncation despite the maxReferences default of 80.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the primary action (finding definitions) before the secondary one (finding usages). No waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with no output schema, the description omits what results look like and how maxReferences truncates them. It covers the core intent but leaves the agent guessing about result shape and truncation behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%, so half the parameters (name, references, maxReferences) are undocumented in the schema. The description partly compensates by clarifying the name format and whole-word reference semantics, but leaves the references toggle and the 80-reference cap unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (find) and resource (symbol definitions and usages), and names the accepted input formats (name or Container.name). It is distinguishable from generic siblings like search_code, though it never distinguishes itself from the very similar lsp_definition/lsp_references/vscode_find_references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus the closely related lsp_definition, lsp_references, vscode_find_references, or search_code siblings. Usage is only implied by the description of what it searches.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_system_infoA
Read-onlyIdempotent

Host info, configured workspaces, default workspace and available integrations. Call this first in a new conversation.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds the 'call first' behavioral cue, which is useful context. It doesn't disclose return format or any other behavioral traits, but with annotations doing the heavy lifting, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences: the first lists the returned information, the second gives the usage directive. It is front-loaded with purpose and wastes no words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only discovery tool with no output schema, the description covers what the tool returns and when to call it. It is nearly complete, though it could briefly mention if the result is cached or if it reflects live state, but that's minor given the simplicity of the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameters for the description to explain. The description's enumeration of returned data gives some sense of scope, but parameter semantics is not applicable. Baseline 4 is correct for zero parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the resource: host info, workspaces, default workspace, and available integrations. It's not a tautology like the name 'get_system_info', and it enumerates what is returned. However, it doesn't distinguish this from sibling tools like list_workspaces or capability_list, which may overlap in scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: 'Call this first in a new conversation.' This is a clear, actionable directive that tells the agent exactly when to invoke the tool, with no ambiguity or exclusions needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_commitB
Destructive

Stage files (all changes by default) and create a commit.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRepository directory (default: default workspace)
filesNoSpecific files to stage; default all
messageYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the agent knows this mutates repository state. The description usefully adds the default staging behavior ('all changes by default'), but says nothing about failure modes, empty-commit handling, or what happens to unstaged/untracked files.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence with the staging behavior front-loaded before the commit action. Nothing wasted, though it is arguably too terse given the mutation's risk profile.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter destructive mutation with no output schema and no annotations covering commit semantics, the description covers the core operation but omits commit-failure behavior, partial-staging implications, and the interaction with the required message parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: path and files are documented in the schema, message is not (though it is self-evident and required). The description adds the meaningful default for the files parameter ('all changes by default'), which reinforces rather than merely repeats the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific compound verb+resource: stage files and create a commit, with the default staging scope ('all changes by default') made explicit. It is clearly distinguishable from read-oriented siblings like git_status, git_diff, and git_log, though it never names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or prerequisite guidance. It doesn't say a commit requires staged changes, that a message is mandatory, or how it relates to git_diff/git_status as the recommended pre-step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_diffA
Read-onlyIdempotent

Full unified diff of the working tree (or staged changes, or against a ref), optionally for one file.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNo
fileNo
pathNoRepository directory (default: default workspace)
stagedNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that the output is a full unified diff, but does not disclose output size, error behavior, or pagination details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the core purpose and parenthetically listing the supported modes. Every phrase carries information and no words are wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still communicates that the return value is a full unified diff, and annotations cover the read-only safety profile. The low schema coverage and lack of ref/file syntax details leave minor gaps, but the definition is nearly complete for a simple read-only git inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, so the description must compensate for three undocumented parameters. It maps to 'staged changes' (staged), 'against a ref' (ref), and 'optionally for one file' (file), adding meaningful semantics, but does not explain ref syntax or how staged and ref interact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Full unified diff of the working tree (or staged changes, or against a ref), optionally for one file.' It clearly distinguishes the tool from siblings like git_status and git_log by naming the exact output type and scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates modes of use: working tree, staged changes, against a ref, optionally one file. However, it does not explicitly compare this tool to alternatives such as git_status or git_log, nor does it state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_logC
Read-onlyIdempotent

Recent commit history.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRepository directory (default: default workspace)
limitNo

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no behavioral context beyond that — no ordering (newest-first?), no pagination semantics for the limit, no indication of what a truncated history means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single four-word fragment with zero filler, so it is maximally brief, but brevity here comes from under-specification rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool this could be adequate, but with no output schema and one undocumented parameter, the description leaves the return shape, ordering, and limit behavior entirely unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 50% schema description coverage, the description is expected to compensate, but it mentions neither parameter. The 'path' parameter is documented in the schema, while 'limit' has no description anywhere, so an agent cannot tell what it caps or how it interacts with the default of 15.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a bare noun phrase, 'Recent commit history,' that essentially restates the tool name git_log. It gives no verb, no scope (which repo, which branch, whole history vs. HEAD), and does nothing to distinguish it from siblings like git_status, git_diff, or git_commit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus git_diff, git_status, or git_commit. The agent must infer that this returns a listing of past commits rather than diffs or working-tree state, with no stated conditions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_statusB
Read-onlyIdempotent

git status (short) plus diff stat against HEAD.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRepository directory (default: default workspace)

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered by structured data. The description adds useful behavioral content in that it reveals output composition (short status format plus a diffstat against HEAD), but says nothing about untracked files, staging areas, or what happens in a non-repo directory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero waste and the key output qualifiers ('short', 'against HEAD') front-loaded. It is arguably under-specified rather than over-long, but it is efficiently structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must convey what is returned; it does partially by naming status-short plus diffstat. However, it omits whether untracked/ignored files appear, staged vs unstaged separation, and failure behavior outside a repository, leaving meaningful gaps for a tool an agent will use to decide next steps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one optional parameter and its schema description ('Repository directory (default: default workspace)') already carries 100% coverage of its meaning. The description adds no syntax or default-behavior detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific git operation and its scope ('status (short) plus diff stat against HEAD'), so an agent knows it reports working-tree state rather than commits or diffs alone. It does not explicitly distinguish itself from siblings like git_diff or git_log, which keeps it out of 5 territory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the obvious alternative git_diff even though the description itself includes diff-stat output. The agent must infer selection context on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_cancelB
Destructive

Cancel a running job (kills its process tree).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true. The description adds useful behavioral context by specifying that cancellation kills the process tree. It does not cover authorization needs, return format, or whether the operation waits for termination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. The key action and its destructive effect are communicated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple destructive tool with rich annotations and no output schema, the description states the effect but leaves the required 'id' parameter unexplained and gives no usage guidance versus alternatives. It is minimally adequate but has clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter 'id' is undocumented in both schema and description. The description does not explain the format, source, or meaning of the job identifier beyond the obvious implication from 'job'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Cancel a running job'. The parenthetical clarifies the scope by noting it kills the process tree. It is clear but does not explicitly differentiate from sibling tools like kill_process or process_stop, which could cause confusion for an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use job_cancel versus alternatives such as kill_process, process_stop, or job_status. Usage is only implied by the tool name and the phrase 'running job'. No exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_listB
Read-onlyIdempotent

Recent background jobs with status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds two small facts beyond that: results are time-scoped ('recent') and include status. It says nothing about ordering, result size, or how the job set is bounded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded fragment with no filler is efficient, though the brevity shades into under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param list tool with no output schema, the description should at least hint at what is returned (fields, ordering, how 'recent' is bounded) and how it differs from job_status. One of those gaps is closed ('with status'), the rest are not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document. Schema coverage is trivially 100% and the baseline for an argument-less tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Recent background jobs with status' names the resource (background jobs) and a scope (recent) but uses no verb, so it reads as a noun caption rather than a statement of what the tool does. It is distinguishable from job_output/job_cancel only by inference, and the relationship to job_status (whose status? how many jobs?) is left unclear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of the obvious alternatives job_status, job_output, or process_list. The agent must guess whether this lists many jobs or reports one job's status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_outputB
Read-onlyIdempotent

Read a job log: tail (default 200 lines), a line range (fromLine/lines), or grep (regex with context lines).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
grepNo
tailNo
linesNo
contextNo
fromLineNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds the default tail size (200 lines), which is useful behavioral context. It does not disclose pagination behavior or log growth considerations, so a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and packs three modes efficiently. It is not overly verbose, though the parenthetical grammar slightly obscures the mode-to-parameter mapping.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six undocumented parameters and no output schema, the description covers the main read modes but leaves id, context, and mode-interaction semantics unexplained. Adequate but with clear gaps for an agent invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema documents none of the six parameters. The description partially compensates by explaining tail default, fromLine/lines range semantics, and grep with context lines, but the meaning of id, the interaction between modes, and the context parameter remain undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: 'Read a job log.' It then enumerates the three access modes (tail, line range, grep), which lets an agent distinguish it from siblings like job_status or process_output. It does not explicitly name those siblings, keeping it just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The three modes imply when each is useful (default tail, explicit line range, regex grep), but there is no explicit when-to-use guidance, no mention of alternatives such as process_output or read_file, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_startA
Destructive

Run a long command (test suite, build, install, benchmark) as a background job. Waits up to waitMs (default 25s): if it finishes you get a smart summary (pass/fail counts, failing tests with context, output tail) instead of the raw log; otherwise poll job_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
titleNo
waitMsNo
commandYes
timeoutMsNoKill after this long (default 1h)
resourceClassNoAdmission class; defaults to automatic command classification
resourceWaitMsNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavior: the wait window, the smart-summary return shape, and the polling fallback when the job outlives the wait. It doesn't disclose that the job keeps running after waitMs elapses, which is implied but not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences: the capability first, then the wait/poll contract and what you get back. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Return behavior is well explained despite no output schema, and the wait/poll loop is clear. But for a 7-parameter destructive mutation tool, five parameters are undocumented in both the description and the schema, leaving an agent guessing about timeout, resource class, and working directory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29%, so the description needs to compensate, but it only restates waitMs and its default (already in the schema). cwd, title, timeoutMs, resourceClass, and resourceWaitMs receive no explanation, and the 50s cap on waitMs is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (run a long command as a background job) and enumerates concrete use cases (test suite, build, install, benchmark). It doesn't explicitly contrast with the similarly named process_start/run_command/pty_start siblings, so an agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: use for long-running commands, then wait up to waitMs, otherwise poll job_status. The fallback path to a named sibling is explicit. It lacks guidance on when to prefer run_command or process_start over a background job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

job_statusA
Read-onlyIdempotent

Status of a job; waitMs (≤50s) waits for it to finish. Finished jobs return the failure summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
waitMsNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and open-world behavior. The description adds useful context beyond that: the wait is bounded to ≤50s, and finished jobs return a failure summary. This extra behavioral detail is helpful, though it omits other traits like retry behavior or status value semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence with a semicolon; purpose is front-loaded, followed by the wait behavior and the return note. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a simple two-parameter status tool with rich annotations and no output schema, the description covers what it does, the optional wait limit, and the failure summary for finished jobs. The only notable gap is the lack of any explanation for the required id parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains waitMs clearly (waits for finish, ≤50s), but the required id parameter is left completely undocumented, leaving the agent to infer that it refers to a job identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource: retrieving the status of a job. The scope is specific enough to differentiate from siblings like job_output and job_cancel, but the description never explicitly names or contrasts with those alternatives, so it misses the top mark.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage by describing the optional waitMs parameter as a way to wait for completion, but gives no explicit when-to-use guidance or alternatives (e.g., use job_output for results, job_list for all jobs). The agent must infer the context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kill_processB
Destructive

Kill an OS process tree by PID.

ParametersJSON Schema
NameRequiredDescriptionDefault
pidYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds useful scope beyond annotations by specifying that it kills the process tree, not just a single process, but does not mention permissions, forcefulness, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single, front-loaded sentence with no wasted words. The core action and targeting mechanism are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter mutation tool with annotations covering the destructive nature and no output schema, the description is minimally adequate. It omits routing guidance against similar stop/cancel siblings and does not describe behavior when the PID is invalid or already terminated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single pid parameter. The description partially compensates by identifying the parameter as a PID, but adds no format, validation, or lookup details beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb, resource, and scope: kill an OS process tree by PID. It is clear what the tool does, but it does not explicitly distinguish itself from nearby siblings like process_stop or pty_stop.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is provided. The agent must infer that this is for terminating an entire process tree rather than using process_stop, pty_stop, or job_cancel.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_directoryB
Read-onlyIdempotent

Tree listing of a directory (skips node_modules, .git, build outputs unless includeIgnored).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoDirectory (default: default workspace)
depthNo
includeIgnoredNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false. The description adds useful context beyond that: it skips node_modules, .git, and build outputs by default unless includeIgnored is set. It still does not describe depth behavior or return format, so it remains adequate rather than rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. The parenthetical adds the key filtering behavior without disrupting readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only directory listing tool whose annotations already cover safety and idempotence, the description is mostly sufficient. However, because there is no output schema and depth is undocumented, an agent still lacks complete context on recursion behavior and result format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate for undocumented parameters. It explains includeIgnored implicitly through the ignore-skipping behavior, but depth is not mentioned at all, and its default and maximum of 10 are left for the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specifies the verb 'Tree listing' and the resource 'directory', so the core action is clear. It does not distinguish itself from nearby filesystem siblings such as find_files or search_code, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states what the tool does but gives no when-to-use guidance, no alternatives, and no exclusions beyond the default ignore behavior. Usage is only weakly implied by the purpose itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_processesA
Read-onlyIdempotent

List OS processes (optionally filtered by name substring).

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds only the filter's substring behavior and does not disclose return format, ordering, or scope beyond OS processes, making it adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero wasted words. The core purpose and the optional filter behavior are communicated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool with one optional parameter, full annotation coverage, and no output schema, the description is nearly complete. A minor gap remains around distinguishing it from process_list and describing ordering or scope, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden for the single optional parameter. It successfully defines 'filter' as a name substring match, which is meaningful semantic detail not present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'List OS processes'. It is clear what the tool does, but it does not explicitly distinguish itself from the sibling tool process_list, leaving some ambiguity for the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that the optional filter is used for a name substring, which implies usage. However, it does not say when to use this tool versus alternatives such as process_list or get_system_info, so guidance beyond parameter usage is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_workspacesA
Read-onlyIdempotent

List configured project workspaces with git branch and number of changed files.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint, idempotentHint and non-destructive behavior, so the safety profile is covered. The description earns credit by disclosing what the listing returns (git branch + changed-file count), which goes beyond the structured fields and is valuable given there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The resource comes before the payload detail, so the agent can stop reading as soon as it has the gist.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only lister with full annotation coverage, the description supplies the essential missing piece: what the result contains. Only the absence of any route to a sibling/alternative tool keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4; there are no parameter semantics to document or omit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb ('List') plus a specific resource ('configured project workspaces') and even the payload ('git branch and number of changed files'). It is distinguishable from the browser/pty/git siblings, though it does not name a specific alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by naming the resource, but gives no explicit when-to-use or when-not-to-use guidance relative to nearby tools such as git_status, project_context, or list_directory. Nothing tells the agent why it would pick this over those.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_callsB
Read-onlyIdempotent

Call hierarchy of the function at file:line: incoming (who calls it) or outgoing (what it calls).

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesSource file
lineYes1-based line
symbolNoIdentifier on that line; its column is found automatically
characterNo1-based column (or give symbol instead)
directionNoincoming

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, openWorld=false, and destructive=false, so the safety profile is covered. The description adds semantic meaning for incoming/outgoing, but does not mention language-server prerequisites, failure modes, or return format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler, and the core operation and direction options are presented efficiently. Every part of the sentence contributes to understanding the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does not explain what the returned call hierarchy looks like or how results are structured. It also omits when-to-use guidance relative to sibling LSP tools, leaving some gaps despite clear annotations and mostly descriptive parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so parameters are largely documented. The description adds useful meaning beyond the enum by defining incoming as 'who calls it' and outgoing as 'what it calls', which clarifies the key choice the caller must make.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation (call hierarchy) on a specific resource (function at file:line) and explains the two direction modes. It is clear but does not explicitly differentiate itself from siblings such as lsp_references or lsp_definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what incoming and outgoing mean but gives no guidance on when to use this tool versus alternatives like lsp_references or lsp_definition. It describes output rather than usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_code_actionsA
Read-onlyIdempotent

List language-server code actions/quick fixes/refactors for a source range without applying them.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesSource file
lineYes1-based line
onlyNoOptional CodeActionKind filter, e.g. quickfix or source.organizeImports
limitNo
symbolNoIdentifier on that line; its column is found automatically
endLineNo
characterNo1-based column (or give symbol instead)
endCharacterNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world, so the safety profile is fully covered; 'without applying them' reinforces this but adds little. The description does not disclose return shape, result limit behavior, or what happens when no actions exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key constraint (non-applying, range-scoped, language-server) front-loaded and zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, no-output-schema tool with 63% coverage, the description is adequate but leaves gaps: it doesn't explain the default limit of 50, the endLine/endCharacter range semantics, or the shape of the returned actions list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 63% schema coverage across 8 params, the schema already documents file, line, only, symbol and character. The description only implies a 'source range' but adds no meaning for limit, endLine, endCharacter, or the CodeActionKind filter, so it neither duplicates nor compensates meaningfully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('language-server code actions/quick fixes/refactors') scoped to a source range, and the qualifier 'without applying them' sharply separates it from any apply/rename/format counterpart among the lsp_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Without applying them' implies the usage context (discovery/preview rather than mutation), but no alternative tool or explicit when-not condition is named, so the agent must infer routing from sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_definitionB
Read-onlyIdempotent

Go to definition (kind: definition | typeDefinition | implementation) of the symbol at file:line (give character or symbol name).

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesSource file
kindNo
lineYes1-based line
symbolNoIdentifier on that line; its column is found automatically
characterNo1-based column (or give symbol instead)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a read-only, idempotent, non-destructive operation, so the safety profile is covered. The description adds the positional requirement (file:line plus a symbol/character marker) but says nothing about the response (a location, possibly multiple at kind=implementation) or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence, front-loaded with the action and immediately followed by the kind options. No wasted words, though the parenthetical phrasing is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters, no output schema, and moderate complexity, the description covers the inputs adequately but omits what a caller gets back and how to disambiguate from the sibling vscode_go_to_definition. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema largely documents the parameters including the kind enum and the 1-based semantics of line/character. The description adds only the mutual-exclusion hint '(give character or symbol name)', which is useful but marginal beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Go to definition') and resource (the symbol at file:line), and enumerates the three supported kinds. It is distinguishable from siblings like lsp_references and lsp_hover, though it does not explicitly differentiate from the near-duplicate vscode_go_to_definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the 'kind' enumeration and the note to supply character or symbol name, but there is no explicit when-to-use vs when-not, and no guidance on choosing between this and vscode_go_to_definition or lsp_references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_diagnosticsA
Read-onlyIdempotent

Compiler/type-checker diagnostics from the real language server (gopls, typescript-language-server, pyright) for files or a directory — no VS Code needed. all: true also returns diagnostics the server reported for other files of the project.

ParametersJSON Schema
NameRequiredDescriptionDefault
allNo
pathNo
pathsNo
waitMsNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that diagnostics come from a live language server and that all:true widens the scope, but omits any note on the 30s waitMs latency, server-not-ready behavior, or result format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with purpose and backends, then the one parameter caveat that matters. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so return shape is unaddressed, and the ambiguous path/paths pair plus waitMs timing semantics are unexplained. Adequate for a read-only tool but short of what an agent needs to call it precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description must carry the burden. It only explains 'all' (returns diagnostics for other project files) and leaves path vs paths selection and the waitMs timeout default undocumented — the latter being important since a 30s wait implies asynchronous server startup.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (compiler/type-checker diagnostics) and names the concrete backends (gopls, typescript-language-server, pyright). The phrase 'no VS Code needed' implicitly distinguishes it from the sibling vscode_diagnostics, so an agent can route correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (real language-server diagnostics without VS Code) but gives no explicit when/when-not against the near-identical sibling vscode_diagnostics, nor guidance on file vs directory scope. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_format_previewB
Read-onlyIdempotent

Preview language-server formatting edits for a file without applying them.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYes
limitNo
tabSizeNo
insertSpacesNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description's only added behavior is that edits are shown and not applied, which reinforces but does not go beyond the read-only annotation; it says nothing about output shape or size limits on the preview.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence, front-loaded with the verb and the non-applying constraint. Nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should convey the shape of the preview, and it only hints at 'formatting edits'. Combined with four undocumented parameters and no clarity on where the preview result goes, the definition leaves real gaps for a formatting tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, and the description only implicitly addresses 'file'. The format-controlling parameters (limit, tabSize, insertSpaces) are never mentioned, so an agent gets no guidance on what they mean or which to set, and with low coverage the description is expected to compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (preview) and resource (language-server formatting edits) and bounds it with 'for a file without applying them', which cleanly distinguishes it from write_file/edit_file/apply_patch. It stops short of naming the sibling that actually applies the edits, so it is clear but not sibling-complete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without applying them' implies the intended workflow (preview before committing a format), but there is no explicit when-to-use/when-not guidance nor any pointer to the apply step. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_hoverB
Read-onlyIdempotent

Type signature and documentation of the symbol at file:line.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesSource file
lineYes1-based line
symbolNoIdentifier on that line; its column is found automatically
characterNo1-based column (or give symbol instead)

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and a closed world, so the safety profile is fully covered by structured data. The description adds some value by disclosing what is returned (type signature plus docs), but says nothing about behavior when no symbol exists at the position or about server availability errors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the location and the returned content are both stated immediately. Nothing could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the important work of naming the return payload (type signature and documentation), which is genuinely necessary. However, it omits edge-case behavior such as empty results when no symbol is under the cursor, leaving a gap for a tool an agent will call speculatively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema itself documents both the file/line anchors and the symbol-vs-character alternative mechanism. The description only echoes 'file:line' and adds no syntax, format, or precedence detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource returned (type signature and documentation) at a precise location (file:line), which is far more informative than a bare 'hover'. It does not explicitly contrast itself with close siblings like lsp_definition or lsp_references, but the returned artifact is distinct enough that an agent can differentiate them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives (lsp_definition, lsp_references, find_symbol), nor any prerequisites such as needing a running LSP server. The agent must infer usage purely from the name and the LSP convention.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_referencesA
Read-onlyIdempotent

All references to the symbol at file:line, resolved by the language server (exact, not text search).

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesSource file
lineYes1-based line
limitNo
symbolNoIdentifier on that line; its column is found automatically
characterNo1-based column (or give symbol instead)
includeDeclarationNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered structurally. The description adds genuine context that results are LSP-resolved (semantic) rather than textual, but it omits return shape, result volume, and the effect of includeDeclaration.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core scope front-loaded and the disambiguating qualifier immediately after. No filler, nothing to cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description says nothing about what is returned (list of locations, ordering, effect of limit). Combined with silent includeDeclaration semantics, the definition is adequate for identification but leaves behavioral gaps for a 6-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, above the 50% baseline, so the schema already documents file, line, symbol, and character. The description adds nothing about the two undocumented/under-documented parameters (limit's default, includeDeclaration's true-default meaning), so it fails to compensate for the remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('all references to the symbol') anchored at file:line, and adds a sharp qualifier ('resolved by the language server (exact, not text search)'). This is clear and self-contained, but it does not distinguish itself from close siblings like lsp_definition or vscode_find_references, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'exact, not text search' clause implicitly tells the agent to prefer this over a text-based search, which is a form of usage context. However, it never names an alternative tool or states when-not-to-use, and the presence of lsp_definition and vscode_find_references as siblings leaves the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_rename_previewA
Read-onlyIdempotent

Preview the exact workspace edits a semantic rename would make. Does not write files; apply reviewed changes through apply_patch/edit_file.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesSource file
lineYes1-based line
limitNo
symbolNoIdentifier on that line; its column is found automatically
newNameYes
characterNo1-based column (or give symbol instead)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so safety is covered structurally; the description's 'does not write files' reinforces but does not extend that. It adds the useful workflow fact that reviewed changes go through apply_patch/edit_file, but says nothing about truncation of the preview (relevant given the limit parameter) or how edits are ordered/grouped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the read-only constraint and the handoff to apply_patch/edit_file are both front-loaded and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description adequately signals the return value type (a set of workspace edits) and the next action. The remaining gap is the limit/truncation behavior, which matters for an agent reading a large preview, but overall the definition is usable as-is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so file/line/symbol/character are documented by the schema itself, but the description adds no meaning to any parameter. The 'limit' parameter (default 200) is undocumented in both places, leaving its truncation semantics unexplained despite being the main behavioral risk of a preview.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Preview the exact workspace edits a semantic rename would make'), which is unambiguous and clearly distinct from the browser/pty/file siblings. An agent knows it produces a dry-run of a rename, not the rename itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Does not write files' and names the follow-up tools (apply_patch/edit_file), which tells the agent both what this tool is not for and how to complete the workflow. It stops short of stating when a preview is required versus optional, so it is not fully prescriptive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_statusA
Read-onlyIdempotent

Which language servers are installed and running (language, project root, pid).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, closed-world behavior. The description adds useful disclosure beyond annotations by naming the returned fields (language, project root, pid), which is important because no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It efficiently communicates the resource and the returned fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only status tool with rich annotations and no output schema, the description is nearly complete. It covers the returned data but could state more about scope or limitations, such as whether results are project-specific.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so parameter semantics baseline is 4. The description does not need to document parameters and correctly focuses on output instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource and result: language servers that are installed and running, with returned fields language, project root, and pid. It is clear enough to distinguish from diagnostic or navigation LSP tools, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit when-to-use guidance, prerequisites, or alternatives among the many sibling tools. Usage is only inferable from the tool name and the returned information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsp_symbolsB
Read-onlyIdempotent

Symbols: of one file (file), or matching a query across the project (query + path).

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNo
pathNo
limitNo
queryNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the scope distinction (one file vs. whole project), which is real behavioral context about cost/breadth, but it says nothing about result volume, truncation, or how the limit interacts with project-wide scans.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the two modes front-loaded and no filler. It is efficient, though the terse 'Symbols:' fragment reads more like a schema note than a description, which slightly hurts first-glance readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a 4-parameter tool with 0% schema coverage and no output schema, so the description bears heavy weight. It maps three of four parameters loosely but omits `limit` entirely, leaves `path` ambiguous, and gives no guidance on output shape or choice among the numerous sibling symbol tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning — and it does clarify that `file` selects single-file mode while `query`+`path` drive project search. It still leaves `path` semantics (project root? relative? filter scope?) and `limit` behavior unstated, so the compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource (symbols) and distinguishes two retrieval modes — single file vs. project-wide query — which is more than a restatement of the name. However, it does not differentiate this tool from close siblings like find_symbol, outline, or repo_map, which appear to overlap in purpose, so an agent still has to guess which symbol tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical param pairings (file) vs (query + path) implicitly tell the agent which argument combination triggers which mode, which is useful. But there is no explicit when-to-use guidance relative to the many sibling symbol/search tools, and no statement of exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

move_pathC
Destructive

Move or rename a file or directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
overwriteNo
destinationYes
allowPartialUndoNoProceed even if the change is too large to snapshot completely (undo would be partial)

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and non-idempotency, so the safety profile is partly covered. The description adds only the rename clarification and says nothing about overwriting an existing destination, snapshot/undo limits, or permissions for a destructive filesystem operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single six-word sentence, front-loaded with the verb and free of filler. Nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive 4-parameter filesystem mutation with no output schema and only 25% parameter documentation, the description is far too thin. It omits overwrite behavior, undo semantics, and the meaning of the undocumented parameters, so an agent cannot call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (just allowPartialUndo is documented), yet the description says nothing about source, destination, or overwrite semantics. It does not compensate at all for the undocumented parameters, leaving overwrite behavior and undo implications opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear specific verb+resource: 'Move or rename a file or directory' tells the agent exactly what operation it performs, and the 'or rename' clause usefully notes that a move doubles as a rename. However, it offers no differentiation from near siblings like copy_path or delete_path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement of prerequisites, and no mention of the copy_path/delete_path alternatives that accomplish related outcomes. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

outlineA
Read-onlyIdempotent

All symbols (functions, classes, methods, types) of a file with line numbers, so you can read_file just the part you need. A directory gives a repo map of it.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is covered. The description adds valuable behavioral detail the annotations can't: it returns symbols with line numbers and switches behavior between file and directory inputs. It leaves pagination/limits unstated, but for a simple read tool that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste. The primary behavior (file symbols with line numbers) is front-loaded, followed by the directory variant and the motivating use case.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only tool with no output schema, the description conveys what is returned, how it varies by input type, and why it's useful. Only minor gaps remain, such as output size or ordering guarantees.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'path' parameter, so the description carries the burden and does so meaningfully: it clarifies that path may be a file (returns symbols) or a directory (returns a repo map). This is a real semantic clarification beyond the bare schema type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: enumerating all symbols (functions, classes, methods, types) of a file with line numbers. It also describes the directory mode. It is clear, but it doesn't explicitly differentiate itself from siblings like repo_map, find_symbol, or lsp_symbols, which overlap in purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'so you can read_file just the part you need' implies a use case and a follow-up workflow, which is helpful context. However, there is no explicit guidance on when to choose this over find_symbol, repo_map, or lsp_symbols, and no when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_inputA
Destructive

Write text to stdin of a process started with process_start (REPLs, prompts, interactive CLIs). close: true writes the data and closes stdin.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
dataYes
closeNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the agent knows this mutates a running process irreversibly. The description adds one genuinely useful behavioral detail beyond the annotations: close:true writes the data and closes stdin. It does not say what happens if the id is stale or the process has already exited.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler. The core action is front-loaded and the close:true caveat follows immediately after.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter, no-output-schema tool, this covers the essential action and the one non-obvious parameter. The remaining gap is minimal: id and data are only implicit, and there is no note on error behavior for a dead process.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the schema gives no help. The description explains the optional close parameter's effect well, and 'text to stdin' implies data while the process_start reference implies id, but neither id nor data is described explicitly, leaving the agent to infer their meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Write text to stdin of a process') and ties it to the sibling that creates the process (process_start), so the agent knows these two go together. It falls short of 5 because it never distinguishes itself from pty_write, which reads and writes to an interactive terminal and could easily be confused with this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete usage context ('REPLs, prompts, interactive CLIs') and a prerequisite ('a process started with process_start'). It does not state any when-not condition or name an alternative tool, so it stops short of explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_listA
Read-onlyIdempotent

List background processes started by this bridge.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint false, so the safety profile is covered. The description adds useful scoping context ('started by this bridge'), but it does not describe return format, pagination, or any behavioral details beyond that scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, and the scope ('started by this bridge') is front-loaded. It is appropriately sized for a zero-parameter listing tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, zero-parameter read-only tool with rich annotations and no output schema, the description covers the essential purpose and scope. The only gap is lack of differentiation from sibling list_processes, but the annotations already handle safety and idempotency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, and the schema is empty with 100% description coverage. The description appropriately adds no parameter details because none are needed; baseline for zero-parameter tools is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('background processes') and adds the scope 'started by this bridge.' This is clear, but it does not distinguish the tool from sibling list_processes or job_list, leaving some ambiguity about what makes this list distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus alternatives like list_processes, process_output, or job_list. The description only states what it does, not the context or conditions that would select it over a sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_outputA
Read-onlyIdempotent

Read new output from a background process since the last read (or all buffered output with all: true).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
allNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds useful context beyond those annotations by disclosing the stateful cursor behavior ('since the last read') and the buffering scope ('all buffered output with all: true'). It does not mention whether the call blocks when no new output is available, but the added cursor semantics are a meaningful contribution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the core action and then qualifies it with the parameter override. No wasted words; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (2 parameters, no output schema), and the description tells the agent what is returned and how to change the read scope. It omits whether the call blocks when no new output exists and does not describe the output format, which is a minor gap given no output schema is present and annotations cover safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the full burden. It explains the optional 'all' boolean's effect (return all buffered output) and implies that 'id' is the background process identifier, but it does not specify the format or type of the id beyond context. Partial compensation for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (Read) and resource (new output from a background process) with a scoping qualifier (since the last read or all buffered output). It implicitly distinguishes itself from pty/job siblings by specifying 'background process', but it does not name an alternative sibling tool by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by explaining the 'since the last read' default and the 'all: true' override, which guides the agent on how to call it. However, it never states when to prefer this tool over siblings like job_output or pty_read, leaving alternatives to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_startA
Destructive

Start a long-running background command (dev server, watcher, long build). Returns an id; read logs with process_output, stop with process_stop.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
nameNo
waitMsNoHow long to wait before returning initial output
commandYes
resourceWaitMsNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive, non-idempotent, and open-world behavior. The description adds lifecycle context: it returns an id and points to process_output for logs and process_stop for termination. It omits waitMs/resourceWaitMs effects, but annotations carry the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, with the primary purpose front-loaded and no filler. Every clause provides actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no output schema and poor parameter descriptions, the description covers the return id and follow-up tools but leaves parameter semantics and wait behavior unexplained. Adequate but incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 20% (waitMs only). The description only implies a 'command' parameter and does not explain cwd, name, or resourceWaitMs. With low coverage, it fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb 'Start' and resource 'long-running background command', with examples (dev server, watcher, long build). It does not explicitly differentiate from run_command, job_start, or pty_start among the many siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides when to use (long-running background commands) and names the follow-up tools process_output and process_stop. It does not give explicit exclusions or compare against run_command, job_start, or pty_start.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_stopA
Destructive

Stop a background process started with process_start (kills its whole process tree).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavioral detail: stopping kills the entire process tree, not just the parent, which an agent could not infer from the schema or annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, 14 words, front-loaded with the action and immediately qualified by the process_start linkage and the tree-kill behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter mutation tool with no output schema and annotations covering the destructive/idempotency profile, the description supplies the essential extras: the originating tool and the tree-kill scope. Only minor gaps remain, such as behavior when the id is stale or already terminated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter 'id' with 0% schema description coverage, so the description is the only source. It implies the id refers to a process previously launched via process_start but does not state whether it is a PID, opaque handle, or whether the id remains valid after the process exits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (stop) plus resource (background process) and it names the originating sibling tool process_start, which anchors the resource type. It does not, however, distinguish itself from the other termination-flavored siblings such as kill_process, pty_stop, or job_cancel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'started with process_start' implies the intended lifecycle pairing, but it never states when to prefer this over kill_process or how it differs from pty_stop/job_cancel. Usage is implied rather than explicitly routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_contextA
Read-onlyIdempotent

Start here for any project: detected stack and build/test/lint commands, git state, top-level layout, project instruction files, README start and saved memory notes.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoProject directory (default: default workspace)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world, so the safety profile is covered. The description then adds what is actually returned (stack, commands, git state, layout, instruction files, memory notes), which is genuine content-level information beyond the annotations. It does not discuss truncation or freshness of the gathered data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the 'start here' positioning and then a tight enumeration of contents. Every clause maps to a distinct piece of returned information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only orientation tool with complete annotations and no output schema, the description enumerates the returned sections well enough for an agent to know what it will get and when to call it first. It could say more about scope limits or ordering, but nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single optional parameter ('path', with default workspace behavior) fully documented in the schema. The description adds no path semantics of its own, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description lists exactly what the tool aggregates: detected stack, build/test/lint commands, git state, top-level layout, instruction files, README start, and memory notes. That distinguishes it from narrower siblings like repo_map, git_status, and project_memory. It lacks an explicit verb ('returns'/'aggregates'), but the resource and scope are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Start here for any project' gives clear, explicit when-to-use guidance and positions it as the entry point before the more specialized siblings. It stops short of naming when not to use it or which sibling to reach for instead once a specific need appears, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_memoryA
Destructive

Durable per-project notes shown by project_context in future conversations. action: list | add (note) | remove (index). Save conventions, gotchas and verified commands — not temporary state.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
pathNoProject directory (default: default workspace)
indexNo
actionNolist

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered. The description adds genuinely useful context beyond them: the notes are 'durable' and resurface via project_context 'in future conversations', and remove operates by index, signaling irreversible per-entry deletion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the persistence purpose, then the action syntax, then the content rule. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, no-output-schema mutation tool, the description covers purpose, action modes, parameter mapping, and the persistence behavior an agent needs. The only gap is the shape of the 'list' return (e.g. whether indexes are shown), which matters since remove relies on index.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only path is documented), so the description must compensate and largely does: it binds 'add' to note and 'remove' to index, revealing the action-to-parameter mapping the schema does not express. path's directory semantics are left to the schema, keeping it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (durable per-project notes) and enumerates the three action modes (list/add/remove), which lets an agent tell it apart from the many browser/process siblings. It does not name the closest alternative by comparison (project_context is mentioned only as the consumer), so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit content guidance — 'Save conventions, gotchas and verified commands — not temporary state' — which is a clear when-to-use plus a when-not-to-use. It stops short of naming a competing tool for ephemeral state, so no full routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pty_listB
Read-onlyIdempotent

List interactive PTY sessions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds nothing beyond that — no note on what a session entry contains, whether stopped sessions are included, or how results are ordered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words, appropriate for a no-argument list operation. It is arguably too terse, but nothing in it is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, non-destructive list tool with no output schema, the description is minimally viable. It omits any indication of what a returned session looks like or how to use the result (e.g., feeding an ID into pty_read or pty_stop), which would help an agent act on the output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. The description correctly implies an unfiltered full listing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (interactive PTY sessions), which is clear enough to distinguish from pty_start/pty_read/pty_write/pty_stop in the sibling set. It does not, however, explicitly differentiate itself from other listing tools like process_list or job_list, which could return overlapping or confusable results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to call this versus alternatives such as process_list or job_list, nor any preconditions. An agent must infer usage purely from the tool name and sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pty_readA
Read-onlyIdempotent

Read new terminal output from a PTY session (or all buffered output with all: true). ANSI escapes are stripped by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
allNo
rawNo
maxCharsNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false. The description adds useful behavioral context beyond annotations: ANSI escapes are stripped by default and all: true returns buffered output. It stops short of explaining truncation, blocking behavior, or error cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence states the core action, and the parenthetical efficiently covers the all option. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should convey return semantics. It indicates that terminal text is returned with ANSI stripped by default, but omits maxChars truncation/default behavior and does not fully explain raw or id. Annotations cover safety, but parameter and return gaps remain for a four-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for four parameters, so the description must carry the burden. It explicitly explains all: true and indirectly hints at raw behavior via 'ANSI escapes are stripped by default', but id and maxChars are undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Read' and resource 'terminal output from a PTY session', and distinguishes scope via 'new' versus 'all buffered output'. This clearly separates it from sibling operations like pty_write and pty_list, even without naming process_output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after starting or writing to a PTY and gives parameter-level guidance for all: true, but it does not state when to choose pty_read over process_output, job_output, or other read tools, nor does it provide any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pty_resizeC
Destructive

Resize a PTY terminal.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
colsNo
rowsNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is a destructive, non-idempotent, non-read-only operation, implying it modifies state. The description doesn't add any behavioral context beyond what annotations already say — it doesn't explain what gets destroyed, whether resizing affects running processes, or any side effects. With annotations covering the safety profile, the description could add value but doesn't. It also doesn't contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the action. It's efficient and not verbose, but it may be too terse given the complexity and missing details. However, it avoids waste, so it scores well on structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters, 0% schema coverage, no output schema, and destructive annotations, the description is completely inadequate. It doesn't explain parameter usage, return values, or behavioral implications. An agent would struggle to invoke this correctly without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for undocumented parameters. It fails to mention the required 'id' parameter or the optional 'cols' and 'rows' parameters, their meanings, formats, or constraints. The agent gets no additional semantic guidance beyond the schema, which itself has no descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Resize) and resource (PTY terminal), so the agent understands the general action. However, it lacks any detail about what resizing entails (e.g., terminal dimensions, cursor behavior) and doesn't differentiate this from other terminal control tools like pty_write or pty_read. It's minimally viable but generic.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as browser_resize or other terminal tools. The description doesn't specify prerequisites, timing, or scenarios where resizing is needed. It simply states the action without context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pty_startB
Destructive

Start a real interactive terminal backed by Windows ConPTY. Use for REPLs, debuggers, SSH/TUI programs and CLIs that require a terminal. By default starts the configured interactive shell; command runs inside it, or file+args starts an executable directly.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
argsNo
colsNo
fileNo
nameNo
rowsNo
waitMsNo
commandNo
resourceWaitMsNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, and openWorldHint=false, so the agent knows this is a local, non-read-only, non-idempotent operation. The description adds that a command runs inside the configured shell or that file+args starts an executable directly, but it does not describe persistence, interaction lifecycle, or other side effects beyond what annotations imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and terminal backing, then moves to usage and invocation details. Every sentence adds distinct information with no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no schema descriptions and no output schema, the description is materially incomplete. It explains the main command/file/args invocation pattern but omits behavioral meaning for most parameters, making it insufficient for reliable parameter selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the description must compensate. It clarifies the relationship between command, file, and args, but leaves cwd, cols, rows, name, waitMs, and resourceWaitMs completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Start a real interactive terminal backed by Windows ConPTY.' It also identifies the use cases (REPLs, debuggers, SSH/TUI programs, terminal-requiring CLIs), which implicitly distinguishes it from non-terminal process and command tools. However, it does not explicitly name or contrast with any sibling tool, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use for REPLs, debuggers, SSH/TUI programs and CLIs that require a terminal' gives clear positive usage guidance. It does not provide when-not conditions or name alternatives such as run_command or process_start, so it is not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pty_stopC
Destructive

Stop and remove an interactive PTY session.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description adds the 'remove' nuance (the session is deleted, not just paused), which is meaningful beyond 'stop', but it omits what happens to running processes, whether an invalid id errors, or whether it blocks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the action front-loaded and no padding. Brevity is appropriate, though it contributes nothing to filling the semantic gaps.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema and an entirely undocumented parameter, the description should explain the identifier source and the consequences of removal. It leaves both to inference, which is insufficient even with annotations carrying the safety signal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required 'id' parameter. The description never states that this is a PTY session identifier obtained from pty_list/pty_start or what format it takes, so nothing compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair and resource: 'Stop and remove an interactive PTY session.' An agent can distinguish it from pty_start/pty_read/pty_write by the verb, though it doesn't explicitly contrast with close siblings like process_stop or kill_process.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. The agent must infer that this is the cleanup counterpart to pty_start and is distinct from process_stop on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pty_writeB
Destructive

Write to an interactive PTY. enter: true appends the Enter key.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
dataYes
enterNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral detail — that enter:true appends the Enter key — but says nothing about whether the write blocks, whether it can fail if the PTY is dead, or how it interacts with pty_read buffering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse sentences, front-loaded with the action and followed by the one non-obvious parameter behavior. Nothing is wasted, though the brevity is partly a symptom of under-specification rather than discipline.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent 3-parameter tool with no output schema and zero schema coverage, the description is thin. It omits the lifecycle dependency on pty_start, the meaning of 'id', and how the caller observes the effect — the agent is left to guess at the basic usage loop.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden. It explains only the optional 'enter' flag; the required 'id' (presumably a PTY session identifier) and 'data' (payload semantics, encoding, newline handling) are entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Write to an interactive PTY'), which clearly separates it from pty_read, pty_start, and pty_stop. It does not, however, clarify what 'id' refers to or how this differs from process_input, so differentiation from the broader process family is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus pty_read, process_input, or browser_type. The description never states that a PTY must already exist (via pty_start) or that output must be retrieved with pty_read, both of which an agent needs to invoke this correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_fileA
Read-onlyIdempotent

Read a text file with line numbers. Relative paths resolve against the default workspace. Use startLine/endLine for large files.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path
endLineNo
startLineNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real value beyond that: relative paths resolve against the default workspace, and output includes line numbers. It doesn't disclose error behavior for missing files, but for a read tool this is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero filler, and the core action is front-loaded. The path-resolution and large-file hints follow in priority order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description conveys the key return characteristic (line numbers) and the path-resolution rule. For a simple 3-param read tool that is close to complete, with only edge-case semantics (out-of-range lines, missing file) left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% — only 'path' is documented as 'File path'. The description compensates partially by naming startLine/endLine and tying them to large files, but it doesn't clarify indexing base, inclusivity of endLine, or what happens when omitted. Baseline 3 fits given the partial compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Read a text file') plus an output trait ('with line numbers'), which cleanly separates it from write_file, edit_file, and read_many_files. It stops short of naming the sibling it differs from (read_many_files for multi-file reads), so it isn't a full 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It offers one usage hint — 'Use startLine/endLine for large files' — which is genuinely actionable, but there is no guidance on when to choose this over read_many_files or file_info, and no stated preconditions. Usage is implied rather than articulated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_many_filesA
Read-onlyIdempotent

Read up to 50 files in one call (strings, or {path,startLine,endLine}). Much faster than several read_file calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesYes
maxTotalCharsNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds one useful behavioral fact, the 50-file batch cap, but says nothing about truncation behavior or what happens when the output exceeds the default maxTotalChars of 150000.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence of capability plus one sentence of motivation, zero filler, with the batch limit front-loaded. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with no output schema, the essential call shape is conveyed, but an agent is left guessing about maxTotalChars and about truncation when many large files are requested. Adequate but with a clear documentation gap on the second parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description has to carry the load: it does document the polymorphic file entry forms (string vs {path, startLine, endLine}), which mirrors the anyOf but gives an agent a usable summary. However, the second parameter maxTotalChars (default 150000) is never mentioned, leaving half the parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) plus resource (files) with an explicit scope qualifier (up to 50 in one call). The contrast with the sibling read_file is implicit in the name and made explicit in the description, so an agent can distinguish the batch reader from the single-file reader immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly positions itself against several read_file calls, which is a genuine routing signal (use this when you'd otherwise call read_file repeatedly). It stops short of stating when-not-to-use it (e.g. for a single file, or when output would exceed maxTotalChars), so it's clear context but no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_clipA
Destructive

Render a video clip from code: pass a complete HTML document (CSS/JS/canvas/SVG animation) or a URL; it is played headlessly at the given size for durationMs and saved as mp4/webm/gif. Default 1080x1920 (vertical, Shorts/TikTok/Reels). If the page defines window.startClip(), it is called when recording begins.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNo
urlNo
htmlNo
nameNoFile name inside the clips folder
widthNo
formatNomp4
heightNo
outputNoExplicit output path (overrides name)
durationMsNo
readySelectorNoWait for this selector before starting the clock
resourceWaitMsNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag this as a destructive, open-world, non-idempotent write, so the safety bar is partly covered. The description adds real behavioral detail: playback is headless, the clip is recorded for durationMs, the pages defines window.startClip() which is invoked at recording start, and readySelector waits before the clock starts. It doesn't disclose errors, rate limits, or the returned handle.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight clauses, front-loaded with the action and input, then output defaults, then the optional startClip hook. Every sentence earns its place, though the input-vs-output sentence is dense and packs several unrelated facts together.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter destructive tool with no output schema and low schema coverage, the description still leaves gaps: fps and resourceWaitMs are never explained, the name vs output resolution is left to the schema, and there is no indication of what is returned (file path/handle) after recording. It is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 27%, so the description must carry weight, and it does: it defines html vs url input, explains width/height defaults as vertical 1080x1920, clarifies format options and durationMs semantics. It omits fps, resourceWaitMs, and the readySelector/name/output relationships, but covers most of the functionally critical parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Render a video clip from code') plus the input forms (HTML document or URL) and the output artifacts (mp4/webm/gif). This clearly separates it from browser_screenshot, browser_record_start/stop, and windows_snapshot in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains what to pass (complete HTML document or a URL) and mentions the window.startClip() hook, which is genuine usage context. However, it never says when to choose this over browser_record_start/browser_record_stop, and it does not note that at least one of html/url is functionally required despite required=0.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_mapB
Read-onlyIdempotent

Compact map of a codebase: the most-referenced files with their key functions/classes/types and line numbers. Understand a large project in one call. focus: file names or symbols to prioritise.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoProject directory (default: default workspace)
focusNo
maxCharsNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds useful behavioral context about the output shape (compact, most-referenced-first, includes line numbers), but says nothing about cost, truncation behavior, or what happens when the map is too large.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the output description before the usage note and the parameter hint. No filler, though the trailing 'focus:' fragment reads as a schema note tacked onto prose rather than integrated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing returns – it does reasonably well (files, symbols, line numbers). However, with a 12k default character budget and truncation likely, an agent gets no guidance on what is dropped or how to react, and one of three params is unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only `path` is documented in the schema), so the description must compensate. It does explain `focus` as 'file names or symbols to prioritise', which adds real meaning, but `maxChars` and its default-truncation implications are left entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: producing a compact map of a codebase showing most-referenced files with key functions/classes/types and line numbers. That is concrete enough to distinguish it from outline or search_code at a glance, though it never names a sibling to sharpen the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Understand a large project in one call" implies the usage context (orientation in a big codebase), but there is no explicit when-to-use vs. when-not, and no mention of alternatives such as outline, project_context, or find_symbol that an agent might otherwise pick for similar goals.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resource_statusA
Read-onlyIdempotent

Current workstation resource pressure and admission-controller slots for commands, builds/tests, LSP, diagnostics, background processes and video work.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false. The description adds useful scope by naming the resource categories it covers, but it does not disclose freshness guarantees, latency, or whether the values are cached or live beyond the word "Current".

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tightly written sentence that front-loads the resource being reported and then enumerates the categories. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter status tool with rich annotations, the description is nearly complete. It identifies what is reported but does not describe the shape or fields of the return value, which would be useful since no output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to clarify. The baseline for a 0-parameter tool is 4, and the description adds no confusion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states what the tool returns: current workstation resource pressure and admission-controller slots across specific work categories. It lacks any explicit differentiation from siblings such as get_system_info or lsp_status, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It does not say when to use this tool, when not to use it, or what alternatives exist. The reader must infer that it is for checking resource status before starting heavy work, but no guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_commandA
Destructive

Run a shell command (/bin/bash) and wait for it to finish. Returns exitCode, stdout, stderr. Use for builds, tests, git, npm, python, ffmpeg, etc. For servers/watchers that never exit use process_start.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory (default: default workspace)
stdinNoOptional text piped to stdin
commandYes
timeoutMsNoDefault 120000
resourceClassNoAdmission class; defaults to automatic command classification
resourceWaitMsNoHow long to wait for CPU/RAM capacity before refusing the command

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the destructive, non-idempotent, open-world profile, so credit goes to what the description adds: blocking-until-completion semantics and the return contract (exitCode, stdout, stderr). It does not explain what happens on timeout (whether the process is killed, partially run, or left alive), which matters for a destructive tool with a 120s default timeout.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and completion behavior, then usage examples, then the routing exclusion. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the return fields and the blocking behavior. The remaining gap is timeout semantics and the resource-admission refusal path (resourceWaitMs), which an agent running heavy/destructive commands would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so the schema already documents cwd, stdin, timeoutMs, resourceClass, and resourceWaitMs. The description adds no parameter-level detail beyond the schema baseline, so the standard 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a shell command (/bin/bash)') plus a key execution trait ('wait for it to finish'), and names the sibling it is not (process_start) for the never-exiting case. An agent can distinguish it from job_start/pty_start/process_start without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete positive examples (builds, tests, git, npm, python, ffmpeg) and an explicit exclusion with the alternative to use instead ('For servers/watchers that never exit use process_start'). This is exactly the when/when-not/alternative pattern.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codeA
Read-onlyIdempotent

Search file contents with a regex (ripgrep when installed, Node fallback otherwise). Returns path:line: text; context adds surrounding lines.

ParametersJSON Schema
NameRequiredDescriptionDefault
globNoFile filter, e.g. "*.ts"
pathNoDirectory to search (default: default workspace)
queryYesRegex (or literal text with literal: true)
contextNo
literalNo
maxResultsNo
caseSensitiveNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuinely new context: the ripgrep-vs-Node backend fallback and the return shape 'path:line: text'. It omits any mention of maxResults truncation or rate/performance behavior, keeping it short of 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action, then the return format. Zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but the description supplies the return format and context behavior. For a 7-param read tool this is nearly complete, with the minor gap being maxResults truncation semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 43%, so the description must compensate. It clarifies the 'query' semantics (regex, or literal with literal: true) and explains 'context' (surrounding lines), which the schema leaves bare. However, maxResults and caseSensitive remain undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Search file contents with a regex.' The phrase 'file contents' implicitly separates it from find_files (filename search), but no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no when-not, and no routing to alternatives like find_files or find_symbol. The agent must infer from the 'file contents' scope that this is content search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_cancelA
Destructive

Cancel a durable task. A running task requires its current lease; queued/verifying/repairing tasks can be cancelled by id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
leaseIdNoOpaque lease returned by task_claim/task_next. Required while a worker owns the task.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already mark the tool as destructive, non-idempotent, and not read-only. The description adds useful operational context about lease requirements for running tasks and which task states can be cancelled by id alone, going beyond what the annotations and schema together provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and followed by the key state-dependent constraint. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation tool with no output schema, the description covers the main operational nuance: when a lease is needed and when cancellation can be done by id. Annotations cover the destructive safety profile, so remaining gaps such as irreversibility details are not critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: leaseId is documented in the schema, while id has no schema description. The description says tasks can be cancelled by id, implying id identifies the task, but it adds no format, origin, or constraints for id and largely repeats the leaseId requirement already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Cancel a durable task.' This clearly distinguishes the operation from sibling task tools such as task_submit, task_claim, task_update, and task_verify. An agent can identify the core action without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives state-dependent usage context: a running task requires its current lease, while queued/verifying/repairing tasks can be cancelled by id. It does not explicitly name alternatives or when not to use the tool, but the cancellation conditions are clear enough to guide invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_claimB
Destructive

Claim one queued/repairing task with an exclusive durable lease. Refuses unmet dependencies, lease conflicts, critical resource pressure, and worker-capacity overflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
workerIdNo
workerRoleNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds meaningful behavioral context beyond annotations: exclusive durable leasing and four specific refusal conditions. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences that front-load the action and then list refusal conditions. Every clause is informative and there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers core behavior and failure modes well, and annotations cover the safety profile. However, with zero parameter description coverage and no output schema, the tool definition still leaves important calling details—especially the meaning and roles of id, workerId, and workerRole—unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has three parameters (id, workerId, workerRole) with 0% description coverage, and the description mentions none of them. It does not explain what id refers to, whether workerId/workerRole are required, or what values are expected, leaving parameter semantics entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Claim one queued/repairing task with an exclusive durable lease.' It also defines scope (queued/repairing) and mechanism (exclusive durable lease), making the action clear. However, it does not explicitly differentiate this tool from similar siblings such as task_next or task_ready.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this tool to claim a queued/repairing task. The description lists refusal conditions—unmet dependencies, lease conflicts, resource pressure, worker-capacity overflow—which are useful preconditions but not explicit when-to-use or when-not-to-use guidance relative to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_contractB
Read-onlyIdempotent

Read the self-contained execution/verification contract for one durable task.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false, and destructiveHint=false, so the safety profile is covered. The description adds only that the contract is "self-contained" and covers execution/verification, but it does not disclose error behavior, missing-task handling, or what the contract contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no filler, and the core action and resource appear immediately. Nothing is repeated from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with clear annotations, the description is minimally viable. However, with no output schema and 0% parameter description coverage, it should explain more about what the contract contains and how to select this tool over the many task_* siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the single required "id" parameter has no schema description. The description implies the id identifies "one durable task," which adds minimal meaning, but it gives no format, example, or lookup semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ("Read") and resource ("self-contained execution/verification contract for one durable task"). It is clear enough to distinguish the general purpose from task_status or task_update, but it does not explicitly name or differentiate itself from the many other task_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to call this tool, when not to call it, or which sibling tool to use instead. The phrase "for one durable task" implies the required id but does not explain the use case versus task_status, task_verify, or task_events.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_eventsB
Read-onlyIdempotent

Read the append-only task audit trail, optionally for one task.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNo
limitNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world behavior. The description adds useful context that the data is an append-only audit trail, implying historical immutability, but omits ordering, pagination/limit behavior, and permission requirements. With annotations covering the safety profile, this is adequate but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence, front-loads the verb and resource, and contains no filler. The size is appropriate for a routine read tool and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no output schema and no schema descriptions, the description is too thin. It does not clarify event time range, ordering, how limit affects pagination, or what the returned audit trail contains, leaving key invocation details underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema itself provides no parameter meaning. The description only loosely hints at the id parameter via 'optionally for one task' and says nothing about id format or the limit parameter's default, maximum, or behavior. It does not compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('append-only task audit trail'), with optional scoping to one task. Clear, but does not explicitly contrast itself with task_list, task_status, or task_stats siblings that also expose task-related data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Only mentions that the tool can optionally be scoped to one task. It gives no when-to-use guidance, no alternatives, and no conditions or exclusions, so the agent must infer when this is preferable to task_list or task_status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_heartbeatA
Idempotent

Renew a running task lease. Long-lived workers should heartbeat before lease expiry.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
leaseIdYesOpaque lease returned by task_claim/task_next. Required while a worker owns the task.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true, readOnlyHint=false, and destructiveHint=false, so the safety profile is covered. The description adds the important timing context of heartbeating before lease expiry, but does not explain what happens on an expired lease or any rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences and front-loads the tool's purpose. Every sentence earns its place by defining the action and giving operational timing guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple lease-renewal tool, the description covers purpose and timing, and annotations cover safety. However, it leaves the required id parameter semantically unexplained and does not describe error behavior when a lease is no longer valid.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: leaseId is well documented in the schema, but id has no description. The description does not compensate by clarifying what id represents, and it adds no new semantic detail beyond the schema's existing leaseId explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: renew a running task lease. This is clear and distinct from sibling task tools like task_claim or task_update, though it does not explicitly name an alternative or contrast itself with siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage condition: long-lived workers should heartbeat before lease expiry. This tells the agent when to call it, but it does not state when not to use it or name an alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_listB
Read-onlyIdempotent

List durable tasks filtered by state/project, ordered by priority then creation time.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statesNo
projectNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this as a read-only, idempotent, non-destructive, closed-world operation, so the safety profile is covered. The description adds useful ordering semantics ('priority then creation time') and filter dimensions, but does not describe pagination behavior, return format, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded with the core action and then adds filtering and ordering constraints. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with annotations covering safety and no output schema, the description is adequate but incomplete. It omits pagination/limit semantics and does not clarify the return shape, though the absence of an output schema lowers the burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It maps to two of the three parameters by mentioning state and project filtering, but omits any explanation of the 'limit' parameter. It also says 'state' singular while the actual parameter is the array 'states', leaving some ambiguity despite the schema enum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List durable tasks.' It also specifies filtering and ordering, which helps distinguish it from task_status/task_stats. However, it does not explicitly name or distinguish itself from the many sibling task_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use guidance and does not mention alternatives such as task_ready, task_next, or task_status. Usage is only loosely implied by the filtering and ordering behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_nextB
Destructive

Agent-loop primitive: atomically claim the highest-priority runnable task for this worker. Multiple model clients can call this concurrently; leases prevent two workers owning the same tree.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
projectNo
workerIdNo
workerRoleNo

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false. The description adds meaningful behavioral context beyond those annotations: atomic claiming and lease-based concurrency protection preventing two workers from owning the same tree. It does not describe lease expiry or exactly what state changes on claim, but it adds real value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and followed by concurrency safety context. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with four undocumented parameters and no output schema, the description leaves the agent without guidance on what arguments to provide. It covers the purpose and concurrency behavior well, but the missing parameter semantics make it incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are 4 parameters with 0% schema description coverage, and the description does not explain cwd, project, workerId, or workerRole. The phrase 'for this worker' loosely hints at workerId but provides no usable semantics for any parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'atomically claim the highest-priority runnable task for this worker.' It clearly distinguishes the core action from generic task listing or status tools, though it does not explicitly contrast with siblings like task_claim or task_ready.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It labels itself an 'Agent-loop primitive' and mentions concurrent callers, which implies a usage context, but there is no explicit when-to-use guidance, no conditions for choosing it over task_claim or task_ready, and no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_readyA
Destructive

Worker handoff: mark implementation ready for verification and release the worker lease. The task keeps its tree reserved while verification runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
resultNo
leaseIdYesOpaque lease returned by task_claim/task_next. Required while a worker owns the task.
evidenceNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and non-idempotent, and the description adds the crucial state effects beyond that: the worker lease is released but the task tree stays reserved during verification. That lease/tree distinction is real behavioral value, though the description never explains why the action is flagged destructive or what a repeated call does.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and followed immediately by the resulting state. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but none is needed: the description already conveys the outcome (lease released, tree still reserved). Since the definition is otherwise complete for a handoff tool, the lack of a return schema is not a material gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% and the description mentions no parameters at all. The required 'id' and the nested 'result'/'evidence' objects are undocumented in both places, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: marking implementation ready for verification plus releasing the worker lease. The handoff framing distinguishes it conceptually from the verifier-side sibling task_verify, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'Worker handoff' — an agent can infer this is called by a worker when implementation is done. There is no explicit statement of when not to call it or which sibling (task_update, task_cancel, task_reconcile) handles adjacent cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_reconcileA
Destructive

Recover a task left interrupted after restart/lease loss. Requeue it, mark it failed, or explicitly reattach a known-live worker with a fresh lease.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
errorNo
evidenceNo
workerIdNo
workerRoleNo
dispositionYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description adds the recovery framing and the disposition→effect mapping (requeue / mark failed / reattach with a fresh lease), but does not say what state is permanently lost or that the operation is non-idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the recovery scenario is front-loaded, then the three outcomes. Every clause carries information and nothing is repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations cover safety and there is no output schema to explain, so the burden is on purpose and parameters. Purpose is covered well, but with 6 undocumented parameters and a destructive hint, the description should say more about the error/evidence inputs and irreversibility.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across six parameters, so the description must carry the load. It clarifies disposition semantics and hints that workerId/workerRole matter for the 'reattach' path, but leaves error, evidence, and id completely unexplained — a significant gap for a destructive mutation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Recover a task left interrupted after restart/lease loss') and enumerates the three recovery modes, which lets an agent distinguish it from task_cancel/task_update at a glance. It does not name a sibling alternative, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering condition — the task was interrupted by a restart or lease loss — which is exactly the context that selects this tool over task_heartbeat or task_claim. It offers no explicit when-not-to-use or routing to siblings, so it is clear context but not 5-level guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_statsA
Read-onlyIdempotent

Read task counts, active workers, configured capacity, and current resource pressure.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds the categories of metrics returned, but it does not explain return format, freshness, auth requirements, or other behavioral details. This is useful context but not rich behavioral disclosure, matching a 3 when annotations carry the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every item in the list earns its place by specifying the readout contents.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only stats tool with rich annotations and no output schema, the description is largely sufficient because it lists the key metrics. However, it leaves minor gaps such as what 'resource pressure' means or whether values are real-time, so it is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description does not need to document parameter semantics, and the empty schema is consistent with a no-argument stats tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and names the exact resource categories returned: task counts, active workers, configured capacity, and resource pressure. It clearly distinguishes itself as an aggregate stats tool rather than a per-task status tool, but it does not explicitly differentiate from siblings like task_status, task_list, or resource_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention task_status, task_list, resource_status, or any other sibling, nor does it give any conditions or exclusions for use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_statusB
Read-onlyIdempotent

Read the current durable state of one task.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description's only added nuance is 'durable' state, implying persisted rather than transient state, which is useful but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, zero filler, front-loaded with the verb and scope. Nothing could be removed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read tool with no output schema, the description is near-adequate, but it omits the error behavior for missing/invalid ids and gives no hint of the returned state fields, leaving gaps for an agent to discover at call time.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'id' parameter is completely undocumented. The description only implies the id selects 'one task' without saying what identifier form is expected (task id vs. name vs. uuid) or what happens on an unknown id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb ('Read') and resource ('current durable state of one task'), and the cardinality ('one task') distinguishes it from list-style siblings like task_list or task_events. It does not explicitly name an alternative or state what 'durable state' contains, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus task_list, task_events, task_verify, or task_update. The agent must infer that this is a point-read keyed by id. No prerequisites or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_submitB
Destructive

Create a durable provider-neutral agent task. Supports dependencies, exclusive tree leases, idempotency, repair limits, and authoritative verification commands.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdYes
titleNo
projectNo
leaseKeyNoTasks with the same leaseKey cannot be active concurrently. Defaults to cwd.
metadataNo
priorityNo
dependsOnNo
objectiveYes
acceptanceNo
allowedPathsNoWorker contract hints. FCA_ALLOWED_ROOTS remains the actual filesystem security boundary.
forbiddenPathsNoWorker contract hints. Use FCA_ALLOWED_ROOTS for hard sandboxing.
idempotencyKeyNo
maxRepairRoundsNo
verificationCommandsNoCommands task_verify runs in cwd. Non-zero exit or timeout fails verification.
requireIndependentVerifierNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is covered structurally. The description adds genuine behavioral context (exclusive tree leases, repair limits, verification commands that task_verify runs), but omits what happens on duplicate submission, whether tasks can be cancelled/reconciled, and the durability/persistence semantics implied by 'durable'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action, no filler. Efficient, though the second sentence is a compressed feature dump rather than structured guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter mutation tool with no output schema and only 27% schema coverage, the description leaves major gaps: what the returned task identifier is, how this relates to task_claim/task_ready/task_verify, and how the many optional contract parameters interact. It is insufficiently complete for correct invocation without reading the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 27% for 15 parameters, so the description must compensate and it does so only loosely: it gestures at dependencies, leases, idempotency, repair limits, and verification commands without mapping them to parameter names or formats. Required fields cwd and objective, plus acceptance, metadata, priority and requireIndependentVerifier, get no elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Create a durable provider-neutral agent task.' An agent can distinguish this as the submission/creation tool within the task_* family. It does not, however, explicitly contrast itself with siblings like task_claim or task_next, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., an active lease/workspace), and no named alternative. The feature list ('Supports dependencies, exclusive tree leases...') describes capabilities, not selection criteria, so the agent must infer that this is the entry point for the task lifecycle.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_updateA
Destructive

Report a worker failure or interruption while it owns the task. Successful work must go through task_ready then task_verify.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
errorNo
actionYes
leaseIdYesOpaque lease returned by task_claim/task_next. Required while a worker owns the task.
evidenceNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and idempotentHint=false, so the safety profile is covered. The description adds useful workflow context: the operation is scoped to lease ownership and is the terminal failure branch rather than the success branch. It does not say what happens to the task afterward (re-queue vs. permanent failure) or whether the lease is released, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the primary purpose front-loaded and the routing constraint immediately after. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no output schema and very low schema coverage, the description establishes the workflow context and correct routing but omits parameter-level detail (error/evidence semantics) and post-call state behavior. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only leaseId is documented). The description loosely maps to the action enum via 'failure or interruption', but adds nothing about the error payload, the evidence object, or the id field, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (report) plus the exact conditions (failure or interruption) and the resource (worker task ownership). It also implicitly distinguishes itself from the success path by naming task_ready/task_verify, so an agent can tell what this tool is not for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it (worker failure or interruption while owning the task) and when not to (successful work must go through task_ready then task_verify). The alternative path is named, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_verifyB
Destructive

Run the task verification gates and apply the verdict durably. Passing gates transition to succeeded; failures transition to repairing until maxRepairRounds is exhausted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
errorNo
verdictNoauto
evidenceNo
timeoutMsNo
verifierIdNoLogical verifier identity. Required when the task requests independent verification.
resourceWaitMsNo
continueOnFailureNo

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=false, but the description adds real behavioral value the annotations cannot: it discloses the resulting state machine transitions and the maxRepairRounds exhaustion path, which is exactly the kind of side-effect an agent needs before calling a destructive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and then the outcome semantics. Little waste, though the word 'durably' is vague and slightly redundant with 'apply the verdict'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with 8 parameters, 13% schema coverage, and no output schema, the description covers only the state-transition behavior. It leaves permissions/auth requirements, the meaning of most inputs, and failure/error handling entirely undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13% across 8 parameters, so the burden falls on the description — yet id, error, evidence, timeoutMs, resourceWaitMs, and continueOnFailure are never mentioned. The one term it cites (maxRepairRounds) is not even a parameter, so it cannot compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource (run verification gates, apply verdict) and clarifies the outcome semantics. It does not, however, differentiate from siblings like task_update or task_reconcile, which also mutate task state, so an agent must infer which to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains what happens after invocation (pass→succeeded, fail→repairing) but never says when to call this versus task_update, task_reconcile, or task_status. No prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vscode_add_breakpointC
Destructive

Add a breakpoint (optionally conditional).

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute file path
lineYes1-based line
conditionNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and non-idempotent, but the description adds nothing: it doesn't say whether breakpoints persist across sessions, whether re-adding the same line is harmless, or that a debug session must be running. With annotations present the bar is lower, yet the description contributes no behavioral context at all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler and the optional aspect front-loaded. It is efficient, though borderline under-specified rather than maximally informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating debug tool with no output schema, the description omits the debug-session prerequisite, where the breakpoint attaches, and any indication of success/failure. An agent could not confidently invoke this without inferring context from outside the definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: 'file' and 'line' are documented in the schema, while 'condition' has no schema description. The description's 'optionally conditional' partially compensates by telling the agent what the parameter is for, but it gives no expression syntax or scope for the condition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add a breakpoint') plus a scope modifier ('optionally conditional'). No sibling performs this action, so differentiation is unnecessary, but the description stops short of saying where the breakpoint lands or its lifecycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites. Critically, it never states that an active debug session is required, or how this relates to vscode_debug_state — the one sibling an agent would plausibly pair with it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vscode_debug_stateC
Read-onlyIdempotent

Active debug session and breakpoints.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. Beyond that the description discloses nothing - no indication of what 'debug state' includes, whether it errors with no active session, or what the response contains (and there is no output schema to fall back on).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is certainly short, but this reads as under-specification rather than tight writing: a single verbless fragment that omits the operation entirely. Brevity here costs clarity rather than buying it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations describing return values, the description is the only place the agent could learn what comes back, and it does not. 'Debug state' is left undefined for a tool that presumably returns session status and the breakpoint list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There is no parameter behavior for the description to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Active debug session and breakpoints' names the resource but contains no verb, so the agent must infer that this is a read operation. It never distinguishes itself from the sibling vscode_add_breakpoint, which it is clearly the read counterpart to. Purpose is guessable but vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus alternatives such as vscode_add_breakpoint or vscode_diagnostics. No prerequisites, no context, no exclusions are given. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vscode_diagnosticsA
Read-onlyIdempotent

Live compiler/linter errors and warnings from VS Code language servers (optionally filtered to one file).

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description contributes the useful 'Live' qualifier, implying diagnostics reflect current editor state rather than a persisted snapshot, but adds nothing about freshness, staleness, or what happens when no server is attached.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the resource first and the scoping qualifier in parentheses. No filler, no restatement of annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter, read-only diagnostic query with full annotation coverage and no output schema, this is nearly sufficient. What is missing is the relationship to lsp_diagnostics and the expected return shape (severity/files grouping), which matters more here than for a trivial read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden, and it does: 'optionally filtered to one file' tells the agent the single 'file' parameter is optional and scoping, not required. It does not specify path format (absolute vs workspace-relative), which is the remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb-resource pair: compiler/linter errors and warnings sourced from VS Code language servers, with an optional file scope. It is distinguishable from file-reading siblings, but it never addresses the near-identical lsp_diagnostics sibling, leaving overlap ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite (e.g. whether the editor/language server must be running), and no direction to alternatives such as lsp_diagnostics for the same information. Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vscode_find_referencesB
Read-onlyIdempotent

All references to the symbol at a position (via the language server).

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute file path
lineYes1-based line
characterYes1-based column

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is well covered. The description adds that the operation uses the language server, which is useful behavioral context beyond the annotations. It does not describe language-server availability requirements, latency, or result completeness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It states the core operation and mechanism efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, read-only language-server tool with 100% schema coverage and rich annotations, the description supplies the essential operation. It implies a list of references without describing output format, but the absence of an output schema makes that a minor omission rather than a critical gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so file, line, and character are already fully documented. The description adds only the concept of a 'symbol at a position,' which maps naturally to the line/character parameters but does not add syntax or format details beyond the schema. The baseline of 3 for high schema coverage is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair: find all references to the symbol at a given position. It also specifies the mechanism (via the language server), which clarifies the semantic nature of the search. However, it does not differentiate this tool from close siblings such as lsp_references or vscode_go_to_definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use guidance, prerequisites, or alternatives. It only implies the use case by describing the operation. With a sibling like lsp_references appearing to perform a similar function, the lack of routing guidance is a meaningful gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vscode_go_to_definitionC
Read-onlyIdempotent

Definition location of the symbol at a position.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute file path
lineYes1-based line
characterYes1-based column

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is covered structurally. The description adds nothing beyond that: no note on what is returned, whether the definition may be in another file, or behavior when no definition exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no waste, and the operative concept (definition location of the symbol) is front-loaded. It is terse to the point of being slightly under-specified rather than verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, fully-schema-documented, read-only navigation tool, the description is minimally sufficient. It omits any handling of the no-definition case and does not clarify its relationship to the overlapping lsp_definition sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear notes that file is absolute and line/character are 1-based. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource (definition location of the symbol) and the positioning input, so an agent knows this resolves a symbol to its definition. However, it is stated as a noun phrase rather than a verb, and it does not differentiate itself from the sibling lsp_definition, which appears to do the same thing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no alternatives named, despite the presence of lsp_definition among siblings that likely overlaps in function. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vscode_open_fileB
Destructive

Open a file in VS Code at a line so the user can see it.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesAbsolute file path
lineNo1-based line

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false), so the description does not need to restate that. It adds modest context by framing the action as making the file visible to the user, but it never explains the surprising destructive flag, what happens if the path does not exist, or any side effects on editor state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the action front-loaded and zero filler; nothing repeats the schema or annotations. The phrasing 'at a line' is slightly loose, which keeps it short of ideal precision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema, the description covers the core contract adequately. It omits failure behavior (missing file, no VS Code instance) and any statement about editor state after the call, which an agent would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – both 'file' (absolute path) and 'line' (1-based) are documented in the schema itself. The description only gestures at the line parameter ('at a line') and adds nothing about path format or line defaults beyond what the schema already says, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (Open) and resource (a file in VS Code) and adds the intent 'so the user can see it', which implicitly separates it from read_file, which returns content to the agent. It does not explicitly name or contrast any of the ~90 sibling tools, so the differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this instead of read_file, vscode_open_files, or file_info, nor any prerequisites (e.g. VS Code must be running, workspace must be open). The only hint is the trailing purpose clause, which implies user-facing display but never says so as guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vscode_open_filesC
Read-onlyIdempotent

Open editor tabs and the active file in VS Code.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds almost nothing beyond that — it does not explain what tabs actually get opened, whether the active file is derived from context, or what side effects (focus stealing, view changes) occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding. It is concise, though its brevity stems partly from under-specification rather than tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no parameters, no output schema, and a near-identical singular sibling, the description should explain what state or context drives the file selection. Instead it leaves the central question — which files and why — unanswered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters there is nothing for the description to disambiguate; the baseline for a no-param tool is 4. The description does not need to compensate for any schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb (open) and resource (editor tabs / active file in VS Code), but the plural name with a zero-parameter schema leaves it unclear which files it opens or how they are chosen. It also fails to distinguish itself from the sibling vscode_open_file, which an agent would naturally confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus vscode_open_file or vscode_debug_state, and no prerequisites or context are given. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

windows_actionA
Destructive

Act on a native Windows UI Automation element from windows_snapshot: invoke, setValue, toggle, select, expand, collapse, focus or scrollIntoView. Re-snapshot after meaningful UI changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
pidNoProcess ID from windows_list
refYes
titleNoTop-level window title substring (alternative to pid)
valueNo
actionYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=false and openWorldHint=false, so the safety profile is largely covered. The description adds the re-snapshot workflow constraint, which is genuine behavioral value, but it does not indicate which of the eight actions are state-mutating (invoke/setValue) versus benign (focus/scrollIntoView), nor how failures surface when a ref goes stale.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the core capability plus its action set is front-loaded before the operational follow-up. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter, non-idempotent, destructive UI-automation tool with no output schema, the description covers the action surface and one workflow rule but omits what happens on success/failure, ref staleness behavior, and any permission or focus prerequisites. It is workable but leaves an agent to infer too much before a destructive call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40%: pid and title are documented in the schema, but ref, value, and action have no field descriptions. The description partially compensates by listing the full action vocabulary and identifying windows_snapshot as where ref comes from, but it says nothing about the ref format or which actions consume value, leaving real gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear resource (a native Windows UI Automation element from windows_snapshot) and enumerates the eight actions it can perform, which concretely defines the tool's scope. It implicitly differentiates itself from the snapshot tool by naming it as the source of the ref, though 'Act on' is a dispatcher-style verb that relies on the action list for precision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The instruction to re-snapshot after meaningful UI changes is a useful workflow cue, and referencing windows_snapshot as the ref source implies the correct sequence. However, it never states when to choose windows_action versus alternatives (browser_*, chrome_*, windows_screenshot) or any preconditions such as needing a valid ref from a fresh snapshot.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

windows_listB
Read-onlyIdempotent

List visible top-level native Windows application windows with process IDs and UI Automation refs.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/closed-world, so the safety profile is covered. The description adds real behavioral detail beyond that: only visible, top-level, native windows are enumerated, and the response carries process IDs and UI Automation refs. It does not state pagination or ordering behavior, but with no output schema the return-shape hint is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the scope front-loaded and zero filler. Every clause (visible, top-level, native, returned IDs/refs) carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only enumeration tool with one optional parameter and no output schema, the description covers what is listed and roughly what comes back. The undocumented 'limit' parameter is the only meaningful omission, and safety semantics are already in the annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'limit' parameter has 0% schema description coverage and is not mentioned anywhere in the description. The agent must guess that it caps the number of returned windows rather than compensating from the text; the description does nothing to fill this gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (visible top-level native Windows application windows) plus the returned artifacts (process IDs, UI Automation refs). The scoping to 'top-level native' implicitly separates it from chrome_list_tabs/browser_tabs and from windows_snapshot, though it never names those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as windows_snapshot or windows_action. The agent must infer that this is the enumeration step preceding an action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

windows_screenshotA
Read-onlyIdempotent

Capture the visible bounds of a native Windows application window selected semantically by pid/title. Returns the PNG to the client and saves it under screenshots/.

ParametersJSON Schema
NameRequiredDescriptionDefault
pidNoProcess ID from windows_list
titleNoTop-level window title substring (alternative to pid)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive). Beyond that, the description discloses two useful traits: the PNG is returned to the client, and a copy is persisted under screenshots/ — a side effect not visible in the schema. It stops short of noting behaviors like occlusion/foreground requirements or failure when a window is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, and the core capability is front-loaded before the return/persistence note. Nothing is repeated from the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and zero required parameters, the description supplies the missing return-channel information (PNG returned and saved to screenshots/). It is nearly complete; the only unaddressed items are error/edge conditions such as a nonexistent pid or a minimized window.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters carry their own descriptions, so the schema does the heavy lifting. The description only echoes the pid/title alternative selection ('selected semantically by pid/title') without adding format or precedence detail beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Capture the visible bounds of a native Windows application window') and names the selection mechanism (pid/title). The qualifier 'native Windows application window' cleanly separates it from browser_screenshot, chrome_screenshot, and windows_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'native' implicitly signals when this tool applies instead of the browser/chrome screenshot siblings, and the pid reference to windows_list hints at a prerequisite. However, it never explicitly says when to choose this over browser_screenshot or record_clip, nor states any exclusion or ordering requirement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

windows_snapshotB
Read-onlyIdempotent

Semantic Windows UI Automation snapshot of a native app window. Returns stable runtime refs, control names/types, supported patterns and bounds. Prefer this over coordinate clicking.

ParametersJSON Schema
NameRequiredDescriptionDefault
pidNoProcess ID from windows_list
titleNoTop-level window title substring (alternative to pid)
maxElementsNo
includeOffscreenNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuine value by disclosing the return contents (stable runtime refs, control names/types, supported patterns, bounds), which no output schema exists to convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the mechanism and purpose, then return values, then the routing hint. Nothing is padded, but the routing sentence is a fragment that could be folded in.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, so listing return contents is the right move. Remaining gaps: no guidance on pid vs title selection strategy, and no explanation of the undocumented maxElements/includeOffscreen parameters. Adequate for invocation basics, incomplete on tuning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: pid and title are documented (including 'alternative to pid'), but maxElements and includeOffscreen have no schema description and the tool description says nothing about parameters at all. The description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Semantic Windows UI Automation snapshot of a native app window') and clarifies the underlying mechanism (UI Automation), which separates it from the pixel-based windows_screenshot. It stops short of naming the sibling it replaces explicitly, only gesturing at 'coordinate clicking'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Prefer this over coordinate clicking' gives a directional preference, but no when-not condition and no named alternative tool (windows_screenshot or windows_action are never mentioned). Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

windows_waitB
Read-onlyIdempotent

Wait for a semantic native Windows control to appear, become enabled, or disappear. Matches by accessible name substring, AutomationId and/or control type.

ParametersJSON Schema
NameRequiredDescriptionDefault
pidNoProcess ID from windows_list
nameNo
stateNopresent
titleNoTop-level window title substring (alternative to pid)
enabledNo
timeoutMsNo
controlTypeNo
maxElementsNo
automationIdNo
includeOffscreenNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds the semantic state model (present/enabled/gone), which is useful, but says nothing about polling behavior, what timeoutMs does on expiry, or what the call returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and immediately followed by the matching criteria. No filler, no repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 10-parameter blocking wait tool with no output schema and 20% schema coverage demands more than two sentences. Timeout semantics, default 15s behavior, match-combination logic, and return shape are all missing, leaving the agent unable to predict behavior on failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (just pid and title), leaving eight parameters undocumented. The description partially compensates by naming the match keys (accessible name substring, AutomationId, control type), but it never explains state (present/gone), enabled, timeoutMs, maxElements, or includeOffscreen, and the 'and/or' matching logic between the three keys is ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Wait for a semantic native Windows control') and enumerates the three wait conditions (appear, become enabled, disappear) plus the matching keys. It is clearly distinguishable from windows_list/windows_snapshot, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (call this before acting on a control that may not yet exist or be ready), but there is no explicit when-to-use or when-not guidance and no alternatives mentioned, e.g. using windows_snapshot to check current state instead of blocking.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_fileA
Destructive

Create or fully overwrite a file (parent folders are created). Prefer apply_patch/edit_file for changes to existing files.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFile path
checkNoRun compilers/linters on the changed files afterwards (default true)
contentYesComplete new file content
expectedSha256NoOptional optimistic precondition from file_info; refuse overwrite if the file changed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavioral context – parent folders are auto-created and the operation is a full overwrite rather than a merge – though it doesn't say what happens on failure or how the SHA precondition rejection surfaces.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste; the destructive overwrite semantics are front-loaded and the alternative-tool guidance follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with destructive annotations, a fully-covered schema, and no output schema, the description covers what it must: overwrite semantics, side effects, and routing to safer alternatives. Only failure behavior is unstated, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all four parameters including check and expectedSha256 are documented in the schema itself. The description adds no parameter-level detail beyond the overwrite scope, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create or fully overwrite a file') and clarifies scope with 'fully overwrite' plus the parent-folder side effect. It also names sibling alternatives, so an agent can distinguish it from apply_patch/edit_file without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Prefer apply_patch/edit_file for changes to existing files,' giving both the preferred alternative and the condition that selects it. This is an explicit when-not-to-use statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 128 tool updatesv3.0.1
    • First observedapply_patch
    • First observedbatch
    • First observedbrowser_click
    • First observedbrowser_close
    • First observedbrowser_console
    • First observedbrowser_dialog
    • First observedbrowser_evaluate
    • First observedbrowser_get_text
    • First observedbrowser_hover
    • First observedbrowser_navigate
    • First observedbrowser_press_key
    • First observedbrowser_record_start
    • First observedbrowser_record_stop
    • First observedbrowser_resize
    • First observedbrowser_screenshot
    • First observedbrowser_scroll
    • First observedbrowser_select_option
    • First observedbrowser_snapshot
    • First observedbrowser_tabs
    • First observedbrowser_type
    • First observedbrowser_upload_file
    • First observedbrowser_wait_for
    • First observedcapability_call
    • First observedcapability_list
    • First observedcheck_files
    • First observedcheckpoint_list
    • First observedcheckpoint_rewind
    • First observedchrome_call_site_tool
    • First observedchrome_click
    • First observedchrome_close_tab
    • First observedchrome_console
    • First observedchrome_dialog
    • First observedchrome_downloads
    • First observedchrome_evaluate
    • First observedchrome_extension_reload
    • First observedchrome_get_text
    • First observedchrome_list_tabs
    • First observedchrome_navigate
    • First observedchrome_new_tab
    • First observedchrome_pairing_setup
    • First observedchrome_press_key
    • First observedchrome_screenshot
    • First observedchrome_scroll
    • First observedchrome_snapshot
    • First observedchrome_status
    • First observedchrome_switch_tab
    • First observedchrome_type
    • First observedcopy_path
    • First observeddelete_path
    • First observededit_file
    • First observedfetch_url
    • First observedfile_info
    • First observedfind_files
    • First observedfind_symbol
    • First observedget_system_info
    • First observedgit_commit
    • First observedgit_diff
    • First observedgit_log
    • First observedgit_status
    • First observedjob_cancel
    • First observedjob_list
    • First observedjob_output
    • First observedjob_start
    • First observedjob_status
    • First observedkill_process
    • First observedlist_directory
    • First observedlist_processes
    • First observedlist_workspaces
    • First observedlsp_calls
    • First observedlsp_code_actions
    • First observedlsp_definition
    • First observedlsp_diagnostics
    • First observedlsp_format_preview
    • First observedlsp_hover
    • First observedlsp_references
    • First observedlsp_rename_preview
    • First observedlsp_status
    • First observedlsp_symbols
    • First observedmove_path
    • First observedoutline
    • First observedprocess_input
    • First observedprocess_list
    • First observedprocess_output
    • First observedprocess_start
    • First observedprocess_stop
    • First observedproject_context
    • First observedproject_memory
    • First observedpty_list
    • First observedpty_read
    • First observedpty_resize
    • First observedpty_start
    • First observedpty_stop
    • First observedpty_write
    • First observedread_file
    • First observedread_many_files
    • First observedrecord_clip
    • First observedrepo_map
    • First observedresource_status
    • First observedrun_command
    • First observedsearch_code
    • First observedtask_cancel
    • First observedtask_claim
    • First observedtask_contract
    • First observedtask_events
    • First observedtask_heartbeat
    • First observedtask_list
    • First observedtask_next
    • First observedtask_ready
    • First observedtask_reconcile
    • First observedtask_stats
    • First observedtask_status
    • First observedtask_submit
    • First observedtask_update
    • First observedtask_verify
    • First observedvscode_add_breakpoint
    • First observedvscode_debug_state
    • First observedvscode_diagnostics
    • First observedvscode_find_references
    • First observedvscode_go_to_definition
    • First observedvscode_open_file
    • First observedvscode_open_files
    • First observedweb_search
    • First observedwindows_action
    • First observedwindows_list
    • First observedwindows_screenshot
    • First observedwindows_snapshot
    • First observedwindows_wait
    • First observedwrite_file

TDQS

C2.9/5.0

Scored across 128 tools

Disambiguation2/5

Several large tool clusters overlap heavily: run_command/process_start/pty_start/job_start all run or manage commands, chrome_* and browser_* duplicate browser automation, and lsp_* and vscode_* duplicate language-server queries. Descriptions distinguish some contexts (real Chrome vs Playwright, background process vs PTY vs durable job), but an agent could easily misselect among them.

Naming Consistency4/5

The dominant convention is snake_case with consistent domain prefixes (browser_, chrome_, process_, task_, lsp_, vscode_), which makes the set mostly predictable. Minor deviations exist, such as list_processes vs process_list, get_system_info, file_info, record_clip, batch, and run_command, preventing a perfect consistency score.

Tool Count1/5

128 tools is an extreme over-provisioning for a single coding-agent server. Even with broad ambitions, the surface is far beyond the typical 3-15 well-scoped range and directly contributes to selection ambiguity and maintenance overhead.

Completeness4/5

The server covers a remarkably broad coding-agent lifecycle: file read/write/patch/edit, search, git status/diff/log/commit, shell, process/job/PTY control, LSP diagnostics and navigation, browser and real-Chrome automation, Windows UI automation, durable tasks, checkpoints, and web access. Some common operations are missing or only reachable via shell, such as advanced git workflows (push/pull/branch/PR) and project scaffolding, so it is strong but not fully complete.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    D
    quality
    C
    maintenance
    The intelligent execution layer for coding agents, exposed as an MCP server for high-stakes engineering projects. It enables AI agents to manage plans, tasks, and integrations via tool calls.
    25
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Secure local development platform that exposes controlled developer capabilities (FS, Git, search, command execution) to AI assistants via MCP with deny-by-default security and audit logging.
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI clients to securely control and interact with a local Windows machine through 218 configurable tools for files, Git, processes, Windows UI, browser automation, WSL, Office, recovery, skills, and child MCP servers.
    11 npm
    MIT