AgentHub
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AgentHubdelegate codex to add input validation to the signup handler and run its tests"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AgentHub

▶ Watch the demo in full quality (MP4, 43 s, with voiceover): a real run, not a mock-up.
Let Claude Code hand work to other coding agents (Codex, Antigravity, Claude, Aider, Goose, or any CLI) without handing them your whole machine.
AgentHub is a small MCP server, CLI and optional HTTP API. Use it to:
Get a second opinion. Send one question to several agents and compare the answers (
compare).Get a cross-model review of your diff before you commit (
review).Delegate safely. An agent works in its own git worktree. You see the diff, then apply or discard it.
Use the subscriptions you already pay for. ChatGPT through Codex, Google through Antigravity, all from one place, with automatic fallback when one runs out of quota.
Run long jobs in the background, wait for them, read readable logs.
Search other agents' past sessions, or continue them.
Every request goes through one policy layer: trusted folders, permission modes, a minimal environment, timeouts, quota tracking and a hash-chained audit log.
Pure Python standard library. No dependencies. Linux and macOS. About 20 MB of RAM.
Quick start
pipx install agenthub-gateway # or: uv tool install agenthub-gateway
# or run without installing: uvx agenthub-gateway doctor
agenthub config --init # writes ~/.agenthub/config.json
$EDITOR ~/.agenthub/config.json # set "trusted_workspaces": ["~/code"]
agenthub doctorThen connect it to Claude Code in one of two ways:
# A) Plugin: MCP server plus slash commands (/second-opinion, /cross-review, /delegate, /agent-tasks)
# Starts the server with `uvx agenthub-gateway mcp`, so it needs uv (https://docs.astral.sh/uv/).
claude plugin marketplace add premanand8800/agenthub-mcp
claude plugin install agenthub@agenthub
# B) MCP server only
agenthub install-claudeRestart Claude Code and try:
/cross-review codex
/delegate codex add input validation to the signup handler and run its tests
Use agenthub to compare codex and claude on: what's the safest way to migrate this table?
Related MCP server: Hydra
How a delegated change works
start_task(isolation="worktree") → agent edits a private checkout (your files are untouched)
wait_task → blocks until done; returns final output, session_id, token usage
get_task_diff → file list + patch
review(task_id=…) → optional: another model reviews that patch
apply_task | discard_task → patch lands in your working tree (uncommitted), or is thrown awayThe worktree starts from your current state, including uncommitted changes to tracked files. Build junk (__pycache__, node_modules, …) is left out of the diff. If your files changed in the meantime, apply_task tries a 3-way merge. If that conflicts, it keeps the worktree so you can resolve it by hand.
Permission modes
Mode | Meaning | Codex | Claude | Antigravity | Aider | Goose |
| Inspect only |
|
|
|
|
|
| Edit files in |
|
|
| no shell commands | – |
| No sandbox, no approvals |
|
|
|
|
|
fullis off until a human sets"allow_full_access": true. A model cannot turn it on.reviewandcomparealways runread-only.If an agent can't enforce a mode, AgentHub refuses the request instead of quietly running with weaker settings.
A delegated Claude runs with
--strict-mcp-config, so it can't use your other MCP servers (email, chat, AgentHub itself).
MCP tools
Tool | What it does | Read-only |
| Installed agents, modes, features, quota health | yes |
| Run a prompt and wait. Returns | no |
| Background job, optionally | no |
| Block until a task finishes (up to 10 min per call) | yes |
| Status, final output, readable logs | yes |
| Kill a task and all its child processes | no |
| What a worktree task changed | yes |
| Land or drop a worktree task's changes | no |
| Read-only review of | yes* |
| Same prompt to 2–5 agents in parallel | yes* |
| Other agents' history | yes |
| Model IDs per agent | yes |
| Message a live session | no |
* These don't change your files, but they do send code to the agents' vendors and use quota.
Options for ask and start_task: workdir, permission_mode, model, and session_id (continue a conversation). Also:
add_dirs: extra directoriesimages: Codex onlyoutput_schema: a JSON Schema; the answer comes back parsed instructuredfallback_agents: other agents to try if this one is out of quota or not installed
Long calls send MCP progress notifications every 10 s when the client asks for them.
To skip permission prompts for tools that only read:
// ~/.claude/settings.json
{ "permissions": { "allow": [
"mcp__agenthub__list_agents", "mcp__agenthub__wait_task", "mcp__agenthub__get_task",
"mcp__agenthub__get_task_logs", "mcp__agenthub__list_tasks", "mcp__agenthub__get_task_diff",
"mcp__agenthub__list_sessions", "mcp__agenthub__search_sessions", "mcp__agenthub__get_transcript",
"mcp__agenthub__list_models"
] } }Reliability
Quota-aware. When a provider reports a quota or rate limit, AgentHub records it, including the reset time when the message gives one ("Resets in 137h"). That agent then fails fast with
quota_exhausteduntil the reset time.fallback_agentsmoves on to the next agent automatically.list_agentsshows each agent's health.No unsafe retries. An agent run that failed partway may already have edited files, so AgentHub never retries it automatically. Falling back to another agent only happens when nothing ran.
Survives restarts. Background tasks keep running if Claude Code closes, and their status still works from any process. In-flight
askcalls are killed when the client disconnects, so nothing keeps spending.Global limits.
max_concurrent_tasksholds across every Claude Code session, enforced with a file lock.
CLI (also built for agents)
Every command is non-interactive, has --json, and uses stable exit codes. Other agents can call AgentHub from their own terminals.
agenthub agents [--json]
agenthub ask codex "explain main.py" -m read-only [--session ID] [--schema s.json] [--fallback claude]
agenthub run codex "add tests" --isolation worktree → prints task ID
agenthub wait ID [--timeout 600] exit 0 succeeded · 1 failed · 3 still running
agenthub diff ID | apply ID | discard ID
agenthub review codex [--base main | --task ID]
agenthub compare codex,claude "which approach is safer?"
agenthub tasks | task ID | logs ID | cancel ID
agenthub audit [--verify] · config [--init] · token [--rotate] · serve · doctor
agenthub approve antigravity "Bash(npm test)" human-onlyExit codes: 0 ok · 1 the agent or task failed · 2 usage error · 3 still running (for wait).
Configuration
~/.agenthub/config.json. Every key is optional. Unknown keys are an error, so typos don't fail silently. Set AGENTHUB_HOME to move the whole directory.
Key | Default | |
|
| Agents may only run inside these. Narrow this. |
|
| Always refused |
|
| |
|
| Allows |
|
| Extra env vars for agents (globs), e.g. |
|
|
|
|
| |
|
| Across all sessions |
|
| |
|
| Stops agent → hub → agent loops |
|
| How long to skip an agent after a quota error with no reset time |
|
|
|
|
| Old tasks and their worktrees are deleted |
|
| Browser origins allowed to call the HTTP API |
Environment: agents get only what they need:
PATH,HOME, locale, proxies, CA bundlestoolchain variables (nvm, cargo, go, java, venv, …)
their own API-key variables (
OPENAI_*for Codex,ANTHROPIC_*for Claude, …)
Cloud credentials such as AWS_* and GH_TOKEN are not passed unless you list them in env_passthrough.
Custom agents
Put a JSON file in ~/.agenthub/agents/ (mode 600). See examples/custom-agent.json.
{
"name": "myagent",
"command": "myagent",
"args": ["run", "--prompt={prompt}"],
"model_args": ["--model", "{model}"],
"add_dir_args": ["--add-dir", "{dir}"],
"output_schema_args": ["--schema", "{schema_file}"],
"json_output": { "reply": "result", "session_id": "session_id", "cost_usd": "cost" },
"env_passthrough": ["MYAGENT_API_KEY"],
"modes": { "read-only": ["--no-write"], "workspace-write": ["--sandbox"] }
}modeslists only what your CLI can really enforce.json_outputtells AgentHub where to find the reply and session ID when your CLI prints JSON.Specs that are group- or world-writable are refused, and so are specs that reuse a built-in name.
agenthub doctorreports why.
HTTP API (optional)
agenthub serve # 127.0.0.1:8765
TOKEN=$(agenthub token)
curl -s localhost:8765/v1/agents/codex/ask -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"prompt":"hi","permission_mode":"read-only"}'Method | Path |
GET |
|
POST |
|
GET |
|
POST |
|
GET |
|
POST |
|
Errors are always {"ok": false, "error": {"code", "message"}}. A quota_exhausted error also includes retry_after_seconds.
Security model
See SECURITY.md. In short:
The model is untrusted. Defaults assume a prompt injection will reach AgentHub. Every agent run is sandboxed, limited to trusted folders, given a minimal environment, time-limited and audited. A model can't enable
fullmode or approve its own tools.No shell, ever. Prompts can't become CLI options.
Private state. All state lives in
~/.agenthub(0700), never in/tmp.Tamper-evident audit log.
agenthub audit --verify.
Development
git clone https://github.com/premanand8800/agenthub-mcp && cd agenthub-mcp
PYTHONPATH=src python -m unittest discover -s tests -t tests -vTests use a fake agent: no API keys, no quota. To add a built-in adapter, subclass agenthub.adapters.base.Adapter, then:
implement
build_command(spec), andparse_outputif your CLI prints JSONdeclare
modesandenv_allowregister it in
adapters/__init__.py
Contributing
Issues and PRs welcome — see CONTRIBUTING.md. Security reports: SECURITY.md.
License
MIT
Available Tools
18 toolsapply_taskApply task changesBDestructive
Apply a worktree-isolated task's changes to the user's working tree (not committed).
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | task_id from start_task | |
| keep_worktree | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is covered. The description adds real context by stating changes land uncommitted in the working tree and come from a worktree-isolated task, but it omits whether existing uncommitted edits are overwritten, whether the worktree survives, and how failures surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the effect on the working tree and its uncommitted status are both stated up front. Nothing is repeated from annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema and one undocumented parameter, the description covers the core effect but not enough: conflict/overwrite behavior, worktree fate, and the keep_worktree choice are all missing. Annotations carry the safety signal, which keeps this from being worse.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% and the description adds nothing about parameters, leaving keep_worktree with no semantics in either place — an agent cannot tell whether it keeps the worktree after applying or deletes it. The word 'worktree-isolated' gestures at the concept but never explains the flag.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (apply) and resource (a worktree-isolated task's changes) plus the exact destination (the user's working tree, not committed), so scope is unambiguous. It does not differentiate itself from siblings such as get_task_diff or discard_task, which would have pushed it to a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given: it never says at what point in the start_task → wait_task → apply_task lifecycle this should be called, nor when to prefer discard_task instead. The only procedural hint is the schema's 'task_id from start_task', which is not in the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
askAsk an agentBDestructive
Run a prompt on an agent and wait for the answer (minutes). Returns reply, session_id, usage.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. | |
| model | No | Model ID from list_models. | |
| images | No | Absolute image paths to attach (agents with supports.images). | |
| prompt | Yes | Instructions for the agent. | |
| workdir | No | Absolute project directory. Default: first trusted workspace. | |
| add_dirs | No | Extra absolute directories the agent may use. | |
| session_id | No | Continue this session (returned by earlier calls). | |
| output_schema | No | Optional JSON Schema; the reply is returned parsed as `structured`. | |
| fallback_agents | No | Agents to try in order if this one is out of quota or unavailable. | |
| permission_mode | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, and non-idempotent, so the agent knows this is a risky, side-effecting operation. The description adds the important blocking/waiting behavior and return fields (reply, session_id, usage), which complements annotations. However, it doesn't explain quota fallback, permission implications, or what makes it destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that conveys the core behavior (run prompt, wait) and return values. No wasted words; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no output schema and destructive/open-world annotations, the description is too thin. It omits session continuation behavior, fallback_agents semantics, output_schema handling, and permission_mode implications—all of which an agent would need to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 90%, so the schema already documents nearly all parameters including fallback_agents, session_id, and output_schema. The description adds no parameter-specific meaning beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource: run a prompt on an agent and wait for the answer. It distinguishes itself from siblings like start_task/wait_task by emphasizing synchronous blocking behavior ('wait for the answer (minutes)'), though it doesn't explicitly mention the async alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus siblings like start_task or send_message. The parenthetical '(minutes)' hints at synchronous blocking but doesn't state when to prefer it over async task tools or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_taskCancel taskBDestructiveIdempotent
Stop a running task and every process it started.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | task_id from start_task |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=false and readOnlyHint=false, so the safety profile is covered. The description adds genuinely useful scope information — the cascade kill of child processes — which the annotations do not convey. It does not say what happens to the task's logs, artifacts, or exit state, or what a repeat call against an already-finished task returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. The most consequential fact — that child processes are killed too — is included without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation tool with full annotation coverage and no output schema, the description covers the core action and its blast radius. It omits behavioral edge cases an agent may need: whether it only works on running tasks, what a repeated call returns given idempotentHint, and what is left behind after cancellation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter with 100% schema coverage; the schema already documents task_id and ties it to start_task. The description adds no extra meaning about the parameter, so the baseline 3 for high-coverage schemas is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Stop') and resource ('task'), and adds scope: the task plus 'every process it started'. That scope detail lets an agent distinguish it from sibling tools like discard_task or wait_task. It stops short of naming an alternative, so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'a running task' weakly implies the precondition that the task must be active, but there is no explicit when-to-use or when-not-to-use guidance. Siblings such as discard_task and wait_task are never mentioned, so the agent must infer which lifecycle tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compareCompare agentsBRead-only
Send one prompt to 2-5 agents in parallel (read-only) and get their answers side by side.
| Name | Required | Description | Default |
|---|---|---|---|
| agents | Yes | 2 to 5 agent names | |
| prompt | Yes | ||
| workdir | No | ||
| output_schema | No | Optional JSON Schema; the reply is returned parsed as `structured`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=true and idempotentHint=false. The description adds concurrency ('in parallel'), the read-only qualifier and the 2-5 bound, which is real context. It does not explain the idempotentHint=false implication or what the side-by-side return looks like, so the added value is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action, the fan-out count, the safety qualifier and the result shape. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Safety is covered by annotations and there is no output schema to explain, but for a 4-parameter tool with a nested `output_schema` object, the description omits what `workdir` controls and how the parallel answers are returned. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: `agents` and `output_schema` are documented in the schema, while `prompt` and `workdir` are not. The description repeats the '2-5 agents' bound already in the schema and says nothing about `workdir` or how `output_schema` changes the response shape, so it does not compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (send one prompt) and resource (2-5 agents), plus the output shape (answers side by side). It implicitly contrasts with the single-agent sibling `ask`, but never names it, so an agent must infer the distinction from the tool list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parallel multi-agent framing implies the use case (comparing several agents on one prompt), but there is no explicit when-to-use/when-not guidance and no mention of `ask` as the single-agent alternative. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discard_taskDiscard task changesADestructiveIdempotent
Throw away a worktree-isolated task's changes (cancels it if running).
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | task_id from start_task |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: the scope is worktree-isolated changes, and the task is also cancelled if it is currently running, which is a non-obvious side effect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence, front-loaded with the core action and parenthetically noting the conditional cancellation. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, destructive tool with no output schema and full annotation coverage, the description conveys what it does, the isolation scope, and the running-task side effect. It could say more about what happens to the worktree or whether the task becomes unretrievable, but it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single task_id parameter is documented as coming from start_task. The description adds nothing about parameter format or sourcing, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (throw away) plus resource (a worktree-isolated task's changes), and it clarifies the interaction with cancellation. It distinguishes itself reasonably from cancel_task and get_task_diff, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or comparison to alternatives like cancel_task, apply_task, or get_task_diff. Usage is only inferable from the phrase 'throw away ... changes'; the agent is left to guess when discarding is preferable to cancelling or applying.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_taskGet taskBRead-only
Task status and, once finished, its final output, session_id and usage.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | task_id from start_task |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds genuine context by explaining the return is state-dependent (output only once finished), which reconciles with idempotentHint=false. It omits what happens for an unknown or still-running task_id, so it adds moderate rather than rich context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight clause with no filler, and the state-conditioned nature of the output is stated up front. It is arguably a touch under-specified rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does carry the burden of describing returns and does enumerate them (status, final output, session_id, usage). Missing only error/edge-case behavior for an unknown or pending task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter and the schema already explains task_id comes from start_task. The description adds nothing about the identifier's format or lifetime, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific resource (task) and enumerates the returned payload (status, final output, session_id, usage), which separates it from siblings like list_tasks and get_task_logs. The verb is only implied by the name, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus wait_task, list_tasks, get_task_logs, or get_task_diff, and no note about whether it is meant for polling. Usage must be inferred from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_task_diffGet task diffBRead-onlyIdempotent
Changes a worktree-isolated task made (file list, stat and patch).
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | task_id from start_task | |
| max_chars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and non-destructive, so the safety profile is covered. The description usefully adds the return shape (file list, stat and patch), which matters since there is no output schema, but it says nothing about truncation via max_chars or the state a task must be in.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the resource and then the payload contents. No filler, though it is arguably too terse given the missing usage and truncation context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does partially cover the return value (file list, stat, patch). However it omits when the tool is applicable and how max_chars affects output, leaving meaningful gaps for a tool that can return large patches.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: task_id is documented in the schema, but max_chars has no description anywhere. The description does not mention the character cap or what happens when the patch exceeds it, so it fails to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (the diff of changes a worktree-isolated task made), and enumerates what the diff contains (file list, stat, patch). This distinguishes it from siblings like get_task and get_task_logs, though it doesn't explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives, nor any precondition such as the task being finished or isolated to a worktree. The reader must infer usage entirely from the phrase 'worktree-isolated task'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_task_logsGet task logsBRead-only
Last lines of a task's log (readable summary of commands, edits and messages).
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | task_id from start_task | |
| tail_lines | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and closed-world scope. The description adds that output is a 'readable summary of commands, edits and messages' rather than raw output, which is genuinely useful, but it says nothing about truncation limits, defaults, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the scope ('last lines') front-loaded and no filler. It is efficient, though the parenthetical could be folded into a clearer statement of return format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry the return-value burden; it does so only loosely by calling the log a 'readable summary'. Default tail size, truncation behavior, and the relationship to other log/transcript tools are left unaddressed, which is a meaningful gap for a 2-param retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: task_id is documented ('task_id from start_task') but tail_lines has no description, no default, and no stated maximum in the description. 'Last lines' hints at tail behavior but does not tell the agent how many lines are returned by default or how tail_lines bounds the result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (a task's log) and its scope (last lines), with a parenthetical clarifying the content type. It is clear what the tool returns, though it does not explicitly distinguish itself from siblings like get_task or get_task_diff.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus get_task, get_task_diff, or get_transcript. The 'last lines' scope implies a tail-style read, but no conditions, prerequisites, or alternatives are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptGet transcriptCRead-onlyIdempotent
Turns (messages, tool calls) of one past session.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. | |
| max_steps | No | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description usefully adds that the payload contains both messages and tool calls, but says nothing about truncation via max_steps or whether the session must be completed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but the terseness comes at the cost of substance: a single verbless fragment carries no structure, no return-format note, and no constraint information. This reads as under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 33% param coverage, the description is the only place return shape and parameter behavior could be conveyed, yet it gives one fragment. An agent cannot tell what a transcript entry looks like or how max_steps affects the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%: only 'agent' is documented in the schema, while session_id and max_steps are bare. The description adds no parameter meaning at all — notably it never explains that max_steps caps the number of steps returned — so it fails to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Turns (messages, tool calls) of one past session' identifies the resource (a session's message/tool-call transcript) but uses no verb like retrieve/get and omits 'transcript' itself, leaving the action implied. It partially distinguishes from list_sessions/search_sessions by scoping to 'one past session', but an agent must infer the operation from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisite (e.g. obtain session_id from list_sessions), and no alternative named such as search_sessions. Usage is only implied by the phrase 'one past session'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agentsList agentsBRead-onlyIdempotent
Installed agents, their permission modes, supported features and quota health.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is fully covered. The description contributes useful context about what is returned (permission modes, features, quota health), but adds nothing about freshness, pagination, or scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight fragment with zero filler, and the most informative content (the returned attributes) is front-loaded. It is under-specified rather than verbose, which is a content problem, not a structural one.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description is the only source of return-value information, and it names the fields but not the structure or ordering. Adequate for a simple enumerator, but an agent still cannot predict the response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. The enumerated fields do hint at the shape of results, which slightly compensates for the absent output schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'Installed agents' names a specific resource and the enumerated fields (permission modes, supported features, quota health) make the retrieval intent clear, distinguishing it from list_models/list_tasks. It lacks an explicit verb, but the name and sibling set make 'list' unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives like list_models or list_tasks, nor any stated precondition (e.g., that it reflects only locally installed agents). Usage is implied by the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList modelsBRead-onlyIdempotent
Model IDs an agent accepts for model.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered by structured data. The description's one added behavioral fact is that the returned IDs are the accepted values for an agent's `model` parameter, which is genuinely useful but thin. It adds nothing about return shape or volatility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with zero padding, which is appropriate for a trivial single-parameter read. It is a sentence fragment rather than a complete statement, which slightly weakens clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description is the only place the return value is described, and 'Model IDs' is adequate but does not confirm the shape (a list of strings) or whether it is ordered or exhaustive. Combined with full schema coverage and complete annotations, this is just enough to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single `agent` parameter, so the schema already documents that it takes an agent name from list_agents. The description adds the linkage between the output and the `model` parameter, but no further format or constraint detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment identifies the resource (model IDs) and the scope (those an agent accepts), which is a specific verb-implied purpose that an agent can act on. It also implicitly distinguishes itself from list_agents by keying on models rather than agents, though the verb 'list/return' is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no naming of alternatives, and no statement of sequencing (e.g., call before setting the `model` parameter). The link to the `model` parameter hints at usage but leaves the agent to infer it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sessionsList sessionsCRead-onlyIdempotent
An agent's recent conversation sessions.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. | |
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description only adds that results are 'recent', hinting at an implicit recency window, but it does not state ordering, default count, or pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence-fragment with no padding; nothing is wasted. It is terse rather than bloated, though its fragment form leaves the action implicit.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations covering return content, the description should describe what a session record contains, its ordering, and how limit behaves. It provides none of these, leaving an agent unable to predict the result set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: 'agent' is documented while 'limit' has no description at all. The description's 'recent' vaguely gestures at bounded results but supplies no default, maximum, or ordering to compensate for the undocumented limit parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('conversation sessions') and scopes it to an agent, but it is a noun fragment with no verb and adds no differentiation from the sibling search_sessions or get_transcript. An agent can infer the action from the tool name, but the description itself doesn't establish a distinct purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative routing is provided. The existence of search_sessions as a sibling is never addressed, so the agent gets no help deciding between browsing and searching.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksList tasksCRead-only
Recent background tasks, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds ordering (newest first) and a recency bias, but omits the default limit, the 500 cap behavior, and whether results are truncated or paginated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the ordering constraint front-loaded and no wasted words. It is efficient, though its brevity is partly the source of the missing detail rather than a virtue in itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 0% schema coverage, an enum parameter, and no output schema, the description should explain filtering by status and the meaning of limit. It does neither, leaving the tool under-specified for a filterable list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two parameters: limit and status. The description never mentions either, so an agent gets no guidance on the status enum values or on how limit interacts with 'recent' results. It fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource (background tasks) and an ordering guarantee (newest first), but the action verb is only implied by the name and there is no differentiation from siblings like get_task, wait_task, or list_agents. An agent can tell it lists something recent, but the exact scope is vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to use this versus get_task, wait_task, or list_agents. There is no mention of whether this is a global list or session-scoped, nor any prerequisite or exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reviewReview changesARead-only
Ask an agent to review a diff (read-only): uncommitted changes vs base (default HEAD), or a task's worktree changes via task_id.
| Name | Required | Description | Default |
|---|---|---|---|
| base | No | git ref, e.g. main | |
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. | |
| model | No | ||
| task_id | No | task_id from start_task | |
| workdir | No | ||
| instructions | No | ||
| output_schema | No | Optional JSON Schema; the reply is returned parsed as `structured`. | |
| fallback_agents | No | Agents to try in order if this one is out of quota or unavailable. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description repeats '(read-only)' and adds only the default base (HEAD) and the task-worktree mode; it says nothing about agent invocation latency, quota/fallback behavior, or what a review returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one front-loaded sentence with no filler. It delivers the tool's action, safety note, and both operating modes compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no output schema and no annotation coverage of return behavior, the description explains the two core modes but omits what the review returns, how agent/fallback selection works, and any constraints around model, workdir, or instructions. It is adequate but leaves important context gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 63%, so the baseline is 3. The description adds meaning for base (default HEAD) and task_id (task worktree changes), but leaves agent, model, workdir, instructions, output_schema, and fallback_agents entirely to the schema or unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource ('review a diff') and scopes it with two modes: uncommitted changes vs base, or a task's worktree changes. It is distinct from simply fetching a diff, but it never names any sibling tool (e.g., get_task_diff or ask) to make the differentiation explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent when to use which mode: base for uncommitted changes (default HEAD) or task_id for a task's worktree. It does not provide exclusions or name an alternative tool for when a simple diff retrieval is sufficient, so it falls short of the 5-level 'when-not/alternatives' standard.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_sessionsSearch sessionsCRead-onlyIdempotent
Search an agent's past sessions by keyword.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. | |
| limit | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false, covering the safety profile. The description adds almost no behavioral context beyond the basic operation—it does not describe matching behavior, result limits, pagination, or what a session search covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It is efficient, though its extreme brevity leaves important details unstated for a 3-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, annotations covering mostly safety, and only 33% schema description coverage, the description is too incomplete. It omits usage guidance, parameter details for limit and query, and any return or pagination context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description must compensate but does not. It vaguely maps 'agent's past sessions' to agent and 'keyword' to query, but says nothing about the limit parameter, its default, its maximum of 100, or how query matching works.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (an agent's past sessions) with the keyword filter. It is clear what the tool does, but it does not distinguish itself from the sibling list_sessions or explain how keyword search differs from listing sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus list_sessions, get_transcript, or other siblings. The description implies a keyword-search use case but provides no conditions, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_messageSend message to sessionCDestructive
Send a message into an agent's live interactive session.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. | |
| message | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, openWorldHint=true, and idempotentHint=false, but the description never explains what destruction is possible, whether the message interrupts or queues behind an in-flight turn, or whether the call blocks for a reply. 'Live interactive session' hints at real-time delivery but adds little beyond the annotation set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler or repetition. It is efficient, though arguably under-specified rather than appropriately sized for a three-parameter mutation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent write with no output schema and two undocumented required parameters, the description omits what the agent should expect: whether the session must be active, whether the call waits, and what happens on failure. It is too thin for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: 'agent' is documented in the schema while 'message' and 'session_id' have no descriptions anywhere. The description adds no format, source, or constraint information to compensate for the two undocumented required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and target: send a message into an agent's live interactive session. It is clearer than a tautology and identifies the resource (interactive session), but it does not distinguish itself from siblings like 'ask' or 'apply_task' which plausibly also deliver input to an agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites (e.g. an existing session from list_sessions), and no mention of when to prefer this over 'ask' or 'apply_task'. The agent must infer the routing from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_taskStart background taskADestructive
Start long agent work in the background; returns task_id. Use isolation='worktree' for code changes.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Agent name from list_agents, e.g. 'codex' or 'claude'. | |
| model | No | Model ID from list_models. | |
| images | No | Absolute image paths to attach (agents with supports.images). | |
| prompt | Yes | Instructions for the agent. | |
| workdir | No | Absolute project directory. Default: first trusted workspace. | |
| add_dirs | No | Extra absolute directories the agent may use. | |
| isolation | No | worktree = private git checkout; review with get_task_diff, then apply_task or discard_task. | |
| session_id | No | Continue this session (returned by earlier calls). | |
| output_schema | No | Optional JSON Schema; the reply is returned parsed as `structured`. | |
| fallback_agents | No | Agents to try in order if this one is out of quota or unavailable. | |
| permission_mode | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, openWorldHint=true and idempotentHint=false, so the safety profile is covered. The description usefully adds that the call is asynchronous and returns a task_id (no output schema exists), but says nothing about what the agent can actually modify, permission_mode effects, or quota/fallback behavior implied by the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core purpose and the async/task_id behavior front-loaded before the parameter tip. Nothing is padded or restated from the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter, destructive, open-world dispatch tool with no output schema, the description is thin: it never warns that the spawned agent may write to the workspace, and it omits the session-continuation and fallback flows. The rich schema compensates for parameter details, but the behavioral gaps keep this at a minimum-viable level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 91%, so the schema already carries parameter meaning; the description only adds a selection cue for the isolation enum ('worktree for code changes'). It adds nothing for the 10 other parameters, including the undocumented permission_mode enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start long agent work in the background') plus the key return value (task_id), which cleanly separates it from read/observation siblings like get_task, list_tasks and wait_task. It stops short of naming a sibling or drawing the explicit contrast with a synchronous 'ask'-style call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for this tool: long-running work that should run in the background rather than block, and a concrete conditional ('isolation=worktree for code changes'). It provides no explicit exclusions or named alternatives, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_taskWait for taskARead-only
Wait until a task finishes (or timeout_seconds passes) and return its result.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | task_id from start_task | |
| timeout_seconds | No | Default 60. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true and destructiveHint=false already covering safety, the description adds genuine context annotations cannot: that the call blocks until completion or a timeout, and that it returns the result rather than just a status. It omits what actually happens when timeout_seconds expires (error, partial result, or status payload), and does not explain the non-idempotentHint=true flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the blocking behavior and the timeout condition, followed by the return value. No filler or redundant restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter blocking tool with no output schema, the entry point (task_id from start_task) and the timeout bound are covered, but the two things an agent most needs for a wait primitive are missing: the outcome when the timeout expires and any hint of the result shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: task_id is documented as coming from start_task and timeout_seconds as having a default of 60 with a 1-600 range. The description only restates that timeout_seconds can elapse, adding no format or usage detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('wait' + 'task') with the terminating conditions stated inline: task finish or timeout_seconds elapsing, plus it returns the result. It is clear what the tool does, but it never names or contrasts with the obvious alternative get_task (polling), so an agent must infer the distinction from the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The blocking semantics imply 'use this instead of repeatedly polling get_task', but that is left to inference. There is no explicit when-to-use, when-not-to-use, or named alternative, so guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v1.2.2- First observed
apply_task - First observed
ask - First observed
cancel_task - First observed
compare - First observed
discard_task - First observed
get_task - First observed
get_task_diff - First observed
get_task_logs - First observed
get_transcript - First observed
list_agents - First observed
list_models - First observed
list_sessions - First observed
list_tasks - First observed
review - First observed
search_sessions - First observed
send_message - First observed
start_task - First observed
wait_task
TDQS
Scored across 18 tools
Most tools have clearly distinct purposes: the task lifecycle set (start/wait/get/logs/list/cancel/diff/apply/discard) is well-partitioned by resource and action, and list/search/get_transcript for sessions are clearly differentiated. Minor overlap risk exists between ask, start_task, and send_message (all send prompts, differing mainly by sync/async/interactive mode), which descriptions largely resolve.
Nearly all names follow a consistent snake_case verb_noun pattern (list_agents, get_task_diff, cancel_task, search_sessions, send_message). A few single-verb names (ask, review, compare) deviate slightly from the pattern but remain readable and unambiguous.
18 tools is on the heavier side but justified: nine tools cover the full background-task lifecycle plus agent, session, and model operations. Each tool has a distinct role, so nothing appears redundant or padded.
Coverage is strong: agent discovery, full task lifecycle (create, monitor, diff, apply, discard, cancel), session browsing/search/transcript, and model listing. Minor gaps exist around agent lifecycle management (list_agents implies installation but there's no install/remove tool) and cross-agent model discovery.
Maintenance
Related MCP Connectors
- ParleyOAuthdev.weldra
Coordination hub for AI coding agents: message teammates, ask humans, audit every event.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Security reviews for coding agents: diffs checked against your org policy and live infrastructure.
Governance layer for AI coding agents: knowledge-graph grounding, session audit, policy controls.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables Claude to delegate tasks to external coding agents (Codex or Antigravity) for independent reviews, separate quota usage, and async processing.6MIT
- AlicenseNot gradedqualityBmaintenanceEnables Codex to delegate bounded engineering jobs to Claude Code CLI in isolated Git worktrees with strict security and allowance pacing.MIT
- FlicenseAqualityBmaintenanceEnables Claude Desktop to delegate specialized tasks like security scanning and code review to external AI agents from multiple providers, with dynamic agent registration, flexible pipelines, and safety controls.9-
- AlicenseNot gradedqualityBmaintenanceEnables delegating coding tasks to sandboxed agents, returning changes as auditable diffs for human review before applying or discarding them.1MIT