Skip to main content
Glama

agentmux

English | 한국어

Use Codex, Claude Code, and Antigravity as each other's subagents — without leaving the coding-agent UI you already use.

npm CI License

agentmux is a local MCP orchestration layer for coding-agent CLIs. Keep Codex, Claude Code, or Antigravity as your control tower, then delegate work to other installed providers while preserving their native sessions.

Quick start

Requirements: Node.js 20+ and at least one supported provider CLI installed (codex, claude, or agy).

Install agentmux as a native plugin in every supported coding-agent host already installed on your machine:

npx -y '@jiho.ko/agentmux@latest' plugins install

Check the result:

npx -y '@jiho.ko/agentmux@latest' plugins status

Then fully exit and relaunch Codex, Claude Code, or Antigravity so the plugin's MCP tools and orchestration skill are loaded. Opening only a new chat/thread may not reload plugin-provided MCP servers.

If you previously registered agentmux as a direct MCP server, migrate to the native plugin and remove the duplicate registration only after plugin installation succeeds:

npx -y '@jiho.ko/agentmux@latest' plugins install --replace-mcp

Related MCP server: all-agents-mcp

Why agentmux?

  • Keep your existing agent UI. No separate multi-agent dashboard is required; Codex, Claude Code, or Antigravity stays in control.

  • Mix providers in one task. A Codex session can delegate to Claude Code or Antigravity, and managed agents can spawn children across providers.

  • Preserve native conversations. Codex thread_id, Claude session_id, and Antigravity conversation_id are retained so work can resume in the same provider-native session.

  • Delegate durable work. By default, jobs run through a detached local broker, while messages, delegations, and orchestration events are persisted across MCP/UI process restarts.

  • Parallelize without immediately colliding on files. Writable agents can be isolated in Git worktrees, with explicit diff/apply control before changes reach the base workspace.

  • Reuse provider authentication. agentmux invokes the installed provider CLIs and does not store provider credentials itself.

  • Install as native plugins. Codex, Claude Code, and Antigravity are all supported through one installer command.

Use cases

Cross-provider control tower

Keep your preferred agent as the supervisor and send specialist work elsewhere.

Use Codex as the control tower.
Ask Claude Code to review the API design.
Ask Antigravity to inspect the implementation for edge cases.
Wait for both and summarize the disagreements.

Parallel independent review

Run multiple providers on the same question before committing to a change.

Spawn one Codex and one Antigravity reviewer for this pull request.
Have them review independently, then compare their findings.

Builder + reviewer

Separate implementation from verification.

Delegate the implementation to Claude Code in an isolated worktree.
Have Codex review the resulting diff before applying it to the main workspace.

Long-running delegated work

Start work from one UI and let the local broker keep ownership if that UI or MCP process exits. Reconnect later through the shared state, job, delegation, and event APIs.

Model name resolution

For Antigravity, agentmux queries the installed CLI with agy models instead of hard-coding a model list. Informal names are resolved before a job is created:

agy 3.8 flash high
Gemini 3.8 Flash High
gemini-3.8-flash-high
        ↓
gemini-3.8-flash-high

If a request matches multiple installed models, agentmux returns the candidates instead of guessing. The MCP models tool can be used to inspect or disambiguate the current catalog. Codex and Claude Code model names currently pass through to their native CLIs.

Supported providers

Provider

Worker CLI

Native plugin

Native session resume

Codex

codex

✓

thread_id

Claude Code

claude

✓

session_id

Antigravity

agy

✓

conversation_id

How it works

          your existing coding-agent UI
     Codex / Claude Code / Antigravity
                     |
                agentmux MCP
                     |
          local orchestration runtime
          /          |           \
       Codex       Claude     Antigravity
       worker      worker        worker

The interactive host can remain the external control tower, or a managed agent can supervise nested children. agentmux supplies the shared session, delegation, messaging, event, execution, and workspace layer rather than introducing another UI.

Status: early alpha. The core runtime and real Codex ↔ Antigravity delegation path are validated, but command surfaces and plugin integration may still evolve.

Current scope

The MCP server exposes a provider-neutral session API:

  • spawn / spawn_many — create one or many agent sessions and start their jobs asynchronously

  • send — continue the same provider-native conversation

  • whoami — identify a managed child agent from inherited runtime context

  • message_send, inbox, message_ack — persisted, attributed agent-to-agent messaging

  • delegate, delegation_list, delegation_accept, delegation_complete, delegation_cancel — explicit tracked work handoff

  • events / events_wait — durable, ordered orchestration history shared across MCP hosts

  • status — inspect an agent and its latest job

  • result / wait — fetch results or wait for multiple jobs in one MCP call

  • list — list local sessions

  • kill — cancel an active job and stop the session

  • team_create, team_status, team_list — group sessions and record supervision

  • providers — show supported runtime adapters

  • models — inspect provider model names; Antigravity models are discovered dynamically from agy models and informal names can be resolved to canonical slugs

  • doctor — report provider install/auth health and MCP-host configuration

Provider sessions are preserved using their native IDs:

Provider

CLI

Native session ID

Codex

codex exec --json

thread_id

Claude Code

claude -p --output-format stream-json --verbose

session_id

Antigravity

agy -p --output-format stream-json

conversation_id

Main vs subagent

agentmux does not hard-code one model as the main agent.

When a team has no supervisorAgentId, the interactive MCP host is the control tower:

You
 |
Codex UI                 <- external supervisor
 |
agentmux team
 |- Claude reviewer
 |- Antigravity implementer
 `- Codex researcher

A managed agent can also supervise children. Provider subprocesses inherit AGENTMUX_AGENT_ID, AGENTMUX_TEAM_ID, AGENTMUX_PARENT_AGENT_ID, and AGENTMUX_ROLE. If that coding agent starts its configured agentmux MCP server, whoami resolves the inherited identity and nested spawn automatically creates children inside the same team.

Managed agents are team-scoped: they can inspect and send work within their team, read only their own inbox, and stop only themselves or descendants. An external Codex/Claude Code/Antigravity UI has no inherited agent ID and remains the unrestricted control tower. This is a coordination boundary, not an OS-level security sandbox.

Multiple agentmux MCP processes on the same machine can share this state safely. State mutations are serialized with an inter-process filesystem lock and committed by atomic replacement.

Provider jobs are owned by a small detached local broker by default, not by the MCP stdio process that happened to launch them. Codex UI, Claude Code UI, Antigravity UI, and nested managed agents therefore share one local execution owner, and a running delegated job can continue if its launching UI or MCP process exits. The broker uses a local Unix socket or Windows named pipe plus a per-state capability token, and shuts itself down after an idle period.

If the broker cannot start, agentmux falls back to MCP-owned execution and reports the mode through runtime_status. Set AGENTMUX_EXECUTION=local to force that fallback behavior explicitly.

Requirements

  • Node.js 20+

  • Provider CLIs are optional. Install only the workers you actually plan to use: codex, claude, and/or agy.

  • Git is needed only for worktree isolation/integration.

A valid installation can therefore be Codex + Antigravity only, Claude Code only, or even agentmux with no provider installed yet.

Installation details

For normal installations, use the Quick start at the top of this README. The sections below cover direct-MCP fallback and development setups.

Direct MCP fallback

If a host does not support or should not use plugins, register agentmux directly as MCP instead:

npx -y '@jiho.ko/agentmux@latest' setup

Or install the CLI globally:

npm install -g '@jiho.ko/agentmux'
agentmux setup

The quotes around the scoped package are intentionally shown so the commands can be pasted unchanged into PowerShell as well as POSIX shells.

Local checkout for development

git clone https://github.com/seaweedsoup98/agentmux.git
cd agentmux
npm install
npm run build
npm link
agentmux setup

Interactive setup detects installed Codex, Claude Code, and Antigravity hosts and asks which ones to configure. Non-interactive examples:

agentmux setup --hosts codex,antigravity
agentmux setup --hosts claude
agentmux setup --yes
agentmux setup --hosts codex,antigravity --dry-run

The setup command uses native host configuration paths:

  • Codex: codex mcp add

  • Claude Code: user-scoped claude mcp add

  • Antigravity: merges only mcpServers.agentmux into ~/.gemini/config/mcp_config.json

Existing Antigravity MCP entries and unrelated JSON keys are preserved. If a native agentmux plugin is already installed for a host, setup does not add a second direct MCP registration.

Verify the machine afterwards:

agentmux doctor
agentmux doctor --json

Provider health and authentication

agentmux doctor separates installation from authentication instead of treating every provider as a required dependency.

Typical states:

  • ready: CLI installed and its non-inference authentication status says it is logged in.

  • auth_required: CLI installed but login is required.

  • installed: CLI installed, but authentication cannot be verified without a real model request.

  • missing: CLI is not installed.

  • unhealthy: the status probe itself failed unexpectedly.

Codex uses codex login status; Claude Code uses claude auth status. Antigravity has no documented zero-cost shell auth-status command, so doctor reports its auth as unknown rather than consuming quota or opening a browser. A real headless Antigravity run is the authoritative check.

If a provider is missing, doctor prints its official installation command rather than installing software implicitly. Provider installation can modify PATH, shell profiles, or system state, so it remains an explicit user action.

If authentication expires, agentmux does not store or repair provider credentials. Re-authenticate with the provider itself:

Codex:       codex login
Claude Code: claude auth login
Antigravity: run agy interactively once and complete sign-in

Runtime failures that look like missing executables or authentication errors include the corresponding remediation hint.

When setup is running from an npx cache, it registers npx -y @jiho.ko/agentmux@latest as the stable MCP launch command rather than pinning an ephemeral cache path.

Native plugins

The repository ships native plugin bundles for Codex/OpenAI, Claude Code, and Antigravity. The recommended installer is:

npx -y '@jiho.ko/agentmux@latest' plugins install

Codex / OpenAI

Manual equivalent:

codex plugin marketplace add seaweedsoup98/agentmux --ref main
codex plugin add agentmux@agentmux

The repository marketplace is .agents/plugins/marketplace.json; the plugin bundle is plugins/codex.

Claude Code

Manual equivalent:

claude plugin marketplace add seaweedsoup98/agentmux@main --scope user
claude plugin install agentmux@agentmux --scope user
claude plugin enable agentmux@agentmux --scope user

The Claude marketplace is .claude-plugin/marketplace.json; the plugin bundle is plugins/claude.

Antigravity

The universal installer stages the npm-bundled plugin automatically. From a repository checkout, the manual equivalent is:

agy plugin install ./plugins/antigravity

The plugin bundle is plugins/antigravity.

All three native plugins bundle the orchestration skill and launch the same published MCP runtime, npx -y @jiho.ko/agentmux@latest. Direct MCP registration and native plugin installation are alternative integration methods; use --replace-mcp when migrating to avoid duplicate tool surfaces.

Development

git clone https://github.com/seaweedsoup98/agentmux.git
cd agentmux
npm install
npm run check

Run the MCP server over stdio:

npm run dev

The server stores local session metadata in ~/.agentmux/state.json. Override that directory with AGENTMUX_HOME.

The state store is shared across local agentmux MCP processes. Reads use fresh snapshots; mutations use a process-safe lock plus atomic file replacement. Windows transient replace failures are retried without falling back to a non-atomic delete-and-rewrite path. A running job records its owner PID/instance so starting another MCP host does not incorrectly recover or overwrite work owned by a live host.

Example

Once the MCP server is registered in your host, you can ask the host agent naturally:

Create a team for this task.
Spawn two Antigravity agents to review this repository independently.
Use one Codex agent to compare their findings, then report the consensus.

The host remains the control tower. agentmux provides the runtime/session layer.

Agent-to-agent messaging

A managed child can discover itself and its team with whoami, then inspect peers with team_status. Direct send is reserved for resuming your own managed session (or for an external control tower); peer-to-peer work must use attributed messages or tracked delegations.

message_send(
  to_agent_id="<peer>",
  message="I changed the repository interface. Rebase your implementation on it.",
  wake=false
)

Messages are persisted before delivery. wake=false leaves the message unread in the peer's inbox. wake=true additionally resumes the peer's provider-native session when that peer is idle and resumable; if it is busy, the wake fails but the message remains in the inbox.

A wake job is owned by the detached broker, so it can outlive the MCP host that launched it. Call wait when the current turn depends on the result; otherwise the delegated work may continue independently and can be observed later through job status or the durable event stream.

inbox(unread_only=true)
message_ack(message_ids=["msg_..."])

This lets agents communicate without requiring a separate agentmux UI.

Tracked delegations

Use messages for coordination and use delegations when work ownership/completion matters.

delegate(
  to_agent_id="<specialist>",
  task="Review the provider adapter and report concrete defects.",
  wake=true
)

A delegation is pending until accepted, active while owned by the target, then completed or canceled. With wake=true, an idle resumable target is woken immediately and the delegation becomes active automatically. If wake fails because the target is busy, the task remains persisted as a pending delegation plus an inbox message.

Managed agents can delegate only within their team. The assigned agent can explicitly accept and complete the work, while the sender or receiver can cancel a non-terminal delegation. Delegation transitions are also emitted into the durable event stream.

Durable orchestration events

Every important lifecycle transition is also appended to a process-safe ordered event stream. Events use a monotonic seq cursor and cover team creation, agent spawning/stopping, job creation/start/completion/cancellation, message delivery/read/wake state, workspace integration, and bounded provider progress.

events(after_seq=0, team_id="<team>")
events_wait(after_seq=42, timeout_ms=30000)

events_wait is a bounded long-poll rather than tight polling. Because the cursor is persisted in the same transactional state store, Codex UI, Claude Code UI, Antigravity UI, and nested managed agents can observe the same orchestration history even when they are backed by different agentmux MCP processes. Managed agents remain restricted to their own team.

The event stream is intentionally metadata-oriented: message bodies, assistant text deltas, command contents, and raw provider stdout are not copied into events. Detailed content remains in inbox/job APIs.

Provider adapters normalize only useful structured progress:

  • Codex exec --json: non-response item start/completion

  • Claude Code stream-json: tool-use and tool-result transitions

  • Antigravity stream-json: non-response step state transitions

These become provider.progress, provider.tool_started, and provider.tool_completed events. Exact duplicates within one second are coalesced, and the durable event journal is bounded rather than growing indefinitely.

Workspace isolation

Each spawned agent accepts workspace: shared | worktree | auto.

  • shared uses the requested working directory directly.

  • worktree creates a detached Git worktree under ~/.agentmux/worktrees/<agent-id>.

  • auto is the default. Read-only agents share the workspace. A single writable agent normally shares it; parallel writable agents in the same spawn_many batch are isolated before they start, and a later writable agent is isolated when another shared writer is already running.

Worktrees are created from Git HEAD, and the exact base commit is recorded on the agent session. To avoid silently dropping local edits, worktree creation refuses a dirty repository; commit/stash first or explicitly choose shared.

The control tower can inspect and integrate isolated writable work explicitly:

workspace_status(agent_id="<agent>")
workspace_diff(agent_id="<agent>")
workspace_apply(agent_id="<agent>")
kill(agent_id="<agent>")
workspace_cleanup(agent_id="<agent>", force=true)

workspace_diff builds one base-relative patch using a temporary Git index, so committed, staged, unstaged, deleted, and untracked files are represented without modifying the agent's real index. workspace_apply is external-control-tower only, refuses a dirty base repository, runs git apply --check, and never performs an automatic merge. Cleanup requires the session to be stopped; dirty isolated worktrees require an explicit force=true.

This keeps parallel writers reproducible while leaving integration authority with the interactive control tower.

Access modes

spawn accepts a provider-neutral access mode. Adapters map it to the nearest native behavior:

agentmux

Codex

Claude Code

Antigravity

read-only

read-only sandbox

plan permission mode

plan + sandbox + read-only tool profile

workspace-write

workspace-write sandbox

acceptEdits

accept-edits + terminal sandbox

full

danger-full-access

skip permission prompts

accept-edits + skip permission prompts

For Antigravity, read-only uses the bundled agentmux-readonly custom agent. It exposes only native repository-reading tools (view_file, list_dir, find_by_name, and grep_search) and excludes run_command plus all file-writing tools. This avoids headless permission prompts for shell-based reads while preserving a true read-only review surface.

These mappings are intentionally conservative and are not identical security models.

Design principles

  1. Keep the existing Codex, Claude Code, or other MCP-host UI.

  2. Treat main vs subagent as a session relationship, not a model property.

  3. Preserve native provider sessions instead of flattening everything into stateless API calls.

  4. Make workspace isolation optional; auto only isolates concurrent writers.

  5. Keep the core small. Worktrees, messaging policy, and richer supervision sit above provider adapters.

Deliberate non-goals

To keep the runtime small, agentmux intentionally does not add:

  • a separate multi-agent UI — the existing Codex, Claude Code, or Antigravity UI stays in control

  • a workflow/task-DAG DSL — nested agents plus tracked delegations cover ownership without another orchestration language

  • a second push/message transport — events_wait is the shared long-poll event primitive

  • automatic worktree merging — the control tower explicitly inspects and applies isolated changes

  • provider-specific workflow abstractions in the core — provider adapters stop at execution/session/progress normalization

Named roles remain lightweight session metadata instead of persistent templates.

Real-provider validation

Normal CI uses deterministic fake provider executables and never consumes Codex, Claude Code, or Antigravity quota. Providers reported as auth_required are excluded from the default real-provider smoke run rather than making an otherwise valid partial installation fail. To validate installed authenticated CLIs explicitly:

npm run e2e:real

The smoke harness detects installed providers, then sequentially verifies initial spawn and native-session resume using read-only prompts. Requests are spaced by 2 seconds by default to avoid aggressive traffic. Select providers explicitly with:

npm run e2e:real -- --providers codex,claude

Run the full heterogeneous parent/child matrix only when desired:

npm run e2e:real:matrix

Matrix mode exercises every selected cross-provider parent -> child pair, nested managed identity, wake/resume delegation, and delegation completion. It writes one machine-readable JSON report to stdout and uses a disposable local state/workspace. It never requests full access.

Tune pacing and timeout with AGENTMUX_E2E_DELAY_MS and AGENTMUX_E2E_TIMEOUT_MS.

License

Apache-2.0

Available Tools

17 tools
doctorA

Check whether each supported provider CLI is installed and report its version.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. The verb 'check' and 'report' imply a non-mutating diagnostic read, which is useful. However, it does not explicitly state side-effect-free behavior, return format, or what happens when no providers are installed. For a parameterless, likely safe tool, the core behavior is covered but not deeply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action ('Check') and clearly states the output ('report its version'). There is zero fluff; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless tool with no output schema, the description is nearly complete. It tells the agent what will happen and what the result will be. The only missing context is clarification of what 'supported providers' means and whether this is a safe/read-only operation, though 'check' strongly implies it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the input schema is empty. Per the baseline for 0-parameter tools, the description need not explain parameters. It correctly references 'each supported provider' as a fixed scope, adding no confusion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Check whether each supported provider CLI is installed') and a concrete result ('report its version'). It is unambiguous and distinguishable from sibling tools like 'providers' or 'status' because it focuses on installation/version diagnostics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, typical scenarios (e.g., diagnosing environment setup), or exclusions. The purpose is clear, but the agent is left to infer when 'doctor' is appropriate among 16 siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inboxA

Read persisted messages. Managed agents can only read their own inbox; external supervisors must specify agent_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNo
mark_readNo
unread_onlyNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility. It discloses the read nature and the permission model, but it does not mention that mark_read can mutate state (mark messages as read). This is a behavioral gap, as an agent might assume a purely read-only operation despite the parameter hinting at a side effect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The primary purpose is stated first, followed by the access rule. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters, no output schema, and no annotations, the description is incomplete. It fails to explain mark_read and unread_only, and does not describe the return format. An agent would be guessing on how to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all parameters. It only explains agent_id (and who needs it), leaving mark_read and unread_only completely undefined. The agent cannot know what these booleans do without external knowledge, which is a significant deficiency.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function ('Read persisted messages') and distinguishes it from sending tools like send and message_send. It also specifies the resource (messages) and the context for who can read what, which sets it apart from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when agent_id must be specified (external supervisors) and implies it's optional for managed agents. This gives concrete usage guidance, though it doesn't explicitly name alternatives or state 'use this when you need to read messages' – but that is clear from the verb 'read'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

killA

Cancel an active job and stop a session. Managed callers may stop only themselves or their descendants.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It states the core action (cancel/stop) and a permission restriction, but does not mention side effects, reversibility, or behavior when the target does not exist. This is a destructive operation, so more transparency is warranted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence that front-loads the action and includes a relevant caveat. It contains no filler or redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sparse for a destructive tool. It does not explain return values (no output schema), error conditions, or what happens to dependent jobs/sessions. The restriction on managed callers is stated but the behavior for unmanaged callers is left ambiguous. Given the one-parameter signature and absence of annotations, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage, and the description does not elaborate on the parameter. However, 'agent_id' is self-explanatory in the context of killing an agent, and the tool name and description imply its purpose. It is a single, simple parameter, so the minimal guidance is acceptable but not compensatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool cancels an active job and stops a session, with a specific verb and resource. It also adds a restriction about who may call it, distinguishing it from monitoring tools like status or wait.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions a constraint for managed callers but does not explicitly state when to use this tool versus alternatives or when not to use it. There is no guidance on checking status before killing or using other tools for related actions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

listA

List visible sessions. Managed agents see only their team; external supervisors see all sessions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It does reveal a non-obvious behavior: visibility is role-dependent (managed agents see only their team; external supervisors see all). Yet it doesn't mention read-only status, output details, ordering, or error behavior, which are relevant but not critical for a simple list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. The core function is front-loaded, and the role-based visibility nuance adds real value without bloating the definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with a straightforward purpose, the description covers what the tool does and who can see what. The lack of an output schema isn't fully compensated by the description, but for a simple list tool this is a minor gap rather than a critical omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so the schema imposes no burden. With zero parameters, the baseline is 4, and the description doesn't need to add parameter-level detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('List') and the resource ('visible sessions'), making the tool's function immediately understandable. The 'visible' qualifier and role-based scope help differentiate it from sibling tools like team_list or status, though it doesn't explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool—when you need to see sessions—and adds useful role-based context about who sees which sessions. However, it does not explicitly state when not to use it or mention any sibling alternatives, leaving some routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message_ackB

Mark one or more inbox messages as read.

ParametersJSON Schema
NameRequiredDescriptionDefault
message_idsYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It states the mutation (mark as read) but does not disclose whether the operation is idempotent, whether it requires authentication, what happens for invalid IDs, or the impact on other state. The word 'mark' implies a state change but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that is front-loaded with the action and object. It contains no filler or redundancy, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple mutation tool, the description provides the core action but omits details like return value, error handling, idempotency, and permission requirements. Given the absence of an output schema and annotations, an agent might not know what to expect after calling it. However, the simplicity of the tool (single parameter, straightforward action) keeps the gaps moderate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate. It implies the 'message_ids' parameter by saying 'one or more inbox messages', which aligns with the schema's minItems=1 and maxItems=100. However, it does not explicitly name the parameter or clarify its format (e.g., string identifiers). The description adds some meaning but not enough to fully cover the parameter's semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Mark') and resource ('inbox messages as read'), which clearly conveys the tool's function. It distinguishes itself from siblings like 'inbox' (which likely lists messages) and 'send'/'message_send' (which create messages) by focusing on the read-state mutation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. It does not mention that it should be used after viewing messages, nor does it exclude cases like marking messages as unread (which might be another tool). The description only states the action, leaving usage context to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message_sendB

Persist an attributed message to another agent. With wake=true, also start a new turn if the recipient is idle and resumable.

ParametersJSON Schema
NameRequiredDescriptionDefault
wakeNo
messageYes
to_agent_idYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the key behavior: the message is persisted, and wake=true conditionally starts a new turn. However, it omits side effects like delivery guarantees, recipient state requirements, or whether failures can occur silently.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no filler. The core action is front-loaded, and the wake behavior is presented as a conditional add-on, making the tool's behavior easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a mutation with no annotations, no output schema, and 0% parameter coverage. The description does not explain return values, error conditions, or how this differs from the sibling 'send' tool, leaving an agent with notable gaps when deciding to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies wake's conditional behavior and the general notion of an attributed message, but gives little explicit semantics for to_agent_id and message beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Persist') and resource ('attributed message to another agent'), making the core action clear. It does not explicitly distinguish itself from the sibling tool 'send', but the persistence and attribution framing separates it at a basic level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the wake=true behavior but gives no guidance on when to use this tool versus alternatives like 'send' or 'message_ack'. There is no mention of prerequisites, conditions, or cases where a sibling would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

providersA

List provider adapters and the local CLI command each adapter expects.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the operation is a list and that the output includes the CLI command expectation, which implies a read-only, non-mutating action. However, it does not explicitly confirm side-effect-free behavior, error conditions, or any rate limits, leaving a minimal but acceptable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, compact sentence front-loads the verb and object and adds the relevant detail about the CLI command. No filler, redundancy, or unnecessary words are present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter, read-only listing tool with no output schema, the description conveys the core purpose and what the output contains. It lacks any mention of usage context or relationship to sibling tools, but the description is functionally sufficient for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the empty input schema is fully covered. The description adds relevant context about the returned data, but since there are no parameters to explain, the baseline of 4 for zero-parameter tools applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('provider adapters') and adds a defining detail about the CLI command each adapter expects. It is clear and specific, but it does not explicitly distinguish itself from the generic sibling 'list' tool, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'list' or other sibling tools. There is no mention of context, prerequisites, or exclusions, leaving the agent without direction on tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resultB

Get an accessible job result by job ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says to 'get' a result, but does not explain behavior for pending, failed, invalid, or inaccessible jobs, nor what the result contains or whether the operation is read-only. The word 'accessible' promises a condition but never elaborates on it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused sentence with no filler, tautology, or redundant detail. It front-loads the action and resource and stays appropriately minimal for a tool with one parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations and no output schema, so the description must compensate. It does not explain the return value, error cases, job lifecycle, or when a result is 'accessible.' Given sibling tools like wait and status, more context is needed for an agent to correctly sequence calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description explicitly references 'by job ID,' which maps directly to the single required job_id parameter. For a one-parameter tool, this is sufficient semantic clarification, even though it does not describe format conventions or value provenance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get an accessible job result by job ID.' It clearly identifies what the tool does and ties it to the only input, distinguishing it from siblings like status, list, or wait. The word 'accessible' hints at a scoping condition, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as status, wait, or list. The only implicit signal is that the caller must already have a job_id, but there is no stated workflow, no exclusion, and no mention of prerequisites or when the result becomes available.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sendC

Continue an existing idle provider-native session. Managed callers are limited to their team.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
agent_idYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose that only idle sessions are targetable and that managed callers are team-limited, but it omits key behavioral traits such as side effects of sending a prompt (e.g., session becoming active, executing actions), error behavior for non-idle sessions, and permission requirements beyond the team restriction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no filler, and it front-loads the core action and constraint. It is appropriately terse, though it sacrifices detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% schema parameter coverage, the description should provide enough context to invoke the tool correctly. It gives a broad purpose and a team restriction, but lacks parameter meanings, return behavior, error conditions, and relationships to sibling tools, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the two parameters. It does not mention 'prompt' or 'agent_id' at all, leaving the agent to infer meaning solely from the parameter names. This fails to provide any added semantic clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Continue') and resource ('existing idle provider-native session'), clearly distinguishing it from sibling tools like 'spawn' (creation) and 'message_send' (likely messaging). It does not explicitly name alternatives but the focus on 'existing idle' sessions makes the purpose fairly unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this tool when an existing idle provider-native session should be continued. However, it does not explicitly state when not to use it or contrast with siblings such as 'message_send' or 'spawn', so the guidance is minimal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spawnC

Spawn a coding-agent session and start its first job asynchronously. Managed callers automatically create children in their own team.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
nameNo
roleNo
modelNo
accessNoworkspace-write
effortNo
promptYes
team_idNo
providerYes
workspaceNoauto
parent_agent_idNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does usefully disclose asynchronous execution and the managed-caller child-team behavior, but it omits other important behavior such as side effects, resource implications, failure modes, or how to retrieve the spawned session's output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tight sentences with no filler, front-loading the core action and the key asynchronous trait. Every clause adds some signal, making it efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—11 parameters, no output schema, no annotations—the description leaves too much undefined for reliable invocation. The managed-caller behavior is a useful detail, but the overall definition is incomplete for an operation that creates a coding-agent session.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no meaning to any of the 11 parameters. 'First job' loosely hints at prompt, but provider, access, workspace, team_id, parent_agent_id, effort, role, and model are entirely unexplained, so the agent gets no parameter-level help.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('spawn'), resource ('coding-agent session'), and the asynchronous start behavior, so an agent can grasp the main action. It does not explicitly contrast with the sibling spawn_many, but the singular 'a coding-agent session' makes the single-session scope clear enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to choose spawn over spawn_many, send, or wait, nor when not to use it. The managed-callers note offers a small usage hint, but there is no explicit routing or alternative-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spawn_manyC

Spawn up to 16 coding-agent sessions. Sessions are started sequentially but their jobs run concurrently.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentsYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior. It covers concurrency (sequential start, concurrent jobs) and the limit of 16, but omits critical aspects: whether it blocks, what happens on partial failure, side effects on the workspace, authentication requirements, or return format. For a tool that spawns multiple agents, these are significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, no wasted words, and states the most important detail (concurrency) early. It is appropriately concise, though a bit more detail would be warranted given the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested array schema, no annotations, and no output schema, the description is woefully incomplete. An agent cannot know what to put in the 'agents' array, what the expected response is, or how failures are handled. It does not even mention that each element is an agent configuration object.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the 'agents' array and its nested fields. It does not mention that each agent requires 'provider' and 'prompt', nor the meaning of optional fields like 'cwd', 'access', 'workspace', or 'model'. The description adds zero parameter meaning beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose: spawning up to 16 coding-agent sessions. The verb 'spawn' with the resource 'coding-agent sessions' is specific. It distinguishes from the sibling 'spawn' by implying multiple sessions, but it does not explicitly name the alternative, so it misses the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the sibling 'spawn' or other session-management tools. The description mentions sequential start and concurrent jobs, which hints at behavior but not usage context. No exclusions or alternatives are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusB

Get one accessible agent and its latest job.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals it is a read operation (get), but does not state whether it requires authentication, whether it might return partial data if the agent is busy, or what happens if the agent is not accessible. No contradictions, but significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence with no fluff. It directly states the action and the object. Perfectly front-loaded and minimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with one param and no output schema, the description is mostly sufficient. However, it lacks context on what 'latest job' means, whether the output includes status fields, or if there are any side effects. Given the tool is read-only, a 3 is fair; it's adequate but missing minor details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the parameter, so the description must compensate. The description does not explain 'agent_id' beyond naming it, but the parameter name is self-explanatory (an ID string). This is mediocre compensation; more detail like 'the ID of the agent to query' would help but isn't critical.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it fetches an accessible agent and its latest job, which is specific and distinguishes it from siblings like 'list' (which likely lists agents) and 'result' (which may fetch job results). However, it could be more specific about what 'accessible' means or what a 'job' entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus siblings like 'status' vs 'doctor' or 'whoami'. It implies it is for checking status, but does not mention alternatives or exclusions. For example, it doesn't say 'use doctor for diagnostics' or 'use whoami for current user'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

team_createC

Create a top-level logical team. This tool is available to external control-tower hosts.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
supervisor_agent_idNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only states the creation of a team, but does not disclose side effects, permissions required, idempotency, or what happens on success or failure. This is minimal for a mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core action, but it is under-specified. While concise, it does not earn its brevity because key information about parameters and usage is missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and zero parameter coverage, the description is severely incomplete. It does not explain parameters, return values, prerequisites, or any operational context. An agent would struggle to call this tool correctly without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does not mention either 'name' or 'supervisor_agent_id' at all, leaving their meaning and relationships unexplained. The description adds no value beyond the parameter names themselves.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb 'Create' and a specific resource 'top-level logical team', distinguishing it from sibling tools like team_status and team_list which are read-only. It also adds context about availability to external control-tower hosts, which is a useful constraint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It mentions availability to external control-tower hosts, which is a constraint rather than usage guidance. No exclusions or explicit conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

team_listA

List visible teams. Managed agents see their own team; external supervisors see all teams.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses a non-obvious behavioral trait: returned teams depend on the caller's role ('Managed agents see their own team; external supervisors see all teams'). For a read-only list with no parameters, this is meaningful transparency, though it does not describe output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose. The role-based detail is delivered efficiently with no filler, and each sentence contributes essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool, the description explains what is listed and the visibility rules. There is no output schema, so a slight gap exists around return format, but the operation is simple enough that the description is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, and the instructions give a baseline of 4 for parameter-free tools. The description adds role-based filtering context, which is relevant to interpreting what the (empty) input means.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('List visible teams'), making the tool's purpose immediately clear. It also adds the role-based scoping distinction (managed agents vs. external supervisors), which differentiates it from generic sibling tools like 'list' and 'team_status'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful context about who sees which teams, implying when the tool is appropriate, but it does not explicitly mention alternatives or when not to use it. Sibling tools like 'team_status' or 'list' are not referenced, leaving the agent to infer the correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

team_statusA

Get a team and its members. Managed agents are limited to their own team.

ParametersJSON Schema
NameRequiredDescriptionDefault
team_idYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It adds the managed agent limitation, which is useful, but it does not explicitly state the operation is read-only (implied by 'Get'), nor does it mention error behavior or response format. The description adds one behavioral constraint but lacks broader transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, consisting of two sentences with the primary purpose stated first. It is front-loaded and contains no extraneous information, making it easy for an agent to quickly understand the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has only one parameter and no output schema, so the description should provide enough context for correct invocation. It explains the main function and the managed agent limitation, but it does not describe the structure of the returned data (e.g., what 'members' includes) or any error conditions. For a simple get tool, it is mostly adequate but missing details that could affect how the agent interprets results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description must compensate for the team_id parameter. It does not explain what team_id refers to, what format it should take, or that it's the identifier for the team. This is a significant gap since the schema only provides a type and minLength.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Get a team and its members' with a specific verb (get) and resource (team plus members). It distinguishes itself from sibling tools like team_list (list teams) and team_create by focusing on retrieving a single team with its members.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions 'Managed agents are limited to their own team,' which provides a usage constraint but does not explicitly guide when to use this tool versus alternatives. It implies that managed agents should use this for their own team, but it doesn't state when not to use it or suggest other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

waitA

Wait for one or more jobs to finish, returning early when all are terminal. Timeout is capped at 60 seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idsYes
timeout_msNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It adds useful details beyond the schema, such as returning early when all jobs are terminal and a 60-second timeout cap. However, it does not disclose what happens when the timeout expires, what the return value contains, or whether waiting has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, front-loading the core action and then adding the key behavioral constraint. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple two-parameter tool, and the description covers the essentials of waiting and early return. However, without an output schema or annotations, the absence of timeout-expiration behavior and return-value semantics leaves a meaningful gap for an agent deciding whether the call succeeded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It maps job_ids to 'one or more jobs' and relates timeout_ms to the 60-second cap, but it leaves out the default timeout, the unit context, and how zero or missing timeout_ms should be interpreted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('wait for'), a clear resource ('one or more jobs'), and a precise completion condition ('returning early when all are terminal'). This distinguishes it from sibling status-checking tools like status and result by making the blocking behavior explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies the use case: block until jobs reach a terminal state rather than polling status or fetching results. It does not explicitly name alternatives or exclusion conditions, but the purpose is unambiguous enough for an agent to select it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiA

Return the managed agent identity inherited by this MCP process, or managed=false for an external control-tower host.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the return value (identity or 'managed=false'), which implies a read-only operation, but it does not explicitly mention that it has no side effects, requires no permissions, or cannot fail. For a simple identity query, this is acceptable but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core function ('Return the managed agent identity') and immediately clarifies the edge case. Every word contributes value, with no redundant phrasing or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description fully specifies what the tool does and what it returns. It covers both the managed and unmanaged cases, leaving nothing missing for an agent to call it correctly. The simplicity of the tool means no further context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema fully covers everything (coverage 100% trivially). Per calibration, a baseline of 4 applies since there are no parameters to document. The description adds no parameter-specific meaning because none exist, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: returns the managed agent identity inherited by the MCP process, with a fallback for external hosts. It uses a specific verb ('Return') and resource ('managed agent identity'), leaving no ambiguity about what it does. This distinguishes it from sibling tools like 'status' or 'list', which handle other concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when you need to know the managed agent identity or whether the host is managed. However, it does not explicitly mention alternatives or when not to use it. Sibling tools like 'doctor' or 'status' might cover similar diagnostics, but no exclusions or comparative guidance is provided, keeping this at a basic clear-context level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv0.1.0
    • First observeddoctor
    • First observedinbox
    • First observedkill
    • First observedlist
    • First observedmessage_ack
    • First observedmessage_send
    • First observedproviders
    • First observedresult
    • First observedsend
    • First observedspawn
    • First observedspawn_many
    • First observedstatus
    • First observedteam_create
    • First observedteam_list
    • First observedteam_status
    • First observedwait
    • First observedwhoami

TDQS

B3.1/5.0

Scored across 17 tools

Disambiguation4/5

Most tools target distinct resources and actions: sessions, jobs, teams, messages, and provider metadata are clearly separated. The main ambiguity is between send and message_send, but their descriptions clarify session continuation versus persisted messaging.

Naming Consistency3/5

Naming is readable but mixes conventions: verb_noun (team_create, team_status), noun_verb (message_send, message_ack), bare verbs (send, wait, list, kill), and nouns (inbox, status, providers). The send/message_send pair is especially inconsistent.

Tool Count3/5

With 17 tools, the server is at the borderline of feeling heavy for an orchestration/mux utility. Each tool covers a distinct concern, but the count is above the typical well-scoped range.

Completeness3/5

The surface covers core session, job, team, and messaging workflows, but there are notable gaps: no way to list all jobs, no team deletion or update, and no session transcript/history retrieval. These gaps can force workarounds but basic orchestration is functional.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables any MCP-compatible client to use existing Claude Code agents from .claude/agents/ directories. Spawns agents in separate CLI sessions for better context optimization and performance across Codex, Gemini CLI, and other AI coding assistants.
    3
    MIT
  • A
    license
    B
    quality
    F
    maintenance
    Enables orchestrating multiple AI CLI agents (Claude Code, Codex, Gemini CLI, Copilot CLI) through a unified MCP interface for task delegation, cross-agent comparison, and specialized tools like code review and debugging.
    14
    7 npm
    14
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables multiple coding agents (Claude Code, Codex, Cursor) to discover each other's sessions, search transcripts, ask questions, and handoff tasks through a shared MCP server.
    5
    8 npm
    MIT