cross-agent
cross-agent is an MCP server that lets a host session run headless Claude, Codex, and Grok specialists as a solo consultant or a full dev team under a configured mode.
Inspect the active mode's loop, roles, workspaces, sandbox defaults, prompts, git policy, and project root (
describe_mode).List roles and their engine/model/effort bindings, including multi-seat roles (
list_roles).Delegate a brief to a specialist role in a working directory, optionally in its own worktree, by seat, or with explicit engine/model/effort/branch/resume/force options (
delegate).Monitor and control tasks: wait for settle/stall/timeout, check status and event tail, fetch full final result, cancel a task and its descendants, and list project tasks by status (
wait,check,result,cancel,list_tasks).Verify a linked worktree and exact branch before operating on it (
verify_worktree).Run journaled, locked git subcommands in a verified worktree (
git_mutate).Run whitelisted git verbs at the project root, including gated
merge --ff-onlywith test and review guards (git_root).Run configured test or setup commands by selector at the root or a worktree, journaling passes (
run_command).Record a review waiver for a branch head to bypass the all-seats-clean-review merge guard (
waive_review).
Owns git on behalf of the team: worktree creation and root repository operations are performed through the server's git_mutate and git_root tools and journaled per task, specialists are never allowed to write git metadata, and merges are gated by guards requiring a passing test suite on the exact branch head plus clean reviews (or a recorded waiver).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cross-agentHave Codex review src/parser.ts and summarize what it finds."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
cross-agent
Run headless claude, codex and grok processes as one team, from whichever of the
three you work in. cross-agent is an MCP server and a launcher skill: your session, the
host, hands work to other engines, each running headless in its own CLI's sandbox on
your own subscription, and a mode decides the team — its roles, the loop its lead runs,
and its git policy.
Mode | The team | The loop runs in |
| one consultant: another engine reads and answers, or takes one change in a worktree of its own | your session |
| planner, plan reviewer, implementer, as many code reviewer seats as you bind, and a resolver; the planners read the project, and each task's change is made and reviewed in a git worktree of its own | your session |
| the same team | a Claude or Codex lead the server launches, so your session stays free |
A git repository with no cross-agent config runs as solo, so a one-off delegation needs
no setup.
How it works
One server, three hosts. The same MCP server and skill attach to Claude Code, Codex and Grok.
.cross-agent/config.jsonbinds each role to an engine, a model and an effort.Authority from process ancestry, not tokens. Your session gets the operator's tools, a lead the loop's, and a specialist a few read tools and never
delegate, so no delegation loop can form.Every specialist in its engine's sandbox. It writes only its own workspace, and it is denied launching
claude,codex,grok, this server or its CLI.Git stays with the server. Specialists never write git metadata; worktree and root git go through the server's
git_mutateandgit_root, journaled per task.Two merge guards. A merge needs the configured test suite to have passed on that exact branch head and, for a team task, a clean review from every reviewer seat, or your recorded waiver.
Related MCP server: Rutherford MCP Server
Requirements
Linux,
git, Node.js 24 or later, and util-linuxflock. There are no other dependencies: Node runs the TypeScript sources directly.The engine CLIs your roles use —
claude,codex,grok— installed and signed in.Each engine's sandbox prerequisites, such as
bwrapandsocatfor Claude: Linux prerequisites.
Install
Claude Code and Codex install it from the agent-plugins marketplace; Grok attaches it per project from a clone of this repository. docs/install.md has every step, check and removal, and the installs from a clone.
Claude Code, in a session:
/plugin marketplace add WSH95/agent-plugins
/plugin install cross-agent@agent-pluginsKeep the default user scope, or use --scope local for one project; never
--scope project, which every Claude specialist would load.
Codex:
codex plugin marketplace add https://github.com/WSH95/agent-plugins
codex plugin add cross-agent@agent-pluginsThen ask Codex: "Use cross-agent to set up its MCP connection once for this host."
The skill configures MCP for the CLI and local Linux desktop Codex. Restart MCP or open
a new session afterwards. Each chat uses its own project; no per-session
CROSS_AGENT_PROJECT or versioned cache path is needed. Install and update the plugin
through the marketplace as usual; the connection follows the installed version.
Grok: clone this repository and name it in the project's own .grok/config.toml:
Install it in Grok.
Quick start
Ask another engine, with no project setup. In any git repository, after attaching the plugin, ask your session, for example: "Use cross-agent to have Codex review src/parser.ts and tell me what it finds."
Run a team. Once per project, write its config — the mode, and the engine, model and
effort of each role. In Claude Code, ask your session to run cross-agent init --mode dev-team in the project; from a terminal, run the installed copy's launcher
(The operator CLI):
cd ~/code/my-project
cross-agent init --mode dev-teamEdit .cross-agent/config.json to bind the roles as you like. Commit that .gitignore
change init made before the team's first task (git add .gitignore && git commit -m '…' -- .gitignore): a loop's first step stops unless git status --porcelain --untracked-files=normal prints nothing at the project's root. Start a new session in the
project — the server reads the mode once, when it starts — and ask it for the work, for
example: "Use the cross-agent dev team to add a --json flag to the export command." Under dev-team your session runs the loop — plan, plan review, implementation
in a worktree, parallel code review, fixes, a tested merge. Under dev-team-engine a lead
runs it and puts its questions to you, which you answer with cross-agent answer.
docs/operator-guide.md covers the team's settings, the merge guards and the waiver, and running teams on several branches at once.
The operator CLI
The CLI ships inside the plugin, so it needs no separate install, but installing the
plugin does not put a cross-agent command on your shell's PATH:
In Claude Code the plugin's
bin/cross-agentis on the session's Bash toolPATH: ask Claude to runcross-agent init --mode dev-team.In Codex ask the cross-agent skill to initialize the project; it finds the bundled CLI relative to its installed skill directory.
From a terminal run the installed copy's launcher, for example
~/.codex/plugins/cache/agent-plugins/cross-agent/<version>/bin/cross-agent, or link a clone'sbin/cross-agentonto yourPATH.
You need init once per project for the team modes. solo works without a config
after the one-time Codex setup. The rest is optional, since your session reaches the same operations
through the plugin's tools:
Command | What it does |
| what the team is doing |
| every task, its outcome and its final message |
| a lead's questions to you |
| stop a task and everything it delegated |
Every verb and its exit codes: The operator CLI.
Documentation
docs/install.md: installing, checking and removing it per host; the Linux prerequisites; installs from a clone.
docs/operator-guide.md: team configuration and the merge guards, every CLI verb, several branches at once.
docs/design.md: the design, which is authoritative, and docs/probes.md: what each engine CLI was observed to do.
Development
npm test # the suite, node:test
node tools/build-dist.mjs # the Claude and Codex payloads, from HEAD, into dist/cross-agent
python3 tools/publish_agent_artifact_pr.py --dry-run # preview the agent-plugins pull requestA release bumps the version in package.json and both plugin manifests (Claude Code keeps
users on their cached copy until it changes), commits, and pushes main. Then
python3 tools/publish_agent_artifact_pr.py --build --commit-message "…" --pr-title "…" --pr-body "…" --keep-temp builds the payloads from that commit and opens the pull request,
which replaces cross-agent/ alone: the marketplace's root entries for cross-agent came with
its first publication.
License
MIT. See LICENSE.
Available Tools
13 toolscancelA
Terminate a task and every task it delegated, leaves first, returning one outcome per task. A lead may cancel only its own.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the cascading destruction of all delegated descendants, the leaves-first ordering, the one-outcome-per-task return, and an authorization constraint (a lead may cancel only its own). It omits whether cancellation is irreversible and what happens to tasks already in flight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler; the cascading scope and ordering are front-loaded, and the ownership constraint closes it out. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description explains the return ('one outcome per task') and the ownership rule, which is the key behavioral fact for a destructive tool. What remains missing is reversibility and failure behavior for non-owned or already-completed tasks.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required parameter, so the description must compensate. It implies the parameter refers to a task ('a task and every task it delegated', 'leaves first'), but adds no id format, source, or lookup detail beyond that inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (terminate) and resource (a task and its entire delegated subtree), plus the cascade order and return shape. Nothing else in the sibling list (delegate, wait, check, result, list_tasks) does tree-wide cancellation, so an agent can distinguish it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the ownership rule ('A lead may cancel only its own') and the cascading scope tell the agent when cancellation is appropriate, but there is no explicit when-to-use guidance versus alternatives like waive_review or check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkC
The status of a task, how long it has been running, and the last lines of its engine's own event stream. Reconciles nothing, and records the stall or the revival its clock reads.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | No | ||
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it falls short. The phrase 'records the stall or the revival its clock reads' hints at possible state mutation or side effects without clarifying them, which is ambiguous and potentially misleading for what appears to be a status read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, but the second is ornamental pseudo-technical prose ('records the stall or the revival its clock reads') that consumes space without adding actionable meaning. The useful content (status, uptime, log tail) is buried in an indirect sentence rather than front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a task-status tool with no output schema, no annotations, and two undocumented parameters, the description is inadequate. It neither documents the parameters nor clarifies the return shape, side effects, or relationship to sibling task tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters (task_id, lines), and the description adds nothing about them. It never explains that task_id identifies the task to check or what 'lines' controls (presumably the number of event-stream lines returned).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description does convey the resource (a task) and what is returned: its status, running duration, and tail of its engine's event stream. However, the tool name 'check' is generic and the second sentence ('Reconciles nothing...') is cryptic prose that obscures rather than clarifies purpose. No differentiation from siblings like list_tasks, result, or wait.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no mention of alternatives, despite siblings (wait, result, list_tasks) that clearly overlap. The only hint is the vague 'Reconciles nothing,' which does not tell an agent when to pick this over a blocking wait or a result fetch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegateC
Launch a specialist for a role on a brief in a working directory, returning its task id. Validates the role, the workspace, its reservation, duplicates and the resume binding first. A role with no binding takes its engine in the call; worktree: true gives a role that works at the project root a writable task worktree of its own instead. A role bound to a list of bindings is delegated by seat, 1-based.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | ||
| role | Yes | ||
| seat | No | ||
| brief | Yes | ||
| force | No | ||
| model | No | ||
| branch | No | ||
| effort | No | ||
| engine | No | ||
| resume | No | ||
| worktree | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full behavioral burden. It does disclose some behavioral traits: it validates role, workspace, reservation, duplicates, and resume binding before launching, and it explains that 'worktree: true' creates a writable task worktree. However, it omits critical details like permissions required, what happens on validation failure, whether the launch is asynchronous, and error handling. The description adds useful context but is incomplete for a complex tool with 11 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, packed into three sentences that front-load the main action and then cover validation and special cases. It avoids filler, and each sentence adds information. However, the structure could be improved with clearer separation between core functionality and parameter explanations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (11 parameters, 0% schema coverage, no annotations, no output schema), the description is incomplete. It fails to explain most parameters, validation outcomes, error behavior, and return details beyond 'task id'. For a tool that launches tasks and interacts with workspaces, this leaves significant gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only explains a few parameters: 'worktree: true' and 'seat' (1-based for bound roles). It does not describe cwd, role, brief, force, model, branch, effort, engine, or resume. The description adds some meaning for worktree and seat but leaves 9 parameters undocumented, which is insufficient at this coverage level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Launch a specialist for a role on a brief in a working directory, returning its task id.' This is clear and distinguishable from siblings like list_roles or run_command, though it doesn't explicitly name alternatives. The core purpose is evident, but the specialized terminology (role, brief, seat) assumes domain familiarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is offered on when to use this tool versus alternatives such as run_command or check. There are no prerequisites or exclusions stated. The description mentions validation but doesn't explain when a delegation is appropriate or when other tools should be used instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_modeA
The active mode's loop text, its roles with their workspace, sandbox default and prompt, its git policy, and projectRoot, the root of the project this server serves. Call this first: it is how a launcher learns the loop, which is served rather than copied. review carries the loop's two settings, limits.planReviewRounds and review.afterResolver.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the informational payload, implying a read, but never states that the call is side-effect free, what happens when no mode is active, or whether the result can error. The 'served rather than copied' remark is useful context but is about the loop, not the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a long run-on list ('its roles with their workspace, sandbox default and prompt, its git policy, and projectRoot') that is hard to parse, and the second sentence stacks an instruction, a rationale, and a note about review settings. Nothing is padding, but the structure resists scanning and buries the call-first instruction mid-paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no annotations, and no params, the description does the right work by enumerating the returned fields (loop text, roles, sandbox, git policy, projectRoot, review limits). It omits error/empty-mode behavior, which is the remaining gap for a zero-input introspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so baseline is 4. The description correctly implies a no-argument call by never referencing inputs, though it could have said explicitly that it takes none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (describe) and resource (the active mode) and enumerates the payload — loop text, roles, git policy, projectRoot — which is more than a tautology. It is distinguishable from siblings like git_root and list_roles, though it never defines what a 'mode' or 'loop' is, leaving the core noun under-specified for a reader who doesn't already know the domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Call this first: it is how a launcher learns the loop, which is served rather than copied" gives an explicit ordering instruction and a rationale, which is genuinely actionable. It stops short of naming alternatives or stating when NOT to call it, so it is clear context rather than full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_mutateB
Run one git subcommand in a verified worktree, under the project's locks and journaled. The only path that writes a worktree's git metadata; the workspace and branch default to the mode's own git policy.
| Name | Required | Description | Default |
|---|---|---|---|
| args | Yes | ||
| path | No | ||
| slug | Yes | ||
| branch | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful traits: worktree must be verified, the call runs under project locks and is journaled, and it mutates git metadata (implying the operation is state-changing). However, it omits auth/permission requirements, reversibility or recovery after the journal, and failure behavior for an unverified worktree — significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the action and scope, with no filler. It is jargon-heavy ('journaled', 'the mode's own git policy') but every clause carries weight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a journaled mutation tool with no annotations, no output schema, and 0% parameter coverage, the description covers behavioral framing but not the practical contract: what slug identifies, how args maps to the 'one subcommand', what path does, and what happens on failure. The agent has to guess at the call signature.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and four parameters exist, so the description must do the explaining. It only touches the branch default ('the workspace and branch default to the mode's own git policy'), leaving slug, args (the actual subcommand + argv), and path entirely undefined. Required parameters slug and args are the most consequential and get no semantic help at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Run one git subcommand in a verified worktree') and stakes out a clear scope claim ('the only path that writes a worktree's git metadata'), which separates it from siblings like git_root and verify_worktree. It stops short of naming those siblings explicitly, which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage through the exclusivity claim — if you need to write git metadata, this is the only path — but never states when not to use it or which sibling handles read-only git inspection. A conditional hint ('the workspace and branch default to the mode's own git policy') gestures at routing but leaves the agent to infer alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
git_rootA
Run one whitelisted git verb at the project root, under the project's git lock and, for a verb that changes the repository, the repository's own lock, and journal the step it completes. A merge --ff-only refuses a head without a tested step (unless testCommand is none), and for a team task in a gating mode, without every seat's clean review or a waiver. The verbs are worktree add -b, worktree remove, branch -d, merge --ff-only, rebase --abort, and the read-only status, log, rev-parse, rev-parse --abbrev-ref HEAD, merge-base, branch --list and worktree list; a verb that journals a step names the slug whose journal it belongs to.
| Name | Required | Description | Default |
|---|---|---|---|
| args | Yes | ||
| slug | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses lock acquisition (project lock plus repository lock for mutating verbs), step journaling, the whitelist enforcement, and the review/test preconditions that cause merge --ff-only to refuse. It stops short of describing the refusal error shape or return format, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first clause and every subsequent clause adds real constraint (locks, journaling, gating, verb whitelist). The first sentence is dense and slightly run-on, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation-capable tool with no annotations, no output schema, and 0% schema description coverage, the description supplies the critical missing context: allowed verbs, locking, journaling, and refusal conditions. Only the response/error shape and the exact args format remain undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with two params, so the description must compensate. It clarifies that a step-journaling verb must name the slug whose journal it belongs to, which explains the slug parameter, and the verb whitelist implicitly defines the args array. However, it never explicitly states that args carries the verb and its arguments, nor the expected ordering/format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: run one whitelisted git verb at the project root, and enumerates the permitted verbs, so an agent knows precisely what this tool executes. It does not, however, differentiate itself from the sibling git_mutate or run_command, leaving overlap unresolved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage constraints (exactly one whitelisted verb per call, specific gating conditions for merge --ff-only, journaling only for step verbs), which is more than nothing. But it never says when to choose this tool over git_mutate or run_command, nor does it state exclusions, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_rolesC
The roles the active mode has: each one's workspace and the sandbox profile it will run under, with the engine, model and effort .cross-agent/config.json binds it to — or binding: null where it binds none, which is the engine a delegate call must name itself. A role bound to a list is answered as seats, each with its own engine, model, effort and sandbox.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it does disclose the return shape in detail: workspace, sandbox profile, engine/model/effort binding, binding:null semantics, and the 'seats' expansion for list-bound roles. It adds genuinely useful context, but it omits any statement of read-only safety or permission requirements, and the sentence structure makes the behavior hard to extract.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense, grammatically awkward run-on ('the engine, model and effort .cross-agent/config.json binds it to') that buries the key fact under convoluted phrasing. It is not front-loaded with a clear statement of what the tool does, and the sentence is hard to parse on first read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must supply the return semantics, and it does attempt to describe the role/binding/seat structure in full. However, the explanation is confusingly worded and omits the read-only nature and any usage context, leaving the agent with an incomplete picture for a tool whose only purpose is its output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify; per the rubric a 0-parameter tool earns a baseline 4. No parameter information is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name list_roles is clear, and the description implies it returns the roles of the active mode along with their workspace, sandbox profile, and engine/model/effort binding. However, it opens with a noun fragment ('The roles the active mode has:') rather than a verb+resource statement, and it never distinguishes this tool from siblings like describe_mode, so the purpose is inferable but not stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no comparison to alternatives such as describe_mode. The only indirect routing hint is the note that a null binding is 'the engine a delegate call must name itself,' which gestures at delegate but does not tell the agent when to call list_roles.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksC
Every task of this project after a reconciliation pass, newest first, with the records no reader could judge.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden; it discloses ordering (newest first), scope (every task of this project), and an unusual inclusion trait ('records no reader could judge'), but omits read-only safety, pagination, filtering behavior, and whether it mutates. It adds some behavioral context but is far from complete for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. However, its compressed jargon ('reconciliation pass', 'records no reader could judge') sacrifices clarity, so it is concise but not optimally structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter list tool with no annotations and no output schema, the description should explain the status filter, return scope, and what 'reconciliation pass' and 'records no reader could judge' mean. It only gives ordering and project scope, leaving an agent unable to invoke it confidently with a status filter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole parameter, status, is never mentioned in the description. Worse, 'Every task' could mislead an agent into ignoring the status enum filter. No meaning is provided for status beyond the bare enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (tasks of this project) and ordering (newest first), so an agent can infer it is a listing tool, but it never states the verb 'list' explicitly and adds obscure qualifiers ('after a reconciliation pass', 'records no reader could judge') that obscure rather than sharpen purpose. It does not distinguish itself from sibling list_roles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no alternatives (e.g., list_roles, check, result), and no exclusions. The only contextual cue is 'after a reconciliation pass,' which is unexplained and does not tell the agent when to choose this tool over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resultB
The final message of a settled task in full, with the engine session it ran under. A task still running answers with its status.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does add real behavior: it returns the message in full including the engine session, and an unsettled task yields status rather than an error. It stops short of permissions, size limits, or output format details, leaving meaningful gaps for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the primary return value front-loaded and the edge-case behavior second. Nothing is wasted, though the compactness is partly under-specification rather than pure efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, single-parameter read tool with no output schema, the description covers what is returned and the running-task fallback, which is the essential context. It omits any explanation of the task_id parameter and return format specifics that would otherwise be carried by an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single task_id parameter is undocumented in the schema. The description never mentions task_id or how to obtain it, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and scope: the full final message of a settled task plus the engine session it ran under. An agent can tell this retrieves a completed task's output, but the description never distinguishes it from siblings like check, wait, or list_tasks, which also concern task state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no alternative tool is named. The clause about running tasks implies a usage condition (call it when you want a result and accept status if unsettled), but that is inference, not instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_commandA
Run this project's configured test or setup command — by selector, never as a command string — at the project root or for a verified worktree, returning the exit code and the last 64 KB of its output. Worktree tests run in a detached checkout of the branch head under .cross-agent/gate/ and journal tested on success. A passing test run at the root after the merge journals the slug's tests-passed step.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | ||
| where | Yes | ||
| which | Yes | ||
| timeout_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers real behavioral detail: the return payload (exit code plus last 64 KB of output), where worktree tests execute (detached checkout of the branch head under .cross-agent/gate/), and a side effect (journaling tested on success). It still omits timeout behavior and whether 'setup' commands mutate anything, so it falls short of fully compensating for the missing annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the selector constraint are front-loaded in the first clause, with the return payload defined before the operational nuances. It is dense and jargon-heavy ('journals the slug's tests-passed step'), but nearly every clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description correctly specifies the return value (exit code + truncated output), and it covers the root-vs-worktree execution model and journaling side effects. It is close to complete, missing only timeout semantics and any prerequisite/permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for most params: 'which' is explained as the test/setup selector, 'where' as project root or verified worktree, and 'slug' is tied to journaling ('the slug's tests-passed step'). Only 'timeout_seconds' is left unexplained, making coverage strong but incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Run this project's configured test or setup command') and immediately scopes it ('by selector, never as a command string'), which distinguishes it from any generic shell-execution sibling. It also names the target ('project root or a verified worktree'), so an agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies context via 'project root or for a verified worktree' and the test/setup selector, but never states when to use this versus alternatives like verify_worktree or git_mutate, nor any prerequisites or when NOT to use it. Usage is inferable but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_worktreeC
Verify a linked worktree and its exact branch, returning canonical Git and worktree paths or a refusal reason.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| branch | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It does disclose a meaningful behavioral trait — that the call can succeed with canonical paths or fail with a refusal reason — but it omits whether the operation is read-only, what permissions it needs, and what conditions trigger a refusal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the verb and outcome front-loaded. Nothing is wasted, though the sentence tries to carry purpose and return contract at once.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, and 0% parameter coverage mean the description must do more than this. It never explains the refusal conditions, path format expectations, or read-only nature of the operation, which are exactly the gaps structured data leaves open.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both required parameters. The description implies 'path' is a worktree path and 'branch' must be 'exact', which adds some meaning, but it never explains format, matching semantics, or what happens when the branch does not match.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (verify) and resource (linked worktree and its exact branch), and even names the outcome shape (canonical paths or refusal reason). It does not, however, distinguish itself from neighboring tools like git_root or check, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as git_root or check, and no prerequisites or preconditions are given. The only guidance is implicit in the word 'verify'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
waitB
Wait for a task to settle, for its engine to go quiet for the configured stall threshold, or for the timeout, and answer with the call to make next. A lead may wait only on a task it delegated.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | ||
| timeout_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful traits: the call blocks until stall threshold or timeout, it returns 'the call to make next' rather than task data, and it is restricted to a lead that delegated the task. It stops short of saying what happens when the timeout elapses (error vs. partial result) or what the default threshold/timeout is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence that front-loads the purpose and ends with the useful return-value note ('answer with the call to make next'). No filler, though the enumerated settle conditions make it slightly hard to parse on first read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a blocking tool with no annotations and no output schema, the description explains the blocking semantics and return nature reasonably well, but omits timeout behavior, default values, and any error/edge-case outcomes, leaving gaps an agent would need to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only obliquely references 'the timeout' and 'a task' without naming task_id or timeout_seconds, their formats, units, or defaults. The two undocumented parameters are effectively left to the agent's guesswork.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (wait) and resource (task) and enumerates the three settle conditions: task settles, engine goes quiet past the stall threshold, or timeout. It does not name the sibling tools (check, result) that an agent might otherwise poll with, so the differentiation from those alternatives is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'A lead may wait only on a task it delegated' gives a real precondition for use, which is useful context. However, it never says when to prefer this over check/result/cancel, so the agent must infer the choice between waiting and polling from the sibling names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
waive_reviewB
Record the waiver of the merge's review guard for a task's branch head, as a review-waived step in its journal: git_root merge then takes that head without every seat's clean review. The commit must name the branch head as it stands. The operator's, and a lead's only under review.afterResolver lead-decides once every seat has finished its review of that head.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | ||
| commit | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the key behavioral consequence that git_root merge then takes the branch head without every seat's clean review, and notes that the waiver is recorded as a journal step. However, it does not clearly state permissions, reversibility, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loads the core action and effect. However, the final authorization sentence is grammatically broken and hard to parse, which weakens the structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0% schema description coverage, the description should be fairly comprehensive. It conveys the critical effect on git_root merge and the commit requirement, but leaves authorization rules and the slug parameter only partially explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds meaning for commit by requiring it to name the branch head as it stands, and implies slug identifies the task's branch head. The meaning of slug is only inferable rather than explicitly defined, so the compensation is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: recording the waiver of the merge's review guard for a task's branch head. It also distinguishes the effect from git_root merge, which then takes the head without every seat's clean review. The purpose is clear despite the dense phrasing, though it never explicitly contrasts itself with named siblings like check or result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers no clear when-to-use guidance or alternatives. The authorization sentence about the operator and a lead under review.afterResolver lead-decides is too garbled to serve as a usable condition or prerequisite.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.1.0- First observed
cancel - First observed
check - First observed
delegate - First observed
describe_mode - First observed
git_mutate - First observed
git_root - First observed
list_roles - First observed
list_tasks - First observed
result - First observed
run_command - First observed
verify_worktree - First observed
wait - First observed
waive_review
TDQS
Scored across 13 tools
Most tools target distinct resource+action pairs, but the task-inspection cluster (wait, check, result, list_tasks) has some conceptual overlap, and git_mutate vs git_root differ mainly by scope rather than action. Descriptions do clarify the boundaries (wait = block until settle, check = status snapshot, result = full final message), so an agent can generally choose correctly.
The set mixes compound verb_noun names (list_roles, list_tasks, verify_worktree, git_mutate, run_command, waive_review) with bare verbs/nouns (delegate, wait, check, cancel, result) and a describe_mode outlier. It remains readable, but there is no single predictable pattern.
13 tools is a reasonable scope for an orchestration server covering mode/role discovery, task lifecycle, git, worktree, and review concerns. Each tool earns its place, though the inspection tools are slightly numerous for their overlap.
Coverage is strong: role/mode discovery, full task lifecycle (delegate, wait, check, result, cancel, list_tasks), worktree verification, git mutation at both root and worktree, test execution, and review waiver. Gaps are minor, e.g. no obvious tool for reading or replying within a running task's stream beyond check.
Maintenance
Related MCP Connectors
The team layer for AI coding agents: shared contracts, collision alerts, E2EE sessions.
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Real-time chat for AI agents. Claude Code, Cursor, Cline and Codex join channels over MCP.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceOrchestrates multiple AI models (Gemini, OpenAI, Claude, local models) within a single conversation context, enabling collaborative workflows like multi-model code reviews, consensus building, and CLI-to-CLI bridging for specialized tasks.-
- AlicenseAqualityAmaintenanceEnables one AI coding agent to delegate tasks to, and build consensus across, multiple other coding CLIs (Claude Code, Codex, etc.) by orchestrating them as headless subprocesses.187MIT
- AlicenseNot gradedqualityAmaintenanceEnables coordinating Claude Code and Codex across separate Git worktrees with shared issue ownership, file reservations, messages, and explicit handoffs.1,876 PyPI4MIT
- AlicenseNot gradedqualityBmaintenanceCoordinates parallel Claude Code agents on shared repositories and orchestrates multiple Claude subscriptions, enabling live session handoff, conflict detection, and remote terminal control.14 npmMIT