throng
throng is an MCP server that hands a task from one AI agent to another (Claude Code, Codex, OpenCode, Gemini CLI) in a real, resumable session, optionally with schema-constrained output.
Start a nested agent —
run_throngletpicks a harness/model/effort (e.g.codex/gpt-6-sol:high), takes a self-contained prompt and acwd, and returns the agent's final message plussession_id, stop reason, usage and duration.Get data, not prose — pass a JSON Schema and the result comes back as
structuredoutput that matches it.Continue a conversation —
send_messageruns the next turn in an existing session, so the nested agent keeps its prior context;steer: trueinterrupts the running turn and delivers the message next, ahead of queued messages.Work in the background —
background: trueon either run returns at once with{session_id, state, queued}; run several agents concurrently.Collect results later —
wait_throngletwaits until the session has no running or queued turn and returns the last result (idempotent); on timeout it returns session state instead of an error.See what's running —
list_throngletslists sessions with description, agent, cwd, state (running/queued/idle/failed), queue length, timestamps and last error.Stop a turn —
cancel_throngletcancels the running turn and drops queued messages, leaving the session idle and usable again.Discover options —
list_harnessesreports installed harnesses with their models and effort levels, unavailable ones with reasons, and server limits.Bound runs — per-call
timeout_slimits wall-clock time (config default 21600s); permissions come only from config, never from tool parameters.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@thronguse codex to add a --verbose flag to src/cli.ts and run the tests in /work/my-app"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
throng-mcp
An MCP server that lets your agent hand work to another one. Any MCP client can call it, and the task goes to any supported harness.
Another model's view. A review by another vendor's model, a design critique, a cheaper model for a mechanical pass.
Second opinions. Hand one agent's result to another: Codex reviews, Claude fixes, Codex checks again, each in its own long-lived session.
Real sessions, not one-shots. The nested agent keeps running; a follow-up goes to the same agent with everything it already knows.
Background work. Several agents run at once while you carry on.
Answers by schema. Pass a JSON Schema and the result comes back as data that fits it, not prose to parse: a list of findings, a verdict, a plan, ready to feed into the next step.
Supported harnesses: Claude Code, Codex, OpenCode (and every model it can reach), Gemini CLI. Setup for each is under Install.
Example
You're in Claude Code and have just changed the payment flow.
Ask Codex to find a way this payment flow could charge someone twice. Don't change any files.
Claude calls run_thronglet with agent: "codex/gpt-6-sol:high", a self-contained prompt, the repository path and a schema, so the findings come back as data: file, line, steps to reproduce.
Give the findings to Opus. Have it fix each one, add a test for it and run the tests.
A second run_thronglet, with agent: "claude/opus". Codex's findings go into its prompt; Opus edits the code and runs the tests.
Now show the same Codex session what changed. Can it still make a customer pay twice?
That is send_message into the first session: Codex still has its findings in context and checks the fixes against them instead of starting over.
Any of these calls takes background: true: it returns at once and wait_thronglet collects the result later, which is how several agents work while you carry on.
Related MCP server: pokeclaw
Install
Three parts, in order: the server, the agents it may run, and the client it is called from. The skill at the end is optional.
The plugin covers parts 1, 3 and 4; the agents from part 2 you still install yourself.
claude plugin marketplace add agent-runbooks/throng-mcp
claude plugin install throng@throng-mcp1. The server
The npm package throng-mcp. Needs node ^22.13 || >=24. Tested with node 24.11.1, claude 2.1.282, codex 0.156.1, opencode 1.18.30.
Two ways to run it:
npx -y throng-mcpstraight in the client config, below. Nothing to install: npx fetches the package on the first start and caches it. The package has no dependencies, so that is one tarball and nothing else.A global install, then the command is
throng-mcp. Starts faster than going through npx.npm i -g throng-mcp
2. The agents to run
Each agent needs its own CLI installed and logged in. throng talks to agents over ACP (Agent Client Protocol); Claude Code and Codex need an ACP adapter, and throng ships none. Install only the agents you want to delegate to.
npm i -g @agentclientprotocol/claude-agent-acp
claude auth statusthrong passes the path of the claude it finds on PATH to the adapter (CLAUDE_CODE_EXECUTABLE), so the nested agent runs your installed, logged-in CLI. If claude auth status says not logged in, put CLAUDE_CODE_OAUTH_TOKEN into the config, see docs/configuration.md.
npm i -g @agentclientprotocol/codex-acp
codex loginThe adapter gets the path of your codex the same way (CODEX_PATH).
OpenCode speaks ACP itself (opencode acp), so there is no adapter. Install it per https://opencode.ai/docs, then:
opencode auth loginCustom providers live in your ~/.config/opencode/opencode.json; throng doesn't touch it.
Gemini CLI speaks ACP itself (gemini --acp), so there is no adapter.
npm i -g @google/gemini-cli
geminiRun gemini once and sign in; the nested agent uses that login. Checked against Gemini CLI 0.61.0 only up to the first model request (handshake, models, modes); a full turn has not been run yet, so expect rough edges and please report them. Limits:
One turn per session: Gemini CLI can't resume a session, so
send_messageto it fails withsession_not_foundbefore anything runs.steer: trueis refused the same way and leaves the running turn alone; to stop a gemini turn, usecancel_throngletand start a new run.list_throngletsshows such a session withaccepts_messages: false.No effort levels: a
:<effort>suffix is ignored with a warning.No usage numbers: the result's
usagestays empty.The
cwdyou give it is trusted for the run (GEMINI_CLI_TRUST_WORKSPACE=true), under every permission policy.
3. The MCP client
The client is the session that calls throng. It may be the same program as one of the agents above, or a different one. The server speaks stdio and the command is npx -y throng-mcp, or throng-mcp after a global install.
claude mcp add --scope user throng -- npx -y throng-mcp--scope user registers it for every project; without it, for the current project only. claude mcp list shows the result.
codex mcp add throng -- npx -y throng-mcpIn ~/.config/opencode/opencode.json:
{
"mcp": {
"throng": {
"type": "local",
"command": ["npx", "-y", "throng-mcp"],
"enabled": true
}
}
}Any other MCP client: register a stdio server with that command.
4. The skill, optional
skills/throng tells the calling agent how to use the server: when to delegate, how to write the prompt, background runs, follow-ups, structured output, what a refused permission means. It is a file for the agent and does not install the server.
npx skills add agent-runbooks/throng-mcp --skill throng -g -a claude-code -y-g installs into the agent's user directory, for every project; without it, into the current project. -a codex or -a opencode for the other agents. npx skills update pulls later changes.
Copy skills/throng into your agent's skills directory: ~/.claude/skills, ~/.codex/skills, ~/.config/opencode/skills. A global install also leaves it under the installed package, $(npm root -g)/throng-mcp/skills/throng.
Agents pick skills up at session start, so open a new session after installing.
Check
In a new session, ask the agent to call list_harnesses. It starts each installed adapter without a prompt (seconds, no tokens) and lists every available harness with its models and effort levels; unavailable names what is missing and how to install it. If a run then fails during the handshake, the usual cause is auth: claude auth status, codex login, opencode auth login, gemini (sign in once). More in troubleshooting.
Before the first run, two things to know. The agent edits the directory you name, with your user's rights; throng adds no isolation and rolls nothing back. Give it only trees you would let an agent edit unattended, and give parallel writers a worktree each.
Using throng
agent names harness, model and effort in one string, <harness>/<model>[:<effort>]:
claude/opus:max
codex/gpt-6-sol:xhigh
opencode/openrouter/z-ai/glm-5.3-flash
gemini/gemini-2.5-proThe model is one of the harness's own values; list_harnesses has the current list. Effort is low | medium | high | xhigh | max; omitted means the harness default.
A call:
{
"agent": "codex/gpt-6-sol:high",
"prompt": "In this repository, add a --verbose flag to the CLI in src/cli.ts and a test for it. Run pnpm test. Leave the changes uncommitted and end with the list of files you changed.",
"cwd": "/work/my-app",
"description": "add --verbose flag"
}The prompt is self-contained: the nested session sees nothing of the calling conversation. The call returns the agent's final message, a session_id for follow-ups, and why the turn stopped.
tool | what it does |
| starts a session and runs the first turn |
| runs the next turn in an existing session; |
| collects the result of a turn started with |
| the sessions and what each is doing |
| stops a session's running turn |
| the installed harnesses, their models and effort levels |
background: true on run_thronglet or send_message returns as soon as the turn is running; wait_thronglet collects the result later, which is how several agents run at once. schema asks for structured output: the agent fills a JSON Schema and the result comes back as structured instead of text.
A failure is an MCP tool error with code, message and, when the session exists, session_id and the partial text. Every field, error code and stop reason of every tool is in DESIGN §3.
Permissions and safety
What a nested agent may do comes from the config file, never from a tool parameter, so the calling model cannot grant itself more than you allowed. Optional ~/.config/throng/config.yaml:
permissions: autoauto, the default: each harness runs in its own auto-approve mode (Claudeauto, Codexagent, OpenCode as configured, Geminiyolo); whatever that mode still asks about, throng refuses.allow_all: every request allowed, once.deny_all: every request refused.elicit: each request is shown to you as a dialog in the MCP client, with the one-time choices the harness offered. Needs a client with elicitation support; Claude Code has it.
throng never answers "always allow", so no rule gets written into the agent's project settings. Per-harness overrides, extra env for an adapter, timeouts and the nesting limit are in docs/configuration.md.
Documentation
docs/configuration.md: the config file, environment variables, permissions in detail, auth, troubleshooting.
DESIGN §3: the exact inputs, results, error codes and stop reasons of every tool.
skills/throng/SKILL.md: working patterns for the calling agent.
docs/development.md: for maintainers; tests, smoke runs against real harnesses, files on disk.
Agent Runbooks: multi-step procedures a session runs through subagents. Any step of a runbook can go to any harness through throng.
License
MIT, see LICENSE.
Available Tools
6 toolscancel_throngletA
Cancels the session's running turn (session/cancel) and drops its queued messages; a pending wait_thronglet returns the cancelled error. The session is idle afterwards and accepts a new send_message. No-op on an idle session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | session_id from run_thronglet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses that queued messages are dropped, that a pending wait_thronglet surfaces a cancelled error, and that the session is idle afterward and accepts new messages. It does not mention authorization or permissions, leaving one gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the primary effect and back-loaded with post-conditions and the no-op case. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no annotations and no output schema, the description covers the key effects, side effects on queued messages and waiting callers, and the resulting state. Only the permission/authorization angle is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single session_id parameter, and the schema already ties it to run_thronglet. The description adds no syntax or format detail beyond that, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (cancels) and resource (the session's running turn) and even names the underlying operation (session/cancel). It also references siblings wait_thronglet and send_message, so an agent can distinguish this tool's role from theirs without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when the tool matters and explicitly notes the no-op case on an idle session, which prevents pointless calls. It stops short of naming an alternative tool or exclusion for other cancel-like scenarios, so it is not a full when/when-not routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_harnessesA
Lists valid agent values for run_thronglet: available harnesses with their models and effort levels, unavailable ones with the reason, and the server limits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full disclosure burden and does well: it reveals that both available and unavailable harnesses are returned, that unavailability comes with a reason, and that server limits are included. It does not explicitly state the operation is read-only, though 'Lists' strongly implies a safe read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The purpose and the returned structure are packed into one clean clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description adequately explains the return payload, covering available, unavailable, and limit categories. Minor gaps remain about exact format or size of the response, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. The description correctly implies no filtering or selection arguments are needed to obtain the full harness list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (harnesses), then enumerates exactly what is returned: available harnesses with models and effort levels, unavailable ones with reasons, and server limits. This clearly separates it from siblings like list_thronglets, since it is specifically about agent values consumed by run_thronglet.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
By framing the output as the 'valid agent values for run_thronglet,' it implies the tool should be called to discover acceptable inputs before invoking run_thronglet. There is no explicit when-not or named alternative, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_throngletsA
Lists thronglet sessions on this machine: description, agent, cwd, state (running | queued | idle | failed), queue length, accepts_messages (false: send_message to it fails), timestamps and the last error. Live state is this server's; a session run by another throng instance shows what its record says.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a good job: it defines the state enum, explains that accepts_messages=false means send_message will fail, and discloses the important caveat that live state belongs to this server while sessions from other throng instances may be stale records. It stops short of stating read-only/non-mutating nature explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the verb and scope come first, followed by the returned fields and the freshness caveat. The long field enumeration is justified because there is no output schema to carry it, though it reads as a run-on list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must document return values itself, and it does so field by field, including the error field and cross-instance freshness semantics. An agent has everything needed to call and interpret this zero-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the schema to describe and the baseline is 4. No parameter-level ambiguity exists for the agent to resolve.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (lists) and resource (thronglet sessions) with a clear scope qualifier ('on this machine'), and enumerates the returned fields. It does not explicitly name the sibling tools it is distinct from, but 'list' versus run/cancel/send/wait is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer this is the discovery/inspection call before acting on a session. There is no explicit when-to-use, when-not-to-use, or pointer to an alternative sibling, so guidance is only adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_throngletA
Runs an agent of an installed harness (see list_harnesses) on a task in cwd and returns its final message as JSON {session_id, text, stop_reason, usage, duration_s, warnings?}; with schema, structured replaces text. description names the thronglet for listings. With background: true the call returns {session_id, state, queued} as soon as the turn runs; collect the result with wait_thronglet.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | Absolute path; the agent works in this tree | |
| agent | Yes | <harness>/<model>[:<effort>], e.g. claude/opus:max, codex/gpt-6-sol:xhigh; valid values: list_harnesses | |
| prompt | Yes | Task for the agent | |
| schema | No | JSON Schema (draft-07 or 2020-12) for structured output: the agent submits a matching result, returned as `structured` instead of `text` | |
| timeout_s | No | Wall-clock limit for the run in seconds; default from config (21600) | |
| background | No | Return as soon as the turn is running (or queued behind the session's current turn); collect the result with wait_thronglet | |
| description | Yes | What this thronglet is for, in a few words; shown in session listings |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does so well: it documents both return shapes ({session_id, text, stop_reason, usage, duration_s, warnings?} and the structured variant), the background contract ({session_id, state, queued}), and the 21600s default timeout. It omits side-effect/risk disclosure (the agent executes commands in the user's tree) and failure/timeout behavior, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return contract, then progressive detail on structured output and background mode. Dense and semicolon-heavy, but nearly every clause carries operational information; the listing-name note is the only mildly incidental line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no output schema and no annotations, the description covers return values for both sync and background modes plus the default timeout, which is what an agent needs to call it. Gaps remain around error/timeout handling and the privilege implications of running an agent in a working tree.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning the schema alone does not: that schema swaps `structured` in place of `text`, that background changes the returned shape and requires wait_thronglet, and that description surfaces the run in session listings. Only `cwd` and `timeout_s` get no added context beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Runs') plus resource ('an agent of an installed harness') and scope ('on a task in cwd'), then distinguishes itself by pointing at list_harnesses for valid agents and wait_thronglet for async collection. An agent can separate this from send_message/cancel_thronglet without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear routing for two cases: fetching valid agents via list_harnesses and retrieving async results via wait_thronglet when background: true. It does not, however, contrast itself with send_message (continuing an existing session) or state when not to launch a new run, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_messageA
Sends the next message into an earlier session (session_id from run_thronglet) and returns the same JSON as run_thronglet. A message to a session whose turn is still running waits for that turn to end: turns on one session never overlap. With background: true the call returns {session_id, state, queued} as soon as the turn runs; collect the result with wait_thronglet. steer: true interrupts the running turn and delivers this message next; queued messages follow it. A session whose harness cannot resume takes no further message, steer included: list_thronglets shows it as accepts_messages: false.
| Name | Required | Description | Default |
|---|---|---|---|
| steer | No | Interrupt the running turn (session/cancel) and run this message as the very next turn, ahead of queued messages. The in-flight tool call is aborted; a half-applied edit may remain. | |
| prompt | Yes | Next message for the agent | |
| schema | No | JSON Schema (draft-07 or 2020-12) for structured output: the agent submits a matching result, returned as `structured` instead of `text` | |
| timeout_s | No | Wall-clock limit for the run in seconds; default from config (21600) | |
| background | No | Return as soon as the turn is running (or queued behind the session's current turn); collect the result with wait_thronglet | |
| session_id | Yes | session_id from run_thronglet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses that turns on one session never overlap, that a message to a busy session blocks until the turn ends, that background mode changes the return shape and requires wait_thronglet to collect, that steer interrupts the running turn and jumps the queue, and that non-resumable sessions reject all messages including steer. These are exactly the traits an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core purpose and the run_thronglet return-equivalence come first, followed by the blocking rule, then background, then steer, then the terminal limitation. Every clause carries information; there is no filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutating tool with nested schema input, no annotations, and no output schema, the description covers the hard parts: concurrency, interrupt semantics, async collection, and the dead-session case. Minor gaps remain — nothing on error/timeout behavior and only a pointer ('same JSON as run_thronglet') for the default return payload — but nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real semantics beyond the schema for two parameters: it explains what background:true actually returns and how to retrieve the result, and clarifies steer's queue-ordering effect relative to pending messages. session_id's provenance is also stated. timeout_s and schema remain schema-only, which keeps this from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Sends the next message into an earlier session') and immediately anchors it to the sibling it depends on ('session_id from run_thronglet'), which cleanly separates it from run_thronglet (start a session) and wait_thronglet (collect a result). An agent can pick this over its siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditional guidance: use background:true when you want an early {session_id, state, queued} return and then wait_thronglet, use steer:true when you must interrupt the in-flight turn, and note that sessions showing accepts_messages: false in list_thronglets cannot receive messages at all. It does not spell out the inverse case (when to prefer run_thronglet over this), so it falls just short of fully explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_throngletA
Waits until a session has no running or queued turn and returns the last turn's result: the same JSON as run_thronglet, or its error. With timeout_s elapsed returns {session_id, state, queued} instead. Idempotent: the result is stored with the session.
| Name | Required | Description | Default |
|---|---|---|---|
| timeout_s | No | How long to wait in seconds; default from config (21600). Elapsed → the session state, not an error | |
| session_id | Yes | session_id from run_thronglet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the timeout behavior (returns {session_id, state, queued} instead of an error), idempotency, and that results are persisted with the session. It stops short of auth requirements or rate limits, but the core behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with zero waste; the primary behavior is front-loaded and the timeout and idempotency caveats follow logically. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain return values, and it does: both the success path (run_thronglet JSON or its error) and the timeout path ({session_id, state, queued}). It leaves the possible 'state' values undefined, but for a wait primitive this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema, and the description's timeout detail (state returned instead of error) mirrors the schema description rather than adding new meaning. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (waits) and resource (session turn), and explicitly distinguishes itself from run_thronglet by noting it returns 'the same JSON as run_thronglet, or its error'. An agent can tell what this does and how it relates to the sibling without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tie to run_thronglet ('session_id from run_thronglet', 'same JSON as run_thronglet') implies this is used to await a previously started run, but there is no explicit when-to-use or when-not guidance, nor a named alternative (e.g., polling run_thronglet). Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.1- Changed
run_thronglet1 field changed- changed
Input schema / properties / agent / descriptionPrevious value: -"<harness>/<model>[:<effort>], e.g. claude/opus[1m]:max, codex/gpt-6-sol:xhigh; valid values: list_harnesses"New value: +"<harness>/<model>[:<effort>], e.g. claude/opus:max, codex/gpt-6-sol:xhigh; valid values: list_harnesses"
6 tool updates
v0.1.0- First observed
cancel_thronglet - First observed
list_harnesses - First observed
list_thronglets - First observed
run_thronglet - First observed
send_message - First observed
wait_thronglet
TDQS
Scored across 6 tools
Most tools target clearly distinct operations (list sessions, list harnesses, cancel, wait, run, send). The main overlap is run_thronglet vs send_message, which both start a turn and return identical JSON, though the descriptions clarify new-session vs existing-session intent.
All six tools use snake_case verb_noun form (list_thronglets, cancel_thronglet, run_thronglet, wait_thronglet, send_message, list_harnesses), which is predictable. Minor inconsistency in pluralization (list_thronglets/harnesses vs singular cancel/run/wait_thronglet) and send_message breaks the 'thronglet' noun pattern.
Six tools is well-scoped for a session-orchestration server; each maps to a distinct lifecycle action (run, send, wait, cancel, list sessions, list harnesses). No filler or redundant tools.
Covers the core lifecycle: discover harnesses, run a session, message it, wait on it, cancel it, and list sessions. Minor gap: no explicit way to fully terminate/delete a session or fetch a single session by id (only cancel the current turn, and list all).
Maintenance
Related MCP Connectors
Build and supervise fleets of agents from Claude Code, Codex or Cursor. Connects over OAuth.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Scoped agent execution. Server-side credentials, policy, budgets and verifiable receipts.
Related MCP Servers
- AlicenseAqualityDmaintenanceWraps Claude Code as tools for MCP clients, enabling autonomous coding tasks via a 4-tool lifecycle with session management, async polling, and permission controls.459 npm20MIT
- AlicenseNot gradedqualityDmaintenanceEnables MCP clients to spawn and control Codex CLI and Claude Code sessions on the host machine, with session management and filesystem access.4MIT
- AlicenseAqualityDmaintenanceEnables AI agents to interact programmatically with Claude Code CLI, managing sessions, streaming outputs, and handling permission requests.711 npm1MIT
- AlicenseBqualityBmaintenanceEnables launching and supervising local Codex and Claude Code agent sessions and one-shot tasks, exposing run, list, stop, and output operations over MCP with session-scoped permissions and sandbox controls.54,779 npmMIT