throng
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@thronguse codex to add a --verbose flag to src/cli.ts and run the tests in /work/my-app"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
throng
An MCP server for delegating coding tasks to other agents. run_thronglet starts Claude Code, Codex or OpenCode over ACP (Agent Client Protocol) in the directory you give it, runs one prompt to completion and returns the agent's final message. send_message sends the next message into that session; with background: true either call returns at once and wait_thronglet collects the result. list_thronglets shows the sessions and what each is doing, cancel_thronglet stops a session's running turn. list_harnesses shows which harnesses are installed, with their models and effort levels.
Security
There is no sandbox, worktree or isolation: the nested agent edits the live tree at
cwdwith your user's rights, and throng rolls nothing back.By default (
permissions: auto) every harness runs in its own auto-approve mode, and whatever that mode still asks about, throng refuses.The other policies are
allow_all(every request allowed once),deny_all(every request refused) andelicit(each request shown to you as a dialog in the MCP client). The policy is set in the config file, never by the calling model; see Permissions.Run throng only on trees you would let an agent edit unattended.
Related MCP server: pokeclaw
Requirements
node ≥ 24 (runs
src/*.tsdirectly via type stripping, no build step) and pnpm.The CLIs of the harnesses you want to use on PATH, logged in:
claude,codex,opencode.The ACP adapters for Claude Code and Codex (OpenCode speaks ACP itself). throng ships no adapters; you install them.
Tested with node 24.11.1, pnpm 11.10, claude 2.1.282, codex 0.156.1, opencode 1.18.30.
Install
git clone https://github.com/Nodge/throng-mcp
cd throng-mcp
pnpm installAdapters:
npm i -g @agentclientprotocol/claude-agent-acp @agentclientprotocol/codex-acpthrong passes the absolute paths of the claude / codex it finds on PATH to the adapters (CLAUDE_CODE_EXECUTABLE / CODEX_PATH), so they run your installed, logged-in CLI.
OpenCode is a single binary with ACP built in (opencode acp): install it per https://opencode.ai/docs.
Register the server in Claude Code (user scope, all projects), from the repo root:
claude mcp add --scope user throng -- node "$(pwd)/src/mcp.ts"
claude mcp listTool names in Claude Code: mcp__throng__run_thronglet, mcp__throng__send_message, mcp__throng__wait_thronglet, mcp__throng__list_thronglets, mcp__throng__cancel_thronglet, mcp__throng__list_harnesses.
Check the setup: in a new Claude Code session ask Claude to call list_harnesses. Each installed adapter is started without a prompt, so it costs no tokens; the answer lists every available harness with its models and efforts, and under unavailable what is missing and how to install it.
Skill
skills/throng is a skill for the calling agent: when to delegate, writing the prompt, background turns and wait_thronglet, follow-ups, steer, housekeeping, structured output and permissions. Install it with the skills CLI:
npx skills add Nodge/throng-mcp --skill throng -g -a claude-code-g installs it for every project (user scope); without -g it goes into the current project only.
You run this yourself; throng itself never touches ~/.claude. Claude Code picks skills up at session start, so start a new session after installing.
Agent spec
agent names harness, model and effort in one string: <harness>/<model>[:<effort>].
claude/opus[1m]:max
codex/gpt-6-sol:xhigh
opencode/openrouter/z-ai/glm-5.3-flashThe first path segment is the harness:
claude,codexoropencode. The rest is the model as the harness names it (for OpenCode that is<provider>/<model>).The model must be one the harness offers.
list_harnesseslists them; they are the harness's own values and change with its versions.The suffix is taken as effort only when it is
low | medium | high | xhigh | max, so a model name with its own:tagstays intact. Omitted, the harness default applies.Effort is mapped to the harness's effort option: claude takes the level as is; codex too, with
maxfalling back toxhighif absent. OpenCode's ACP adapter exposes no effort option (1.18.31), so a suffix there only produces a warning. An effort the harness doesn't offer is a warning, not an error.
Example
Ask Claude to delegate a task, and it calls run_thronglet:
{
"agent": "codex/gpt-6-sol:high",
"prompt": "In this repository, add a --verbose flag to the CLI in src/cli.ts and a test for it. Run pnpm test. Leave the changes uncommitted and end with the list of files you changed.",
"cwd": "/work/my-app",
"description": "add --verbose flag"
}The prompt is self-contained: the nested session sees nothing of the calling conversation. When the turn ends, the call returns the agent's final message with the session id, why the turn stopped, and what it cost:
{
"session_id": "019a4c2e-7d1b-7f40-9a2c-5e8b1d3f6a90",
"text": "Added --verbose to src/cli.ts and a test in src/cli.test.ts; pnpm test passes. Changed: src/cli.ts, src/cli.test.ts.",
"stop_reason": "end_turn",
"usage": { "input_tokens": 48210, "output_tokens": 3120 },
"duration_s": 94.3
}A follow-up goes into the same session with send_message; the agent still has the context of its first turn, and the result has the same shape and the same session_id:
{
"session_id": "019a4c2e-7d1b-7f40-9a2c-5e8b1d3f6a90",
"prompt": "Also document the flag in README.md."
}A failure is an MCP tool error carrying code, message and, when the session exists, session_id and the partial text. Every field, error code and stop reason of every tool is in DESIGN §3.
Configuration
Optional: ~/.config/throng/config.yaml.
# Permission policy: auto | allow_all | deny_all | elicit (see Permissions below).
permissions: auto
# Per-harness overrides, all optional.
harnesses:
# claude:
# env: { CLAUDE_CODE_OAUTH_TOKEN: "..." } # extra adapter env; e.g. auth when the standalone `claude` isn't logged in
# codex:
# permissions: allow_all # per-harness policy override
# opencode:
# command: /opt/opencode # adapter outside PATH: absolute path, or a name looked up on PATH
# args: [acp] # replaces the default args
# env: { X: "1" }
limits:
timeout_s: 21600 # default turn timeout (6 h)
handshake_s: 60 # adapter start + session setup
elicitation_s: 600 # policy elicit: how long a permission dialog waits for an answer
max_concurrency: 10 # parallel turns per server process
max_depth: 2 # nested throng → harness → throng → … levelsUnknown keys are rejected. A broken config is logged on server start and shows up in every list_harnesses unavailable reason; run_thronglet refuses to run until it's fixed (it never falls back to defaults, which might be less strict than what you meant).
Environment variables of the server:
variable | meaning |
| config path instead of |
| cache dir instead of |
| nesting depth; set by throng for its children, you don't set it by hand |
Permissions
The policy comes from the config only, never from a tool parameter, so the calling model can't grant itself more than the config allows. throng answers every permission request with a one-time option, never "always allow" (Claude would write that rule into the project settings).
auto(default): each harness runs in its own auto-approve mode (Claudeauto, Codexagent, OpenCode as configured); whatever it still asks about is rejected.allow_all: the harness runs in its asking mode (Claudedefault, Codexread-only, OpenCode with every permission set toask) and every request is allowed once.deny_all: the same asking mode; every request is rejected.elicit: the same asking mode; each request is shown to you as a dialog in the MCP client (the tool title, kind, input truncated to 2 KB, paths) with the one-time choices the harness offered. Your answer goes to the agent; Decline rejects, dismissing the dialog or no answer withinlimits.elicitation_scancels the request, and the agent carries on either way. Background turns ask the same way. Needs an MCP client that supports elicitation (Claude Code does); otherwise the call fails withelicitation_unsupportedbefore anything starts. A throng call inside a thronglet usually fails that way, since its client is the harness.
throng's own submit_result (structured output) is allowed under every policy. Cancelling a run answers its pending requests cancelled and closes any open dialog.
Auth: the nested harness uses whatever login its CLI has. If claude auth status says not logged in, put CLAUDE_CODE_OAUTH_TOKEN (or ANTHROPIC_API_KEY) into harnesses.claude.env as above. Codex and OpenCode use their own logins: codex login, opencode auth login. OpenCode custom providers live in your ~/.config/opencode/opencode.json; throng doesn't touch it.
Troubleshooting
harness_unavailable: the adapter isn't on PATH, and the message carries the install command; or the config is broken (message starts withconfig error:): fix the yaml, throng won't run on defaults.elicitation_unsupported:permissions: elicit(global orharnesses.<harness>.permissions) but the MCP client has no elicitation support; the message names the key. Use a client that has it or pick another policy.A warning
permission mode "auto" not applied: the agent switched to "acceptEdits": Claude Code has no auto mode for that model (haiku, for one) and falls back to accept-edits; file edits are still auto-approved, anything else the harness asks about is rejected by throng (autonever widens into allow-all). Pick another model if you need the real auto mode.model_rejected: the model isn't one of the harness's values. Calllist_harnessesfor the current list; they are the harness's own option values and change with harness versions.handshake_timeout/spawn_failed/handshake_failed: the message includes the adapter's stderr. Usual cause is auth: checkclaude auth status(or setCLAUDE_CODE_OAUTH_TOKENin config),codex login,opencode auth login. Slow first start: raiselimits.handshake_s.empty_result: the agent ended its turn without saying anything; the harness's own session log shows what it did.timeout: raisetimeout_sfor the call orlimits.timeout_s. The payload keepssession_idand any partialtext.Anything else: the server log (stderr of the server; the MCP client decides where it ends up), then the harness's session log by
session_id. throng keeps no transcripts.
Links
DESIGN §3 "External contract": the exact inputs, results, error codes and stop reasons of every tool, background turns and
wait_thronglet, the per-session queue andsteer.skills/throng/SKILL.md: working patterns for the calling agent.
docs/development.md: for maintainers: tests, smoke runs against real harnesses, files on disk.
LICENSE: MIT.
Available Tools
6 toolscancel_throngletA
Cancels the session's running turn (session/cancel) and drops its queued messages; a pending wait_thronglet returns the cancelled error. The session is idle afterwards and accepts a new send_message. No-op on an idle session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | session_id from run_thronglet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses that queued messages are dropped, that a pending wait_thronglet surfaces a cancelled error, and that the session is idle afterward and accepts new messages. It does not mention authorization or permissions, leaving one gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the primary effect and back-loaded with post-conditions and the no-op case. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no annotations and no output schema, the description covers the key effects, side effects on queued messages and waiting callers, and the resulting state. Only the permission/authorization angle is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single session_id parameter, and the schema already ties it to run_thronglet. The description adds no syntax or format detail beyond that, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (cancels) and resource (the session's running turn) and even names the underlying operation (session/cancel). It also references siblings wait_thronglet and send_message, so an agent can distinguish this tool's role from theirs without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when the tool matters and explicitly notes the no-op case on an idle session, which prevents pointless calls. It stops short of naming an alternative tool or exclusion for other cancel-like scenarios, so it is not a full when/when-not routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_harnessesA
Lists valid agent values for run_thronglet: available harnesses with their models and effort levels, unavailable ones with the reason, and the server limits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full disclosure burden and does well: it reveals that both available and unavailable harnesses are returned, that unavailability comes with a reason, and that server limits are included. It does not explicitly state the operation is read-only, though 'Lists' strongly implies a safe read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The purpose and the returned structure are packed into one clean clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description adequately explains the return payload, covering available, unavailable, and limit categories. Minor gaps remain about exact format or size of the response, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. The description correctly implies no filtering or selection arguments are needed to obtain the full harness list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (harnesses), then enumerates exactly what is returned: available harnesses with models and effort levels, unavailable ones with reasons, and server limits. This clearly separates it from siblings like list_thronglets, since it is specifically about agent values consumed by run_thronglet.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
By framing the output as the 'valid agent values for run_thronglet,' it implies the tool should be called to discover acceptable inputs before invoking run_thronglet. There is no explicit when-not or named alternative, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_throngletsA
Lists thronglet sessions on this machine: description, agent, cwd, state (running | queued | idle | failed), queue length, timestamps and the last error. Live state is this server's; a session run by another throng instance shows what its record says.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a non-obvious behavioral caveat: live state reflects this server while sessions run by another throng instance show their recorded state. That is genuinely useful, but read-only safety, ordering, and staleness limits are not spelled out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the action and scope, followed by the field list and the cross-instance caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully enumerates the return fields (description, agent, cwd, state, queue length, timestamps, last error), which is exactly the compensation needed. Ordering and pagination behavior are the only gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter syntax to document; baseline 4 applies. Nothing is missing on this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (lists) and resource (thronglet sessions) with the returned fields enumerated, so the agent knows exactly what comes back. It is clearly distinguishable from siblings like list_harnesses and the mutating run_thronglet/cancel_thronglet tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and the 'on this machine' scope, but there is no explicit statement of when to call this instead of wait_thronglet or run_thronglet, and no exclusions or prerequisites. Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_throngletA
Runs a coding agent (Claude Code, Codex, OpenCode) on a task in cwd and returns its final message as JSON {session_id, text, stop_reason, usage, duration_s, warnings?}; with schema, structured replaces text. description names the thronglet for listings. With background: true the call returns {session_id, state, queued} as soon as the turn runs; collect the result with wait_thronglet.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | Absolute path; the agent works in this tree | |
| agent | Yes | <harness>/<model>[:<effort>], e.g. claude/opus[1m]:max, codex/gpt-6-sol:xhigh; valid values: list_harnesses | |
| prompt | Yes | Task for the agent | |
| schema | No | JSON Schema (draft-07 or 2020-12) for structured output: the agent submits a matching result, returned as `structured` instead of `text` | |
| timeout_s | No | Wall-clock limit for the run in seconds; default from config (21600) | |
| background | No | Return as soon as the turn is running (or queued behind the session's current turn); collect the result with wait_thronglet | |
| description | Yes | What this thronglet is for, in a few words; shown in session listings |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it discloses the JSON return shape ({session_id, text, stop_reason, usage, duration_s, warnings?}), the structured-output substitution, and the background return contract ({session_id, state, queued}). It implies filesystem impact via "the agent works in this tree" but never states permissions, reversibility, or that the run mutates cwd.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Everything is front-loaded into one dense paragraph with no filler, and the return contract comes first. The semicolon-chained clauses make it slightly harder to parse than a clean multi-sentence structure, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by documenting both the sync and background return shapes plus the structured-output path, which is exactly what an agent needs. It stops short of noting cancel_thronglet as a lifecycle alternative or any failure/timeout behavior beyond the default config value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description earns above that by tying parameters to behavior the schema doesn't state — schema replaces text with structured, background returns early, and description feeds session listings. It adds meaning beyond the field-level docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ("Runs a coding agent") with named harnesses and the execution context (cwd), and it distinguishes this tool from the sibling wait_thronglet by explaining the async handoff. An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear conditional usage for the background: true path and names wait_thronglet as the collection mechanism. It doesn't explicitly state when to prefer this over alternatives like send_message to an existing thread, but the sync/async guidance is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_messageA
Sends the next message into an earlier session (session_id from run_thronglet) and returns the same JSON as run_thronglet. A message to a session whose turn is still running waits for that turn to end: turns on one session never overlap. With background: true the call returns {session_id, state, queued} as soon as the turn runs; collect the result with wait_thronglet. steer: true interrupts the running turn and delivers this message next; queued messages follow it.
| Name | Required | Description | Default |
|---|---|---|---|
| steer | No | Interrupt the running turn (session/cancel) and run this message as the very next turn, ahead of queued messages. The in-flight tool call is aborted; a half-applied edit may remain. | |
| prompt | Yes | Next message for the agent | |
| schema | No | JSON Schema (draft-07 or 2020-12) for structured output: the agent submits a matching result, returned as `structured` instead of `text` | |
| timeout_s | No | Wall-clock limit for the run in seconds; default from config (21600) | |
| background | No | Return as soon as the turn is running (or queued behind the session's current turn); collect the result with wait_thronglet | |
| session_id | Yes | session_id from run_thronglet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden and does so richly: turn serialization ('turns on one session never overlap'), the waiting behavior, the background return shape {session_id, state, queued}, and steer semantics including interruption of the in-flight call. This is exactly the kind of behavioral disclosure annotations would otherwise provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with the core action and session_id provenance, then waiting semantics, then background, then steer. No filler; every clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Comprehensive for a 6-param mutation tool with no output schema: it covers return shape, concurrency, background and steer flows. It does not mention timeout_s behavior or the structured-output schema param, both of which are documented in the schema but the description stays silent on their runtime effect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds non-obvious semantics: background returns early and needs wait_thronglet, and steer interrupts the running turn and jumps ahead of queued messages. That sequencing relationship between steer and queued messages is not derivable from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (sends) and resource (next message into an earlier session), and ties the session_id directly to the sibling run_thronglet. An agent can tell this apart from cancel/wait/list siblings immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly positions this as the continuation mechanism after run_thronglet and routes to wait_thronglet for background collection. Does not explicitly say when-not-to-use, but the alternative and its trigger conditions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_throngletA
Waits until a session has no running or queued turn and returns the last turn's result: the same JSON as run_thronglet, or its error. With timeout_s elapsed returns {session_id, state, queued} instead. Idempotent: the result is stored with the session.
| Name | Required | Description | Default |
|---|---|---|---|
| timeout_s | No | How long to wait in seconds; default from config (21600). Elapsed → the session state, not an error | |
| session_id | Yes | session_id from run_thronglet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the timeout behavior (returns {session_id, state, queued} instead of an error), idempotency, and that results are persisted with the session. It stops short of auth requirements or rate limits, but the core behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with zero waste; the primary behavior is front-loaded and the timeout and idempotency caveats follow logically. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain return values, and it does: both the success path (run_thronglet JSON or its error) and the timeout path ({session_id, state, queued}). It leaves the possible 'state' values undefined, but for a wait primitive this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema, and the description's timeout detail (state returned instead of error) mirrors the schema description rather than adding new meaning. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (waits) and resource (session turn), and explicitly distinguishes itself from run_thronglet by noting it returns 'the same JSON as run_thronglet, or its error'. An agent can tell what this does and how it relates to the sibling without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tie to run_thronglet ('session_id from run_thronglet', 'same JSON as run_thronglet') implies this is used to await a previously started run, but there is no explicit when-to-use or when-not guidance, nor a named alternative (e.g., polling run_thronglet). Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
cancel_thronglet - First observed
list_harnesses - First observed
list_thronglets - First observed
run_thronglet - First observed
send_message - First observed
wait_thronglet
TDQS
Scored across 6 tools
Most tools are clearly distinct: cancel, list, run, send, and wait target different operations. However, run_thronglet and send_message both initiate turns and return the same result shape, which could cause slight confusion without careful reading of descriptions. The overlap is manageable but not perfectly clean.
All tool names use snake_case and follow a consistent verb_noun pattern (cancel_thronglet, list_harnesses, list_thronglets, send_message, run_thronglet, wait_thronglet). There are no mixed conventions or vague verbs. The naming is predictable and readable.
Six tools is well-scoped for a coding agent session manager. Each tool covers a distinct lifecycle operation without redundancy or bloat. The count is neither too thin nor too heavy for the apparent purpose.
The set covers core session lifecycle operations: start (run_thronglet), continue (send_message), wait (wait_thronglet), cancel (cancel_thronglet), and list sessions/harnesses. Minor gaps exist, such as no explicit tool to delete or archive a session, and no dedicated single-session status retrieval beyond wait_thronglet with a timeout. These are workaround-able but not fully complete.
Maintenance
Related MCP Connectors
Build and supervise fleets of agents from Claude Code, Codex or Cursor. Connects over OAuth.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Scoped agent execution. Server-side credentials, policy, budgets and verifiable receipts.
Related MCP Servers
- AlicenseAqualityDmaintenanceWraps Claude Code as tools for MCP clients, enabling autonomous coding tasks via a 4-tool lifecycle with session management, async polling, and permission controls.459 npm20MIT
- AlicenseNot gradedqualityDmaintenanceEnables MCP clients to spawn and control Codex CLI and Claude Code sessions on the host machine, with session management and filesystem access.4MIT
- AlicenseAqualityDmaintenanceEnables AI agents to interact programmatically with Claude Code CLI, managing sessions, streaming outputs, and handling permission requests.711 npm1MIT
- AlicenseBqualityBmaintenanceEnables launching and supervising local Codex and Claude Code agent sessions and one-shot tasks, exposing run, list, stop, and output operations over MCP with session-scoped permissions and sandbox controls.54,779 npmMIT