severally
This server lets you start and manage background peer consultations with AI coding agents, record and export the results as decision records.
Start consultations (
consult_start) with one target or fan out to up to 3 consultants at once (targets), with an immediatejob_id/group_id.Choose a mode:
explore(withhold your proposal),review(get weaknesses on a proposal), ordebate(test disputed points and counterpoints).Provide context: facts, artifacts, constraints, success criteria, counterpoints, and an optional pre-registered prediction (never sent to the consultant).
Follow up on a consultation (
followup_to) to continue a chain with one named consultant.Check results (
consult_get) with optional blockingwait_ms; completed jobs include structured findings, unknowns, next checks, and failure details like timeout/auth/usage_limit.Cancel (
consult_cancel) a running job or whole fan-out group.Record verdicts (
consult_record) against findings/unknowns/checks (confirmed, not_applicable, unverifiable, unverified), plus a reflection on what the answer added.Export (
consult_export) a consultation chain or fan-out group as Markdown for committing to the repository.List (
consult_list) recent consultations from the session.Current environment caveat: no consultant CLI is installed, so
consult_startwill refuse targets until one is available; the schema also enforces safety limits (max rounds, no task support, read-only consultants).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@severallyAsk Codex, Claude, and Antigravity for independent opinions on this refactoring plan."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
severally
Independent opinions, returned severally. The verdict is yours.
An MCP server and Skill for the moment a coding agent asks you "can I go ahead with this plan?". The agent asks another CLI (Codex, Claude Code, Antigravity running Gemini, or OpenCode running GLM) for its opinion, checks the findings that come back in its own repository, and only then asks you for a decision again.
The name comes from the legal phrase jointly and severally — each party bound on its own. Ask several consultants and their answers are not merged: each comes back separately, and the reader decides which to take.
Example: before publishing a tool
I had written a small tool that rewrites AI CLI conversation histories when a project folder moves. Before putting it on GitHub and PyPI, I asked three consultants at once "what must be fixed before this is published?". About two minutes later their answers came back side by side (excerpt):
codex do_not_proceed no handling for a failure partway through moving the folder and updating several config files
antigravity do_not_proceed no LICENSE; --force means both "ignore the running-session check" and "merge folders"
claude-code proceed fine to publish, once 5 items such as adding a LICENSE and removing the author's local paths are doneThe blockers the three raised overlap heavily, yet the bottom lines split. The difference is "fix it, then publish" versus "publish once it is fixed". Merge the answers into one and that difference disappears.
The agent that asked (the lead) does not stop at reading the findings. It checks them in its own repository.
codex finding needs handling for a failure partway through
checked if the destination is inside the source, it crashes with a traceback.
it also turned up a bug: after rolling back partway, the tool still exits 0
decision added a fix to the pre-publication work: count failures, exit non-zero, tell the user to re-run
antigravity finding reads large history files entirely into memory, so it runs out of memory
checked the largest file on this machine is 43.8 MB. the reading will be fixed, but the severity was loweredWhat comes back to you is not a vote of "1 for, 2 against". It is what was checked, the recommendation, and a list of what has not been checked yet. The checked results stay next to the findings and can be exported as Markdown into your repository.
Related MCP server: Senior Consult MCP
Compared with pasting into another terminal
You can get something similar by pasting a brief into another terminal. severally adds three things:
Only the brief goes across. Each consultant starts in a fresh child session and never sees your conversation history. In
exploremode, which asks without showing your plan, the field for the plan (proposal) cannot be usedEvery answer has the same shape. Bottom line, findings with grounds, missing information, conditions that would change the judgement, and how to check. If no answer arrives, the reason (rate limit, auth failure, timeout). Ask several and nothing is summarised
Checked results can be written back next to each finding. For each finding, record "confirmed / not applicable / unverifiable / unverified" and its effect on the decision, then export it as Markdown
Whichever of Codex, Claude Code, Antigravity, or OpenCode you use, you can ask the other three.
How it works
Every consultation starts a dedicated child session; it never attaches to an existing one. The consultant
receives only the brief you wrote, never your conversation history. One request can send the identical brief to
up to four consultants and collect the answers under one group_id. What consultants may and may not do is
listed under "Before you use it".
Claude Code ──(skill: severally)──> mcp: severally ──> codex exec | agy | opencode run
Codex ──(skill: severally)──> mcp: severally ──> claude -p | agy | opencode run
Antigravity ──(skill: severally)──> mcp: severally ──> codex exec | claude -p | opencode run
OpenCode ──(skill: severally)──> mcp: severally ──> codex exec | claude -p | agyClaude Code and Codex get it as a plugin; Antigravity and OpenCode are registered directly by the installer (Antigravity has no verified plugin path, OpenCode has no plugin mechanism). Neither needs a public marketplace.
Install
Requires Node.js 20.10 or newer. Windows runs natively from PowerShell or Command Prompt; WSL is not required. Install and authenticate the consultant CLIs you want to use first.
npm install # fetch dependencies
npm test # offline checks (no real API calls)
npm run build # regenerate dist/ inside the plugin (committed, so usually not needed)
node scripts/install.mjs # --dry-run prints the plan without changing anythingVerify:
claude plugin details severally # Skills (1) / MCP servers (1)
codex plugin list # severally@severally-local installed, enabled
agy mcp list # severally stdio enabled
opencode mcp list # severally connectedRestart the clients (a running session does not reload plugins).
After updating the plugin, run npm run build && node scripts/install.mjs again.
Windows (PowerShell)
npm.cmd install
npm.cmd test
npm.cmd run build
node scripts/install.mjsOn Windows the installer defaults to manual mode for all four clients. It copies the
standalone server (including its dependencies) to ~/.severally/runtime/severally-mcp.mjs,
registers that file with the absolute path to node.exe, and copies each client's Skill.
After successful registration, the source checkout can be moved or deleted. Node.js must remain
installed. Global npm installation is unnecessary; --skip-global is still accepted but is optional
on Windows. Add --dry-run to inspect the changes without writing anything.
Upgrading an earlier checkout-based installation: run node scripts/install.mjs --force
to switch the existing MCP registrations to the dedicated runtime. Without --force, existing
registrations are preserved and may still point to the checkout. Updates replace the runtime
after checking the new bundle's syntax and backing up the previous file. To update later, obtain
a new checkout, run npm.cmd install, npm.cmd run build, and the installer again, then restart
the clients. The temporary checkout is no longer needed after registration succeeds.
Verify with claude mcp get severally, codex mcp get severally, agy mcp list, and opencode mcp list,
then restart the clients. Tool names in this mode are mcp__severally__*.
CLI detection supports .exe, .cmd, and .bat through PATH/PATHEXT, including paths with
spaces. For a custom executable path in config.json, use forward slashes
("bin": "C:/Tools/claude.exe") or escaped backslashes ("bin": "C:\\Tools\\claude.exe").
The test suite uses local stand-in CLIs; authentication and live consultations still depend on
the installed client versions and accounts.
Linux/macOS: installation independent of the checkout
After a successful installation, the source checkout can be moved or deleted on Linux and macOS too:
Claude Code receives a copy of the bundled plugin in
~/.claude/skills/severally/.The standalone server lives in
~/.severally/runtime/severally-mcp.mjs.Codex launches the global
severally-mcpcommand, installed from that persistent runtime. Its marketplace and plugin files are copied into~/.severally/marketplace/.Antigravity and OpenCode (and all clients in
--manualmode) launch Node.js with the copied runtime directly; OpenCode is registered into~/.config/opencode/opencode.jsonand its Skill is copied to~/.config/opencode/skills/severally/.
To migrate an existing installation, run:
npm install
npm run build
node scripts/install.mjs --forceThis refreshes the Codex marketplace location and switches existing direct MCP registrations to the
copied runtime. Restart the clients after registration succeeds; the checkout is then disposable.
Node.js and the installed runtime must remain. For updates, obtain a fresh checkout and run the same
commands. Previous runtime and marketplace files are backed up under ~/.severally/backups/.
Plugin mode requires npm's global executable directory on PATH. --skip-global is accepted only
when severally-mcp already resolves to the copied runtime; a command linked to an old checkout is
rejected. In --manual mode no global install is needed and --skip-global is optional.
Two install modes
Plugin mode (default on macOS/Linux) | Manual mode ( | |
Claude Code | places the plugin in |
|
Codex |
|
|
Antigravity |
| same |
OpenCode |
| same |
MCP tool names |
|
|
The installer does not break existing settings. It changes client settings only through each CLI's own commands
(plugin add / mcp add). Before it does anything, it backs up ~/.claude.json, ~/.codex/config.toml and any
existing Skill directory to ~/.severally/backups/<timestamp>/, naming each backup after its original path.
Switching modes backs up and then removes the duplicate registration left by the other mode.
No public marketplace listing is needed. Claude Code works without a marketplace, and the Codex marketplace file
.agents/plugins/marketplace.json is copied from this repository into the managed marketplace directory.
Usage
The Skill starts on requests such as "ask Codex", "have Claude review this" or "I want a second opinion", or in these situations:
important design decisions and hard-to-reverse choices (architecture, data migration, concurrency, security, pre-publication checks)
options that stay neck and neck however long you think about them
two or more failed attempts at the same bug with no new information
For "ask everyone" or "consult all of them", the same brief goes out at once to the other three CLIs and to a fresh session of your own CLI. Answers from your own CLI carry a "same lineage" note.
When the agent asks you to approve a hard-to-reverse change, it offers "consult first?" as one of the choices. You decide whether to start.
Not for small fixes. One consultation costs a few minutes and real quota.
A routine consultation is small. One consultant. The brief is one paragraph of plan, a few lines of facts it rests on, and one relevant code excerpt. The Skill checks the one finding that would change the decision, reports what it checked, what it did not adopt and what is still unverified, and saves the checked result next to the finding. Starting a consultation returns immediately, so the agent keeps working while it waits.
For decisions you will have to justify later (interfaces, migrations, security, concurrency), the Skill adds steps: it pins down the question and success criteria, separates imposed constraints from its own assumptions, writes down a prediction before consulting, and exports the record as Markdown into the repository.
The Skill tells the lead how to write the brief, which mode to pick (explore / review / debate), and how to
read the results.
Configuration
No configuration needed by default. At startup the server checks whether codex / claude / agy /
opencode are on PATH and drops the ones that are missing. To disable a consultant, change a model, or point
at a specific executable, put a single ~/.severally/config.json in place. A template can be generated for
your machine:
npm run init-config # writes ~/.severally/config.json (never overwrites an existing one)Every key is explained in CONFIG.md; config.example.json is a working example that exercises each one. Precedence is environment variables > config file > auto-detection > defaults. The config is read once at server startup, so restart the client after changing it.
Before you use it
Consultation round limit
A consultation chain defaults to 5 total rounds (one initial consultation and up to four follow-ups).
Set SEVERALLY_MAX_ROUNDS in the MCP server's environment to change the limit to 1–20 total rounds;
values above 20 are capped at 20. For example, SEVERALLY_MAX_ROUNDS=20 allows the initial consultation
plus 19 follow-ups. Restart the client/server after changing this setting. This is an environment setting,
not a key in config.json. The server reports the active budget in rounds_remaining.
What never reaches a consultant: your conversation history. What no consultant can do:
write
use the network (web search and browsing are the only exception)
use MCP
consult anyone else
Reading is not blocked. Per child session:
Consultant | Reads the disk | Shell | Write / execute |
Codex | yes | read-only | no |
Claude Code | yes ( | none | no |
Antigravity | yes ( | none | no |
OpenCode | yes ( | none | no |
Every consultant starts in an empty working directory, and the server does not tell it where your repository
is. Paste whatever it should see into the brief — or, when a consultant needs to explore rather than read what
you picked out, name absolute paths in context.expose_paths. Those files and directories are copied read-only
into the consultant's working directory (under ./workspace) and are then the only part of your repository it
has. At most 20 entries, 500 files and 5 MB in total; symlinks are skipped rather than followed; text files are
credential-masked exactly as the brief is, binary files are copied unchanged. The copy is deleted with the job,
and the history keeps only which paths were shown, not their contents. The Claude Code consultant's whole-disk
read scope is dropped for a consultation that uses it.
A consultant that failed (
usage_limit/auth/timeout…) and one that answered on thin grounds come back as different things. A failure is not "no problems found"History is kept in
~/.severally/history/. Each round is one file holding the brief that was sent, the consultant's answer, and the verdicts the lead wrote per finding (withconsult_record).consult_exportturns it into Markdown for the repository. Verdicts can be added later, from another sessionCredentials are masked in the brief that is sent and in the results that come back
You can also consult a different model of your own CLI (a Fable lead asking Opus,
target: "claude:opus"). The answer carries a "same lineage" note. Declare your own model withcaller_modeland the record keeps who asked whomThe record also keeps whether you asked for the consultation or accepted the agent's offer (self-declared). Offers you declined are written one per line to
~/.severally/history/offers.jsonl. Neither restricts anythingConsultations about code decisions return mostly checks the lead can run itself. Consultations about project policy return checks that depend on other people, and those come back to you unrun
Uninstall
Windows (the default manual installation):
claude mcp remove severally -s user
codex mcp remove severally
agy mcp remove severally
# opencode has no `mcp remove`; delete the "severally" entry from
# the "mcp" object in $HOME/.config/opencode/opencode.json
Remove-Item -LiteralPath "$HOME/.claude/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.codex/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.gemini/config/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.config/opencode/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.severally/runtime" -Recurse -Force
# Only if you previously installed the global command:
npm.cmd uninstall -g severally-mcpIf CODEX_HOME is set, use that directory instead of $HOME/.codex for the Codex Skill.
macOS/Linux:
Plugin mode:
rm -rf ~/.claude/skills/severally # Claude Code
codex plugin remove severally --marketplace severally-local
codex plugin marketplace remove severally-local
agy mcp remove severally # Antigravity (registered directly in both modes)
rm -rf ~/.gemini/config/skills/severally
# OpenCode (registered directly in both modes; opencode has no `mcp remove`,
# so delete the "severally" entry from the "mcp" object by hand):
# edit ~/.config/opencode/opencode.json(c)
rm -rf ~/.config/opencode/skills/severally
npm uninstall -g severally-mcp
rm -rf ~/.severally/runtime ~/.severally/marketplaceManual mode:
claude mcp remove severally -s user
codex mcp remove severally
agy mcp remove severally
# opencode: delete the "severally" entry from the "mcp" object in
# ~/.config/opencode/opencode.json(c) by hand
rm -rf ~/.claude/skills/severally ~/.codex/skills/severally ~/.gemini/config/skills/severally \
~/.config/opencode/skills/severally
# Only if a global command was previously installed:
npm uninstall -g severally-mcp
rm -rf ~/.severally/runtime ~/.severally/marketplaceEither way, history and backups stay in ~/.severally/ (delete it if you no longer need them).
Available Tools
6 toolsconsult_cancelCancel a running consultationA
Stop a running consultation and kill the consultant process and everything it spawned. Pass group_id instead of job_id to stop every consultant in a fan-out.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | ||
| group_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the destructive nature ('kill the consultant process and everything it spawned') and the group behavior, which is significant and goes beyond what would be assumed. However, it does not address reversibility, side effects on data, or permission requirements, but for a cancel operation, the key behaviors are adequately covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—two sentences—with the primary action front-loaded. It includes the essential parameter guidance without extraneous detail, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema and no annotations, the description covers the purpose, the parameter semantics, and the key behavioral trait (killing spawned processes). It does not address edge cases like providing both parameters or neither, but these are not critical for typical usage and the description is sufficiently complete for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clearly explains the semantic difference between job_id (single job) and group_id (fan-out group), which is critical for correct usage. It does not specify that both are optional or what happens if neither is provided, but the core distinction is well conveyed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear, specific action: stop a running consultation and kill the consultant process. It explicitly mentions the spawned processes, making it distinct from the sibling tools (consult_get, consult_start, consult_record, consult_export, consult_list), which all serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides guidance on parameter usage ('Pass group_id instead of job_id to stop every consultant in a fan-out'), but it does not explicitly address when to use this tool versus alternatives. It does not mention any sibling tools or conditions for selection, so the when-to-use context is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consult_exportExport a consultation as a record to commitA
Render one consultation (chain_id) or one fan-out (group_id) as Markdown: the brief as it was sent, each consultant's answer as it came back, and the verdicts recorded against each point -- with the ones nobody checked marked as unchecked. Nothing is summarised across consultants and nothing is scored. The text is returned, not written: put it wherever the decision belongs in the repository (a decision record next to the code it is about), which is the only place a teammate will find it. Reads the on-disk history, so a consultation from an earlier session can still be exported.
| Name | Required | Description | Default |
|---|---|---|---|
| chain_id | No | ||
| group_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure and succeeds: it states the output format (Markdown), what is included, what is excluded ('Nothing is summarised... nothing is scored'), that the operation is non-writing ('returned, not written'), and that it reads on-disk history. This is far beyond a minimal description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and then adds details in a logical order; every sentence contributes either behavioral scope, content expectations, or usage guidance. Despite being three sentences, it is efficient for the information it must convey without annotations or output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given two simple optional-looking parameters, no annotations, and no output schema, the description covers output format, contents, non-mutation, and data source. The main missing piece is an explicit statement of the required/forbidden parameter combination (exactly one of chain_id/group_id), which prevents full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain chain_id and group_id. It does: chain_id selects a consultation, group_id selects a fan-out. It stops short of stating whether exactly one is required or what happens if both/neither are provided, which leaves some ambiguity for invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Render one consultation (chain_id) or one fan-out (group_id) as Markdown.' It details the exact contents of the output (brief, answers, verdicts, unchecked marks) and distinguishes itself from write-oriented siblings by saying 'The text is returned, not written.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when the tool is appropriate: to export a consultation/fan-out as a decision record in the repository, and it notes that on-disk history lets earlier sessions be exported. It does not explicitly name alternative sibling tools or call out when-not-to-use conditions, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consult_getGet consultation status or resultA
Fetch the state of a consultation. Pass wait_ms to block until it finishes (capped at 45000 ms, which stays under the request timeout MCP clients apply) instead of polling in a tight loop; a typical consultation takes one to five minutes, so expect to call this several times, and do something else in between. A completed job carries the structured answer; a failed one carries failure.kind (timeout, auth, usage_limit, model_unavailable, invalid_output, cli_error, spawn_error) — that is "no advice was obtained", which is different from advice that arrived with thin evidence (see quality.evidence_basis). Pass group_id instead of job_id to fetch a fan-out; wait_ms then waits for every consultant in it.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | ||
| wait_ms | No | ||
| group_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it explains blocking behavior, the 45-second cap, the expectation of multiple calls, failure kinds, and the distinction between failure and thin evidence. It does not cover auth or rate limits, but the core behavior is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph, but every sentence adds value—blocking, failure kinds, fan-out. It is front-loaded with the core purpose and then elaborates. Slightly long, but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and three parameters, the description covers the essential behavioral and semantic ground. It mentions the structured answer and failure.kind, hinting at response fields. It could add more about edge cases (e.g., invalid job_id) or the exact response shape, but it is adequate for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It does: wait_ms is detailed (blocking, cap, purpose), group_id is explained as fan-out, and job_id is implied as the primary identifier. This goes beyond the bare schema and compensates for the lack of parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches the state of a consultation, which is a specific verb-resource pair. It does not explicitly contrast with sibling consult_list, but the purpose is unambiguous and distinct from starting or canceling consultations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete guidance on when to use wait_ms (to block instead of polling), how long consultations take, and when to pass group_id for fan-out. It does not explicitly name alternative tools, but the situational guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consult_listList recent consultationsA
Recent consultations from this session, newest first, with their status and one-line summary.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses ordering (newest first) and that results include status and summary, which is useful. However, it doesn't disclose whether this is a read-only operation, whether it reflects only the current session, or what happens when there are no consultations. The 'from this session' scoping is a meaningful behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that conveys the resource, scope, ordering, and included fields with zero waste. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional parameter and no output schema, the description is nearly complete. It could mention that the output is a list of consultations with status and summary, but that's already implied. The main gap is not explaining the 'limit' parameter, though the schema covers it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It doesn't explain the 'limit' parameter at all, but the schema itself defines it as an integer with min/max, which is self-explanatory. The description adds no meaning beyond the schema for the parameter, but the parameter is simple enough that this is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a clear resource ('recent consultations from this session'), and adds ordering ('newest first') plus included fields ('status and one-line summary'). This clearly distinguishes it from siblings like consult_get (single consultation retrieval) and consult_start (initiating a consultation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is for browsing recent session consultations, which is distinct from consult_get (retrieving a specific one) and consult_cancel (cancelling). It doesn't explicitly state when not to use it or name alternatives, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consult_recordRecord what checking a point showedA
Write your own verdict against one or more points of an answer, after you have checked them in the repository. Ids come from the result: findings are f1, f2 ..., unknowns u1 ..., next_checks c1 ... . verdict says what checking showed -- "confirmed" (it holds here), "not_applicable" (true in general, not for this codebase), "unverifiable" (cannot be settled with what you can reach), "unverified" (not checked yet, and say in effect why not). It does not say whether you adopted the point. effect is what it changed about your decision; note is the evidence you used. Recording the same id again replaces that entry. This server stores what you write and counts the verdicts; it never infers one, and never decides a consultation was worth it. The entry is saved beside the answer and the brief in ~/.severally/history, which is what makes the decision readable a month from now; the job is read back from that history, so a consultation from an earlier session can still be recorded against. Pass reflection to record what the answer added over what you already expected, when the consultation was started with a prediction. A prediction itself cannot be written here: it goes in consult_start, before the consultant runs, which is the only thing that makes it a prediction.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| entries | No | ||
| reflection | No | written after you have read the answer. There is deliberately no hit/miss label: a point you predicted can still arrive with the evidence that settles it, and a surprise can still be wrong. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly. It discloses that the server stores entries, overwrites on duplicate ids, never infers verdicts, and persists to ~/.severally/history, enabling cross-session recording. This goes well beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense. Each sentence contributes to understanding the tool's behavior, id conventions, verdict meanings, and storage details. It is front-loaded with the core purpose and avoids fluff, though it could be slightly tightened without losing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (nested entries, reflection object, multiple verdicts, no output schema), the description covers all necessary details: id format, verdict semantics, overwrite behavior, persistence location, cross-session support, and reflection conditions. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only reflection has a description). The description compensates fully by explaining what entries contain, the meaning of each verdict value, what effect and note represent, and when reflection should be used. It adds meaning that the schema omits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (write a verdict), the resource (points of an answer), and the context (after checking them in the repository). It differentiates from siblings by explicitly noting that predictions belong in consult_start, not here.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the reflection parameter (when started with a prediction) and explicitly excludes writing predictions here, pointing to consult_start. It also indicates the id source from the result. It doesn't explicitly contrast with consult_get or consult_export, but the usage context is fairly clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consult_startStart a peer consultationA
Start a consultation with another agent (or several, via targets). Returns a job_id (or a group_id for several) immediately; the work runs in the background.
target: which consultant to ask. This machine can reach: (none -- no consultant CLI is installed). The everyday names work too: gpt/chatgpt/openai, claude/anthropic, gemini/agy/google. A consultant that is not in that list is refused up front, so do not retry it -- say which ones are available instead. Consulting your own CLI is allowed but is a fresh-context check rather than an independent opinion, and the result says so. targets: ask up to 3 consultants the same question at once (mutually exclusive with target, no duplicates). Every member gets the byte-identical brief and one group_id; poll it with consult_get({ group_id }). A follow-up (followup_to) always names one consultant -- fan-out is never available on a follow-up. A consultant may name the model to run it on as a suffix: "claude:claude-opus-5". What each consultant is allowed to run is set by the operator, and this server currently allows:
Only pass a model when the user asked for one; a name outside the list is refused before the
consultation starts, and a name matching two of them is refused rather than guessed.
caller: optional -- the CLI you are running in ("codex" / "claude-code" / "antigravity"), so the server can annotate a same-vendor consultation. caller_model: optional -- the model you are running on (e.g. "claude-opus-5"), self-declared and never checked. With it, a same-vendor caveat can say "same lineage, different model" and the record keeps who asked whom; it never changes which model the consultant runs. mode: explore - hand over objective/constraints/facts and withhold your own preferred solution, to get independent options, alternative problem framings and blind spots. context.proposal MUST be empty on round 1. review - hand over your current proposal AND the reasoning behind it, to get weaknesses, counterexamples, concrete improvements and the conditions under which it holds. context.proposal is required. debate - hand over the disputed point, your position (context.proposal), the other side's claims (context.counterpoints) and the extra evidence, to get what would change the judgement, how to test it, and which disagreements remain. Both are required.
The consultant starts in an empty working directory and is not told where your repository is: put every fact it needs into context.facts and paste the relevant passages into context.artifacts. Model, permissions, round count, timeout and size caps are fixed by this server and cannot be raised from a request.
prediction: optional -- { expected, worry }: the bottom line you expect back and, in one sentence, what you are most worried about. Stored with the consultation and NEVER sent to the consultant. It can only be written here, before the consultant runs, so that afterwards you cannot rewrite what you thought beforehand; consult_record takes the other half (reflection) once you have read the answer.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that work runs asynchronously, the consultant starts in an empty working directory and is not told the repo location, model/permissions/timeouts are fixed, same-vendor calls are annotated with a caveat, and prediction is never sent to the consultant and cannot be rewritten afterward.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the tool is genuinely complex with nested parameters, three modes, fan-out, model suffixes, and privacy behavior. The structure is front-loaded with the key contract, then organized by parameter, and nearly every sentence carries operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the high parameter complexity, no output schema, and no annotations, the description is remarkably complete. It covers return values, how to poll via consult_get, mode preconditions, unavailable targets, same-vendor caveats, prediction privacy, and environment constraints that would otherwise be invisible to an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the parameter documentation burden falls entirely on the description. It compensates thoroughly: target/targets semantics, followup_to restrictions, caller/caller_model intent, prediction shape and immutability, and mode-specific requirements for context.proposal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Start a consultation with another agent (or several, via targets)' and states the immediate return contract (job_id/group_id, background execution). It clearly distinguishes consult_start from its siblings by framing it as the initiation step and explicitly referencing consult_get and consult_record for follow-up actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance for each mode: explore requires empty proposal, review requires proposal, debate requires both proposal and counterpoints. It also names alternatives and exclusions: targets is mutually exclusive with target, follow-ups never fan out, unavailable consultants are refused up front and should not be retried, and a model suffix should only be passed when the user asked for one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.1.0- First observed
consult_cancel - First observed
consult_export - First observed
consult_get - First observed
consult_list - First observed
consult_record - First observed
consult_start
TDQS
Scored across 6 tools
Each tool targets a distinct lifecycle stage: start, get, cancel, record, export, list. No two tools overlap in purpose; consult_get and consult_list both read state but one is for a specific job/group and the other is a session overview, which is clear from descriptions.
All tools follow a consistent consult_verb pattern: consult_start, consult_get, consult_cancel, consult_record, consult_export, consult_list. The verb is always second and snake_case is used throughout.
Six tools cover the full consultation workflow without redundancy: start, poll/get, cancel, record verdicts, export, and list. This is a well-scoped set for a single-purpose server.
The lifecycle is complete: create, read, cancel, record, export, list. The only minor gap is no explicit delete/cleanup tool for history, but that is not essential to the core workflow and can be handled outside the server.
Maintenance
Related MCP Connectors
Second opinion before an irreversible agent action; signed proofs, free verify, public ledger.
A second opinion for AI agents: one prompt across several live Gonka models + roles, one call.
- ParleyOAuthdev.weldra
Coordination hub for AI coding agents: message teammates, ask humans, audit every event.
Multi-LLM council: 25+ frontier models in parallel, consensus scoring, verdict-first code review.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables multiple AI agents to share and read each other's responses to the same prompt, allowing them to reflect on what other LLMs said to the same question.2MIT
- AlicenseBqualityDmaintenanceEnables AI agents to consult expert models (Claude, GPT, Gemini, DeepSeek, Z.ai) for technical guidance, code reviews, and architectural advice without switching context.420 npm4MIT
- AlicenseAqualityAmaintenanceEnables one AI coding agent to delegate tasks to, and build consensus across, multiple other coding CLIs (Claude Code, Codex, etc.) by orchestrating them as headless subprocesses.1829 PyPI6MIT

LLM Council MCPofficial
AlicenseNot gradedqualityDmaintenanceEnables Claude Code to consult external LLMs (GPT, Gemini) through multi-turn sessions for second opinions, parallel consultations, and web-grounded research.MIT