Skip to main content
Glama

severally

Independent opinions, returned severally. The verdict is yours.

An MCP server and Skill for the moment a coding agent asks you "can I go ahead with this plan?". The agent asks another CLI (Codex, Claude Code, Antigravity running Gemini, or OpenCode running GLM) for its opinion, checks the findings that come back in its own repository, and only then asks you for a decision again.

The name comes from the legal phrase jointly and severally — each party bound on its own. Ask several consultants and their answers are not merged: each comes back separately, and the reader decides which to take.

Example: before publishing a tool

I had written a small tool that rewrites AI CLI conversation histories when a project folder moves. Before putting it on GitHub and PyPI, I asked three consultants at once "what must be fixed before this is published?". About two minutes later their answers came back side by side (excerpt):

codex        do_not_proceed  no handling for a failure partway through moving the folder and updating several config files
antigravity  do_not_proceed  no LICENSE; --force means both "ignore the running-session check" and "merge folders"
claude-code  proceed         fine to publish, once 5 items such as adding a LICENSE and removing the author's local paths are done

The blockers the three raised overlap heavily, yet the bottom lines split. The difference is "fix it, then publish" versus "publish once it is fixed". Merge the answers into one and that difference disappears.

The agent that asked (the lead) does not stop at reading the findings. It checks them in its own repository.

codex finding        needs handling for a failure partway through
  checked            if the destination is inside the source, it crashes with a traceback.
                     it also turned up a bug: after rolling back partway, the tool still exits 0
  decision           added a fix to the pre-publication work: count failures, exit non-zero, tell the user to re-run

antigravity finding  reads large history files entirely into memory, so it runs out of memory
  checked            the largest file on this machine is 43.8 MB. the reading will be fixed, but the severity was lowered

What comes back to you is not a vote of "1 for, 2 against". It is what was checked, the recommendation, and a list of what has not been checked yet. The checked results stay next to the findings and can be exported as Markdown into your repository.

Related MCP server: Senior Consult MCP

Compared with pasting into another terminal

You can get something similar by pasting a brief into another terminal. severally adds three things:

  • Only the brief goes across. Each consultant starts in a fresh child session and never sees your conversation history. In explore mode, which asks without showing your plan, the field for the plan (proposal) cannot be used

  • Every answer has the same shape. Bottom line, findings with grounds, missing information, conditions that would change the judgement, and how to check. If no answer arrives, the reason (rate limit, auth failure, timeout). Ask several and nothing is summarised

  • Checked results can be written back next to each finding. For each finding, record "confirmed / not applicable / unverifiable / unverified" and its effect on the decision, then export it as Markdown

Whichever of Codex, Claude Code, Antigravity, or OpenCode you use, you can ask the other three.

How it works

Every consultation starts a dedicated child session; it never attaches to an existing one. The consultant receives only the brief you wrote, never your conversation history. One request can send the identical brief to up to four consultants and collect the answers under one group_id. What consultants may and may not do is listed under "Before you use it".

Claude Code ──(skill: severally)──> mcp: severally ──> codex exec | agy | opencode run
Codex       ──(skill: severally)──> mcp: severally ──> claude -p  | agy | opencode run
Antigravity ──(skill: severally)──> mcp: severally ──> codex exec | claude -p | opencode run
OpenCode    ──(skill: severally)──> mcp: severally ──> codex exec | claude -p | agy

Claude Code and Codex get it as a plugin; Antigravity and OpenCode are registered directly by the installer (Antigravity has no verified plugin path, OpenCode has no plugin mechanism). Neither needs a public marketplace.

Install

Requires Node.js 20.10 or newer. Windows runs natively from PowerShell or Command Prompt; WSL is not required. Install and authenticate the consultant CLIs you want to use first.

npm install                 # fetch dependencies
npm test                    # offline checks (no real API calls)
npm run build               # regenerate dist/ inside the plugin (committed, so usually not needed)
node scripts/install.mjs    # --dry-run prints the plan without changing anything

Verify:

claude plugin details severally      # Skills (1) / MCP servers (1)
codex  plugin list                   # severally@severally-local  installed, enabled
agy    mcp list                      # severally  stdio  enabled
opencode mcp list                    # severally  connected

Restart the clients (a running session does not reload plugins). After updating the plugin, run npm run build && node scripts/install.mjs again.

Windows (PowerShell)

npm.cmd install
npm.cmd test
npm.cmd run build
node scripts/install.mjs

On Windows the installer defaults to manual mode for all four clients. It copies the standalone server (including its dependencies) to ~/.severally/runtime/severally-mcp.mjs, registers that file with the absolute path to node.exe, and copies each client's Skill. After successful registration, the source checkout can be moved or deleted. Node.js must remain installed. Global npm installation is unnecessary; --skip-global is still accepted but is optional on Windows. Add --dry-run to inspect the changes without writing anything.

Upgrading an earlier checkout-based installation: run node scripts/install.mjs --force to switch the existing MCP registrations to the dedicated runtime. Without --force, existing registrations are preserved and may still point to the checkout. Updates replace the runtime after checking the new bundle's syntax and backing up the previous file. To update later, obtain a new checkout, run npm.cmd install, npm.cmd run build, and the installer again, then restart the clients. The temporary checkout is no longer needed after registration succeeds.

Verify with claude mcp get severally, codex mcp get severally, agy mcp list, and opencode mcp list, then restart the clients. Tool names in this mode are mcp__severally__*.

CLI detection supports .exe, .cmd, and .bat through PATH/PATHEXT, including paths with spaces. For a custom executable path in config.json, use forward slashes ("bin": "C:/Tools/claude.exe") or escaped backslashes ("bin": "C:\\Tools\\claude.exe"). The test suite uses local stand-in CLIs; authentication and live consultations still depend on the installed client versions and accounts.

Linux/macOS: installation independent of the checkout

After a successful installation, the source checkout can be moved or deleted on Linux and macOS too:

  • Claude Code receives a copy of the bundled plugin in ~/.claude/skills/severally/.

  • The standalone server lives in ~/.severally/runtime/severally-mcp.mjs.

  • Codex launches the global severally-mcp command, installed from that persistent runtime. Its marketplace and plugin files are copied into ~/.severally/marketplace/.

  • Antigravity and OpenCode (and all clients in --manual mode) launch Node.js with the copied runtime directly; OpenCode is registered into ~/.config/opencode/opencode.json and its Skill is copied to ~/.config/opencode/skills/severally/.

To migrate an existing installation, run:

npm install
npm run build
node scripts/install.mjs --force

This refreshes the Codex marketplace location and switches existing direct MCP registrations to the copied runtime. Restart the clients after registration succeeds; the checkout is then disposable. Node.js and the installed runtime must remain. For updates, obtain a fresh checkout and run the same commands. Previous runtime and marketplace files are backed up under ~/.severally/backups/.

Plugin mode requires npm's global executable directory on PATH. --skip-global is accepted only when severally-mcp already resolves to the copied runtime; a command linked to an old checkout is rejected. In --manual mode no global install is needed and --skip-global is optional.

Two install modes

Plugin mode (default on macOS/Linux)

Manual mode (--manual, default on Windows)

Claude Code

places the plugin in ~/.claude/skills/severally/ (severally@skills-dir)

claude mcp add --scope user + copy the Skill on its own

Codex

codex plugin add from the copied marketplace

codex mcp add + copy the Skill on its own

Antigravity

agy mcp add + copy the Skill on its own (there is no plugin route, so both modes are the same)

same

OpenCode

opencode mcp add severally -- … + copy the Skill on its own (no plugin mechanism, so both modes are the same)

same

MCP tool names

mcp__plugin_severally_severally__*

mcp__severally__*

The installer does not break existing settings. It changes client settings only through each CLI's own commands (plugin add / mcp add). Before it does anything, it backs up ~/.claude.json, ~/.codex/config.toml and any existing Skill directory to ~/.severally/backups/<timestamp>/, naming each backup after its original path. Switching modes backs up and then removes the duplicate registration left by the other mode.

No public marketplace listing is needed. Claude Code works without a marketplace, and the Codex marketplace file .agents/plugins/marketplace.json is copied from this repository into the managed marketplace directory.

Usage

The Skill starts on requests such as "ask Codex", "have Claude review this" or "I want a second opinion", or in these situations:

  • important design decisions and hard-to-reverse choices (architecture, data migration, concurrency, security, pre-publication checks)

  • options that stay neck and neck however long you think about them

  • two or more failed attempts at the same bug with no new information

For "ask everyone" or "consult all of them", the same brief goes out at once to the other three CLIs and to a fresh session of your own CLI. Answers from your own CLI carry a "same lineage" note.

When the agent asks you to approve a hard-to-reverse change, it offers "consult first?" as one of the choices. You decide whether to start.

Not for small fixes. One consultation costs a few minutes and real quota.

A routine consultation is small. One consultant. The brief is one paragraph of plan, a few lines of facts it rests on, and one relevant code excerpt. The Skill checks the one finding that would change the decision, reports what it checked, what it did not adopt and what is still unverified, and saves the checked result next to the finding. Starting a consultation returns immediately, so the agent keeps working while it waits.

For decisions you will have to justify later (interfaces, migrations, security, concurrency), the Skill adds steps: it pins down the question and success criteria, separates imposed constraints from its own assumptions, writes down a prediction before consulting, and exports the record as Markdown into the repository.

The Skill tells the lead how to write the brief, which mode to pick (explore / review / debate), and how to read the results.

Configuration

No configuration needed by default. At startup the server checks whether codex / claude / agy / opencode are on PATH and drops the ones that are missing. To disable a consultant, change a model, or point at a specific executable, put a single ~/.severally/config.json in place. A template can be generated for your machine:

npm run init-config            # writes ~/.severally/config.json (never overwrites an existing one)

Every key is explained in CONFIG.md; config.example.json is a working example that exercises each one. Precedence is environment variables > config file > auto-detection > defaults. The config is read once at server startup, so restart the client after changing it.

Before you use it

Consultation round limit

A consultation chain defaults to 5 total rounds (one initial consultation and up to four follow-ups). Set SEVERALLY_MAX_ROUNDS in the MCP server's environment to change the limit to 1–20 total rounds; values above 20 are capped at 20. For example, SEVERALLY_MAX_ROUNDS=20 allows the initial consultation plus 19 follow-ups. Restart the client/server after changing this setting. This is an environment setting, not a key in config.json. The server reports the active budget in rounds_remaining.

What never reaches a consultant: your conversation history. What no consultant can do:

  • write

  • use the network (web search and browsing are the only exception)

  • use MCP

  • consult anyone else

Reading is not blocked. Per child session:

Consultant

Reads the disk

Shell

Write / execute

Codex

yes

read-only

no

Claude Code

yes (Read / Glob / Grep)

none

no

Antigravity

yes (read_file)

none

no

OpenCode

yes (read / glob / grep)

none

no

Every consultant starts in an empty working directory, and the server does not tell it where your repository is. Paste whatever it should see into the brief — or, when a consultant needs to explore rather than read what you picked out, name absolute paths in context.expose_paths. Those files and directories are copied read-only into the consultant's working directory (under ./workspace) and are then the only part of your repository it has. At most 20 entries, 500 files and 5 MB in total; symlinks are skipped rather than followed; text files are credential-masked exactly as the brief is, binary files are copied unchanged. The copy is deleted with the job, and the history keeps only which paths were shown, not their contents. The Claude Code consultant's whole-disk read scope is dropped for a consultation that uses it.

  • A consultant that failed (usage_limit / auth / timeout …) and one that answered on thin grounds come back as different things. A failure is not "no problems found"

  • History is kept in ~/.severally/history/. Each round is one file holding the brief that was sent, the consultant's answer, and the verdicts the lead wrote per finding (with consult_record). consult_export turns it into Markdown for the repository. Verdicts can be added later, from another session

  • Credentials are masked in the brief that is sent and in the results that come back

  • You can also consult a different model of your own CLI (a Fable lead asking Opus, target: "claude:opus"). The answer carries a "same lineage" note. Declare your own model with caller_model and the record keeps who asked whom

  • The record also keeps whether you asked for the consultation or accepted the agent's offer (self-declared). Offers you declined are written one per line to ~/.severally/history/offers.jsonl. Neither restricts anything

  • Consultations about code decisions return mostly checks the lead can run itself. Consultations about project policy return checks that depend on other people, and those come back to you unrun

Uninstall

Windows (the default manual installation):

claude mcp remove severally -s user
codex mcp remove severally
agy mcp remove severally
# opencode has no `mcp remove`; delete the "severally" entry from
# the "mcp" object in $HOME/.config/opencode/opencode.json
Remove-Item -LiteralPath "$HOME/.claude/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.codex/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.gemini/config/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.config/opencode/skills/severally" -Recurse -Force
Remove-Item -LiteralPath "$HOME/.severally/runtime" -Recurse -Force
# Only if you previously installed the global command:
npm.cmd uninstall -g severally-mcp

If CODEX_HOME is set, use that directory instead of $HOME/.codex for the Codex Skill.

macOS/Linux:

Plugin mode:

rm -rf ~/.claude/skills/severally                       # Claude Code
codex plugin remove severally --marketplace severally-local
codex plugin marketplace remove severally-local
agy   mcp remove severally                              # Antigravity (registered directly in both modes)
rm -rf ~/.gemini/config/skills/severally
# OpenCode (registered directly in both modes; opencode has no `mcp remove`,
# so delete the "severally" entry from the "mcp" object by hand):
#   edit ~/.config/opencode/opencode.json(c)
rm -rf ~/.config/opencode/skills/severally
npm uninstall -g severally-mcp
rm -rf ~/.severally/runtime ~/.severally/marketplace

Manual mode:

claude mcp remove severally -s user
codex  mcp remove severally
agy    mcp remove severally
# opencode: delete the "severally" entry from the "mcp" object in
# ~/.config/opencode/opencode.json(c) by hand
rm -rf ~/.claude/skills/severally ~/.codex/skills/severally ~/.gemini/config/skills/severally \
       ~/.config/opencode/skills/severally
# Only if a global command was previously installed:
npm uninstall -g severally-mcp
rm -rf ~/.severally/runtime ~/.severally/marketplace

Either way, history and backups stay in ~/.severally/ (delete it if you no longer need them).

Available Tools

6 tools
consult_cancelCancel a running consultationA

Stop a running consultation and kill the consultant process and everything it spawned. Pass group_id instead of job_id to stop every consultant in a fan-out.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idNo
group_idNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the destructive nature ('kill the consultant process and everything it spawned') and the group behavior, which is significant and goes beyond what would be assumed. However, it does not address reversibility, side effects on data, or permission requirements, but for a cancel operation, the key behaviors are adequately covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—two sentences—with the primary action front-loaded. It includes the essential parameter guidance without extraneous detail, making it efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema and no annotations, the description covers the purpose, the parameter semantics, and the key behavioral trait (killing spawned processes). It does not address edge cases like providing both parameters or neither, but these are not critical for typical usage and the description is sufficiently complete for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It clearly explains the semantic difference between job_id (single job) and group_id (fan-out group), which is critical for correct usage. It does not specify that both are optional or what happens if neither is provided, but the core distinction is well conveyed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear, specific action: stop a running consultation and kill the consultant process. It explicitly mentions the spawned processes, making it distinct from the sibling tools (consult_get, consult_start, consult_record, consult_export, consult_list), which all serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides guidance on parameter usage ('Pass group_id instead of job_id to stop every consultant in a fan-out'), but it does not explicitly address when to use this tool versus alternatives. It does not mention any sibling tools or conditions for selection, so the when-to-use context is only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consult_exportExport a consultation as a record to commitA

Render one consultation (chain_id) or one fan-out (group_id) as Markdown: the brief as it was sent, each consultant's answer as it came back, and the verdicts recorded against each point -- with the ones nobody checked marked as unchecked. Nothing is summarised across consultants and nothing is scored. The text is returned, not written: put it wherever the decision belongs in the repository (a decision record next to the code it is about), which is the only place a teammate will find it. Reads the on-disk history, so a consultation from an earlier session can still be exported.

ParametersJSON Schema
NameRequiredDescriptionDefault
chain_idNo
group_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral disclosure and succeeds: it states the output format (Markdown), what is included, what is excluded ('Nothing is summarised... nothing is scored'), that the operation is non-writing ('returned, not written'), and that it reads on-disk history. This is far beyond a minimal description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then adds details in a logical order; every sentence contributes either behavioral scope, content expectations, or usage guidance. Despite being three sentences, it is efficient for the information it must convey without annotations or output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given two simple optional-looking parameters, no annotations, and no output schema, the description covers output format, contents, non-mutation, and data source. The main missing piece is an explicit statement of the required/forbidden parameter combination (exactly one of chain_id/group_id), which prevents full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain chain_id and group_id. It does: chain_id selects a consultation, group_id selects a fan-out. It stops short of stating whether exactly one is required or what happens if both/neither are provided, which leaves some ambiguity for invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Render one consultation (chain_id) or one fan-out (group_id) as Markdown.' It details the exact contents of the output (brief, answers, verdicts, unchecked marks) and distinguishes itself from write-oriented siblings by saying 'The text is returned, not written.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context for when the tool is appropriate: to export a consultation/fan-out as a decision record in the repository, and it notes that on-disk history lets earlier sessions be exported. It does not explicitly name alternative sibling tools or call out when-not-to-use conditions, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consult_getGet consultation status or resultA

Fetch the state of a consultation. Pass wait_ms to block until it finishes (capped at 45000 ms, which stays under the request timeout MCP clients apply) instead of polling in a tight loop; a typical consultation takes one to five minutes, so expect to call this several times, and do something else in between. A completed job carries the structured answer; a failed one carries failure.kind (timeout, auth, usage_limit, model_unavailable, invalid_output, cli_error, spawn_error) — that is "no advice was obtained", which is different from advice that arrived with thin evidence (see quality.evidence_basis). Pass group_id instead of job_id to fetch a fan-out; wait_ms then waits for every consultant in it.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idNo
wait_msNo
group_idNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does well: it explains blocking behavior, the 45-second cap, the expectation of multiple calls, failure kinds, and the distinction between failure and thin evidence. It does not cover auth or rate limits, but the core behavior is well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph, but every sentence adds value—blocking, failure kinds, fan-out. It is front-loaded with the core purpose and then elaborates. Slightly long, but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and three parameters, the description covers the essential behavioral and semantic ground. It mentions the structured answer and failure.kind, hinting at response fields. It could add more about edge cases (e.g., invalid job_id) or the exact response shape, but it is adequate for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It does: wait_ms is detailed (blocking, cap, purpose), group_id is explained as fan-out, and job_id is implied as the primary identifier. This goes beyond the bare schema and compensates for the lack of parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fetches the state of a consultation, which is a specific verb-resource pair. It does not explicitly contrast with sibling consult_list, but the purpose is unambiguous and distinct from starting or canceling consultations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete guidance on when to use wait_ms (to block instead of polling), how long consultations take, and when to pass group_id for fan-out. It does not explicitly name alternative tools, but the situational guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consult_listList recent consultationsA

Recent consultations from this session, newest first, with their status and one-line summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses ordering (newest first) and that results include status and summary, which is useful. However, it doesn't disclose whether this is a read-only operation, whether it reflects only the current session, or what happens when there are no consultations. The 'from this session' scoping is a meaningful behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that conveys the resource, scope, ordering, and included fields with zero waste. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with one optional parameter and no output schema, the description is nearly complete. It could mention that the output is a list of consultations with status and summary, but that's already implied. The main gap is not explaining the 'limit' parameter, though the schema covers it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It doesn't explain the 'limit' parameter at all, but the schema itself defines it as an integer with min/max, which is self-explanatory. The description adds no meaning beyond the schema for the parameter, but the parameter is simple enough that this is a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a clear resource ('recent consultations from this session'), and adds ordering ('newest first') plus included fields ('status and one-line summary'). This clearly distinguishes it from siblings like consult_get (single consultation retrieval) and consult_start (initiating a consultation).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is for browsing recent session consultations, which is distinct from consult_get (retrieving a specific one) and consult_cancel (cancelling). It doesn't explicitly state when not to use it or name alternatives, but the context is clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consult_recordRecord what checking a point showedA

Write your own verdict against one or more points of an answer, after you have checked them in the repository. Ids come from the result: findings are f1, f2 ..., unknowns u1 ..., next_checks c1 ... . verdict says what checking showed -- "confirmed" (it holds here), "not_applicable" (true in general, not for this codebase), "unverifiable" (cannot be settled with what you can reach), "unverified" (not checked yet, and say in effect why not). It does not say whether you adopted the point. effect is what it changed about your decision; note is the evidence you used. Recording the same id again replaces that entry. This server stores what you write and counts the verdicts; it never infers one, and never decides a consultation was worth it. The entry is saved beside the answer and the brief in ~/.severally/history, which is what makes the decision readable a month from now; the job is read back from that history, so a consultation from an earlier session can still be recorded against. Pass reflection to record what the answer added over what you already expected, when the consultation was started with a prediction. A prediction itself cannot be written here: it goes in consult_start, before the consultant runs, which is the only thing that makes it a prediction.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
entriesNo
reflectionNowritten after you have read the answer. There is deliberately no hit/miss label: a point you predicted can still arrive with the evidence that settles it, and a surprise can still be wrong.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly. It discloses that the server stores entries, overwrites on duplicate ids, never infers verdicts, and persists to ~/.severally/history, enabling cross-session recording. This goes well beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense. Each sentence contributes to understanding the tool's behavior, id conventions, verdict meanings, and storage details. It is front-loaded with the core purpose and avoids fluff, though it could be slightly tightened without losing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (nested entries, reflection object, multiple verdicts, no output schema), the description covers all necessary details: id format, verdict semantics, overwrite behavior, persistence location, cross-session support, and reflection conditions. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only reflection has a description). The description compensates fully by explaining what entries contain, the meaning of each verdict value, what effect and note represent, and when reflection should be used. It adds meaning that the schema omits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (write a verdict), the resource (points of an answer), and the context (after checking them in the repository). It differentiates from siblings by explicitly noting that predictions belong in consult_start, not here.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the reflection parameter (when started with a prediction) and explicitly excludes writing predictions here, pointing to consult_start. It also indicates the id source from the result. It doesn't explicitly contrast with consult_get or consult_export, but the usage context is fairly clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consult_startStart a peer consultationA

Start a consultation with another agent (or several, via targets). Returns a job_id (or a group_id for several) immediately; the work runs in the background.

target: which consultant to ask. This machine can reach: (none -- no consultant CLI is installed). The everyday names work too: gpt/chatgpt/openai, claude/anthropic, gemini/agy/google. A consultant that is not in that list is refused up front, so do not retry it -- say which ones are available instead. Consulting your own CLI is allowed but is a fresh-context check rather than an independent opinion, and the result says so. targets: ask up to 3 consultants the same question at once (mutually exclusive with target, no duplicates). Every member gets the byte-identical brief and one group_id; poll it with consult_get({ group_id }). A follow-up (followup_to) always names one consultant -- fan-out is never available on a follow-up. A consultant may name the model to run it on as a suffix: "claude:claude-opus-5". What each consultant is allowed to run is set by the operator, and this server currently allows:

    Only pass a model when the user asked for one; a name outside the list is refused before the
    consultation starts, and a name matching two of them is refused rather than guessed.

caller: optional -- the CLI you are running in ("codex" / "claude-code" / "antigravity"), so the server can annotate a same-vendor consultation. caller_model: optional -- the model you are running on (e.g. "claude-opus-5"), self-declared and never checked. With it, a same-vendor caveat can say "same lineage, different model" and the record keeps who asked whom; it never changes which model the consultant runs. mode: explore - hand over objective/constraints/facts and withhold your own preferred solution, to get independent options, alternative problem framings and blind spots. context.proposal MUST be empty on round 1. review - hand over your current proposal AND the reasoning behind it, to get weaknesses, counterexamples, concrete improvements and the conditions under which it holds. context.proposal is required. debate - hand over the disputed point, your position (context.proposal), the other side's claims (context.counterpoints) and the extra evidence, to get what would change the judgement, how to test it, and which disagreements remain. Both are required.

The consultant starts in an empty working directory and is not told where your repository is: put every fact it needs into context.facts and paste the relevant passages into context.artifacts. Model, permissions, round count, timeout and size caps are fixed by this server and cannot be raised from a request.

prediction: optional -- { expected, worry }: the bottom line you expect back and, in one sentence, what you are most worried about. Stored with the consultation and NEVER sent to the consultant. It can only be written here, before the consultant runs, so that afterwards you cannot rewrite what you thought beforehand; consult_record takes the other half (reflection) once you have read the answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that work runs asynchronously, the consultant starts in an empty working directory and is not told the repo location, model/permissions/timeouts are fixed, same-vendor calls are annotated with a caveat, and prediction is never sent to the consultant and cannot be rewritten afterward.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool is genuinely complex with nested parameters, three modes, fan-out, model suffixes, and privacy behavior. The structure is front-loaded with the key contract, then organized by parameter, and nearly every sentence carries operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high parameter complexity, no output schema, and no annotations, the description is remarkably complete. It covers return values, how to poll via consult_get, mode preconditions, unavailable targets, same-vendor caveats, prediction privacy, and environment constraints that would otherwise be invisible to an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the parameter documentation burden falls entirely on the description. It compensates thoroughly: target/targets semantics, followup_to restrictions, caller/caller_model intent, prediction shape and immutability, and mode-specific requirements for context.proposal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Start a consultation with another agent (or several, via targets)' and states the immediate return contract (job_id/group_id, background execution). It clearly distinguishes consult_start from its siblings by framing it as the initiation step and explicitly referencing consult_get and consult_record for follow-up actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance for each mode: explore requires empty proposal, review requires proposal, debate requires both proposal and counterpoints. It also names alternatives and exclusions: targets is mutually exclusive with target, follow-ups never fan out, unavailable consultants are refused up front and should not be retried, and a model suffix should only be passed when the user asked for one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv1.1.0
    • First observedconsult_cancel
    • First observedconsult_export
    • First observedconsult_get
    • First observedconsult_list
    • First observedconsult_record
    • First observedconsult_start

TDQS

A4.4/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct lifecycle stage: start, get, cancel, record, export, list. No two tools overlap in purpose; consult_get and consult_list both read state but one is for a specific job/group and the other is a session overview, which is clear from descriptions.

Naming Consistency5/5

All tools follow a consistent consult_verb pattern: consult_start, consult_get, consult_cancel, consult_record, consult_export, consult_list. The verb is always second and snake_case is used throughout.

Tool Count5/5

Six tools cover the full consultation workflow without redundancy: start, poll/get, cancel, record verdicts, export, and list. This is a well-scoped set for a single-purpose server.

Completeness4/5

The lifecycle is complete: create, read, cancel, record, export, list. The only minor gap is no explicit delete/cleanup tool for history, but that is not essential to the core workflow and can be handled outside the server.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers