mcp-devils-advocate
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-devils-advocateRun a devil's advocate on my plan to launch a new product."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-devils-advocate
MCP server that stress-tests reasoning: devil's advocate, premortem analysis, assumption audits, steelmanning and an all-in-one gauntlet, run as enforced, structured protocols.
Why
LLMs agree too easily. Ask one whether your plan is good and you get a polite yes with three bullet points. This server makes it harder to shortcut adversarial thinking. It never writes content itself. It is a state machine that makes the client LLM finish each phase properly: minimum counts, categories, severity ratings and length floors. It checks every submission and returns errors the model can act on. It will not move to the next phase until the current one is complete. At the end it writes a report with a deterministic assessment.
Length floors alone are easy to game, so every submission also goes through a quality gate (see below). It rejects padding ("x" * 30), near-duplicate items, "counterarguments" that only restate the claim, and rebuttals, mitigations, tests or responses copied from the item they answer. Nothing is saved until the whole batch passes. The gate uses deterministic text heuristics, not a semantic judge. It stops lazy and degenerate submissions. It cannot prove that an argument is good.
Five modes:
Mode | Protocol |
| ≥3 counterarguments (categorized, severity 1–5, ≥2 distinct categories) → honest rebuttal of every severity ≥3 counterargument ( |
| set a failure horizon → ≥4 failure causes (likelihood × impact) → concrete mitigation for every cause scoring ≥9, with residual risk → verdict |
| ≥4 assumptions ( |
| ≥3 strongest points for the OPPOSING position → honest |
| all four lenses on the same claim, in this order: devil's advocate → assumptions → premortem → steelman (9 phases). Each lens gets its own verdict and a combined verdict is computed from them |
Related MCP server: Steelmind MCP
Tools
Tool | Arguments | Returns |
|
| New |
|
| Validates the whole batch at once. Returns |
|
| Only when all phases are complete. The compiled report: claim, every item organized by phase (rebuttals/mitigations/tests/responses attached to their targets), aggregate risk score, and the assessment with its rules. Gauntlet reports add one sub-verdict per lens under |
|
| A completed review as a shareable decision memo (Markdown, see below) or as JSON ( |
|
| Where a review stands: status, step (e.g. |
| — | All reviews, newest first: id, claim snippet, mode, status, current phase, timestamps. Files in the data directory that are not valid reviews are listed under |
|
| Marks an unfinished review as abandoned. It stays listed for the record and accepts no more submissions. A completed review cannot be abandoned because its verdict is final |
Every validation error reaches the model word for word under both SDK generations, for example item 1: 'text' is a near-duplicate of item 0 (similarity 1.00) — each counterargument must make a distinct point. Under mcp 2.x a plain exception in a tool shows up as just Error executing tool submit, so the server converts every validation error into the SDK's ToolError. The end-to-end tests check this.
Resources & prompts
Kind | Name | What it gives |
Resource |
| The review listing as JSON (same content as |
Resource template |
| The Markdown decision memo of a completed review, or a progress page (phase, step, what is missing) for an active or abandoned one |
Prompt |
| Step-by-step instructions for running the whole protocol in the chosen mode. A client like Claude Desktop can offer it from the prompt picker. An unknown mode returns an invalid-params error that lists the valid modes |
Quality gate
Text is compared as sets of content words. The text is lowercased, accents are stripped and punctuation is removed. A small English + Spanish stopword and filler list is dropped (negations included, so "we should not do X" adds nothing to "we should do X"). A tiny stemmer then makes rewrite/rewriting and service/services count as the same word. The thresholds are constants in quality.py, and every phase's instructions list them under quality_checks.
Check | Applies to | Rejected when |
Low information | every free-text field | fewer than 4 distinct content words (counterarguments, steelman points) or 3 (everything else). Keyboard-mash tokens like |
Near-duplicate | every item against the others in the same phase, both in this batch and already saved | Jaccard word overlap ≥ 0.8 |
Restating the claim | counterarguments, steelman points, rebuttal justifications, | ≥ 80% of the content words come from the claim, or fewer than 2 new ones |
Parroting the target | rebuttals, mitigations, tests, responses (both | ≥ 80% of the content words are copied from the item being answered, or fewer than 2 new ones |
All four shortcuts found in the 0.1.0 audit now fail with a message that names the item, and nothing is saved: 30 × x, one counterargument pasted under three categories, counter responses that repeat the claim, and justifications copied from the counterargument. Each of them used to end in "claim survives scrutiny". tests/test_quality.py keeps them as regression tests, next to the README example below, which still passes unchanged.
Assessment rules (deterministic)
Mode | Risk score |
|
|
|
| # counterarguments that | ≥2 hold, or any severity-5 holds | exactly 1 holds, or ≥2 partially hold | otherwise |
| average likelihood × impact (1–25) | average > 12 | average > 6, or any mitigation with | otherwise |
| # unverified load-bearing assumptions | ≥2 load-bearing with evidence | exactly 1 with | otherwise |
| # opposing points conceded | every point conceded | concessions ≥ counters | counters outnumber concessions |
| 2 per refuting lens + 1 per lens needing revision (0–8) | ≥2 lenses refute | exactly 1 lens refutes, or ≥2 lenses need revision | otherwise |
In a gauntlet, each lens is assessed with its own rule above, applied to the same review.
How it works
flowchart TD
S[start_review claim + mode] --> M{mode}
M -->|devils_advocate| C["counterarguments<br/>≥3, ≥2 categories, severity 1–5"]
C --> D{any severity ≥ 3?}
D -->|yes| R["rebuttals<br/>one per severe counterargument"]
D -->|no| V
R --> V[all phases complete]
M -->|premortem| P1[setup: horizon] --> P2["failure_causes<br/>≥4, likelihood × impact"]
P2 --> P3{any score ≥ 9?}
P3 -->|yes| P4["mitigations<br/>action + residual risk"] --> V
P3 -->|no| V
M -->|assumptions| A1["assumptions<br/>≥4, load_bearing + evidence"]
A1 --> A2{unverified load-bearing?}
A2 -->|yes| A3["tests<br/>cheapest verification"] --> V
A2 -->|no| V
M -->|steelman| T1["strongest_case<br/>≥3 points for the opposing side"]
T1 --> T2["responses<br/>concede or counter each"] --> V
M -->|gauntlet| GA["devil's advocate phases"] --> GB["assumptions phases"] --> GC["premortem phases"] --> GD["steelman phases"] --> V
Q[[quality gate<br/>junk · duplicates · restating · parroting]] -. checks every submit .-> C
V --> G["get_verdict<br/>report + risk score + assessment"]
G --> E["export_report / review://id<br/>Markdown decision memo"]Every submit is checked in two layers, and the whole batch passes or fails together. First the fields are validated (types, enums, ranges, length floors, target indices). Then the quality gate runs. The server only moves on when the phase's requirements are met. It skips dependent phases that have no targets, for example when no counterargument reached severity 3. State is saved as one JSON file per review in ~/.mcp-devils-advocate/ (override it with the DEVILS_ADVOCATE_DIR environment variable). Reviews written by 0.1.0 still load.
Quickstart
Not on PyPI yet. The
publish.ymlworkflow uploads to PyPI (and the MCP Registry entry inserver.jsonpoints there) when av*tag is pushed. No release has been tagged yet, sopip install mcp-devils-advocate/uvx mcp-devils-advocatedo not work. Until then, install from GitHub:
Claude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"devils-advocate": {
"command": "uvx",
"args": ["--from", "git+https://github.com/AleBrito124356/mcp-devils-advocate", "mcp-devils-advocate"]
}
}
}Claude Code:
claude mcp add devils-advocate -- uvx --from git+https://github.com/AleBrito124356/mcp-devils-advocate mcp-devils-advocatepip (permanent install, then use mcp-devils-advocate as the command):
pip install "git+https://github.com/AleBrito124356/mcp-devils-advocate"With no arguments, mcp-devils-advocate runs the stdio MCP server, the same as in 0.1.0, so existing client configs keep working.
Example session
User: We're considering rewriting our backend in Rust. Play devil's advocate before we commit.
The assistant calls start_review(claim="We should rewrite our backend in Rust", mode="devils_advocate") and receives (abridged):
{
"review_id": "rev-k4d7",
"status": "active",
"phases": ["counterarguments", "rebuttals"],
"instructions": {
"phase": "counterarguments",
"step": "1/2",
"claim": "We should rewrite our backend in Rust",
"goal": "Attack the claim as a devil's advocate. Generate the strongest, most specific counterarguments you can — no strawmen.",
"item_format": {
"text": "str, >= 30 characters — the counterargument, concrete and falsifiable",
"category": "one of: evidence, incentives, base_rates, alternatives, second_order, scope",
"severity": "int 1-5 — how damaging to the claim if true (5 = fatal)"
},
"rules": [
"Submit at least 3 counterarguments (send them all in one call).",
"Use at least 2 distinct categories.",
"Every counterargument with severity >= 3 must be rebutted in the next phase — assign severity honestly."
],
"quality_checks": [
"'text' needs at least 4 distinct content words (...) — junk padding is rejected.",
"Each counterargument must be distinct: a Jaccard word-overlap of >= 0.8 with another counterargument in this phase is rejected as a near-duplicate.",
"A counterargument that only restates the claim (>= 80% of its content words from the claim, or fewer than 2 new ones) is rejected."
]
}
}The assistant does the actual thinking and calls submit:
submit("rev-k4d7", [
{"text": "Full rewrites of working systems fail or massively overrun far more often than they succeed; incremental strangler migrations have much better base rates.", "category": "base_rates", "severity": 4},
{"text": "The team has zero production Rust experience; hiring and ramp-up costs will land exactly while feature delivery is frozen.", "category": "evidence", "severity": 4},
{"text": "Profiling shows the bottleneck is the database layer, not CPU — a rewrite optimizes the part that isn't slow.", "category": "alternatives", "severity": 5},
{"text": "Engineers pushing the rewrite are the ones who want Rust on their CV — incentives are not aligned with the business case.", "category": "incentives", "severity": 2}
])The server validates the batch, marks the phase complete, and returns the rebuttals phase with three targets (the severity ≥3 items). The assistant rebuts each one honestly:
submit("rev-k4d7", [
{"index": 0, "verdict": "partially_holds", "justification": "True in general, but we can scope the rewrite to the two stateless services first, which is effectively a strangler migration."},
{"index": 1, "verdict": "holds", "justification": "No honest rebuttal: nobody on the team has shipped Rust, and the hiring market for it is thin in our region."},
{"index": 2, "verdict": "holds", "justification": "The profiling data is real — the p99 latency lives in Postgres queries. A Rust rewrite does not touch that."}
])All phases are complete, so get_verdict("rev-k4d7") compiles the report:
{
"review_id": "rev-k4d7",
"claim": "We should rewrite our backend in Rust",
"mode": "devils_advocate",
"risk_score": {"value": 2, "scale": "0-3 counterarguments that held after rebuttal"},
"assessment": "claim refuted",
"assessment_reason": "Of 3 rebutted counterargument(s): 2 hold, 1 partially hold, 0 refuted; a severity-5 counterargument holds. Rules: refuted if >=2 hold or any severity-5 holds; ..."
}export_report("rev-k4d7") then returns the decision memo shown in the demo below, with a Next actions checklist the team can paste into an issue.
Assistant: The claim did not survive scrutiny. Two counterarguments held, including a severity-5 one: our bottleneck is the database, not CPU, so a Rust rewrite attacks the wrong problem — and we have no Rust experience in-house. Recommendation: fix the query layer first; if CPU ever becomes the bottleneck, migrate one stateless service as a pilot.
CLI
The same command works on the review files without an LLM client and without importing the MCP SDK:
Command | What it does |
| Run the stdio MCP server (the default) |
| Every review, newest first. Malformed files are reported as warnings |
| Status, progress, what is missing, and the current phase's instructions |
| Export a completed review as a decision memo or JSON |
| Offline scripted review through the real |
--data-dir DIR (before or after the command) overrides DEVILS_ADVOCATE_DIR. python -m mcp_devils_advocate ... is equivalent. For demo, --data-dir keeps the review so that list/show/report can open it afterwards. Otherwise it runs in a temporary directory and deletes it at the end. Try demo --mode gauntlet --memo-only for the full four-lens memo.
This is the real output of mcp-devils-advocate demo. A test checks it against the code, so it cannot go stale:
Offline demo: no LLM, no network. A scripted client plays the model and runs a
'devils_advocate' review through the real ReviewStore.
start_review(claim='We should rewrite our backend in Rust', mode='devils_advocate', context=...)
-> rev-ess4, phases: counterarguments -> rebuttals
== Phase 1/2: counterarguments ==
Goal: Attack the claim as a devil's advocate. Generate the strongest, most specific counterarguments you can — no strawmen.
Item format:
text: str, >= 30 characters — the counterargument, concrete and falsifiable
category: one of: evidence, incentives, base_rates, alternatives, second_order, scope
severity: int 1-5 — how damaging to the claim if true (5 = fatal)
Rules:
- Submit at least 3 counterarguments (send them all in one call).
- Use at least 2 distinct categories.
- Every counterargument with severity >= 3 must be rebutted in the next phase — assign severity honestly.
Quality checks:
- 'text' needs at least 4 distinct content words (stopwords and filler like 'good', 'bad', 'really' do not count), and once it has 6+ content words no single word may make up more than 50% of them — junk padding is rejected.
- Each counterargument must be distinct: a Jaccard word-overlap of >= 0.8 with another counterargument in this phase is rejected as a near-duplicate.
- A counterargument that only restates the claim (>= 80% of its content words from the claim, or fewer than 2 new ones) is rejected.
A lazy client pastes one counterargument under three categories:
REJECTED — Invalid submission — nothing was saved. Fix these problems and resubmit:
- item 1: 'text' is a near-duplicate of item 0 (similarity 1.00) — each counterargument must make a distinct point
- item 2: 'text' is a near-duplicate of item 0 (similarity 1.00) — each counterargument must make a distinct point
The real client does the thinking instead.
submit(rev-ess4, 4 item(s)):
[base_rates, severity 4] Full rewrites of working systems fail or massively overrun far more often than they succeed; in…
[evidence, severity 4] The team has zero production Rust experience; hiring and ramp-up costs will land exactly while …
[alternatives, severity 5] Profiling shows the bottleneck is the database layer, not CPU — a rewrite optimizes the part th…
[incentives, severity 2] Engineers pushing the rewrite are the ones who want Rust on their CV — incentives are not align…
-> phase_complete; next phase: rebuttals
== Phase 2/2: rebuttals ==
Goal: Rebut each high-severity counterargument honestly. If you cannot refute one, admit that it holds.
Targets: 3 (#0, #1, #2)
Item format:
index: int — index of the counterargument being rebutted (see 'targets')
verdict: one of: holds, partially_holds, refuted — does the counterargument survive your rebuttal?
justification: str, >= 20 characters — why that verdict
Rules:
- Provide exactly one rebuttal per target (3 pending).
- 'holds' means the counterargument stands and damages the claim. Do not mark 'refuted' without an actual refutation.
Quality checks:
- 'justification' needs at least 3 distinct content words (stopwords and filler like 'good', 'bad', 'really' do not count), and once it has 6+ content words no single word may make up more than 50% of them — junk padding is rejected.
- Each rebuttal must be distinct: a Jaccard word-overlap of >= 0.8 with another rebuttal in this phase is rejected as a near-duplicate.
- A rebuttal that only restates the claim (>= 80% of its content words from the claim, or fewer than 2 new ones) is rejected.
- A rebuttal that parrots the counterargument it answers (>= 80% of its content words copied, or fewer than 2 new ones) is rejected.
submit(rev-ess4, 3 item(s)):
#0 partially_holds: True in general, but we can scope the rewrite to the two stateless services first, which is eff…
#1 holds: No honest rebuttal: nobody on the team has shipped Rust, and the hiring market for it is thin i…
#2 holds: The profiling data is real — the p99 latency lives in Postgres queries. A Rust rewrite does not…
-> complete: All phases complete. Call get_verdict('rev-ess4') to compile the final report.
get_verdict(rev-ess4) -> claim refuted (risk score 2: 0-3 counterarguments that held after rebuttal)
export_report(rev-ess4, format='markdown'):
# Decision memo: We should rewrite our backend in Rust
- **Review:** `rev-ess4` · mode `devils_advocate` (Devil's advocate)
- **Assessment:** **claim refuted**
- **Risk score:** 2 — 0-3 counterarguments that held after rebuttal
- **Started:** 2026-01-15T09:00:00+00:00 · **Completed:** 2026-01-15T09:02:00+00:00
**Context:** 9-year-old Python monolith, 14 engineers, enterprise customers complaining about p99 latency.
## Verdict
**claim refuted** — Of 3 rebutted counterargument(s): 2 hold, 1 partially hold, 0 refuted; a severity-5 counterargument holds. Rules: refuted if >=2 hold or any severity-5 holds; revision if exactly 1 holds or >=2 partially hold; survives otherwise.
## Counterarguments
- **#0 · base_rates · severity 4** — Full rewrites of working systems fail or massively overrun far more often than they succeed; incremental strangler migrations have much better base rates.
- Rebuttal (*partially holds*): True in general, but we can scope the rewrite to the two stateless services first, which is effectively a strangler migration.
- **#1 · evidence · severity 4** — The team has zero production Rust experience; hiring and ramp-up costs will land exactly while feature delivery is frozen.
- Rebuttal (*holds*): No honest rebuttal: nobody on the team has shipped Rust, and the hiring market for it is thin in our region.
- **#2 · alternatives · severity 5** — Profiling shows the bottleneck is the database layer, not CPU — a rewrite optimizes the part that isn't slow.
- Rebuttal (*holds*): The profiling data is real — the p99 latency lives in Postgres queries. A Rust rewrite does not touch that.
- **#3 · incentives · severity 2** — Engineers pushing the rewrite are the ones who want Rust on their CV — incentives are not aligned with the business case.
- _severity below 3 — no rebuttal required_
## Next actions
- [ ] Resolve counterargument #0 (partially holds, severity 4): Full rewrites of working systems fail or massively overrun far more often than they succeed; incremental strangler migrations have much better base rates.
- [ ] Resolve counterargument #1 (holds, severity 4): The team has zero production Rust experience; hiring and ramp-up costs will land exactly while feature delivery is frozen.
- [ ] Resolve counterargument #2 (holds, severity 5): Profiling shows the bottleneck is the database layer, not CPU — a rewrite optimizes the part that isn't slow.
<sub>Generated by mcp-devils-advocate from review `rev-ess4`.</sub>Compatibility
Tested | |
Python | 3.10+ (developed and tested on 3.14) |
MCP Python SDK | 1.5.0, 1.9.0, 1.30.0 ( |
The dependency is mcp>=1.5.0,<3. SDKs before 1.3 write stdout in the Windows code page, so any non-ASCII text (an em dash, a Spanish accent) corrupts the JSON-RPC stream. The 1.3/1.4 stdio client hangs on Windows, which is why 1.5.0 is the lowest version the end-to-end suite can verify. The upper bound keeps a future 3.x from breaking installs silently, as 2.0 did to 0.1.0.
Development
git clone https://github.com/AleBrito124356/mcp-devils-advocate
cd mcp-devils-advocate
python -m venv .venv && .venv/bin/pip install -e ".[dev]" # Windows: .venv\Scripts\pip
.venv/bin/python -m pytestRun the server from the source tree with python -m mcp_devils_advocate (or python -m mcp_devils_advocate.server).
The suite has 140+ tests, all offline:
test_core.py: phase machine, validation, verdict rules and persistence. It also covers the lifecycle fixes: complete reviews can't be abandoned, malformed files are skipped, same-second ordering is stable, context reaches every phase, and 0.1.0 files still load.test_quality.py: every quality check on both sides of its threshold, the four audit exploits, and the README example as a regression test.test_gauntlet.py: the nine-phase flow, auto-skipped phases, and every branch of the combined rule.test_memo.py: the memo has every section in all five modes, and Next actions matches the open work.test_cli.py:demo,list,show,reportand error exits, run as real subprocesses.test_server_e2e.py: starts the real server (module and console script) and drives it over stdio with the official SDK client. It covers tools, resources, the prompt, a full flow, a gauntlet, and error messages reaching the model word for word.
mcp_devils_advocate/core.py, quality.py and memo.py use only the standard library, and every test except test_server_e2e.py runs without mcp (that file is skipped when the SDK is missing). To check the other SDK generation, install it into a second venv: pip install -e . "mcp<2" pytest, then run python -m pytest.
Related MCP servers
Part of a family of small, dependency-light MCP servers:
mcp-decision-lab — weighted decision matrices with sensitivity analysis
mcp-secret-sentinel — scan code for exposed secrets, always redacted
mcp-git-historian — churn hotspots, blame summaries, bus factor
mcp-memory-vault — persistent memory with SQLite FTS5 search
License
MIT
Available Tools
7 toolsabandon_reviewA
Abandon an unfinished review that is no longer worth finishing.
Abandoned reviews stay listed for the record but accept no further submissions and produce no verdict. Completed reviews cannot be abandoned — their verdict is final.
Args: review_id: The review session id (e.g. "rev-x9k2"). reason: Why the review is being abandoned (required, non-empty).
Returns: dict confirming the abandonment.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| review_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and handles the key behavioral facts: abandoned reviews remain listed, accept no further submissions, and produce no verdict; completed reviews cannot be abandoned. It omits author/permission requirements and whether the action is reversible, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by state effects and the exclusion, then a compact Args/Returns block. No sentence is wasted or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no annotations and no output schema, the description covers behavior, preconditions, and parameter constraints adequately, and hints at the return shape. Missing authorization/prerequisite context is the only notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: review_id is given a concrete format example ('rev-x9k2') and reason is flagged as required and non-empty, a constraint absent from the schema. It does not describe length limits or valid reason content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (abandon) and resource (unfinished review) and scopes it to reviews that are 'no longer worth finishing.' The second sentence distinguishes this lifecycle transition from sibling operations like start_review or get_verdict by clarifying what an abandoned review is and is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly covers when to use it (unfinished review not worth finishing) and when not to (completed reviews, whose verdict is final). This is a genuine exclusion that prevents a wrong call, matching the sibling-routing standard.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_reportA
Export a completed review as a shareable decision memo.
'markdown' (default) gives a memo with the claim, context, assessment and its rule, every item with its rebuttal/mitigation/test/response, and a 'Next actions' checklist (counterarguments that still hold, mitigations to carry out, tests to run, conceded points). 'json' gives the raw verdict report.
Args: review_id: The review session id (e.g. "rev-x9k2"). format: "markdown" (alias "md") or "json".
Returns: dict with review_id, format, assessment and the rendered 'content'.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | markdown | |
| review_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden, and it does well: it discloses the two output shapes, details what the markdown memo contains, and specifies the return dict keys. It omits permission/auth requirements and whether the review must be in a particular state beyond 'completed', but the behavioral profile is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence followed by structured Args/Returns sections. The markdown-format enumeration is somewhat long but each item conveys output content an agent would otherwise have to discover.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no output schema, the description supplies purpose, param semantics, and a Returns section describing the dict. Complete enough to invoke correctly; only state preconditions and permissions are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it documents review_id with a concrete example ('rev-x9k2') and format including the 'md' alias and the markdown default that the schema only shows as a default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Export a completed review as a shareable decision memo.' The 'completed review' qualifier implicitly distinguishes it from start_review/get_verdict, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'completed review' (you export after the review finishes), but there is no explicit when-to-use/when-not or routing to alternatives such as get_verdict or get_review.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_reviewA
Where a review stands — use it to resume after losing track.
Args: review_id: The review session id (e.g. "rev-x9k2").
Returns: dict with status, current phase and step, items per phase, what the phase is still missing, and (if active) the current phase's full instructions.
| Name | Required | Description | Default |
|---|---|---|---|
| review_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden, and it does reasonably well: it discloses that the response contains status, current phase and step, per-phase item counts, what the phase is missing, and the active phase's full instructions. It never states whether a missing/invalid review_id errors or returns empty, nor any permission or rate-limit behavior, which leaves a gap on a bare tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose line is front-loaded and the Args/Returns structure is scannable with no filler. It is slightly verbose in the Returns enumeration, but every listed field is decision-relevant for a resumption flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description correctly spends its budget explaining the return payload, which is the main unknown for this tool. It is nearly complete; only error behavior and any auth expectations are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it names review_id as the review session id and supplies a concrete format example ('rev-x9k2'), which the bare string schema does not provide. With only one required parameter, this is adequate without being exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first line states a specific resource and outcome — the current standing of a review — with a clear retrieval verb implied and a named purpose ('resume after losing track'). It is distinguishable from siblings like get_verdict and list_reviews, though the description never explicitly contrasts itself with them (e.g. 'get_verdict returns only the final decision').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete triggering condition: use it to resume a review when you have lost track of where things stand. That is clear context for invocation. It does not state when NOT to use it or name an alternative for adjacent needs (e.g. list_reviews for enumeration), so it stops short of the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_verdictA
Compile the final report for a completed review.
Only available once every phase is complete — otherwise it raises an error saying what is still missing. The report contains the claim, all items organized by phase (with rebuttals/mitigations/tests/responses attached to their targets), an aggregate risk score, and a deterministic assessment: "claim survives scrutiny", "claim needs revision", or "claim refuted" (rules documented in assessment_reason). Gauntlet reviews also include one sub-verdict per lens under 'lenses'.
Args: review_id: The review session id (e.g. "rev-x9k2").
Returns: dict — the compiled verdict report.
| Name | Required | Description | Default |
|---|---|---|---|
| review_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the blocking precondition, the error behavior, and the exact report contents (claims, phases, attached rebuttals/tests, aggregate risk score, deterministic assessment plus gauntlet lenses). It stops short of stating whether the call mutates state or is purely read-only, but the 'compile report' framing implies a read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: the action and the precondition come first, then report contents, with Args/Returns separated. The multi-line detail about report structure is dense but each element describes a concrete part of the output, so it largely earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by enumerating the returned report's structure and possible assessment strings. Combined with the precondition and error disclosure, an agent has what it needs to invoke and interpret the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it explains review_id as the review session id and supplies a concrete example format ("rev-x9k2"). That is useful beyond the bare string type, though it is a single simple parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Compile) and resource (final report/verdict) and scopes it to a completed review, which is meaningfully more specific than the bare name 'get_verdict'. It does not explicitly distinguish itself from the sibling export_report, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear precondition ("Only available once every phase is complete") and describes what happens if that is violated, which tells the agent when this tool is callable. It stops short of routing between siblings like export_report vs get_review.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_reviewsA
List every review session (active, complete, and abandoned).
Returns: dict with 'count', 'reviews' (id, claim snippet, mode, status, current phase, timestamps; newest first), the data directory, and 'skipped' for any file in it that is not a valid review.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load, and it delivers real behavioral detail: all statuses are included, ordering is newest-first, the returned keys are enumerated, and it discloses a 'skipped' bucket for invalid files. It does not mention permissions or side effects, but for a pure read this is good disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded single-sentence purpose followed by a structured Returns block; no filler. Slightly verbose in the return detail but every element earns its place given the absence of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so describing the return shape (count, reviews fields, data directory, skipped) is exactly the right compensation. No annotations or params to cover. It stops short of stating guarantees like pagination or permission requirements, but is complete enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document; baseline is 4. The description correctly implies no filtering or argument is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List every review session') and scopes it with 'active, complete, and abandoned', which implicitly separates it from get_review (single) and the mutating siblings start_review/abandon_review. It does not name a sibling outright, but the enumeration of statuses makes the breadth clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: an agent can infer this is the enumeration entry point versus get_review for a single session, but there is no explicit when-to-use/when-not or alternative routing. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_reviewA
Start a structured stress-test of a claim or decision.
You (the client) do all the thinking; this server enforces the protocol, validates every submission, and refuses to advance until each phase is genuinely complete. Pick the mode that fits the question:
devils_advocate: generate counterarguments (categorized, severity-rated), then rebut the severe ones honestly.
premortem: imagine the decision failed at a chosen horizon, list failure causes (likelihood x impact), then mitigate the high-risk ones.
assumptions: audit what the claim silently relies on (load-bearing? evidence level?), then design cheap tests for the unverified ones.
steelman: build the strongest honest case for the OPPOSING position, then concede or counter each point.
gauntlet: all four lenses on the same claim, one after another (devil's advocate -> assumptions -> premortem -> steelman), with a combined verdict. Use it for decisions that really matter.
Submissions are quality-gated: junk padding, near-duplicate items, restating the claim and copying the item you answer are rejected.
Args: claim: The claim or decision under review, stated plainly (e.g. "We should rewrite our backend in Rust"). mode: One of: devils_advocate, premortem, assumptions, steelman, gauntlet. context: Optional background that matters for the review (constraints, stakes, prior discussion). It is repeated in every phase's instructions.
Returns: dict with the new review_id, the phase list, and the exact instructions for the first phase: what to send, the item format, the minimums and the quality checks enforced.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| claim | Yes | ||
| context | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that the server enforces a protocol and refuses to advance until phases are complete, that submissions are quality-gated against padding, near-duplicates, restating the claim, and copied items, and what the initial response contains. This is rich behavioral context an agent would otherwise have to discover by trial and error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then a scannable mode list, then args and returns. The mode descriptions are long but each earns its place by defining a distinct mode. Slight verbosity in the preamble, but no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the Returns block explains the shape of the response (review_id, phase list, first-phase instructions with format, minimums, and quality checks). Combined with the mode semantics and protocol disclosure, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (bare 'Mode', 'Claim', 'Context' titles with no descriptions), so the description must compensate, and it does: it enumerates the exact valid mode strings with meanings, gives a concrete claim example, and explains context's purpose including that it is repeated in every phase. This fully compensates for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start a structured stress-test of a claim or decision') and then enumerates five named modes, making it unmistakably distinct from siblings like submit, get_verdict, or export_report. An agent knows immediately this is the entry point that creates a review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete mode-selection guidance for each of the five modes, including an explicit escalation rule ('gauntlet ... Use it for decisions that really matter'). It does not spell out when to choose start_review over other lifecycle siblings, but the creation-vs-follow-up distinction is strongly implied by the description and Returns block.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submitA
Submit items for the current phase of a review.
Each item must match the item_format from the latest instructions. Validation is atomic: if any item is invalid (wrong fields, or it fails a quality check), nothing is saved and the error lists every problem to fix. When the phase's requirements are met the review advances and the next phase's instructions are returned; otherwise the response lists exactly what is still missing.
Args: review_id: The review session id (e.g. "rev-x9k2"). items: List of dicts matching the current phase's item_format.
Returns: dict with status "in_progress" (plus a 'missing' list), "phase_complete" (plus 'next_phase' instructions), or "complete" (call get_verdict next).
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| review_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses atomic validation ('if any item is invalid... nothing is saved'), that errors enumerate every problem, and the three-state transition behavior. What it omits is anything about authentication, side effects beyond the review state, or how items are persisted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then validation semantics, then Args and Returns sections. Every sentence conveys actionable behavior with no filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description appropriately documents the three return states including the 'missing' list and next_phase instructions. It is nearly complete for this workflow tool; only the concrete item schema (deferred to external instructions) is left undefined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: review_id is documented with a format example and items is described as dicts matching the current phase's item_format. The reference to 'latest instructions' is a deferred contract the agent may not possess, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Submit items') scoped to 'the current phase of a review', which is distinct from siblings like start_review, get_verdict, and get_review. It does not explicitly name the sibling it is not, so it stops short of a 5, but an agent can identify the tool's role without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operative context: items must match the latest instructions' item_format, validation is atomic, and on 'complete' the agent should call get_verdict next. This is explicit routing guidance, though there is no explicit when-not-to-use statement or contrast with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.2.0- First observed
abandon_review - First observed
export_report - First observed
get_review - First observed
get_verdict - First observed
list_reviews - First observed
start_review - First observed
submit
TDQS
Scored across 7 tools
Each tool serves a distinct step in the review workflow: starting, submitting, checking status, getting verdict, exporting, listing, and abandoning. get_verdict and export_report both produce output but are clearly differentiated (compilation vs. shareable export). No overlapping purposes exist.
All names use snake_case and follow an imperative verb pattern. The one exception is 'submit', which lacks a noun, but its purpose is clear from context and the rest are consistently verb_noun.
Seven tools perfectly cover the review lifecycle without redundancy. Each tool is necessary and earns its place, from initiation through final export and cleanup.
The tool set provides full CRUD-like lifecycle coverage for structured reviews: start, submit phase items, track progress, compile verdict, export results, list sessions, and abandon unfinished work. No obvious gaps remain for the stated purpose.
Maintenance
Related MCP Connectors
Stress-test a decision through named thinkers' lenses: assumptions, counter-arguments, receipt.
Adversarial behavioural-bias engine — audits your decisions for cognitive biases via your own AI.
Devil's-advocate QC API for AIs: post a decision, get strongest counter-argument. 0.1 USDT/call
Writes adversarial test suites for AI-built code. Your agent's test engineer.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables structured, iterative reasoning for complex problem-solving with features like confidence tracking, revision mechanisms, and branching support. Provides flexible validation and multiple output formats for systematic analysis and decision-making tasks.134 npm72MIT
- AlicenseAqualityDmaintenanceProvides structured thinking with step-by-step reasoning and steel-manning verification for AI agents, backed by cognitive science research.240 npm1MIT
- AlicenseAqualityDmaintenanceEnables users to stress-test decisions and plans with structured contrarian analysis, surfacing blind spots, hidden assumptions, and failure scenarios through multiple modes such as counter, probe, redteam, and premortem.139 npm2Apache 2.0
- AlicenseNot gradedqualityDmaintenanceProvides five rigorous reasoning protocols (debate, red team, audit_argument, threat_model, check_study) that run on the AI you're already using, requiring no extra API keys or costs.MIT