mcp-devils-advocate
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| DEVILS_ADVOCATE_DIR | No | Directory for storing review state (default: ~/.mcp-devils-advocate/) | ~/.mcp-devils-advocate/ |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| start_reviewA | Start a structured stress-test of a claim or decision. You (the client) do all the thinking; this server enforces the protocol, validates every submission, and refuses to advance until each phase is genuinely complete. Pick the mode that fits the question:
Submissions are quality-gated: junk padding, near-duplicate items, restating the claim and copying the item you answer are rejected. Args: claim: The claim or decision under review, stated plainly (e.g. "We should rewrite our backend in Rust"). mode: One of: devils_advocate, premortem, assumptions, steelman, gauntlet. context: Optional background that matters for the review (constraints, stakes, prior discussion). It is repeated in every phase's instructions. Returns: dict with the new review_id, the phase list, and the exact instructions for the first phase: what to send, the item format, the minimums and the quality checks enforced. |
| submitA | Submit items for the current phase of a review. Each item must match the item_format from the latest instructions. Validation is atomic: if any item is invalid (wrong fields, or it fails a quality check), nothing is saved and the error lists every problem to fix. When the phase's requirements are met the review advances and the next phase's instructions are returned; otherwise the response lists exactly what is still missing. Args: review_id: The review session id (e.g. "rev-x9k2"). items: List of dicts matching the current phase's item_format. Returns: dict with status "in_progress" (plus a 'missing' list), "phase_complete" (plus 'next_phase' instructions), or "complete" (call get_verdict next). |
| get_verdictA | Compile the final report for a completed review. Only available once every phase is complete — otherwise it raises an error saying what is still missing. The report contains the claim, all items organized by phase (with rebuttals/mitigations/tests/responses attached to their targets), an aggregate risk score, and a deterministic assessment: "claim survives scrutiny", "claim needs revision", or "claim refuted" (rules documented in assessment_reason). Gauntlet reviews also include one sub-verdict per lens under 'lenses'. Args: review_id: The review session id (e.g. "rev-x9k2"). Returns: dict — the compiled verdict report. |
| export_reportA | Export a completed review as a shareable decision memo. 'markdown' (default) gives a memo with the claim, context, assessment and its rule, every item with its rebuttal/mitigation/test/response, and a 'Next actions' checklist (counterarguments that still hold, mitigations to carry out, tests to run, conceded points). 'json' gives the raw verdict report. Args: review_id: The review session id (e.g. "rev-x9k2"). format: "markdown" (alias "md") or "json". Returns: dict with review_id, format, assessment and the rendered 'content'. |
| get_reviewA | Where a review stands — use it to resume after losing track. Args: review_id: The review session id (e.g. "rev-x9k2"). Returns: dict with status, current phase and step, items per phase, what the phase is still missing, and (if active) the current phase's full instructions. |
| list_reviewsA | List every review session (active, complete, and abandoned). Returns: dict with 'count', 'reviews' (id, claim snippet, mode, status, current phase, timestamps; newest first), the data directory, and 'skipped' for any file in it that is not a valid review. |
| abandon_reviewA | Abandon an unfinished review that is no longer worth finishing. Abandoned reviews stay listed for the record but accept no further submissions and produce no verdict. Completed reviews cannot be abandoned — their verdict is final. Args: review_id: The review session id (e.g. "rev-x9k2"). reason: Why the review is being abandoned (required, non-empty). Returns: dict confirming the abandonment. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| stress_test | Stress-test a claim with one of the protocols: devils_advocate, premortem, assumptions, steelman or gauntlet. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| reviews | Every review (id, claim, mode, status, phase), newest first, as JSON. |
TDQS
Scored across 7 tools
Each tool serves a distinct step in the review workflow: starting, submitting, checking status, getting verdict, exporting, listing, and abandoning. get_verdict and export_report both produce output but are clearly differentiated (compilation vs. shareable export). No overlapping purposes exist.
All names use snake_case and follow an imperative verb pattern. The one exception is 'submit', which lacks a noun, but its purpose is clear from context and the rest are consistently verb_noun.
Seven tools perfectly cover the review lifecycle without redundancy. Each tool is necessary and earns its place, from initiation through final export and cleanup.
The tool set provides full CRUD-like lifecycle coverage for structured reviews: start, submit phase items, track progress, compile verdict, export results, list sessions, and abandon unfinished work. No obvious gaps remain for the stated purpose.