thinking-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@thinking-mcphelp me reason through why our API latency spiked, deep mode"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
thinking-mcp
An MCP server that forces the host AI to reason through a structured protocol — without calling a single LLM.
No API key. No GPU. No inference cost. ~1.2k LOC you can audit in one sitting.
A CAPABLE AI != AN AI THAT THINKS WELL
The problem is not IQ — it is discipline and structure.The problem
Claude 4 and GPT-5 are already smart enough. What they do when answering freely is skip steps:
Failure mode | Symptom |
Anchoring | Jumps on the first hypothesis |
Overconfidence |
|
Sunk cost | Never backtracks after being shown wrong |
Confirmation bias | Only searches for supporting evidence |
Symptom solving | Fixes the leaf, not the root |
Single hypothesis | Offers one possibility and stops |
No inversion | Never attacks its own idea |
thinking-mcp does not make the model smarter. It makes the model not skip steps.
Related MCP server: cs-agent-mcp
What it is / isn't
Is: an MCP server that renders a scaffold, validates structure, and formats output. Isn't: an LLM proxy, a model aggregator, a caching layer, a silver bullet, or "an IQ boost".
┌────────────────────────────────────────────────────────────┐
│ HOST AI (Claude / GPT / Agnes in Cursor) — THE BRAIN │
│ • Reasoning · creativity · context understanding │
└───────────────────────┬────────────────────────────────────┘
│ MCP (stdio)
▼
┌────────────────────────────────────────────────────────────┐
│ thinking-mcp — THE COACH │
│ • Scaffold protocol · validate · format │
│ • 0 API calls · 0 API keys │
└────────────────────────────────────────────────────────────┘The three modes
QUICK | STANDARD | DEEP | |
When | Simple questions | Most questions | Hard problems |
Steps | 3 | 5 | 11 |
Scaffold tokens (measured, rendered) | ~170 | ~550 | ~1,140 |
For | Factual, lookup | Debugging, design, review | Architecture, strategy, root cause |
The mode is auto-selected from the question text; the caller can always override with
mode="deep". Every mode names its own breaking point: the scaffold tells the host AI
when to switch modes instead of grinding through the wrong depth.
DEEP walks through: Axioms ([TRUTH]/[ASSUMPTION] labels) → Decompose → Root Cause (5 Whys with a multi-cause guard: single-cause linear gets one 5-link chain, multi-cause gets an issue tree with one why-chain per branch) → Explore → Verify (evidence labelled [MEASURED]/[INFERRED]) → Red Team (inversion + tripwires) → System Map → Bias Check (+ base-rate) → Synthesize (bootstrapping pass for estimates) → Self-Critique → Finalize.
Why the guards, and where they come from
The scaffolds carry several evidence-graded guard rails adapted from thinking-framework-skills (an Apache-2.0 library of 63 frameworks graded S/M/P/V/A/C/X) and other reference repos:
Guard | Why | Source |
5 Whys caveat + issue-tree branching | Five Whys is reliable only for simple single-cause linear failures and misleads beyond them (Card 2017) — graded X in the reference library. The scaffold says so and branches multi-cause problems instead. | thinking-framework-skills (Problem Framing) |
| Separate verified facts from inherited conventions so assumptions actually get challenged. | first-principles-thinking-skill |
| Never present an inference as a measurement. | Evidence vs Inference Sort (thinking-framework-skills) |
Red-team tripwires | Premortem's kill criteria: each attack names the observable signal that it is coming true. | Premortem, S/M tier (thinking-framework-skills) |
"When NOT to use" per mode | Every method names where it misleads — mode-mismatch wastes tokens both ways. | thinking-framework-skills (every skill carries one) |
Bootstrapping pass on estimates | Averaging a first estimate with a re-estimate from flipped assumptions measurably beats a single guess (Herzog & Hertwig 2009). | Dialectical Bootstrapping, M tier (thinking-framework-skills) |
Natural-frequency framing | "3 in 1,000" beats "0.3%" for conditional reasoning (Gigerenzer & Hoffrage 1995) — the calibrator warns on bare small percentages. | Natural-Frequency Bayesian Framing, S tier (thinking-framework-skills) |
Not adopted on purpose: ACH (controlled trials found it raises confidence without accuracy — graded X) and Six Thinking Hats-style role lenses (P tier, but the red-team/inversion step already covers the adversarial move). Chain of Draft and token-budget prompting were skipped: the scaffold is a fixed ~1.1k-token cost, and compressing it trades away the very structure being enforced.
Install
Not published to PyPI yet. Install from source:
git clone https://github.com/vipera-iso/thinking-mcp.git thinking-mcp && cd thinking-mcp
uv venv .venv && uv pip install -e .One dependency: mcp. Python ≥ 3.10.
Once published to PyPI, replace the command/args below with "command": "uvx", "args": ["thinking-mcp"] — no clone needed.
Integrate with coding agents
The server speaks stdio (how every MCP host launches local servers). Point your agent at the binary and it gets the think tools:
claude mcp add thinking -- /path/to/thinking-mcp/.venv/bin/thinking-mcp{
"mcpServers": {
"thinking": {
"command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp"
}
}
}{
"mcp": {
"thinking": {
"type": "local",
"command": ["/path/to/thinking-mcp/.venv/bin/thinking-mcp"],
"enabled": true
}
}
}{
"mcpServers": {
"thinking": {
"command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp",
"args": ["--log-level", "error"]
}
}
}It also serves Streamable HTTP for clients without stdio support:
thinking-mcp --http --port 8000 (endpoint :8000/mcp).
Drop .cursorrules into your project so the host AI knows when it must
call think.
Tool surface
Tool | What it does |
| Opens a session and returns the protocol scaffold (not the answer) |
| Takes the JSON response → validates → formats |
| Inspects session state for debugging |
| Lists modes and their scaffold cost |
| Builds independent verification prompts for a conclusion |
| Takes the refutation / samples back → verdict + agreement |
The /think <problem> prompt drives the flow: think → reason → think_submit.
Retry flow
think_submit -> validate -- OK --> format + return result
|
FAIL --> { status: "invalid", errors: [...], retry_count: 1 }
| host AI fixes it -> submits again
+-- after 2 failures --> "best_effort" with validation_errorsAfter the retries run out the server does not pretend the result is valid: the status
becomes best_effort and the structural errors are returned verbatim.
Independent verification (think_verify)
Structure enforcement is not the same as being right. think_verify adds two passes that
make the model do independent work:
Pass | What it does |
A. Refutation | A fresh-context pass that never sees the reasoning chain gets only the conclusion and its evidence, and is told to break it. The fresh context is the mechanism — anchoring is the failure mode targeted. |
B. Self-consistency | N answers to the raw problem, majority-voted, so you can see whether the answer is even reproducible. |
Why the host runs the prompts, not the server
MCP was supposed to let the server initiate these calls itself through sampling
(session.create_message). That does not work here: on mcp 2.3.0 both the stdio and the
Streamable HTTP transports raise
NoBackChannelError: Cannot send 'sampling/createMessage':
this transport context has no back-channel for server-initiated requests.— neither transport carries a channel for server-initiated requests, and sampling is deprecated as of 2026-07-28 (SEP-2577). So the server builds the prompts and parses the results, and the host executes them. No API key is involved either way.
The honest weakness
The server cannot enforce that the host ran the refutation in a fresh context. A host
that answers inline from its existing chain gets a weaker version of pass A. The prompt
says so, and the returned output says so too (cannot_enforce).
Neither pass is a judge: A is a model opinion that can reject a correct conclusion, and B
measures reproducibility, not correctness — a confidently wrong answer repeated N times is
still wrong. Both caveats are attached to every result, and when nothing usable comes back
the status is verification_failed with the conclusion labelled UNVERIFIED, never a
pretend-success.
Validation measures structure, not quality
The host AI can still write a fake counter. The validator cannot tell — and deliberately does not try, because there is no LLM judge here.
Validation does not measure quality. It measures structure. The host AI can still write a fake counter. But without structural enforcement, it skips the step entirely.
DEEP, for example, requires: ≥3 assumptions each labelled [TRUTH]/[ASSUMPTION], a
root_cause object with single_cause + branches (1 branch × 5 links if linear,
≥2 branches × 3 links if an issue tree), ≥2 red-team attacks each with a
what_would_change_my_mind and a tripwire, evidence labelled
[MEASURED]/[INFERRED], ≥3 bias checks, and ≥1 piece of disconfirming evidence.
Calibration only warns, never rewrites: confidence: 0.95 with thin evidence returns
a warning and leaves 0.95 untouched — rewriting it would invent a number.
Measured results
agnes-3.0-flash over its OpenAI-compatible API, all 26 tasks in eval/tasks.jsonl,
concurrency 2. Raw output is committed in eval/results/.
.venv/bin/python eval/compliance.py --concurrency 2Two complete runs are committed: the baseline (20261006-123628, 10-step deep scaffold) and the guard-rails run (20261006-143122, 11-step deep with all the guards above). The tables show both where they differ; where only one number is given it is the guard-rails run.
Compliance — does the host AI follow the protocol?
Metric | Baseline | Guard rails |
Valid on first submission | 18 / 26 (69%) | 13 / 26 (50%) |
Valid after retry (final) | 25 / 26 (96%) | 25 / 26 (96%) |
Ran out of retries | 1 / 26 (4%) | 1 / 26 (4%) |
Non-JSON replies | 5, across 4 tasks | 4 |
By mode:
mode | n | first-try (base → rails) | final ok (base → rails) |
quick | 3 | 100% → 100% | 100% → 100% |
standard | 16 | 69% → 50% | 100% → 94% |
deep | 7 | 57% → 29% | 86% → 100% |
Correctness — on the 8 tasks with an objective answer
Baseline | Guard rails | |
Protocol correct | 8 / 8 | 7 / 8 |
Direct correct | 8 / 8 | 8 / 8 |
The regression is math-widgets, and the raw log shows exactly why: the model submitted
confidence: 0.9+ backed by one piece of evidence, was rejected twice by the pre-existing
">0.9 with thin evidence" rule, and exhausted its retry budget. The failure is a retry
mechanics problem, not a guard misfiring — the rule worked as designed, the model just
repeated the same mistake instead of adding evidence or lowering confidence. The flipped
task also swapped: baseline exhausted on debug-useeffect-leak (deep), guard rails run
recovers it. n=8 is far too small to read either flip as signal.
Cost
Metric | Baseline | Guard rails |
Protocol tokens | 90,003 | 113,879 |
Protocol host calls | 37 | 41 |
Direct tokens | 19,058 | 18,120 |
Ratio | 4.72× | 6.28× |
Average latency | 50.1 s / task | 61.2 s / task |
Average scaffold | 539 tokens | 715 tokens |
What this shows — and what it does not
Shows: adding structure costs compliance. The 11-step deep scaffold with mandatory labels and the issue-tree shape cut first-try validity from 69% to 50% overall, and from 57% to 29% in deep — the model needs more retries to produce the richer shape. The retry loop still lands at 96% final validity, and deep final validity went 86% → 100%.
Shows: the new calibrator warnings fire on real output — e.g.
arch-monolith-splitwas told to restate "1%" as "10 in 1,000". They are advisory, so they cost tokens but cannot fail a submission.Does not show a correctness change. 7/8 vs 8/8 on 8 regex-scored tasks is one retry mechanics failure, not evidence either way. The honest headline is unchanged: this measurement supports structure enforcement, not accuracy improvement, and the extra structure is paid for in tokens (6.28×) and latency (61 s/task).
Known limitations
❌ Does not make the AI smarter — only stops it skipping steps
❌ Cannot verify semantic quality (a genuine counter vs a fake one)
❌ Cannot force 100% compliance — a model that ignores protocols gets nothing from this
❌ Guard rails have a compliance price: first-try validity fell 69% → 50% (deep 57% → 29%) because the richer shape is harder to produce on the first attempt
❌ Deep mode does not suit every question (scaffold token cost)
⚠️ The eval needs a real host AI to measure compliance; the server itself produces no answer
⚠️
think_verifycannot force a fresh context, and its two passes are opinions, not a judge
Tests
uv venv .venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest -q # 166 tests, including e2e over real stdioThe end-to-end tests spawn the real server over stdio and drive it with mcp.Client —
the same path Cursor / Claude Desktop take.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
GodPrompt MCP + Agent Skill for coding agents with TDD, debugging, verification, and task routing.
Persistent memory, hybrid search and a goal graph for AI agents, over stdio or remote HTTP.
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
Lints + auto-fixes how AI coding agents discover any new product. 24 rules, 6 tools, score 0-100.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAutonomous spec-to-product coding-agent CLI. Its MCP server exposes 34 tools over stdio: project state and task-queue ops, memory retrieve/store, code search, quality and verification reports, repo hotspots/co-changes, and structured findings/learnings.13,524 npm1,088Business Source 1.1
- AlicenseAqualityBmaintenanceA local stdio MCP service that unifies coding agents like Codex and Claude into cs_agent_* tools, enabling the root agent to create, invoke, and manage child agents with recursive delegation.1438 npm3MIT
- AlicenseBqualityAmaintenanceLocal-first Agent OS that wraps Claude Code, Codex CLI, and other coding agents in a replayable Seed → Ledger → Runtime contract, driven by an interview → seed → execute → evaluate → evolve workflow loop.3418,092 PyPI6,193MIT
- AlicenseAqualityAmaintenanceZero-dependency token optimizer, test-failure triage gate, and System 1.5 semantic guardrail for AI coding agents. Intercepts compiler errors, missing dependencies, and doom loops in < 500µs.6309 PyPI11MIT