thinking-mcp
by vipera-iso
README.md
# thinking-mcp
**An MCP server that forces the host AI to reason through a structured protocol — without calling a single LLM.**
No API key. No GPU. No inference cost. ~1.2k LOC you can audit in one sitting.
```
A CAPABLE AI != AN AI THAT THINKS WELL
The problem is not IQ — it is discipline and structure.
```
---
## The problem
Claude 4 and GPT-5 are already smart enough. What they do when answering freely is
skip steps:
| Failure mode | Symptom |
|---|---|
| Anchoring | Jumps on the first hypothesis |
| Overconfidence | `confidence: 0.9` with one piece of evidence |
| Sunk cost | Never backtracks after being shown wrong |
| Confirmation bias | Only searches for supporting evidence |
| Symptom solving | Fixes the leaf, not the root |
| Single hypothesis | Offers one possibility and stops |
| No inversion | Never attacks its own idea |
thinking-mcp does not make the model smarter. It makes the model **not skip steps**.
## What it is / isn't
**Is:** an MCP server that renders a scaffold, validates structure, and formats output.
**Isn't:** an LLM proxy, a model aggregator, a caching layer, a silver bullet, or "an IQ boost".
```
┌────────────────────────────────────────────────────────────┐
│ HOST AI (Claude / GPT / Agnes in Cursor) — THE BRAIN │
│ • Reasoning · creativity · context understanding │
└───────────────────────┬────────────────────────────────────┘
│ MCP (stdio)
▼
┌────────────────────────────────────────────────────────────┐
│ thinking-mcp — THE COACH │
│ • Scaffold protocol · validate · format │
│ • 0 API calls · 0 API keys │
└────────────────────────────────────────────────────────────┘
```
## The three modes
| | **QUICK** | **STANDARD** | **DEEP** |
|---|---|---|---|
| When | Simple questions | Most questions | Hard problems |
| Steps | 3 | 5 | 11 |
| Scaffold tokens (measured, rendered) | ~170 | ~550 | ~1,140 |
| For | Factual, lookup | Debugging, design, review | Architecture, strategy, root cause |
The mode is auto-selected from the question text; the caller can always override with
`mode="deep"`. Every mode names its own breaking point: the scaffold tells the host AI
when to switch modes instead of grinding through the wrong depth.
**DEEP** walks through: Axioms ([TRUTH]/[ASSUMPTION] labels) → Decompose → Root Cause
(5 Whys **with a multi-cause guard**: single-cause linear gets one 5-link chain,
multi-cause gets an issue tree with one why-chain per branch) → Explore → Verify
(evidence labelled [MEASURED]/[INFERRED]) → **Red Team (inversion + tripwires)** →
System Map → Bias Check (+ base-rate) → Synthesize (**bootstrapping pass** for
estimates) → Self-Critique → Finalize.
### Why the guards, and where they come from
The scaffolds carry several evidence-graded guard rails adapted from
[thinking-framework-skills](https://github.com/product-on-purpose/thinking-framework-skills)
(an Apache-2.0 library of 63 frameworks graded S/M/P/V/A/C/X) and other reference repos:
| Guard | Why | Source |
|---|---|---|
| 5 Whys caveat + issue-tree branching | Five Whys is reliable only for simple single-cause linear failures and misleads beyond them (Card 2017) — graded **X** in the reference library. The scaffold says so and branches multi-cause problems instead. | thinking-framework-skills (Problem Framing) |
| `[TRUTH]`/`[ASSUMPTION]` labels | Separate verified facts from inherited conventions so assumptions actually get challenged. | first-principles-thinking-skill |
| `[MEASURED]`/`[INFERRED]` evidence labels | Never present an inference as a measurement. | Evidence vs Inference Sort (thinking-framework-skills) |
| Red-team tripwires | Premortem's kill criteria: each attack names the observable signal that it is coming true. | Premortem, S/M tier (thinking-framework-skills) |
| "When NOT to use" per mode | Every method names where it misleads — mode-mismatch wastes tokens both ways. | thinking-framework-skills (every skill carries one) |
| Bootstrapping pass on estimates | Averaging a first estimate with a re-estimate from flipped assumptions measurably beats a single guess (Herzog & Hertwig 2009). | Dialectical Bootstrapping, M tier (thinking-framework-skills) |
| Natural-frequency framing | "3 in 1,000" beats "0.3%" for conditional reasoning (Gigerenzer & Hoffrage 1995) — the calibrator warns on bare small percentages. | Natural-Frequency Bayesian Framing, S tier (thinking-framework-skills) |
Not adopted on purpose: **ACH** (controlled trials found it raises confidence without
accuracy — graded X) and **Six Thinking Hats**-style role lenses (P tier, but the
red-team/inversion step already covers the adversarial move). Chain of Draft and
token-budget prompting were skipped: the scaffold is a fixed ~1.1k-token cost, and
compressing it trades away the very structure being enforced.
## Install
Not published to PyPI yet. Install from source:
```bash
git clone https://github.com/vipera-iso/thinking-mcp.git thinking-mcp && cd thinking-mcp
uv venv .venv && uv pip install -e .
```
One dependency: `mcp`. Python ≥ 3.10.
Once published to PyPI, replace the `command`/`args` below with `"command": "uvx", "args": ["thinking-mcp"]` — no clone needed.
## Integrate with coding agents
The server speaks stdio (how every MCP host launches local servers). Point your agent at the binary and it gets the `think` tools:
<details>
<summary><b>Claude Code</b></summary>
```bash
claude mcp add thinking -- /path/to/thinking-mcp/.venv/bin/thinking-mcp
```
</details>
<details>
<summary><b>Cursor / Claude Desktop (mcpServers JSON)</b></summary>
```json
{
"mcpServers": {
"thinking": {
"command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp"
}
}
}
```
</details>
<details>
<summary><b>OpenCode (~/.config/opencode/opencode.json)</b></summary>
```json
{
"mcp": {
"thinking": {
"type": "local",
"command": ["/path/to/thinking-mcp/.venv/bin/thinking-mcp"],
"enabled": true
}
}
}
```
</details>
<details>
<summary><b>Generic MCP client (.agents/mcp.json and similar)</b></summary>
```json
{
"mcpServers": {
"thinking": {
"command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp",
"args": ["--log-level", "error"]
}
}
}
```
</details>
It also serves Streamable HTTP for clients without stdio support:
`thinking-mcp --http --port 8000` (endpoint `:8000/mcp`).
Drop [`.cursorrules`](.cursorrules) into your project so the host AI knows when it must
call `think`.
## Tool surface
| Tool | What it does |
|---|---|
| `think` | Opens a session and returns the protocol scaffold (**not** the answer) |
| `think_submit` | Takes the JSON response → validates → formats |
| `think_status` | Inspects session state for debugging |
| `think_list_templates` | Lists modes and their scaffold cost |
| `think_verify` | Builds independent verification prompts for a conclusion |
| `think_verify_submit` | Takes the refutation / samples back → verdict + agreement |
The `/think <problem>` prompt drives the flow: `think` → reason → `think_submit`.
### Retry flow
```
think_submit -> validate -- OK --> format + return result
|
FAIL --> { status: "invalid", errors: [...], retry_count: 1 }
| host AI fixes it -> submits again
+-- after 2 failures --> "best_effort" with validation_errors
```
After the retries run out the server does **not** pretend the result is valid: the status
becomes `best_effort` and the structural errors are returned verbatim.
## Independent verification (`think_verify`)
Structure enforcement is not the same as being right. `think_verify` adds two passes that
make the model do **independent work**:
| Pass | What it does |
|---|---|
| **A. Refutation** | A fresh-context pass that never sees the reasoning chain gets only the conclusion and its evidence, and is told to break it. The fresh context is the mechanism — anchoring is the failure mode targeted. |
| **B. Self-consistency** | N answers to the raw problem, majority-voted, so you can see whether the answer is even reproducible. |
### Why the host runs the prompts, not the server
MCP was supposed to let the server initiate these calls itself through **sampling**
(`session.create_message`). That does not work here: on `mcp` 2.3.0 both the stdio and the
Streamable HTTP transports raise
```
NoBackChannelError: Cannot send 'sampling/createMessage':
this transport context has no back-channel for server-initiated requests.
```
— neither transport carries a channel for server-initiated requests, and sampling is
deprecated as of 2026-07-28 (SEP-2577). So the server builds the prompts and parses the
results, and the **host** executes them. No API key is involved either way.
### The honest weakness
The server **cannot enforce** that the host ran the refutation in a fresh context. A host
that answers inline from its existing chain gets a weaker version of pass A. The prompt
says so, and the returned output says so too (`cannot_enforce`).
Neither pass is a judge: A is a model opinion that can reject a correct conclusion, and B
measures reproducibility, not correctness — a confidently wrong answer repeated N times is
still wrong. Both caveats are attached to every result, and when nothing usable comes back
the status is `verification_failed` with the conclusion labelled **UNVERIFIED**, never a
pretend-success.
## Validation measures structure, not quality
The host AI can still write a fake counter. The validator cannot tell — and deliberately
does not try, because there is no LLM judge here.
> **Validation does not measure quality. It measures structure.**
> The host AI can still write a fake counter. But without structural enforcement, it skips the step entirely.
DEEP, for example, requires: ≥3 assumptions each labelled `[TRUTH]`/`[ASSUMPTION]`, a
`root_cause` object with `single_cause` + `branches` (1 branch × 5 links if linear,
≥2 branches × 3 links if an issue tree), ≥2 red-team attacks each with a
`what_would_change_my_mind` **and a `tripwire`**, evidence labelled
`[MEASURED]`/`[INFERRED]`, ≥3 bias checks, and ≥1 piece of disconfirming evidence.
Calibration **only warns**, never rewrites: `confidence: 0.95` with thin evidence returns
a warning and leaves 0.95 untouched — rewriting it would invent a number.
## Measured results
`agnes-3.0-flash` over its OpenAI-compatible API, all 26 tasks in `eval/tasks.jsonl`,
concurrency 2. Raw output is committed in [`eval/results/`](eval/results/).
```bash
.venv/bin/python eval/compliance.py --concurrency 2
```
Two complete runs are committed: the **baseline** (20261006-123628, 10-step deep scaffold)
and the **guard-rails run** (20261006-143122, 11-step deep with all the guards above). The
tables show both where they differ; where only one number is given it is the guard-rails
run.
### Compliance — does the host AI follow the protocol?
| Metric | Baseline | Guard rails |
|---|---|---|
| Valid on first submission | **18 / 26 (69%)** | **13 / 26 (50%)** |
| Valid after retry (final) | 25 / 26 (96%) | 25 / 26 (96%) |
| Ran out of retries | 1 / 26 (4%) | 1 / 26 (4%) |
| Non-JSON replies | 5, across 4 tasks | 4 |
By mode:
| mode | n | first-try (base → rails) | final ok (base → rails) |
|---|---|---|---|
| quick | 3 | 100% → 100% | 100% → 100% |
| standard | 16 | 69% → 50% | 100% → 94% |
| deep | 7 | **57% → 29%** | **86% → 100%** |
### Correctness — on the 8 tasks with an objective answer
| | Baseline | Guard rails |
|---|---|---|
| Protocol correct | 8 / 8 | **7 / 8** |
| Direct correct | 8 / 8 | 8 / 8 |
The regression is `math-widgets`, and the raw log shows exactly why: the model submitted
`confidence: 0.9+` backed by one piece of evidence, was rejected twice by the **pre-existing**
">0.9 with thin evidence" rule, and exhausted its retry budget. The failure is a retry
mechanics problem, not a guard misfiring — the rule worked as designed, the model just
repeated the same mistake instead of adding evidence or lowering confidence. The flipped
task also swapped: baseline exhausted on `debug-useeffect-leak` (deep), guard rails run
recovers it. n=8 is far too small to read either flip as signal.
### Cost
| Metric | Baseline | Guard rails |
|---|---|---|
| Protocol tokens | 90,003 | 113,879 |
| Protocol host calls | 37 | 41 |
| Direct tokens | 19,058 | 18,120 |
| Ratio | **4.72×** | **6.28×** |
| Average latency | 50.1 s / task | 61.2 s / task |
| Average scaffold | 539 tokens | 715 tokens |
### What this shows — and what it does not
- **Shows:** adding structure costs compliance. The 11-step deep scaffold with mandatory
labels and the issue-tree shape cut first-try validity from 69% to 50% overall, and from
57% to 29% in deep — the model needs more retries to produce the richer shape. The retry
loop still lands at 96% final validity, and **deep final validity went 86% → 100%**.
- **Shows:** the new calibrator warnings fire on real output — e.g. `arch-monolith-split`
was told to restate "1%" as "10 in 1,000". They are advisory, so they cost tokens but
cannot fail a submission.
- **Does not show** a correctness change. 7/8 vs 8/8 on 8 regex-scored tasks is one retry
mechanics failure, not evidence either way. The honest headline is unchanged: this
measurement supports *structure enforcement*, not accuracy improvement, and the extra
structure is paid for in tokens (6.28×) and latency (61 s/task).
## Known limitations
- ❌ Does not make the AI smarter — only stops it skipping steps
- ❌ Cannot verify semantic quality (a genuine counter vs a fake one)
- ❌ Cannot force 100% compliance — a model that ignores protocols gets nothing from this
- ❌ Guard rails have a compliance price: first-try validity fell 69% → 50% (deep 57% → 29%)
because the richer shape is harder to produce on the first attempt
- ❌ Deep mode does not suit every question (scaffold token cost)
- ⚠️ The eval needs a real host AI to measure compliance; the server itself produces no answer
- ⚠️ `think_verify` cannot force a fresh context, and its two passes are opinions, not a judge
## Tests
```bash
uv venv .venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest -q # 166 tests, including e2e over real stdio
```
The end-to-end tests spawn the real server over stdio and drive it with `mcp.Client` —
the same path Cursor / Claude Desktop take.
## License
MIT
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues