Skip to main content
Glama
vipera-iso
by vipera-iso
README.md
# thinking-mcp

**An MCP server that forces the host AI to reason through a structured protocol — without calling a single LLM.**

No API key. No GPU. No inference cost. ~1.2k LOC you can audit in one sitting.

```
A CAPABLE AI != AN AI THAT THINKS WELL
The problem is not IQ — it is discipline and structure.
```

---

## The problem

Claude 4 and GPT-5 are already smart enough. What they do when answering freely is
skip steps:

| Failure mode | Symptom |
|---|---|
| Anchoring | Jumps on the first hypothesis |
| Overconfidence | `confidence: 0.9` with one piece of evidence |
| Sunk cost | Never backtracks after being shown wrong |
| Confirmation bias | Only searches for supporting evidence |
| Symptom solving | Fixes the leaf, not the root |
| Single hypothesis | Offers one possibility and stops |
| No inversion | Never attacks its own idea |

thinking-mcp does not make the model smarter. It makes the model **not skip steps**.

## What it is / isn't

**Is:** an MCP server that renders a scaffold, validates structure, and formats output.
**Isn't:** an LLM proxy, a model aggregator, a caching layer, a silver bullet, or "an IQ boost".

```
┌────────────────────────────────────────────────────────────┐
│  HOST AI (Claude / GPT / Agnes in Cursor) — THE BRAIN      │
│  • Reasoning · creativity · context understanding          │
└───────────────────────┬────────────────────────────────────┘
                        │ MCP (stdio)
                        ▼
┌────────────────────────────────────────────────────────────┐
│  thinking-mcp — THE COACH                                  │
│  • Scaffold protocol · validate · format                   │
│  • 0 API calls · 0 API keys                                │
└────────────────────────────────────────────────────────────┘
```

## The three modes

| | **QUICK** | **STANDARD** | **DEEP** |
|---|---|---|---|
| When | Simple questions | Most questions | Hard problems |
| Steps | 3 | 5 | 11 |
| Scaffold tokens (measured, rendered) | ~170 | ~550 | ~1,140 |
| For | Factual, lookup | Debugging, design, review | Architecture, strategy, root cause |

The mode is auto-selected from the question text; the caller can always override with
`mode="deep"`. Every mode names its own breaking point: the scaffold tells the host AI
when to switch modes instead of grinding through the wrong depth.

**DEEP** walks through: Axioms ([TRUTH]/[ASSUMPTION] labels) → Decompose → Root Cause
(5 Whys **with a multi-cause guard**: single-cause linear gets one 5-link chain,
multi-cause gets an issue tree with one why-chain per branch) → Explore → Verify
(evidence labelled [MEASURED]/[INFERRED]) → **Red Team (inversion + tripwires)** →
System Map → Bias Check (+ base-rate) → Synthesize (**bootstrapping pass** for
estimates) → Self-Critique → Finalize.

### Why the guards, and where they come from

The scaffolds carry several evidence-graded guard rails adapted from
[thinking-framework-skills](https://github.com/product-on-purpose/thinking-framework-skills)
(an Apache-2.0 library of 63 frameworks graded S/M/P/V/A/C/X) and other reference repos:

| Guard | Why | Source |
|---|---|---|
| 5 Whys caveat + issue-tree branching | Five Whys is reliable only for simple single-cause linear failures and misleads beyond them (Card 2017) — graded **X** in the reference library. The scaffold says so and branches multi-cause problems instead. | thinking-framework-skills (Problem Framing) |
| `[TRUTH]`/`[ASSUMPTION]` labels | Separate verified facts from inherited conventions so assumptions actually get challenged. | first-principles-thinking-skill |
| `[MEASURED]`/`[INFERRED]` evidence labels | Never present an inference as a measurement. | Evidence vs Inference Sort (thinking-framework-skills) |
| Red-team tripwires | Premortem's kill criteria: each attack names the observable signal that it is coming true. | Premortem, S/M tier (thinking-framework-skills) |
| "When NOT to use" per mode | Every method names where it misleads — mode-mismatch wastes tokens both ways. | thinking-framework-skills (every skill carries one) |
| Bootstrapping pass on estimates | Averaging a first estimate with a re-estimate from flipped assumptions measurably beats a single guess (Herzog & Hertwig 2009). | Dialectical Bootstrapping, M tier (thinking-framework-skills) |
| Natural-frequency framing | "3 in 1,000" beats "0.3%" for conditional reasoning (Gigerenzer & Hoffrage 1995) — the calibrator warns on bare small percentages. | Natural-Frequency Bayesian Framing, S tier (thinking-framework-skills) |

Not adopted on purpose: **ACH** (controlled trials found it raises confidence without
accuracy — graded X) and **Six Thinking Hats**-style role lenses (P tier, but the
red-team/inversion step already covers the adversarial move). Chain of Draft and
token-budget prompting were skipped: the scaffold is a fixed ~1.1k-token cost, and
compressing it trades away the very structure being enforced.

## Install

Not published to PyPI yet. Install from source:

```bash
git clone https://github.com/vipera-iso/thinking-mcp.git thinking-mcp && cd thinking-mcp
uv venv .venv && uv pip install -e .
```

One dependency: `mcp`. Python ≥ 3.10.

Once published to PyPI, replace the `command`/`args` below with `"command": "uvx", "args": ["thinking-mcp"]` — no clone needed.

## Integrate with coding agents

The server speaks stdio (how every MCP host launches local servers). Point your agent at the binary and it gets the `think` tools:

<details>
<summary><b>Claude Code</b></summary>

```bash
claude mcp add thinking -- /path/to/thinking-mcp/.venv/bin/thinking-mcp
```

</details>

<details>
<summary><b>Cursor / Claude Desktop (mcpServers JSON)</b></summary>

```json
{
  "mcpServers": {
    "thinking": {
      "command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp"
    }
  }
}
```

</details>

<details>
<summary><b>OpenCode (~/.config/opencode/opencode.json)</b></summary>

```json
{
  "mcp": {
    "thinking": {
      "type": "local",
      "command": ["/path/to/thinking-mcp/.venv/bin/thinking-mcp"],
      "enabled": true
    }
  }
}
```

</details>

<details>
<summary><b>Generic MCP client (.agents/mcp.json and similar)</b></summary>

```json
{
  "mcpServers": {
    "thinking": {
      "command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp",
      "args": ["--log-level", "error"]
    }
  }
}
```

</details>

It also serves Streamable HTTP for clients without stdio support:
`thinking-mcp --http --port 8000` (endpoint `:8000/mcp`).

Drop [`.cursorrules`](.cursorrules) into your project so the host AI knows when it must
call `think`.

## Tool surface

| Tool | What it does |
|---|---|
| `think` | Opens a session and returns the protocol scaffold (**not** the answer) |
| `think_submit` | Takes the JSON response → validates → formats |
| `think_status` | Inspects session state for debugging |
| `think_list_templates` | Lists modes and their scaffold cost |
| `think_verify` | Builds independent verification prompts for a conclusion |
| `think_verify_submit` | Takes the refutation / samples back → verdict + agreement |

The `/think <problem>` prompt drives the flow: `think` → reason → `think_submit`.

### Retry flow

```
think_submit -> validate -- OK --> format + return result
                 |
                 FAIL --> { status: "invalid", errors: [...], retry_count: 1 }
                 |        host AI fixes it -> submits again
                 +-- after 2 failures --> "best_effort" with validation_errors
```

After the retries run out the server does **not** pretend the result is valid: the status
becomes `best_effort` and the structural errors are returned verbatim.

## Independent verification (`think_verify`)

Structure enforcement is not the same as being right. `think_verify` adds two passes that
make the model do **independent work**:

| Pass | What it does |
|---|---|
| **A. Refutation** | A fresh-context pass that never sees the reasoning chain gets only the conclusion and its evidence, and is told to break it. The fresh context is the mechanism — anchoring is the failure mode targeted. |
| **B. Self-consistency** | N answers to the raw problem, majority-voted, so you can see whether the answer is even reproducible. |

### Why the host runs the prompts, not the server

MCP was supposed to let the server initiate these calls itself through **sampling**
(`session.create_message`). That does not work here: on `mcp` 2.3.0 both the stdio and the
Streamable HTTP transports raise

```
NoBackChannelError: Cannot send 'sampling/createMessage':
this transport context has no back-channel for server-initiated requests.
```

— neither transport carries a channel for server-initiated requests, and sampling is
deprecated as of 2026-07-28 (SEP-2577). So the server builds the prompts and parses the
results, and the **host** executes them. No API key is involved either way.

### The honest weakness

The server **cannot enforce** that the host ran the refutation in a fresh context. A host
that answers inline from its existing chain gets a weaker version of pass A. The prompt
says so, and the returned output says so too (`cannot_enforce`).

Neither pass is a judge: A is a model opinion that can reject a correct conclusion, and B
measures reproducibility, not correctness — a confidently wrong answer repeated N times is
still wrong. Both caveats are attached to every result, and when nothing usable comes back
the status is `verification_failed` with the conclusion labelled **UNVERIFIED**, never a
pretend-success.

## Validation measures structure, not quality

The host AI can still write a fake counter. The validator cannot tell — and deliberately
does not try, because there is no LLM judge here.

> **Validation does not measure quality. It measures structure.**
> The host AI can still write a fake counter. But without structural enforcement, it skips the step entirely.

DEEP, for example, requires: ≥3 assumptions each labelled `[TRUTH]`/`[ASSUMPTION]`, a
`root_cause` object with `single_cause` + `branches` (1 branch × 5 links if linear,
≥2 branches × 3 links if an issue tree), ≥2 red-team attacks each with a
`what_would_change_my_mind` **and a `tripwire`**, evidence labelled
`[MEASURED]`/`[INFERRED]`, ≥3 bias checks, and ≥1 piece of disconfirming evidence.

Calibration **only warns**, never rewrites: `confidence: 0.95` with thin evidence returns
a warning and leaves 0.95 untouched — rewriting it would invent a number.

## Measured results

`agnes-3.0-flash` over its OpenAI-compatible API, all 26 tasks in `eval/tasks.jsonl`,
concurrency 2. Raw output is committed in [`eval/results/`](eval/results/).

```bash
.venv/bin/python eval/compliance.py --concurrency 2
```

Two complete runs are committed: the **baseline** (20261006-123628, 10-step deep scaffold)
and the **guard-rails run** (20261006-143122, 11-step deep with all the guards above). The
tables show both where they differ; where only one number is given it is the guard-rails
run.

### Compliance — does the host AI follow the protocol?

| Metric | Baseline | Guard rails |
|---|---|---|
| Valid on first submission | **18 / 26 (69%)** | **13 / 26 (50%)** |
| Valid after retry (final) | 25 / 26 (96%) | 25 / 26 (96%) |
| Ran out of retries | 1 / 26 (4%) | 1 / 26 (4%) |
| Non-JSON replies | 5, across 4 tasks | 4 |

By mode:

| mode | n | first-try (base → rails) | final ok (base → rails) |
|---|---|---|---|
| quick | 3 | 100% → 100% | 100% → 100% |
| standard | 16 | 69% → 50% | 100% → 94% |
| deep | 7 | **57% → 29%** | **86% → 100%** |

### Correctness — on the 8 tasks with an objective answer

| | Baseline | Guard rails |
|---|---|---|
| Protocol correct | 8 / 8 | **7 / 8** |
| Direct correct | 8 / 8 | 8 / 8 |

The regression is `math-widgets`, and the raw log shows exactly why: the model submitted
`confidence: 0.9+` backed by one piece of evidence, was rejected twice by the **pre-existing**
">0.9 with thin evidence" rule, and exhausted its retry budget. The failure is a retry
mechanics problem, not a guard misfiring — the rule worked as designed, the model just
repeated the same mistake instead of adding evidence or lowering confidence. The flipped
task also swapped: baseline exhausted on `debug-useeffect-leak` (deep), guard rails run
recovers it. n=8 is far too small to read either flip as signal.

### Cost

| Metric | Baseline | Guard rails |
|---|---|---|
| Protocol tokens | 90,003 | 113,879 |
| Protocol host calls | 37 | 41 |
| Direct tokens | 19,058 | 18,120 |
| Ratio | **4.72×** | **6.28×** |
| Average latency | 50.1 s / task | 61.2 s / task |
| Average scaffold | 539 tokens | 715 tokens |

### What this shows — and what it does not

- **Shows:** adding structure costs compliance. The 11-step deep scaffold with mandatory
  labels and the issue-tree shape cut first-try validity from 69% to 50% overall, and from
  57% to 29% in deep — the model needs more retries to produce the richer shape. The retry
  loop still lands at 96% final validity, and **deep final validity went 86% → 100%**.
- **Shows:** the new calibrator warnings fire on real output — e.g. `arch-monolith-split`
  was told to restate "1%" as "10 in 1,000". They are advisory, so they cost tokens but
  cannot fail a submission.
- **Does not show** a correctness change. 7/8 vs 8/8 on 8 regex-scored tasks is one retry
  mechanics failure, not evidence either way. The honest headline is unchanged: this
  measurement supports *structure enforcement*, not accuracy improvement, and the extra
  structure is paid for in tokens (6.28×) and latency (61 s/task).

## Known limitations

- ❌ Does not make the AI smarter — only stops it skipping steps
- ❌ Cannot verify semantic quality (a genuine counter vs a fake one)
- ❌ Cannot force 100% compliance — a model that ignores protocols gets nothing from this
- ❌ Guard rails have a compliance price: first-try validity fell 69% → 50% (deep 57% → 29%)
  because the richer shape is harder to produce on the first attempt
- ❌ Deep mode does not suit every question (scaffold token cost)
- ⚠️ The eval needs a real host AI to measure compliance; the server itself produces no answer
- ⚠️ `think_verify` cannot force a fresh context, and its two passes are opinions, not a judge

## Tests

```bash
uv venv .venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest -q          # 166 tests, including e2e over real stdio
```

The end-to-end tests spawn the real server over stdio and drive it with `mcp.Client` —
the same path Cursor / Claude Desktop take.

## License

MIT