Skip to main content
Glama
vipera-iso
by vipera-iso

thinking-mcp

An MCP server that forces the host AI to reason through a structured protocol — without calling a single LLM.

No API key. No GPU. No inference cost. ~1.2k LOC you can audit in one sitting.

A CAPABLE AI != AN AI THAT THINKS WELL
The problem is not IQ — it is discipline and structure.

The problem

Claude 4 and GPT-5 are already smart enough. What they do when answering freely is skip steps:

Failure mode

Symptom

Anchoring

Jumps on the first hypothesis

Overconfidence

confidence: 0.9 with one piece of evidence

Sunk cost

Never backtracks after being shown wrong

Confirmation bias

Only searches for supporting evidence

Symptom solving

Fixes the leaf, not the root

Single hypothesis

Offers one possibility and stops

No inversion

Never attacks its own idea

thinking-mcp does not make the model smarter. It makes the model not skip steps.

Related MCP server: cs-agent-mcp

What it is / isn't

Is: an MCP server that renders a scaffold, validates structure, and formats output. Isn't: an LLM proxy, a model aggregator, a caching layer, a silver bullet, or "an IQ boost".

┌────────────────────────────────────────────────────────────┐
│  HOST AI (Claude / GPT / Agnes in Cursor) — THE BRAIN      │
│  • Reasoning · creativity · context understanding          │
└───────────────────────┬────────────────────────────────────┘
                        │ MCP (stdio)
                        ▼
┌────────────────────────────────────────────────────────────┐
│  thinking-mcp — THE COACH                                  │
│  • Scaffold protocol · validate · format                   │
│  • 0 API calls · 0 API keys                                │
└────────────────────────────────────────────────────────────┘

The three modes

QUICK

STANDARD

DEEP

When

Simple questions

Most questions

Hard problems

Steps

3

5

11

Scaffold tokens (measured, rendered)

~170

~550

~1,140

For

Factual, lookup

Debugging, design, review

Architecture, strategy, root cause

The mode is auto-selected from the question text; the caller can always override with mode="deep". Every mode names its own breaking point: the scaffold tells the host AI when to switch modes instead of grinding through the wrong depth.

DEEP walks through: Axioms ([TRUTH]/[ASSUMPTION] labels) → Decompose → Root Cause (5 Whys with a multi-cause guard: single-cause linear gets one 5-link chain, multi-cause gets an issue tree with one why-chain per branch) → Explore → Verify (evidence labelled [MEASURED]/[INFERRED]) → Red Team (inversion + tripwires) → System Map → Bias Check (+ base-rate) → Synthesize (bootstrapping pass for estimates) → Self-Critique → Finalize.

Why the guards, and where they come from

The scaffolds carry several evidence-graded guard rails adapted from thinking-framework-skills (an Apache-2.0 library of 63 frameworks graded S/M/P/V/A/C/X) and other reference repos:

Guard

Why

Source

5 Whys caveat + issue-tree branching

Five Whys is reliable only for simple single-cause linear failures and misleads beyond them (Card 2017) — graded X in the reference library. The scaffold says so and branches multi-cause problems instead.

thinking-framework-skills (Problem Framing)

[TRUTH]/[ASSUMPTION] labels

Separate verified facts from inherited conventions so assumptions actually get challenged.

first-principles-thinking-skill

[MEASURED]/[INFERRED] evidence labels

Never present an inference as a measurement.

Evidence vs Inference Sort (thinking-framework-skills)

Red-team tripwires

Premortem's kill criteria: each attack names the observable signal that it is coming true.

Premortem, S/M tier (thinking-framework-skills)

"When NOT to use" per mode

Every method names where it misleads — mode-mismatch wastes tokens both ways.

thinking-framework-skills (every skill carries one)

Bootstrapping pass on estimates

Averaging a first estimate with a re-estimate from flipped assumptions measurably beats a single guess (Herzog & Hertwig 2009).

Dialectical Bootstrapping, M tier (thinking-framework-skills)

Natural-frequency framing

"3 in 1,000" beats "0.3%" for conditional reasoning (Gigerenzer & Hoffrage 1995) — the calibrator warns on bare small percentages.

Natural-Frequency Bayesian Framing, S tier (thinking-framework-skills)

Not adopted on purpose: ACH (controlled trials found it raises confidence without accuracy — graded X) and Six Thinking Hats-style role lenses (P tier, but the red-team/inversion step already covers the adversarial move). Chain of Draft and token-budget prompting were skipped: the scaffold is a fixed ~1.1k-token cost, and compressing it trades away the very structure being enforced.

Install

Not published to PyPI yet. Install from source:

git clone https://github.com/vipera-iso/thinking-mcp.git thinking-mcp && cd thinking-mcp
uv venv .venv && uv pip install -e .

One dependency: mcp. Python ≥ 3.10.

Once published to PyPI, replace the command/args below with "command": "uvx", "args": ["thinking-mcp"] — no clone needed.

Integrate with coding agents

The server speaks stdio (how every MCP host launches local servers). Point your agent at the binary and it gets the think tools:

claude mcp add thinking -- /path/to/thinking-mcp/.venv/bin/thinking-mcp
{
  "mcpServers": {
    "thinking": {
      "command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp"
    }
  }
}
{
  "mcp": {
    "thinking": {
      "type": "local",
      "command": ["/path/to/thinking-mcp/.venv/bin/thinking-mcp"],
      "enabled": true
    }
  }
}
{
  "mcpServers": {
    "thinking": {
      "command": "/path/to/thinking-mcp/.venv/bin/thinking-mcp",
      "args": ["--log-level", "error"]
    }
  }
}

It also serves Streamable HTTP for clients without stdio support: thinking-mcp --http --port 8000 (endpoint :8000/mcp).

Drop .cursorrules into your project so the host AI knows when it must call think.

Tool surface

Tool

What it does

think

Opens a session and returns the protocol scaffold (not the answer)

think_submit

Takes the JSON response → validates → formats

think_status

Inspects session state for debugging

think_list_templates

Lists modes and their scaffold cost

think_verify

Builds independent verification prompts for a conclusion

think_verify_submit

Takes the refutation / samples back → verdict + agreement

The /think <problem> prompt drives the flow: think → reason → think_submit.

Retry flow

think_submit -> validate -- OK --> format + return result
                 |
                 FAIL --> { status: "invalid", errors: [...], retry_count: 1 }
                 |        host AI fixes it -> submits again
                 +-- after 2 failures --> "best_effort" with validation_errors

After the retries run out the server does not pretend the result is valid: the status becomes best_effort and the structural errors are returned verbatim.

Independent verification (think_verify)

Structure enforcement is not the same as being right. think_verify adds two passes that make the model do independent work:

Pass

What it does

A. Refutation

A fresh-context pass that never sees the reasoning chain gets only the conclusion and its evidence, and is told to break it. The fresh context is the mechanism — anchoring is the failure mode targeted.

B. Self-consistency

N answers to the raw problem, majority-voted, so you can see whether the answer is even reproducible.

Why the host runs the prompts, not the server

MCP was supposed to let the server initiate these calls itself through sampling (session.create_message). That does not work here: on mcp 2.3.0 both the stdio and the Streamable HTTP transports raise

NoBackChannelError: Cannot send 'sampling/createMessage':
this transport context has no back-channel for server-initiated requests.

— neither transport carries a channel for server-initiated requests, and sampling is deprecated as of 2026-07-28 (SEP-2577). So the server builds the prompts and parses the results, and the host executes them. No API key is involved either way.

The honest weakness

The server cannot enforce that the host ran the refutation in a fresh context. A host that answers inline from its existing chain gets a weaker version of pass A. The prompt says so, and the returned output says so too (cannot_enforce).

Neither pass is a judge: A is a model opinion that can reject a correct conclusion, and B measures reproducibility, not correctness — a confidently wrong answer repeated N times is still wrong. Both caveats are attached to every result, and when nothing usable comes back the status is verification_failed with the conclusion labelled UNVERIFIED, never a pretend-success.

Validation measures structure, not quality

The host AI can still write a fake counter. The validator cannot tell — and deliberately does not try, because there is no LLM judge here.

Validation does not measure quality. It measures structure. The host AI can still write a fake counter. But without structural enforcement, it skips the step entirely.

DEEP, for example, requires: ≥3 assumptions each labelled [TRUTH]/[ASSUMPTION], a root_cause object with single_cause + branches (1 branch × 5 links if linear, ≥2 branches × 3 links if an issue tree), ≥2 red-team attacks each with a what_would_change_my_mind and a tripwire, evidence labelled [MEASURED]/[INFERRED], ≥3 bias checks, and ≥1 piece of disconfirming evidence.

Calibration only warns, never rewrites: confidence: 0.95 with thin evidence returns a warning and leaves 0.95 untouched — rewriting it would invent a number.

Measured results

agnes-3.0-flash over its OpenAI-compatible API, all 26 tasks in eval/tasks.jsonl, concurrency 2. Raw output is committed in eval/results/.

.venv/bin/python eval/compliance.py --concurrency 2

Two complete runs are committed: the baseline (20261006-123628, 10-step deep scaffold) and the guard-rails run (20261006-143122, 11-step deep with all the guards above). The tables show both where they differ; where only one number is given it is the guard-rails run.

Compliance — does the host AI follow the protocol?

Metric

Baseline

Guard rails

Valid on first submission

18 / 26 (69%)

13 / 26 (50%)

Valid after retry (final)

25 / 26 (96%)

25 / 26 (96%)

Ran out of retries

1 / 26 (4%)

1 / 26 (4%)

Non-JSON replies

5, across 4 tasks

4

By mode:

mode

n

first-try (base → rails)

final ok (base → rails)

quick

3

100% → 100%

100% → 100%

standard

16

69% → 50%

100% → 94%

deep

7

57% → 29%

86% → 100%

Correctness — on the 8 tasks with an objective answer

Baseline

Guard rails

Protocol correct

8 / 8

7 / 8

Direct correct

8 / 8

8 / 8

The regression is math-widgets, and the raw log shows exactly why: the model submitted confidence: 0.9+ backed by one piece of evidence, was rejected twice by the pre-existing ">0.9 with thin evidence" rule, and exhausted its retry budget. The failure is a retry mechanics problem, not a guard misfiring — the rule worked as designed, the model just repeated the same mistake instead of adding evidence or lowering confidence. The flipped task also swapped: baseline exhausted on debug-useeffect-leak (deep), guard rails run recovers it. n=8 is far too small to read either flip as signal.

Cost

Metric

Baseline

Guard rails

Protocol tokens

90,003

113,879

Protocol host calls

37

41

Direct tokens

19,058

18,120

Ratio

4.72×

6.28×

Average latency

50.1 s / task

61.2 s / task

Average scaffold

539 tokens

715 tokens

What this shows — and what it does not

  • Shows: adding structure costs compliance. The 11-step deep scaffold with mandatory labels and the issue-tree shape cut first-try validity from 69% to 50% overall, and from 57% to 29% in deep — the model needs more retries to produce the richer shape. The retry loop still lands at 96% final validity, and deep final validity went 86% → 100%.

  • Shows: the new calibrator warnings fire on real output — e.g. arch-monolith-split was told to restate "1%" as "10 in 1,000". They are advisory, so they cost tokens but cannot fail a submission.

  • Does not show a correctness change. 7/8 vs 8/8 on 8 regex-scored tasks is one retry mechanics failure, not evidence either way. The honest headline is unchanged: this measurement supports structure enforcement, not accuracy improvement, and the extra structure is paid for in tokens (6.28×) and latency (61 s/task).

Known limitations

  • ❌ Does not make the AI smarter — only stops it skipping steps

  • ❌ Cannot verify semantic quality (a genuine counter vs a fake one)

  • ❌ Cannot force 100% compliance — a model that ignores protocols gets nothing from this

  • ❌ Guard rails have a compliance price: first-try validity fell 69% → 50% (deep 57% → 29%) because the richer shape is harder to produce on the first attempt

  • ❌ Deep mode does not suit every question (scaffold token cost)

  • ⚠️ The eval needs a real host AI to measure compliance; the server itself produces no answer

  • ⚠️ think_verify cannot force a fresh context, and its two passes are opinions, not a judge

Tests

uv venv .venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest -q          # 166 tests, including e2e over real stdio

The end-to-end tests spawn the real server over stdio and drive it with mcp.Client — the same path Cursor / Claude Desktop take.

License

MIT

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Autonomous spec-to-product coding-agent CLI. Its MCP server exposes 34 tools over stdio: project state and task-queue ops, memory retrieve/store, code search, quality and verification reports, repo hotspots/co-changes, and structured findings/learnings.
    13,524 npm
    1,088
    Business Source 1.1
  • A
    license
    B
    quality
    A
    maintenance
    Local-first Agent OS that wraps Claude Code, Codex CLI, and other coding agents in a replayable Seed → Ledger → Runtime contract, driven by an interview → seed → execute → evaluate → evolve workflow loop.
    34
    18,092 PyPI
    6,193
    MIT