deep-think-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@deep-think-mcpthink through the pros and cons of remote work"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
The Problem
Ask a capable model a hard question and it will often produce a fluent answer that sounds reasoned but skipped the hard parts — it assumed something it never checked, leaned on a weak analogy, ignored a stakeholder, or committed to the first framing that came to mind. The usual fixes ("think step by step," "critique your answer") work unevenly and leave nothing behind: the reasoning evaporates with the context window.
Three concrete gaps:
Reasoning is ephemeral. Once the conversation scrolls away, the chain of thought is gone. You can't revisit why a conclusion was reached, or resume a half-finished analysis tomorrow.
Self-critique is unstructured. "Critique yourself" gives a model too much latitude — it critiques what's easiest, not what's load-bearing. Nothing guarantees it stress-tests its evidence, its assumptions, and its blind spots in turn.
Local models make both worse. A 7B/8B model asked to run a multi-step reasoning protocol and remember where it is in that protocol and emit clean JSON at each step will drop one of those balls.
Related MCP server: Sequential Thinking MCP Server
The Solution
deep-think-mcp is an MCP server that externalizes the reasoning protocol into a state machine the server runs on the model's behalf. A problem is worked in explicit stages (Problem Definition → Research → Analysis → Synthesis → Conclusion), and within each stage the model either sharpens one line of reasoning through rounds of structured self-critique, or spins up competing specialist perspectives that are scored and converged. Every intermediate step is scored on a shared 7-dimension utility matrix and saved to disk.
Because it targets local models — small context, weak instruction-following, no reliable JSON mode — every tool response is short, flat, and directive: it tells the model exactly which tool to call next. A single tool, next_action, answers "what do I do now?" from any state, so the model never has to hold the protocol in its head.
┌─ Problem Definition ─┐ ┌─ Analysis ─┐ ┌─ Synthesis ─┐ ┌─ Conclusion ─┐
│ draft ─▶ critique │ │ specialist │ │ specialist │ │ commit ─▶ │
│ ─▶ refine ─▶ score │ ▶ │ candidates │ ▶ │ candidates │ ▶ │ finalize ─▶ │
│ ─▶ (converged?) ─▶ │ │ ─▶ score │ │ ─▶ score │ │ move/keep │
│ commit │ │ ─▶ winner │ │ ─▶ winner │ │ │
└──────────────────────┘ └─────────────┘ └─────────────┘ └──────────────┘
every step scored, persisted to disk, and resumableFeatures
Two reasoning modes, one schema
Every session picks serial (one line of reasoning, sharpened by rotating critique lenses) or subagent (competing specialist perspectives, scored and converged) — fixed for the life of the session. Both emit the same stage machine, thoughts, and 7-dim utility scores, so you can run a question through both and compare.
Structured self-critique
Serial mode ships 8 bundled critique lenses — overconfidence, weak_evidence, missing_perspective, unstated_assumption, scope_creep, alternative_framing, steel_man, first_principles — each a directive prompt that hunts one specific failure mode. Drop your own .md lenses in to add or override by name.
Persistent by default
One JSON file per session, written under a Portalocker lock with a crash-safe .bak protocol, tracked in a central index. Finalize prompts you to move the artifact anywhere (a project folder, a synced drive) and it stays fully resumable there.
Built for weak models
Flat tool signatures, short directive responses, and next_action as an authoritative "what next?" resolver. Every input is accepted as JSON or tolerant plaintext (scores="correctness: 0.8, clarity: 0.7"). Nothing ever raises a traceback — failures return a retry_with_clarification directive naming the fix.
Local-first, offline-capable
Serial mode and the endpoint-free manual subagent engine need no GPU, no API key, and no network. Point the optional engines at any OpenAI-compatible endpoint (Ollama, llama.cpp, vLLM) only if you want to.
Honest hybrid engine
Subagent mode has two engines: necort drives a vendored Nash-equilibrium core against an endpoint; manual is endpoint-free, where the calling model plays each specialist and self-scores all 7 dimensions for real. (See the honest NECoRT story — most of the upstream PR turned out to be filler.)
Quick Start
Requires Python ≥ 3.11 and uv. The vendored NECoRT core is a git submodule, so clone recursively:
git clone --recurse-submodules <this-repo-url> deep-think-mcp
cd deep-think-mcp
uv sync # core deps; add --extra autopilot for the optional autopilot feature
uv run pytest # confirm a healthy install (tests never touch your real home dir)(Already cloned without submodules? git submodule update --init. The submodule is only needed for [subagent] engine = "necort"; everything else works without it.)
Launch the stdio server:
uv run python -m deep_think_mcp.serverThis is a dev-checkout tool — it reads config/default.toml from the repo root, so every client config points --directory at your clone (see docs/wiring.md).
Drive a serial session (every response carries a message and a next_tool — when unsure, call next_action(session_id)):
start_session(question="Should we cache API responses at the edge or origin?")
→ { "mode_required": true, "next_tool": "set_session_mode", "session_id": "…" }
set_session_mode(session_id, mode="serial")
begin_thought(session_id, content="Cache at the edge: lower latency for users…")
critique_current_thought(session_id) # server picks a stage-appropriate lens
→ { "lens": "weak_evidence", "draft_content": "…", "lens_template": "…", "next_tool": "submit_critique" }
submit_critique(session_id, text="No numbers back the latency claim…")
refine_current_thought(session_id, new_content="Cache at the edge (CDN PoPs) when…")
score_current_thought(session_id, scores="correctness: 0.8, clarity: 0.8, evidence: 0.7, …")
→ { "converged": false, "next_tool": "critique_current_thought" } # loop until converged or max_rounds
commit_thought(session_id)
advance_stage(session_id) # … repeat through the stages …
finalize_session(session_id) # → prompts you to move or keep the saved artifactNew here?
docs/GUIDE.mdis a complete, self-contained teaching document — the concepts, the architecture, both modes in depth, every tool and config key, and how to extend the system. This README is the map; the guide is the tutorial.
The Two Modes
A session's mode is chosen once at creation and is immutable — to use the other mode, start a new session. Creating a session without a mode returns a mode_required directive rather than silently defaulting, forcing the choice to surface.
Serial — one line of reasoning, critiqued
Within a stage, a thought cycles begin → critique → submit → refine → score and repeats with a new lens until it converges. Four convergence rules are checked in precedence order:
fixed_point— the refinement barely changed the text (normalized edit distance< edit_distance_epsilon, default0.05).diminishing_returns— two rounds in a row each improved the score by< score_threshold(default0.05).max_rounds— the round cap (default3) is hit.Otherwise keep going with the next lens.
Natural convergence outranks the ceiling, so you learn why it stopped. Lenses rotate through stage-appropriate defaults first (e.g. Analysis → weak_evidence, overconfidence), then the rest of the library.
Subagent — competing perspectives, converged
Specialists (default roster: Analysis, Creativity, Skeptic) propose competing candidates scored on the 7-dim matrix; the strongest wins. Two engines, same four tools (begin_subagent_thought, advance_subagent_round, inspect_utility_matrix, commit_subagent_thought):
|
| |
Needs an endpoint? | No — fully local & offline | Yes — any OpenAI-compatible |
Who plays the specialists? | The calling model itself | The vendored Nash core |
Utility scoring | All 7 dims, real self-scores | 3 dims real ( |
Commit gate | 7-dim mean ≥ | winner's |
Selection | highest mean wins, ties → earliest | Nash equilibrium |
With engine = "necort" but no endpoint configured (the shipped default), begin_subagent_thought doesn't fail opaquely — it returns a directive pointing at the endpoint-free manual path.
The honest NECoRT story
The original design imagined subagent mode as a full port of PhialsBasement/Chain-of-Recursive-Thoughts PR #7 — specialist agents, a native 7-dim utility matrix, bias detection, continuous learning. A code recon during the build found that most of that PR is disconnected filler: the files advertising those features are never imported, make zero LLM calls, and several aren't even valid Python. The one part that works is NashEquilibriumRecursiveChat. So this project vendors PR #7 in full (a faithful, re-pinnable submodule mirror) but imports only those two working files, wrapped by a single adapter (necort_adapter.py) that shims a real crash, a hardcoded endpoint, and a stdout-corrupts-the-transport bug — without editing a vendored line. Because a single blended Nash rating can honestly inform only 3 of 7 dimensions, genuine multi-perspective diversity comes from the second, from-scratch manual engine instead. The lesson is baked in: verify third-party code against reality before building on its advertised behavior.
Data & the Finalize/Move Lifecycle
Everything lives under one data root, ~/deep-think-mcp/ by default (override with DEEP_THINK_HOME):
~/deep-think-mcp/
├── config.toml seeded from config/default.toml on first use; edit freely
├── index.json session_id → { path, mode, status, created_at, updated_at }
├── sessions/ one JSON file per session
├── lenses/ optional: drop-in .md critique lenses (override by name)
└── logs/ reserved directory (unused in v1)finalize_session returns a human_prompt offering to relocate the artifact; move_session moves it atomically (write → verify → unlink, won't clobber without force) and keep_here records the decline. Sessions moved outside the root stay fully functional — list_sessions / resume_session find them via the index's absolute paths, and move_history tracks every hop.
Configuration
Layered, lowest to highest precedence: packaged defaults (config/default.toml) → user config (<root>/config.toml, seeded on first use) → per-session overrides (start_session(overrides={…})). Key settings:
Section | Key | Default | Notes |
|
|
| Overridden by |
|
|
| The convergence knobs. |
|
| the 8 bundled lens names | Rotation order after stage defaults. |
|
|
|
|
|
|
| Round cap and commit gate. |
|
|
| Specialist roster. |
|
|
| NECoRT engine target. Empty endpoint → the manual-path directive. |
|
|
| Per-session overridable via |
|
|
| Off by default; when off, no network code path is reachable. |
The full table with every key lives in docs/GUIDE.md.
Tolerant input. Every structured parameter accepts JSON or plaintext (tags="a, b, c", scores="correctness: 0.8, clarity: 0.7"). Unparseable input returns a retry_with_clarification payload naming the parameter, expected shape, and an example — never a raw error.
Autopilot (optional). With [autopilot].enabled = true (and uv sync --extra autopilot), two extra tools let the server drive a whole stage internally against a configured endpoint, stopping cleanly with a resumable partial-progress directive on any fault. Off by default, it imports zero networking code.
Tool Surface
25 tools always registered; 27 with autopilot enabled. All responses are flat objects with a message and usually a next_tool.
Group | Tools |
Session lifecycle |
|
Stage cursor |
|
Serial loop |
|
Subagent loop |
|
Meta / guidance / I-O |
|
Autopilot (when enabled) |
|
Full signatures, return fields, and every directive/error code are in docs/GUIDE.md.
Wiring Into an MCP Client
Copy-pasteable config for Claude Desktop, Claude Code, Cursor, Continue, and LibreChat is in docs/wiring.md. The mcpServers-style shape:
{
"mcpServers": {
"deep-think": {
"command": "uv",
"args": ["--directory", "/absolute/path/to/deep-think-mcp", "run", "python", "-m", "deep_think_mcp.server"],
"env": { "DEEP_THINK_HOME": "/absolute/path/to/your/data-root" }
}
}
}Sharing one server between clients (or tools keep vanishing from a long-lived host)? You can instead run deep-think as a single always-live Streamable HTTP daemon that multiple clients reach over a URL (
http://127.0.0.1:8182/mcp) rather than each spawning its own stdio process. This is also the fix when an agent host intermittently drops the tools from its cached schema. Seedocs/http-transport.md.
Documentation
Document | What it is |
The complete teaching guide — concepts, architecture, both modes in depth, full tool/config/directive/data-model references, extension, FAQ, glossary. | |
Exact client config for Claude Desktop, Claude Code, Cursor, Continue, LibreChat. | |
Running as a Streamable HTTP daemon — one always-live server shared by multiple clients (e.g. an agent host + a DAG), the systemd unit, security posture, and the fix for hosts that drop stdio tools from a cached schema. | |
Agent-runnable A/B/C test — does driving a model through the tool beat answering directly? Self-contained prompt, rubric, judge instructions, and report template. | |
The original design document (the "why" behind the architecture). | |
The task-by-task build breakdown with global constraints. | |
Why | |
How to re-pin the vendored NECoRT submodule. |
Architecture
The system is layered: a dispatch layer (server.py) that registers the tools, gates wrong-mode calls, parses tolerant input, and turns storage faults into directives; the engines (serial_engine, subagent_engine, manual_engine, necort_adapter, optional autopilot) that do the thinking; and a domain + persistence layer (session, stages, lens_loader, store, index, lifecycle, config, prompts, tolerant). Two invariants hold the design together: all model-facing wording lives in prompts.py, and necort_adapter.py is the only file that imports vendored code — the entire third-party surface is quarantined behind one boundary. Full diagram in the guide.
Testing
uv run pytest # full suite (423 tests)
uv run pytest -q -W error # the CI bar: pristine, warnings are errorsThe suite drives the real MCP SDK's in-memory client against the real server for every tool contract, plus one subprocess test that speaks real stdio MCP to the launched server. Every test injects a tmp_path data root, so running the suite never touches your real home directory.
How it was built. Implemented task-by-task with a fresh-implementer → adversarial spec+quality review → fix-loop discipline, closed out by a whole-branch multi-lens review with adversarial verification of every finding (including two real security fixes: import path traversal and credential exfiltration). Design docs are docs/build-plan.md and docs/execution-plan.md.
Benchmarks
Not yet run. A head-to-head of serial vs. subagent on three canonical prompts is planned but requires blind human rating to be meaningful, and is deliberately deferred rather than shipped as a self-graded number.
License
MIT — see LICENSE. This project vendors third-party source code (vendor/necort/, a git submodule of PhialsBasement/Chain-of-Recursive-Thoughts PR #7) under its own MIT license; see LICENSE-NOTICES for full attribution.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ClintMoody/deep-think-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server