Skip to main content
Glama

Thinking Agent MCP

License: MIT Node.js TypeScript MCP

Extend your thinking model's chain of thought via MCP tools โ€” a Model Context Protocol server that exposes chat_agent, create_branch, and get_branch_details tools, enabling thinking models to offload subtasks to non-thinking models and build tree-structured multi-perspective analysis.


Features

  • ๐Ÿง  Chain of Thought Extension โ€” Thinking models can delegate reasoning subtasks to non-thinking models via chat_agent, extending effective reasoning depth beyond single-model token limits

  • ๐ŸŒณ Tree-Structured Thinking โ€” create_branch enables recursive, multi-perspective exploration with four branch types (drill down / verify / explore / stash)

  • ๐Ÿ” Full Traceability โ€” get_branch_details retrieves the complete raw reasoning process of any created branch

  • ๐Ÿ›ก๏ธ Context Isolation โ€” Tools are stateless and self-contained; all context must be packed into input_text. No conversation history dependency

  • ๐ŸŽ›๏ธ Parameter Control โ€” Fine-grained control over tool model output via temperature, top_p, seed, stop, and max_tokens

  • ๐Ÿ”Œ Dual API Support โ€” Works with both DeepSeek official API (recommended) and SiliconFlow API


Table of Contents


Quick Start

# Clone and install
git clone https://github.com/ScarletLilith/DeepSeekV4Flash_Thinking_TreeMCP.git
cd meditatorMCP
npm install

# Configure API (see Configuration section below)
# edit test/config.json or set environment variables

# Start the server
npm run build
npm start

# Or run development mode
npm run dev

Configuration

Configuration is loaded with the following priority: Environment variables > test/config.json

export DEEPSEEK_API_KEY=sk-your-key
export DEEPSEEK_BASE_URL=https://api.deepseek.com
export DEEPSEEK_MODEL=deepseek-v4-pro

Note: DeepSeek's thinking mode uses thinking: {type: "enabled"} (not enable_thinking: true).

Option 2: SiliconFlow API (Fallback)

export SILICONFLOW_API_KEY=sk-your-key
export SILICONFLOW_BASE_URL=https://api.siliconflow.cn/v1
export SILICONFLOW_MODEL=deepseek-ai/DeepSeek-V4-Flash

Config File

Create test/config.json (gitignored automatically):

{
  "baseUrl": "https://api.deepseek.com",
  "model": "deepseek-v4-pro",
  "apiKey": "sk-xxx"
}

Tools

chat_agent

Calls a non-thinking model to execute an independent subtask, extending the thinking model's chain of thought.

Parameters

Parameter

Type

Default

Description

input_text

string

required

Complete, self-contained task description with all context

system_prompt

string

optional

System prompt for role/behavior constraints

temperature

number

0.7

Sampling temperature (0.0โ€“2.0). Low = precise, high = creative

top_p

number

0.9

Nucleus sampling threshold (0.0โ€“1.0)

max_tokens

number

4096

Maximum output tokens (enforced server-side via API)

stop

string[]

[]

Stop sequences; empty array = natural completion

seed

number

optional

Random seed for reproducible output (with low temperature)

Parameter Strategies

Verification:  temperature=0.1, top_p=0.1,  max_tokens=2048, seed=42
Exploration:   temperature=1.2, top_p=0.95, max_tokens=4096
Balanced:      temperature=0.5, top_p=0.8,  max_tokens=4096

create_branch

Creates a thinking branch node with recursive nesting support for deep multi-perspective analysis.

Parameters

Parameter

Type

Default

Description

session_id

string

required

Session ID, consistent within a single reasoning session

input_text

string

required

Self-contained subtask description (โ‰ฅ30 characters)

call_type

string

drill_down

Branch type: drill_down / verify / explore / stash

parent_node_id

string

trunk

Parent node ID for tree nesting

Four Branch Types:

Type

Temperature

Purpose

drill_down

0.2

Deep-dive into a subproblem with focused precision

verify

0.0

Verify a conclusion or hypothesis with maximal determinism

explore

1.0

Divergent thinking from different angles with high creativity

stash

0.6

Temporarily record intermediate thoughts for later reference

Response

{
  "status": "success",
  "node_id": "n_a1b2c3d4",
  "conclusion": "The extracted conclusion text...",
  "confidence": 0.85,
  "remaining_quota": 12,
  "suggestions": [
    "ๅ‘ๆ•ฃๆŽข็ดขๅฎŒๆˆ๏ผŒๅฏๅฏนๆœ‰ไปทๅ€ผ็š„ๆ–นๅ‘็”จ drill_down ๆทฑๅ…ฅ",
    "่ฟ˜ๅฏๅˆ›ๅปบ 12 ไธชๅˆ†ๆ”ฏ๏ผŒๅปบ่ฎฎ็ปง็ปญๅคš่ง’ๅบฆๆŽข็ดข"
  ]
}

get_branch_details

Retrieves the complete raw reasoning process of a previously created branch node.

Parameters

Parameter

Type

Default

Description

session_id

string

required

Session ID

node_id

string

required

Branch node ID returned by create_branch

Response

{
  "status": "success",
  "node_id": "n_a1b2c3d4",
  "raw_process": "The complete raw reasoning output from the model..."
}

Error Handling

Tools return structured errors with type and action fields for the thinking model to make informed decisions:

{
  "success": false,
  "type": "api",
  "action": "report",
  "error": "API authentication failed (401)",
  "status_code": 401
}

Error Type

action

Trigger

network

retry

DNS resolution failure, connection refused

api

backoff

429 rate limited

api

report

401 authentication failure

api

retry

5xx server errors

validation

fix_input

Empty input_text

config

report

Missing API Key / Model configuration

The server-side retry mechanism uses exponential backoff with jitter (1sโ†’3sโ†’7s, max 3 retries) for 429 and 5xx errors. Network errors (ENOTFOUND, ECONNREFUSED, ECONNRESET) are not automatically retried.


MCP Client Setup

Claude Desktop

{
  "mcpServers": {
    "thinking-agent": {
      "command": "node",
      "args": ["path/to/meditatorMCP/dist/index.js"],
      "env": {
        "DEEPSEEK_API_KEY": "sk-your-key",
        "DEEPSEEK_BASE_URL": "https://api.deepseek.com",
        "DEEPSEEK_MODEL": "deepseek-v4-pro"
      }
    }
  }
}

Any MCP-compatible Client

Configure stdio transport to point to node dist/index.js in the project directory, with the required environment variables set.


Testing

The project includes both interactive and automated test frameworks:

# Interactive CLI (with tools mode)
npm run test:with-tool

# Interactive CLI (pure thinking, no tools)
npm run test:without-tool

# Automated comparison test (runs both scenarios + generates report)
npm run test:comparison

# Batch end-to-end tests
npm run test:batch

Test Scripts

Script

Description

test/testFramework.ts

Interactive CLI test framework

test/comparisonTest.ts

Automated A/B comparison (with-tool vs without-tool)

test/runA.js

Scenario A: thinking model + tools (standalone, DeepSeek)

test/runB.js

Scenario B: pure thinking model (standalone, DeepSeek)

test/batchTest.ts

Batch end-to-end tests

Scoring: Each question is evaluated against 10 objective checkpoints (50 total). Evaluation is done by human reviewers, not automated scripts.

Note: The test/config.json file contains your API key and is automatically gitignored.


Project Structure

โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ index.ts           # MCP Server entry point
โ”‚   โ”œโ”€โ”€ chatAgentTool.ts   # Tool implementations (chat_agent, create_branch, get_branch_details)
โ”‚   โ”œโ”€โ”€ gatekeeper.ts      # Input validation and quota enforcement
โ”‚   โ”œโ”€โ”€ strategyEngine.ts  # Parameter strategy mapping (call_type โ†’ temperature/top_p)
โ”‚   โ”œโ”€โ”€ nodeStore.ts       # Branch node storage and conclusion extraction
โ”‚   โ”œโ”€โ”€ schemas.ts         # Zod validation schemas and TypeScript types
โ”‚   โ”œโ”€โ”€ logger.ts          # Structured logging to stderr
โ”‚   โ””โ”€โ”€ polyfill.ts        # Node 14 fetch polyfill
โ”œโ”€โ”€ test/
โ”‚   โ”œโ”€โ”€ comparisonTest.ts  # A/B comparison test
โ”‚   โ”œโ”€โ”€ testFramework.ts   # Interactive CLI test framework
โ”‚   โ”œโ”€โ”€ batchTest.ts       # Batch testing
โ”‚   โ”œโ”€โ”€ runA.js            # Scenario A test (DeepSeek)
โ”‚   โ”œโ”€โ”€ runB.js            # Scenario B test (DeepSeek)
โ”‚   โ””โ”€โ”€ config.json        # API configuration (gitignored)
โ”œโ”€โ”€ .env.example           # Environment variable template
โ”œโ”€โ”€ blueprint.md           # Project design blueprint (Chinese)
โ”œโ”€โ”€ package.json
โ”œโ”€โ”€ tsconfig.json
โ””โ”€โ”€ README.md

Development

# Build TypeScript
npm run build

# Start production server
npm run start

# Development mode (ts-node, no build step)
npm run dev

Design Philosophy

  1. Self-Contained Task Descriptions โ€” All context must be packed into input_text; tools never rely on conversation history

  2. Context Isolation โ€” Each tool call is stateless and independent, preventing context explosion in the main chain

  3. Token Cost Optimization โ€” Context is consumed by the cheaper non-thinking model's input tokens, not the thinking model's output tokens

  4. Tree-Structured Reasoning โ€” Complex problems are decomposed into independent branches, each analyzed separately, then synthesized


Benchmark: MCP Tools Impact on Output Quality

We conducted a controlled experiment comparing 3 approaches across 5 challenging engineering problems (distributed consensus, service mesh, RTOS kernel, columnar storage engine, multi-modal AI agent framework).

Test Groups

Group

Model

API

Tools

A

GLM-5.2

SiliconFlow

None

B

DeepSeek-V4-Flash

DeepSeek Official

chat_agent + create_branch

C

DeepSeek-V4-Flash

DeepSeek Official

None

Key Results

Metric

A (GLM-5.2)

B (DS + Tools)

C (DS Pure)

Total Output

38,888 chars

169,484 chars ๐Ÿ†

57,565 chars

Total Time

705s

1,516s

239s ๐Ÿ†

Total Tokens

28,459

210,482

29,184

Total Cost

ยฅ0.75

ยฅ0.21

ยฅ0.06 ๐Ÿ†

Avg Output/Question

7,778 chars

33,897 chars (4.4x) ๐Ÿ†

11,513 chars

Tool Calls

0

30 ๐Ÿ†

0

Cache Hit Rate

0%

up to 68% ๐Ÿ†

0%

What We Found

  • With MCP tools, DeepSeek-V4-Flash produced 4.4x more detailed engineering solutions โ€” including complete Go-style Raft consensus implementations, assembly-level RTOS scheduler code, and production-ready service mesh configurations

  • Tree-structured thinking (create_branch) enabled the model to explore 5-6 levels deep on complex problems, creating subtrees for architecture, implementation, testing, and verification

  • Cost comparison: GLM-5.2 costs 13x more than DeepSeek-V4-Flash pure thinking mode (ยฅ0.75 vs ยฅ0.06) for comparable output quality. With tools enabled, DeepSeek-V4-Flash cost increased to ยฅ0.21 due to deeper exploration (3.7x more tokens), but remained 3.6x cheaper than GLM-5.2.

  • Cache hit rates reached 68% during multi-round tool calls, dramatically reducing effective input costs via DeepSeek's prefix caching

Full experiment results and data: results/comparison/report.md


License

MIT