Skip to main content
Glama
README.md
# DeepSeek Delegate MCP

A tiny local [MCP](https://modelcontextprotocol.io) server that exposes
**DeepSeek V4** to Codex as a single tool: `call_deepseek_sub_agent`.

It turns Codex into a hybrid agent: **GPT Sol stays the architect** (planning,
reviewing, integrating) and **DeepSeek V4 Flash does the cheap execution**
(boilerplate files, test-suite generation, bulk text transformation) at
**$0.14 per 1M input tokens**.

## How it works

```
┌────────────────────────┐   call_deepseek_sub_agent   ┌────────────────────────────┐
│  GPT Sol (Codex agent) │ ───────────────────────────▶ │  local MCP server (node)   │
│  plans / reviews /     │ ◀─────────────────────────── │  "deepseek-delegate"       │
│  integrates            │        output text          └─────────────┬──────────────┘
└────────────────────────┘                                          │ POST /responses
                                                                     ▼
                                                    DeepSeek API (api.deepseek.com)
                                                    deepseek-v4-flash / deepseek-v4-pro
```

Codex runs one model per session, so instead of switching providers mid-task
you give Sol a delegation tool. Sol crafts a precise prompt, DeepSeek answers
with plain text, and Sol reviews + applies the result. No file access is granted
to DeepSeek — it only ever sees the exact prompt you send.

## Features

- Single MCP tool: `call_deepseek_sub_agent` (prompt required; model, system, temperature, max_output_tokens, reasoning_effort optional)
- Uses DeepSeek's native [Responses API](https://api-docs.deepseek.com/guides/responses_api)
- Zero extra key setup if you already use DeepSeek's official Codex integration
- Works with Codex CLI, the ChatGPT desktop app, and the Codex IDE extension

## Requirements

- [Codex](https://developers.openai.com/codex/) (CLI or desktop app)
- A DeepSeek API key from the [DeepSeek Platform](https://platform.deepseek.com)
- Node.js 18+

## Quick start

### 1. Clone and install

```bash
git clone https://github.com/trixmix821/deepseek-delegate-mcp.git
cd deepseek-delegate-mcp
npm install
```

### 2. Register the MCP server

Add this to `~/.codex/config.toml` (use the absolute path from step 1):

```toml
[mcp_servers.deepseek-delegate]
command = "node"
args = ["/absolute/path/to/deepseek-delegate-mcp/deepseek-mcp-server.mjs"]
env_vars = ["DEEPSEEK_API_KEY"]
tool_timeout_sec = 600
default_tools_approval_mode = "auto"
```

### 3. Provide your API key

The server looks for the key in this order:

1. `DEEPSEEK_API_KEY` environment variable
2. `experimental_bearer_token` under `[model_providers.deepseek]` in `~/.codex/config.toml`
   (this is what DeepSeek's [official Codex setup script](https://api-docs.deepseek.com/quick_start/agent_integrations/codex/) writes — if you've run it, you're already done)

```bash
export DEEPSEEK_API_KEY=sk-...
```

### 4. Teach Sol to delegate

Append to `~/.codex/custom_instructions.md`:

```markdown
## Hybrid Architect: Sol (planner) + DeepSeek (executor)

You are the master architect (GPT Sol). You own the high-level plan, design,
and project structure. DeepSeek is your cheap execution layer.

For massive boilerplate files, extensive test-suite generation, bulk/repetitive
text transformation, or well-specified mechanical subtasks, invoke the
`call_deepseek_sub_agent` tool instead of doing the work in your own context.
Craft a precise, self-contained prompt: exact file paths, signatures, language,
framework, constraints, and expected output format. Never delegate open-ended
design decisions. Review DeepSeek's output, fix logic/interface mismatches, then
integrate. If you are already running as a DeepSeek model, do not delegate.
```

### 5. Verify

```bash
node test-deepseek.mjs "Reply with exactly: OK"
```

Expected output: the model replies `OK`, and a usage line
(`[deepseek-delegate] model=... input_tokens=... output_tokens=...`) is printed
to stderr.

Then **restart Codex** so it loads the new MCP server (in the TUI, `/mcp` shows
active servers), and ask something like: *"delegate the test-suite boilerplate
to DeepSeek."*

## Tool reference

### `call_deepseek_sub_agent`

| Parameter | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `prompt` | string | yes | — | The exact coding instruction or context to process |
| `model` | string | no | `deepseek-v4-flash` | `deepseek-v4-flash` or `deepseek-v4-pro` |
| `system` | string | no | — | Optional system prompt for the sub-agent |
| `temperature` | number | no | — | 0.0–2.0 (no effect in thinking mode) |
| `max_output_tokens` | number | no | `8192` | Maximum output tokens |
| `reasoning_effort` | string | no | — | `low`, `high`, or `max` |
| `thinking` | boolean | no | off | Force thinking mode on/off |

**Which model?** `deepseek-v4-flash` is the default and right for nearly all
delegated work (boilerplate, tests, transformations). Escalate to
`deepseek-v4-pro` only when a single small task genuinely needs stronger
reasoning — it costs ~3x more.

## Configuration

Environment variables (all optional):

| Variable | Default | Description |
| --- | --- | --- |
| `DEEPSEEK_API_KEY` | — | DeepSeek API key (falls back to your Codex config) |
| `DEEPSEEK_MODEL` | `deepseek-v4-flash` | Default model for the tool |
| `DEEPSEEK_BASE_URL` | `https://api.deepseek.com` | API base URL |
| `DEEPSEEK_WIRE_API` | `responses` | `responses` or `chat` (OpenAI-style chat completions) |
| `DEEPSEEK_THINKING` | `disabled` | `enabled` or `disabled` thinking mode default (disabled = fast/cheap) |
| `DEEPSEEK_MAX_PROMPT_CHARS` | `30000` | Reject prompts larger than this to prevent giant-payload delegation |

## Costs

DeepSeek V4 pricing (per 1M tokens, [source](https://api-docs.deepseek.com/quick_start/pricing)):

| Model | Input (cache miss) | Input (cache hit) | Output |
| --- | --- | --- | --- |
| `deepseek-v4-flash` | $0.14 | $0.0028 | $0.28 |
| `deepseek-v4-pro` | $0.435 | $0.003625 | $0.87 |

Context window is 1M tokens; max output is 384K.

## Security notes

- The server only calls the DeepSeek API. It never reads or writes your files,
  and DeepSeek only sees the text you put in the `prompt`.
- DeepSeek's official setup stores your key in `~/.codex/config.toml` in
  plaintext. Prefer `export DEEPSEEK_API_KEY=...` and keep the token out of the
  file.
- `default_tools_approval_mode = "auto"` lets Sol call the tool without a
  prompt; change it to `prompt` if you want to approve every delegation.

## Limitations

- The tool only receives the prompt text — it has no access to your repository,
  so delegated prompts must be self-contained.
- Delegation is only worth it for SMALL, mechanical, self-contained tasks.
  Prompts over 30,000 chars are rejected (configurable via
  `DEEPSEEK_MAX_PROMPT_CHARS`) — split them into smaller subtasks instead.
- Thinking mode is off by default for speed; enable it via the `thinking` tool
  parameter only when the task genuinely needs reasoning.
- Delegation is a judgment call by the agent. Explicitly asking for it
  ("delegate X to DeepSeek") makes it deterministic.
- If your session is already running on DeepSeek, delegation is redundant.

## License

MIT

TDQS

A4.2/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusing it with others. The tool's purpose is clearly defined as delegating to DeepSeek V4.

Naming Consistency5/5

The tool name follows a clear verb_noun pattern ('call' + 'deepseek_sub_agent'). With only one tool, there are no inconsistencies or mixed conventions.

Tool Count4/5

The server has a single tool, which is slightly below the typical well-scoped range. However, it is appropriate for the narrow purpose of delegating to DeepSeek, making the count reasonable.

Completeness5/5

The tool provides a generic delegation capability for any text processing or code generation task, covering the entire stated domain. There are no obvious missing operations for a simple delegate service.

Maintenance

ActivitySlowing
ResponsivenessNo issues