multimodels-mcp
With this server, you can delegate tasks from your main AI agent to other models across providers. Capabilities include:
List models: view all enabled models with IDs and provider status.
Delegate tasks: send self-contained tasks to GPT (via Codex), DeepSeek, z.ai/GLM, OpenRouter, LM Studio (local), or "with hands" lanes (models that can read files, grep, and run npm test/builds but not edit). Control effort per task and optionally set a project workdir.
Execution modes: run tasks in background (receiving an immediate task ID) or synchronously with
wait: true.Check results: retrieve task output, track live progress for "with hands" lanes, and recover partial results on failures.
Resilience: automatic retries on network errors/rate limits, per-provider concurrency queuing, configurable timeouts.
Cost & safety: use subscription-based models without extra API cost, local models for free; manufacturer rule prevents delegating to same-vendor models to avoid unnecessary costs; with-hands lanes are read-only.
Configuration: manage providers, API keys, and models via
config/models.json,.env, or a local control panel.
Allows delegating tasks to OpenAI models via OpenAI-compatible APIs or the Codex CLI, with configurable reasoning effort and support for GPT-series models.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@multimodels-mcpdelegate code review to DeepSeek Pro"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
multimodels-mcp
Delegate tasks from Claude Code to other companies' models — without leaving the app.
This is a small MCP (Model Context Protocol) server that acts as a "waiter" between your main coding agent and every other model you have access to. Claude Code stays the orchestrator; the waiter takes an order to whichever kitchen you point at:
Codex CLI → GPT-5.6 Sol / Terra / Luna via your ChatGPT subscription (no API cost)
DeepSeek (DS4 Flash / Pro) via direct API
z.ai (GLM 5.2) via the coding-plan subscription
OpenRouter → anything in their catalog
LM Studio → local models on your machine or another box on your LAN, for free
"With hands" lanes → GLM (z.ai), DeepSeek and Kimi (Moonshot) piloting a disposable headless Claude Code pointed at their Anthropic-compatible endpoint — plus Claude itself on your subscription (
claude-maos, no API key) — the model reads your files, greps the code and actually runsnpm test/npm run build, but cannot edit anything
The same pattern works for any MCP-capable agent — nothing here is Claude-specific except where it's registered.
Tools exposed
Tool | What it does |
| Returns the menu: every enabled model with its exact id and provider status (missing key, offline local server, etc.) |
| Sends a self-contained task to the chosen model. Backgrounds it by default and returns a task id immediately; |
| With |
Delegation niceties, all born from the benchmarks below: pick the Codex model per call (codex:gpt-5.6-luna), set reasoning effort per call (effort works for Codex, claude-maos, z.ai, DeepSeek and OpenRouter) or as a per-model default picked in the panel, per-provider concurrency queues (z.ai and LM Studio silently choke on parallel calls — the server now queues them), configurable per-provider timeouts, and automatic retry on network drops / 429 / 5xx (the answer footer says repescada 1× when the second attempt saved the day).
Related MCP server: Zen MCP Server
Background by default
A delegation to a "with hands" lane or to a local box can take 20–30 minutes. Blocking the caller's session for that long is the whole reason this exists, so delegate_task runs in the background unless you ask otherwise:
It answers in milliseconds with a human-readable ticket (
tarefa-1,tarefa-2, …), the model, and how to collect the result. The tool description tells the calling agent explicitly to go do other work and check back later rather than polling in a loop.check_taskcollects it. The stored result includes the same origin/token footer a synchronous call would have printed."wait": truerestores the old synchronous behaviour — right call for a twenty-second task.Everything that can be refused is refused before the task is created (unknown/disabled model, the manufacturer rule,
efforton a with-hands lane), so a refusal is immediate rather than a task that fails later.Background tasks go through the same provider code paths, so per-provider concurrency queues (
maxConcurrent) still apply: five background delegations to LM Studio queue up exactly like five synchronous ones.
Progress signals and salvaged partials (with-hands lanes only)
A backgrounded delegation used to be a black box: check_task could only say "still running". Since 0.12.0 the claude-cli ("with hands") lanes report what is happening while it happens, and hand back the work in progress if the run dies. Codex, Gemini and the openai-compat API lanes are untouched — they have no equivalent signal, so check_task tells you honestly that this lane sends no progress.
How it's read. The engine no longer runs claude -p --output-format json and parses one document at the end. It runs:
--output-format stream-json --include-partial-messages --verbosewhich emits one JSON event per line as things happen. --verbose is not optional: with --print, the CLI refuses stream-json without it (When using --print, --output-format=stream-json requires --verbose). Event types actually observed on this machine (2026-08-01): system (init, status, hook_started, hook_response, thinking_tokens, post_turn_summary), stream_event (wrapping the raw API events message_start, content_block_start, content_block_delta with text_delta / thinking_delta / input_json_delta / signature_delta, content_block_stop, message_delta, message_stop), assistant (one per closed content block), user (tool results), rate_limit_event, and result last.
The final text is unchanged. The result event carries an identical key set to the single document the old --output-format json produced, so the extraction rules and their error messages are the same code as before, just fed from the stream. A test proves it: it runs the old parser and the new reader over the same recorded run and asserts the two strings are equal.
What the signals are. Steps (trips to the model, grouped by request_id — the CLI emits several assistant events per trip, so counting events would inflate it), tools used with counts, and output tokens (summed from message_delta, which matches the total the result event reports). Progress is persisted at most once every 3 seconds, plus a guaranteed final write — a single run emits hundreds of events and writing each one would hammer the disk.
Partials. When the process is killed by the deadline or by the 10 MB output cap, the text accumulated so far is attached to the error (a dedicated ErroComParcial) instead of being thrown away. It is stored on the task as a parcial field inside the erro state, not as a new state: the task did fail, and inventing a third state would force the list, the housekeeping pass and the orphan warning to learn about it. Anything displaying it must wrap it in an unmissable warning — check_task prints the warning before and after the draft, with explicit start/end markers, and the list shows erro (com rascunho incompleto). "wait": true gets the same treatment inline, since in synchronous mode there's no ticket to store it on.
Deliberately not implemented: live partial text during a normal run. A reasoning model passes through mid-way conclusions it later revises away; exposing those invites acting on something the model itself has already discarded. The draft surfaces only when the task dies — the one case where it is all that's left. The reader also never mixes thinking_delta into the accumulated text, for the same reason.
Unknown events and junk lines are ignored in silence. A future CLI update adding a field or a type, or stray non-JSON noise on stdout, must never take a delegation down.
Where results live
.multimodels/tarefas/<id>.json in the project root (gitignored) — one JSON file per task holding its state, model id, a ~200-character summary of the request, start/end timestamps, that lane's deadline, the result or the error, plus (with-hands lanes) a progresso block (passos, ferramentas, tokensSaida, atualizadoEm) and a parcial field when a dead run left a draft behind. On disk, so a result outlives the session that ordered it.
Ids are claimed exclusively. The next number comes from the files already present, and the file is created with the
wxflag; onEEXISTthe creator tries the next number. Two sessions running at once can never land on the same id.50 tasks are kept. A housekeeping pass runs on every creation and deletes the oldest beyond that, so the folder can't grow forever. A task still running inside its deadline is never deleted — something is still going to write to it.
Orphan tasks are reported honestly, not guessed away. If the session that started a task is closed, nobody will ever update that file: it stays
rodandoforever. There's no reliable way to detect this, so once a running task passes its own deadline plus a 2-minute grace period, it is presented as "provavelmente interrompida — a sessão que iniciou essa tarefa foi fechada". The file itself is never rewritten; if the answer does land, the warning disappears by itself.
Quick start
git clone https://github.com/dpmadsen/multimodels-mcp.git
cd multimodels-mcp
npm install
npm run build
# copy the key template and fill in what you use
cp .env.example .env
# register in Claude Code (user scope = available in every project)
claude mcp add --scope user multimodels -- node "$(pwd)/dist/index.js"Then ask Claude: "use the list_models tool and show me the menu".
Configuring providers
config/models.json— which providers exist, their base URLs, and which models are enabled. Adding an OpenAI-compatible provider is one JSON block; enabling a model is one line in itsmodelsarray..env— API keys only. Never in models.json, never in code. The server reads models.json fresh on every call (edit and it applies immediately);.envis read at startup (restart the server after adding a key).Local control panel —
npm run panelopens a localhost page (http://127.0.0.1:4747) to manage keys and toggle models. Keys are shown last-4-only; the panel binds to localhost.Codex lane — needs the Codex CLI installed and logged in. It uses whatever model your
~/.codex/config.tomlsets (the CLI accepts-m gpt-5.6-lunaetc.).Reasoning effort, per model — an OpenAI-compatible provider that declares
effortStyle("openai"→ top-levelreasoning_effort;"openrouter"→reasoning: { effort }) can also declareeffortOptions(the vendor's own level names, which the panel offers as-is) anddefaultEffortByModel(a{ model: level }map you edit from the panel). Precedence when building the request:effortpassed todelegate_task→defaultEffortByModel[model]→ the provider'sdefaultEffort→ nothing sent, so the vendor's own default applies. Shipped levels: z.aihigh/max(vendor defaultmax), OpenRouterlow/medium/high, DeepSeeklow/high/max(vendor defaulthigh). Aclaude-cli("with hands") lane joins the same mechanism by declaringeffortOptions— same cascade, same panel selector, same{ model: level }map — except the level is passed as--effort <level>on the CLI instead of in a request body. Who gets effort control is declarative, not hardcoded: an OpenAI-compatible provider needseffortStyle, a with-hands lane needseffortOptions. Everything else (LM Studio, Codex, Gemini, the vendor with-hands lanes) has no effort control — the panel shows no selector, and asking for one returns a friendly error.DeepSeek
lowcaveat — ondeepseek-v4-prothe vendor currently treatslowashigh, so pickinglowthere changes the bill, not the thinking. DeepSeek expects this to change in August 2026;deepseek-v4-flashhonours all three levels.Effort on the vendor "with hands" lanes is not controllable — tested, don't retry. The reason is the engine on the other end, not the protocol: DeepSeek documents that it discards the field, and states it auto-raises effort to max for agentic clients anyway, so the lane already runs at the level you'd pick. Measured against the live endpoint (2026-07-31): sending
reasoning_effortin the body changed nothing (low→ 1140/916 output tokens,max→ 393/863), andbudget_tokenswas overrun and undershot at will (500 → 773, 12000 → 549). The real lever on these lanes is the model id (provsflash). (Correction, 0.11.0: an earlier version of this note claimed "the Anthropic wire protocol has no effort concept, only a thinking budget". That was wrong — effort is part of the Anthropic API (output_config.effort) and the CLI exposes--effort. The measured conclusion held; the explanation didn't, and the wrong explanation is exactly what would stop anyone from trying it on the Anthropic lane.)z.ai gotcha — coding-plan subscription keys only work on the coding endpoint (
https://api.z.ai/api/coding/paas/v4). On the generic endpoint they fail with a misleading "insufficient balance". The default config already points at the right one."With hands" lanes — need the Claude Code CLI on your PATH; each lane is a
claude-cliprovider with write tools blocked. The convention: aclaude-clilane withbaseUrl+envKeyis a third-party engine (that address, that key, run under a throwawayCLAUDE_CONFIG_DIR); a lane with neither is a subscription lane, run under your real Claude Code login with no key at all. One engine, two configurations. Shipped vendor lanes:glm-maos:glm-5.2(z.ai,ZAI_API_KEY),deepseek-maos:deepseek-v4-pro/deepseek-maos:deepseek-v4-flash(https://api.deepseek.com/anthropic,DEEPSEEK_API_KEY, pay-per-use) andkimi-maos:kimi-k3(https://api.moonshot.ai/anthropic,MOONSHOT_API_KEY). The panel shows a key field on these cards (last-4-only, as always).claude-maos— Anthropic's own models, on your subscription. Ids:claude-maos:claude-fable-5,claude-maos:claude-opus-5,claude-maos:claude-opus-4-8,claude-maos:claude-sonnet-5. Identical hands to every other lane (Read/Glob/Grep +npm test/npm run build, no Edit/Write). No key, so no key field in the panel. Three things worth knowing before you reach for it:This lane does take reasoning effort — measured, and it works. Levels:
low,medium,high,xhigh,max(the five theclaudeCLI accepts), passed through as--effort <level>, picked per call viadelegate_task'seffortor as a per-model default in the panel. Measured on this machine (2026-08-01,claude-sonnet-5, same short reasoning puzzle,--output-format json):low→ 168 output tokens, 6.0 s, $0.2629;max→ 1321 output tokens, 16.3 s, $0.2803. So ~7.9× the output tokens and ~2.7× the wall clock for ~6.6% more cost — the cost barely moves because on this lane the ~43k tokens of inherited global configuration dominate the bill, not the thinking. Both answers were correct;maxwas slightly better structured. NodefaultEffortis shipped: with nothing picked,--effortisn't sent at all and the CLI's own default applies.A throwaway config dir does NOT authenticate a subscription — measured, don't "fix" this. With
CLAUDE_CONFIG_DIRpointed at an empty temp folder and no key, the CLI returns{"is_error":true,...,"result":"Not logged in · Please run /login"}(2026-08-01, this machine). That is why this lane deliberately runs against the real~/.claude. The consequence: it inherits your global user configuration, roughly 31k tokens of baggage per delegation before the task is even read. Rewiring it to a disposable identity breaks the lane outright.Safety lock. Because it runs on the real config, the child process is spawned with
ANTHROPIC_API_KEY,ANTHROPIC_AUTH_TOKENandANTHROPIC_BASE_URLdeleted from the inherited environment. Any one of them surviving would silently divert the delegation to metered API billing (or to another vendor's engine) instead of the subscription. And "free of API charge" still means it eats the same subscription allowance your main session is spending.
The manufacturer rule — a lane is hidden from the host that made it. This server exists to cross the border: call GPT/Gemini/GLM from inside Claude Code, or call Claude from inside Codex. Delegating to a lane from the same vendor as the calling program is an expensive detour — it spawns a second agent process that reloads the whole global configuration (measured: ~31k tokens per delegation on the subscription lane) and eats the same subscription allowance the session is already spending. The native subagent does it cheaper. So the rule is code, not discipline:
The host identifies itself in the MCP handshake (
getClientVersion(), read at call time — it does not exist yet at startup). Its name is matched by substring, case-insensitively, because the exact name drifts between versions:claude→anthropic,codex→openai,gemini→google.A provider declaring
"fabricante": "<vendor>"inconfig/models.jsonis dropped fromlist_modelsand refused bydelegate_task(the id can be typed by hand, so both doors are locked) when it matches the host's vendor. Shipped marks:codex→openai,gemini→google,claude-maos→anthropic. A provider without the field is never hidden — including the with-hands lanes of other engines (glm-maos,deepseek-maos,kimi-maos), which merely use Claude Code as a chassis while another vendor's model answers.Never silent: when something is omitted, the menu ends with a line naming what was dropped and why. Fails open: an unknown or unidentified host hides nothing.
Escape hatch —
MULTIMODELS_ANFITRIAOoverrides the detection:nenhum(ornone) disables the rule entirely, any other value is forced as the host's vendor. It's a cost optimisation, not a security boundary, so a wrong guess must never block work.The first tool call of each session logs one line to stderr (stdout belongs to the protocol) with the client name, version and deduced vendor — that's how you verify detection against a real host.
Moonshot key — the Kimi-with-hands lane needs an official Moonshot key from platform.kimi.ai, stored as
MOONSHOT_API_KEYin.env. It is not the OpenRouter key used by the text-onlyopenrouter:moonshotai/kimi-k3lane.DeepSeek model names — tested, no gotcha. DeepSeek's docs suggest an unrecognized model name is silently downgraded to
deepseek-v4-flash. Verified against the live API (2026-07-31): a bogus name returns400 The supported API model names are deepseek-v4-pro or deepseek-v4-flash, and each valid name is served by that exact model (checked via the run's usage report, not the model's self-report). Both names stay enabled.Cost figures printed by the Claude Code CLI are wrong for these lanes — they're computed with Anthropic's price table, not the vendor's. Read the real bill on the vendor's dashboard.
The benchmark: who can you actually trust with delegated work?
The benchmark/ folder contains a full evaluation run through this server: 6 stations × 11 models × 3 rounds = 198 runs, graded by hidden test suites written before any model saw the tasks. Stations: build-from-spec, find-and-fix-a-bug, code review with seeded bugs, strict JSON extraction, a long compound deliverable, and honesty under missing context.

Highlights:
The GPT-5.6 Codex family (including Luna at $1/M input) went 54/54 perfect runs, and verified 9/9 times that a phantom file didn't exist instead of hallucinating a fix.
Sonnet 5 and Haiku 4.5 failed the same cent-distribution contract in 2 of 3 rounds each — while every cheap delegate passed 9/9.
Strict JSON extraction: 33/33 across all models. Solved problem.
Single-run benchmarks lied in both directions; three rounds changed half the conclusions.
Everything needed to reproduce is in the folder: station prompts (benchmark/estacoes/, in Portuguese), automated graders (benchmark/corretores/), and every raw response (benchmark/respostas/).

Round 2 — a real task instead of synthetic stations
Seven implementers (Claude, GPT-5.6 and GLM lanes, agentic and text-only) built the same real feature of this very server, each on an isolated git branch, judged by 12 hidden acceptance checks: benchmark/rodada2-implementacao/. Sonnet 5 won on fine-grained review; the text-only lanes revealed their two blind spots (context and verification).
Round 3 — the knowledge-cutoff round
Designed by the Reddit comment section: 13 lanes × 2 stations × 3 rounds, with reasoning effort controlled and a station built against the actually installed zod v4: benchmark/rodada3-esforco-e-cutoff/. The cheap models didn't fail at reasoning — they failed at knowing what year it is (0/14 nine-for-nine on the trap, 18/18 on pure reasoning). Only two defenses exist: file access, or fresh training data.
Round 4 (partial) — the newcomers
Two requested lanes on the same two stations: benchmark/rodada4-raias-novas/. Kimi K3 (text-only, via OpenRouter) ran and became the second text-only lane ever to beat the cutoff trap — 14/14 on the zod v4 station from memory alone, joining Grok 4.5 in the "fresh memory" club. It went 5 of 6 perfect; the one blemish is a systematic failure mode (it reasons to completion but never emits the final answer, 3× identically on the same cell). It's the slowest and heaviest reasoner in the study — 6-12 min per task, ~$0.20 per delivered task ($3/$15 per M). The two Gemini lanes (3.1 Pro high and 3.6 Flash high) are pending — the Google subscription quota ran out; that window resets ~Jul 29.

There's also an interactive decision report (in Portuguese) consolidating all rounds: benchmark/relatorio-decisao.html.
Repo notes
This project is built entirely through vibecoding, in Portuguese. The originals stay in Portuguese as part of how it's made, and every document has an English version: CLAUDE.en.md (working instructions), CHANGELOG.en.md (project diary), benchmark/README.md (benchmark guide) and benchmark/estacoes/en/ (station prompts).
The benchmark ran with the Portuguese prompts; the raw model responses in
benchmark/respostas/are untranslated on purpose — they're the evidence. The graders are language-independent.Tests:
npm test.
License
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- FlicenseCqualityFmaintenanceAllows Claude Code to offload AI coding tasks to Aider, reducing costs and enabling more control over which models handle specific coding tasks.2305
- -license-quality-maintenanceGives Claude access to multiple AI models (Gemini, OpenAI, OpenRouter, Ollama) for enhanced development capabilities including extended reasoning, collaborative development, code review, and advanced debugging.
- Alicense-qualityDmaintenanceLets Claude Code query multiple AI models (Gemini, Grok, ChatGPT, DeepSeek) for diverse perspectives, code reviews, debates, and more.MIT
- AlicenseAqualityDmaintenanceEnables Claude Code to delegate tasks to OpenAI's Codex CLI (GPT-5.4) with structured execution traces, parallel execution, session persistence, and adversarial code review.15MIT
Related MCP Connectors
Real-time chat hub for AI agents — Claude Code, Cursor, Cline, Codex over MCP or REST.
Stop copy-pasting between Claude Chat and Claude Code.
Coding agents from Claude Code, Cursor and Codex claim jobs and lock files on one shared board.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/dpmadsen/multimodels-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server