llm-routing
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@llm-routingEstimate cost for routing 500 math queries on the wide ladder."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
LLM Routing: a measured benchmark, and the router it argues for
A cost-aware LLM routing service (LangGraph + MCP), and the 417-task benchmark that decides its policy.
Answer at the cheapest model that can be verified to have got it right, and escalate only when verification fails. Whether that beats simply paying for the best model is not a matter of opinion — it depends on the models you are choosing between, and this repository measures it on three real ladders.
The finding, in one sentence: cascade when the top rung is genuinely better and verification is cheap — no price-ratio threshold gets all three ladders right. The shipped router computes that verdict per ladder from the committed measurements, and refuses to answer for a ladder it has no data on.
What runs | A LangGraph state machine — |
What decides its policy | 417 tasks (MBPP+ code, MATH-500 level 5), 9 policies, 3 price ladders, all measured on real models: cost–quality frontiers, exact McNemar, paired bootstrap. |
Built with | Python 3.10–3.13 · LangGraph · MCP · Anthropic + DeepSeek APIs · pytest (268 tests) · GitHub Actions. The research core is pure standard library — no dependency can change a benchmark number. |
Evidence | 5,075 real model responses, committed. $8.51 spent. Every figure and table regenerates offline, with no API key, for $0.00. |
Quickstart
pip install -e ".[agent]"
python -m llm_routing.build_taskset
python -m router_agent.cli --demo # real model output, no API key, $0.00No account, no key, no money: the responses were bought once and committed, so the router replays genuine model output rather than simulating it.
python scripts/demo.py prints the three canonical traces — the cascade winning
at the cheap rung, the cascade paying twice, and the code case where
verification is exact and free. The first of them:
1. The cascade's win - verified at the cheap rung
--------------------------------------------------------------------------
query: Let f(x) = x^3 - 3x + 1. Find the sum of the squares of all real roots. [...]
classify domain=math, start=cheap, verifier=self_consistency
answer cheap (deepseek-v4-flash) answered
verify self_consistency -> ACCEPT, confidence=1.00
finalize done: verified
answered by deepseek-v4-flash
verified True (self_consistency)
cost $0.000315 backend $0.000000Those four lines are a walk through the graph below, which is parsed out of
router_agent/graph.py rather than drawn — the escalation edge loops back into
answer, and that loop is what makes this a cascade rather than a router.
Three independent draws from DeepSeek all gave the right answer, so the cascade accepted and never called Opus 5 — roughly 27x cheaper than routing straight to the top rung. When verification fails the cascade escalates and pays for both rungs. Whether that trade is worth making is what the rest of this repository measures.
The finding
Pre-registered comparison against simply always paying for the best model. Exact McNemar over paired outcomes, n=209 held-out tasks per ladder.
ladder | rungs | cascade | always-expensive | Δ acc | p | Δ cost/task |
| v4-flash → Opus 5 | 95.7% | 92.3% | +3.3% | 0.039 | −$0.00307 |
| Haiku 4.5 → Sonnet 5 → Opus 5 | 96.7% | 92.3% | +4.3% | 0.012 | +$0.00097 |
| v4-flash → v4-pro | 86.6% | 83.7% | +2.9% | 0.070 | −$0.00000 |
On wide the cascade is more accurate and four times cheaper. On claude
it buys the accuracy at a premium — verification is not free when the cheap
rung is Haiku and the maths half draws five samples from it. The ladder decides
the sign, which is why the router below reads it rather than assuming it.
Three further results, each with its numbers and its caveats in docs/RESULTS.md:
Predictive routing does not beat a coin flip — six comparisons out of six. Neither an LLM-as-router nor RouteLLM's pretrained BERT beats a cost-matched random null on any ladder, while the cascade beats both on every one. The distinction is when the decision is made: a predictive router commits before seeing an attempt, a cascade decides after verifying one. → the six comparisons, and the frontier AUC behind them
Accuracy hides what a router actually did. Two policies can reach the same accuracy by escalating the right ten tasks or by escalating everything.
always_expensiveescalates 201 tasks to buy 27 rescues, burning $0.71 on escalations that could not improve the answer;cascadegets 24 of those rescues and wastes $0.084. → the per-policy scorecardEvery policy is a curve, not a point. Each router here has a knob trading accuracy for money, so comparing two at one setting each lets whoever set the knobs pick the winner.
frontier.pysweeps each knob across its whole range and compares the resulting curves. → the frontiers, and why the price ratio does not decide it
Each of those has a figure in figures/, which lists what every chart
claims and which artefact in runs/ it was drawn from.
The benchmark ships its own conclusion
A benchmark that ends in a table leaves the reader to apply it. This one ends in
a function. findings.ratio_verdict(ladder) reads that ladder's committed
frontier and returns the verdict for it. Same query, two ladders, opposite
answers:
$ llm-router --estimate "prove that sqrt(2) is irrational" --ladder wide
recommended policy cascade (measured on the wide ladder)
cascade vs always-best, at matched accuracy -83.1%
$ llm-router --estimate "prove that sqrt(2) is irrational" --ladder claude
recommended policy route (measured on the claude ladder)
cascade vs always-best, at matched accuracy +11.7%That flip is the finding, and the router reads it rather than assuming it — and
declines for a ladder it has no data on. The CLI, the MCP explain_routing
tool and the RouterConfig defaults all call the same function, so changing what
the benchmark measured changes what the router recommends. There is no constant
to drift out of date — there used to be one, and two of its three verdicts were
backwards.
Layout
llm_routing/ the experiment — 16 modules, standard library only
router_agent/ the product — LangGraph cascade + MCP server
cache/ 5,075 real model responses — what makes replay free
runs/ every derived artefact: results, frontiers, scorecards
data/ docs/ figures/ scripts/ tests/ archive/The two halves share one model client, one price table and one response cache,
which is what makes a dollar figure from the router mean the same thing as a
dollar figure in the tables. The arrow runs one way — router_agent imports
llm_routing, never the reverse — and CI has a job whose only purpose is to
keep it that way. Module by module:
docs/ARCHITECTURE.md.
Use it from an MCP client
The router is an MCP server: five tools (route_query, resume_routing,
estimate_cost, compare_policies, explain_routing), four read-only
resources under routing://, and one prompt that walks a client through
choosing a policy.
A .mcp.json is committed, so Claude Code picks the server up on pip install -e ".[agent,mcp]" and nothing else. For Claude Desktop or any other client,
the same block registers it by hand:
{
"mcpServers": {
"llm-routing": {
"command": "python",
"args": ["-m", "router_agent.mcp_server"],
"env": {"ROUTER_LADDER": "wide", "ROUTER_MODE": "replay",
"ROUTER_K": "3", "ROUTER_AGREEMENT": "1.0"}
}
}
}ROUTER_MODE=replay is the safe registration: the server answers from the
committed responses and cannot spend money, at the price of only serving
prompts that were actually paid for — anything else comes back as a structured
no_cached_response rather than a fabricated answer. ROUTER_K=3 is pinned to
match the parameters those responses were bought under; the default of 5 would
ask the cache for samples nobody purchased. ROUTER_MODE=real with a key
serves arbitrary queries and bills them.
Approving an escalation
Unset by default. Add ROUTER_APPROVAL_USD and an escalation projected dearer
than it suspends the graph instead of spending: route_query returns
stop_reason: awaiting_approval with a thread_id and an interrupted
payload naming the model and the price, and resume_routing carries the
human's answer back in.
"env": {"ROUTER_LADDER": "wide", "ROUTER_MODE": "replay",
"ROUTER_K": "3", "ROUTER_AGREEMENT": "1.0",
"ROUTER_APPROVAL_USD": "0.001"}Approval is per escalation — the escalate node clears it on the way
through — so a three-rung ladder asks twice, and a client has to resume until
stop_reason is something else. The checkpoint is an InMemorySaver living in
the server process, so both calls must reach the same running server: a client
that spawns one per call, scripts/mcp_call.py included, can never resume what
the previous one paused. A thread_id that no longer exists comes back as
no_suspended_run rather than a KeyError from inside LangGraph.
Seeing the whole surface at once
python scripts/demo_mcp.pyA scripted walkthrough of the server over a real stdio client session — what it advertises and which of its own calls spend, a resource, the ladder flip, a free projection, a routed answer, and the approval loop answered both ways. No key, no spend; it ends by printing what the queries would have cost in production against what actually left the account.
It starts two servers, and the reason is the point of ROUTER_K:
self-consistency samples are cached per sample index, so k is pinned at
start-up to what the responses were bought under — k=3 for the query that
verifies at the cheap rung, k=4 for the one whose fourth draw disagrees and
triggers the escalation. demo.py shows what the router does; this shows what
the server does.
Driving it from a terminal
scripts/mcp_call.py is a one-shot MCP client — it starts the server, does the
handshake, calls one tool and prints the result:
python scripts/mcp_call.py --listpython scripts/mcp_call.py explain_routing ladder=widepython scripts/mcp_call.py --resource routing://findings/probePiping JSON-RPC in by hand does not work, and the failure is quiet: the server
takes stdin EOF as shutdown and exits without draining its queue, so
echo '...' | python -m router_agent.mcp_server prints the initialize reply,
drops the tool call and exits 0. A client holds the pipe open.
To spend real money, name the mode — this is a genuine DeepSeek call, routed and priced through MCP:
ROUTER_MODE=real ROUTER_LADDER=deepseek ROUTER_K=3 ROUTER_AGREEMENT=1.0 python scripts/mcp_call.py route_query query="What is 17 * 23? Give the final answer in \boxed{}." domain=math answered by deepseek-v4-flash (cheap)
verified True via self_consistency
cost $0.000068 backend $0.000068
classify domain=math, start=cheap, verifier=self_consistency
answer cheap (deepseek-v4-flash) answered
verify self_consistency -> ACCEPT, confidence=1.00Three HTTP calls — one greedy answer and two more to check it against itself —
accepted unanimously at the cheap rung, so v4-pro was never touched. Run it a
second time and backend_cost_usd is $0.00 while cost_usd is unchanged:
the responses were cached on the way out, which is the same mechanism that lets
the benchmark replay 5,075 of them for free. The two figures are separate on
purpose — one is what serving costs in production, the other is what left the
account.
What a served query buys lands in cache/serving.<ladder>.jsonl, not in the
benchmark's cache/raw_calls.<ladder>.jsonl. Both hold real paid responses, but
only one is evidence: the benchmark's file is the closed set every published
table is computed from, and letting an arbitrary query append to it would move
the response count and the total spend quoted below. Serving still reads the
benchmark cache, which is what makes --demo free.
Verifying it
python scripts/check_mcp_server.pyTwo phases, and the second is the one that matters. It lists and calls the
tools in-process, then launches the server as a subprocess and speaks JSON-RPC
to it by hand — because on stdio, stdout is the protocol, and a single
stray print beneath a tool corrupts the frame while every in-process test
still passes. That is not hypothetical: response_cache warned about stale
keys on stdout, on a code path only route_query reaches, so the server
listed its tools perfectly and then returned a mangled answer to the first
real call.
Reproducing everything
Replay mode reruns the published analysis against the committed responses — no key, no network, $0.00:
ROUTER_MODE=replay python scripts/run_all_ladders.py --ladders wide # ~30 minDrop --ladders wide for all three, about 75 minutes. The published figures
were produced exactly that way after deleting every derived artefact: 0 calls
reached a backend, 0 rows are simulated, and every regenerated file came back
byte-identical to the committed one.
Replay needs nothing installed at all — pure standard library, offline,
byte-deterministic down to the figures — and it is the default, so none of the
commands above name a mode. Real mode needs a key and spends money. There is a
third mode, mock, which fabricates responses for the test suite and which
every analysis module refuses to run in. All three, plus every analysis entry
point and the order to buy data in, are in
docs/METHOD.md.
Documentation
file | read it when |
you want the plain-language version, no familiarity with routing assumed — start here | |
you want every finding, with the numbers and what they cost | |
you want the method: task set, why these datasets, ladders, policies, verifiers, the degradation experiment, how to run it for real, and the bugs this project found in itself | |
you want to know how the benchmark and the serving layer fit together, module by module | |
you want what bounds the claims |
What bounds the claims, and what is new here
Stated on the front page rather than buried: the verifier that produces the signal is not the verifier that ships. The code half is graded by executing the tests MBPP+ supplies, and a deployed router does not have them.
That gap is priced rather than noted, and pricing it is what this repository
adds to the literature. FrugalGPT (2305.05176)
is the cascade baseline, and it and AutoMix take their verifier as given;
Dekoninck et al. (2410.10347) identify
quality-estimator accuracy as the factor deciding whether any of this works, but
test it by injecting synthetic noise. Here sweep_degraded.py instead
degrades a real verifier by a controlled amount on objectively-graded tasks,
holding domain, models, prompts and grader fixed — so shipping a proxy verifier
is a move along a measured curve rather than a step into the unknown.
Every other bound is stated once in docs/LIMITATIONS.md with what would settle it, and the full bibliography is in docs/METHOD.md.
License
MIT — see LICENSE.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
Agent Cost Allocator MCP — multi-tenant LLM cost attribution for chargeback billing. Companion to
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/APantov/llm-routing-comparison'
If you have feedback or need assistance with the MCP directory API, please join our Discord server