Skip to main content
Glama

Cheapest-LLM Router (CLR)

Route any prompt to the cheapest reachable free/cheap LLM — automatically. Reuses the free-model channels of Free & Cheap Tokens (FACT) (Kimi K2.6 · Qwen · DeepSeek · Cloudflare Workers AI · Groq · Gemini …). A zero-dependency MCP server (runs over stdio, no npm install) with 4 tools: route · cost_compare · list_models · cache_route.

License: MIT Node MCP Pay-Per-Event

🌐 English · 简体中文 · 繁體中文

It lives at

Link

MCP endpoint (hosted)

https://neeenja--cheapest-llm-router.apify.actor/mcp

Apify Store

https://apify.com/neeenja/cheapest-llm-router

Source

https://github.com/PanStories/cheapest-llm-router


English

Cheapest-LLM Router is an MCP server that, given a prompt, automatically picks the cheapest model that can actually serve it — free models first, then the lowest-cost paid fallback — and returns the cost in both USD and CNY plus a fallback chain and plain-English reasoning. It does not call any model; it only does the routing math, so it costs nothing to run and nothing to call except the tiny Apify Pay-Per-Event fee when hosted.

Why this exists

  • Zero new learning — it reuses the curated model list and free channels already built for FACT.

  • Zero marginal cost — routing is pure local computation; no paid API is ever called.

  • 100% brand synergy — it complements FACT's "free/cheap tokens" positioning and cross-sells to the same users.

  • Monetization — advanced routing / caching / the cost-comparison report layer is the hosted paid tier.

  • Cold-start revenue estimate — ¥0–¥500/month via FACT-user conversion.

Related MCP server: token-scout

What you get (MCP tools)

Tool

What it does

route

prompt → cheapest reachable model + USD/CNY cost estimate + fallback chain + reasoning

cost_compare

ranks every reachable model by cost and reports the max savings vs the most expensive

list_models

filtered registry listing (capability / region / free-only)

cache_route

same as route, but demonstrates the in-process cache (cached: true on repeat hits)

route input

{
  "prompt": "Analyze the core risks of this earnings report",
  "max_output_tokens": 512,
  "required_capabilities": ["chinese", "reasoning"],
  "region": "CN",
  "priority": "cost",
  "include_paid": true
}

route output (excerpt)

{
  "ok": true,
  "chosen": { "name": "Qwen3-Plus (via Aliyun Bailian)", "free": true, "cost_usd": 0 },
  "fallback_chain": [ "Kimi K2 (Moonshot direct trial)", "DeepSeek-V3 (via SiliconFlow)" ],
  "estimated_cost_usd": 0,
  "estimated_cost_cny": 0,
  "reasoning": "Routed by lowest cost. Estimated 18 input + 512 output tokens. Cheapest reachable is Qwen3-Plus ... free. No token cost."
}

Quick start

1. Run locally (stdio, zero install)

git clone https://github.com/PanStories/cheapest-llm-router.git
cd cheapest-llm-router
node src/server.mjs          # plain node — no npm install needed

2. Verify (no dependencies)

node --test tests/                     # routing core unit tests
node scripts/mcp-smoke.mjs             # end-to-end MCP protocol smoke test
node scripts/http-e2e-test.mjs         # Streamable HTTP transport test (needs deps)
node scripts/showcase.mjs              # generate showcase.html from real routing output

3. Host remotely (Apify Pay-Per-Event)

npm install                            # pulls express + MCP SDK + apify (hosted variant only)
apify login && apify push              # requires Apify KYC first

The hosted MCP endpoint is then reachable at https://neeenja--cheapest-llm-router.apify.actor/mcp (Bearer-token auth).

Connect it to your client

stdio (mcp.json) — for local desktop clients:

{
  "mcpServers": {
    "cheapest-llm-router": {
      "command": "node",
      "args": ["/abs/path/cheapest-llm-router/src/server.mjs"]
    }
  }
}

hosted (remote, via mcp-remote) — Apify's gateway requires a per-request Bearer token:

{
  "mcpServers": {
    "cheapest-llm-router": {
      "command": "npx",
      "args": [
        "mcp-remote",
        "https://neeenja--cheapest-llm-router.apify.actor/mcp",
        "--header", "Authorization: Bearer <YOUR_APIFY_TOKEN>"
      ]
    }
  }
}

Pricing (Pay-Per-Event)

No subscription, no monthly fee. You pay only when a tool actually runs:

Event

Price (USD)

initialize / tools/list / report-issue

free

route

$0.0005

cost_compare

$0.001

list_models / cache_route

$0.0005

Free events are never billed, so agents can connect and discover tools at zero cost. Apify also grants ~$5/month of free platform credits (≈ thousands of calls).

How routing works

  1. Token estimate — CJK ≈ 1.6 tokens/char, other text ≈ 0.25 tokens/char (no external tokenizer).

  2. Reachability filter — required_capabilities ⊆ model capabilities and region match (CN = mainland-accessible without a VPN).

  3. Cost — free = $0; paid = (in×pin + out×pout) / 1e6.

  4. Ranking — cost (default: free first, then by free-quota then latency; paid by cost) · latency · quality.

  5. Fallback chain — the next 3 reachable models.

⚠️ Prices are indicative (USD per 1M tokens) and drift with providers; verify on each provider's pricing page before production.

What it does NOT do (honest boundaries)

  1. It does not send real inference requests — it only selects a route.

  2. It does not cache/store your prompts (a persistent per-account cache is a hosted paid feature).

  3. It does not guarantee free quotas are live (provider-controlled; may return 429).

  4. Prices change; always trust the provider's official page.

Development

Script

Purpose

node src/server.mjs

stdio MCP server (zero dependency)

node src/http.mjs

HTTP MCP server (PORT=3000) for local/dev

node --test tests/

routing core unit tests

node scripts/mcp-smoke.mjs

stdio MCP protocol smoke test

node scripts/http-e2e-test.mjs

HTTP transport e2e (PORT=3100)

node scripts/showcase.mjs

build showcase.html from real output

Project layout:

cheapest-llm-router/
├── data/models.json        curated model registry (reuses FACT's verified list)
├── src/core/router.mjs     routing engine (zero-dependency, single source of truth)
├── src/server.mjs          zero-dependency stdio MCP server (default, verified)
├── src/http.mjs            self-managed HTTP MCP server (Apify Standby)
├── src/handler.mjs         MCP server factory (shared by both transports)
├── src/billing.mjs         Apify Pay-Per-Event billing (hosted only)
├── tests/router.test.mjs   unit tests (node --test)
├── scripts/                smoke / http-e2e / showcase scripts
└── .actor/                 Apify actor.json + Dockerfile + schemas

License

MIT — fork, self-host, self-modify freely.


简体中文

Cheapest-LLM Router 是一个 MCP 服务器:给定一条 prompt,它会自动选出当下最便宜且能真正服务它的模型——优先免费模型,其次最低成本的付费兜底——并以美元和人民币双币种返回成本、兜底链与推理说明。它不会真正调用任何模型,只做路由计算,因此本地运行零成本,托管后除极低的 Apify 按事件计费外也无其他开销。

为什么做这个

  • 零新学习——直接复用为 FACT 策展的模型清单与免费通道。

  • 零边际成本——路由是纯本地计算,从不调用任何付费 API。

  • 100% 品牌协同——与 FACT「免费/廉价 token」定位天然互补,可向同一批用户交叉转化。

  • 变现点——高级路由 / 缓存 / 成本对比报表层即托管的付费能力。

  • 冷启动月收入预估——¥0–¥500(靠 FACT 用户转化)。

你得到什么(MCP 工具)

工具

作用

route

给定 prompt → 最便宜可达模型 + 美元/人民币成本估算 + 兜底链 + 推理说明

cost_compare

把所有可达模型按成本排序,并给出相比最贵模型的最大节省额

list_models

按 capability / region / free-only 过滤的模型清单

cache_route

同 route,但演示进程内缓存(重复命中返回 cached: true)

route 入参

{
  "prompt": "分析这份财报的核心风险",
  "max_output_tokens": 512,
  "required_capabilities": ["chinese", "reasoning"],
  "region": "CN",
  "priority": "cost",
  "include_paid": true
}

route 出参(节选)

{
  "ok": true,
  "chosen": { "name": "Qwen3-Plus (via Aliyun Bailian)", "free": true, "cost_usd": 0 },
  "fallback_chain": [ "Kimi K2 (Moonshot direct trial)", "DeepSeek-V3 (via SiliconFlow)" ],
  "estimated_cost_usd": 0,
  "estimated_cost_cny": 0,
  "reasoning": "Routed by lowest cost. Estimated 18 input + 512 output tokens. Cheapest reachable is Qwen3-Plus ... free. No token cost."
}

快速开始

1. 本地跑(stdio,零安装)

git clone https://github.com/PanStories/cheapest-llm-router.git
cd cheapest-llm-router
node src/server.mjs          # 直接跑,无需 npm install

2. 验证(无需依赖)

node --test tests/                     # 路由核心单元测试
node scripts/mcp-smoke.mjs             # MCP 协议端到端冒烟测试
node scripts/http-e2e-test.mjs         # Streamable HTTP 传输测试(需装依赖)
node scripts/showcase.mjs              # 用真实路由输出生成 showcase.html

3. 远程托管(Apify 按事件计费)

npm install                            # 仅托管变体需要:express + MCP SDK + apify
apify login && apify push              # 需先完成 Apify KYC

托管后的 MCP 端点:https://neeenja--cheapest-llm-router.apify.actor/mcp(Bearer token 鉴权)。

接入你的客户端

stdio(mcp.json)——本地桌面客户端:

{
  "mcpServers": {
    "cheapest-llm-router": {
      "command": "node",
      "args": ["/绝对路径/cheapest-llm-router/src/server.mjs"]
    }
  }
}

托管(远程,借助 mcp-remote)——Apify 网关要求每次请求带 Bearer token:

{
  "mcpServers": {
    "cheapest-llm-router": {
      "command": "npx",
      "args": [
        "mcp-remote",
        "https://neeenja--cheapest-llm-router.apify.actor/mcp",
        "--header", "Authorization: Bearer <你的_APIFY_TOKEN>"
      ]
    }
  }
}

定价(按事件计费)

无订阅、无月费,只在工具真正运行时付费:

事件

价格(美元)

initialize / tools/list / report-issue

免费

route

$0.0005

cost_compare

$0.001

list_models / cache_route

$0.0005

免费事件永不计费,Agent 可零成本连接与发现工具。Apify 另送约 $5/月的免费平台额度(≈ 数千次调用)。

路由逻辑

  1. Token 估算:CJK 字符 ≈ 1.6 token,其他文本 ≈ 0.25 token(无外部 tokenizer)。

  2. 可达性过滤:required_capabilities ⊆ 模型能力 且 region 匹配(CN = 大陆免 VPN 可达)。

  3. 成本计算:免费模型 = $0;付费 = (in×pin + out×pout) / 1e6。

  4. 排序:cost(默认:免费优先 → 免费内按免费额度再按延迟;付费按成本升序)· latency · quality。

  5. 兜底链:次优 3 个可达模型。

⚠️ 价格为指示性(美元/百万 token),随厂商变动;生产前请以各厂商定价页为准。

本工具不做什么(诚实边界)

  1. 不替你发起真实推理请求——只选路由。

  2. 不缓存/存储你的 prompt(hosted 持久缓存为付费能力)。

  3. 不保证免费额度实时可用(额度由厂商控制,可能 429)。

  4. 价格随厂商变动,请以官方为准。

开发

脚本

用途

node src/server.mjs

stdio MCP 服务器(零依赖)

node src/http.mjs

HTTP MCP 服务器(PORT=3000),本地/开发用

node --test tests/

路由核心单元测试

node scripts/mcp-smoke.mjs

stdio MCP 协议冒烟测试

node scripts/http-e2e-test.mjs

HTTP 传输端到端(PORT=3100)

node scripts/showcase.mjs

用真实输出生成 showcase.html

项目结构:

cheapest-llm-router/
├── data/models.json        策展模型注册表(复用 FACT 已核验清单)
├── src/core/router.mjs     路由引擎(零依赖,唯一事实源)
├── src/server.mjs          零依赖 stdio MCP 服务器(默认、已验证)
├── src/http.mjs            自托管 HTTP MCP 服务器(Apify Standby)
├── src/handler.mjs         MCP 服务器工厂(两个传输共用)
├── src/billing.mjs         Apify 按事件计费(仅托管)
├── tests/router.test.mjs   单元测试(node --test)
├── scripts/                冒烟 / HTTP e2e / 展示脚本
└── .actor/                 Apify actor.json + Dockerfile + schema

许可证

MIT——可 fork、自部署、自托管、自修改。


繁體中文

Cheapest-LLM Router 是一個 MCP 伺服器:給定一條 prompt,它會自動選出當下最便宜且能真正服務它的模型——優先免費模型,其次最低成本的付費兜底——並以美元與人民幣雙幣種回傳成本、兜底鏈與推理說明。它不會真正呼叫任何模型,只做路由計算,因此本地執行零成本,託管後除極低的 Apify 按事件計費外也無其他開銷。

為什麼做這個

  • 零新學習——直接複用為 FACT 策展的模型清單與免費通道。

  • 零邊際成本——路由是純本地計算,從不呼叫任何付費 API。

  • 100% 品牌協同——與 FACT「免費/廉價 token」定位天然互補,可向同一批用戶交叉轉化。

  • 變現點——進階路由 / 快取 / 成本對比報表層即託管的付費能力。

  • 冷啟動月收入預估——¥0–¥500(靠 FACT 用戶轉化)。

你得到什麼(MCP 工具)

工具

作用

route

給定 prompt → 最便宜可達模型 + 美元/人民幣成本估算 + 兜底鏈 + 推理說明

cost_compare

把所有可達模型按成本排序,並給出相比最貴模型的最大節省額

list_models

按 capability / region / free-only 過濾的模型清單

cache_route

同 route,但示範程序內快取(重複命中回傳 cached: true)

route 入參

{
  "prompt": "分析這份財報的核心風險",
  "max_output_tokens": 512,
  "required_capabilities": ["chinese", "reasoning"],
  "region": "CN",
  "priority": "cost",
  "include_paid": true
}

route 出參(節選)

{
  "ok": true,
  "chosen": { "name": "Qwen3-Plus (via Aliyun Bailian)", "free": true, "cost_usd": 0 },
  "fallback_chain": [ "Kimi K2 (Moonshot direct trial)", "DeepSeek-V3 (via SiliconFlow)" ],
  "estimated_cost_usd": 0,
  "estimated_cost_cny": 0,
  "reasoning": "Routed by lowest cost. Estimated 18 input + 512 output tokens. Cheapest reachable is Qwen3-Plus ... free. No token cost."
}

快速開始

1. 本地執行(stdio,零安裝)

git clone https://github.com/PanStories/cheapest-llm-router.git
cd cheapest-llm-router
node src/server.mjs          # 直接執行,無需 npm install

2. 驗證(無需依賴)

node --test tests/                     # 路由核心單元測試
node scripts/mcp-smoke.mjs             # MCP 協定端到端冒煙測試
node scripts/http-e2e-test.mjs         # Streamable HTTP 傳輸測試(需裝依賴)
node scripts/showcase.mjs              # 用真實路由輸出生成 showcase.html

3. 遠端託管(Apify 按事件計費)

npm install                            # 僅託管變體需要:express + MCP SDK + apify
apify login && apify push              # 需先完成 Apify KYC

託管後的 MCP 端點:https://neeenja--cheapest-llm-router.apify.actor/mcp(Bearer token 鑑權)。

接入你的客戶端

stdio(mcp.json)——本地桌面客戶端:

{
  "mcpServers": {
    "cheapest-llm-router": {
      "command": "node",
      "args": ["/絕對路徑/cheapest-llm-router/src/server.mjs"]
    }
  }
}

託管(遠端,借助 mcp-remote)——Apify 閘道要求每次請求帶 Bearer token:

{
  "mcpServers": {
    "cheapest-llm-router": {
      "command": "npx",
      "args": [
        "mcp-remote",
        "https://neeenja--cheapest-llm-router.apify.actor/mcp",
        "--header", "Authorization: Bearer <你的_APIFY_TOKEN>"
      ]
    }
  }
}

定價(按事件計費)

無訂閱、無月費,只在工具真正執行時付費:

事件

價格(美元)

initialize / tools/list / report-issue

免費

route

$0.0005

cost_compare

$0.001

list_models / cache_route

$0.0005

免費事件永不计費,Agent 可零成本連線與發現工具。Apify 另送約 $5/月的免費平台額度(≈ 數千次呼叫)。

路由邏輯

  1. Token 估算:CJK 字元 ≈ 1.6 token,其他文字 ≈ 0.25 token(無外部 tokenizer)。

  2. 可達性過濾:required_capabilities ⊆ 模型能力 且 region 匹配(CN = 大陸免 VPN 可達)。

  3. 成本計算:免費模型 = $0;付費 = (in×pin + out×pout) / 1e6。

  4. 排序:cost(預設:免費優先 → 免費內按免費額度再按延遲;付費按成本升序)· latency · quality。

  5. 兜底鏈:次優 3 個可達模型。

⚠️ 價格為指示性(美元/百萬 token),隨廠商變動;生產前請以各廠商定價頁為準。

本工具不做什么(誠實邊界)

  1. 不替你發起真實推理請求——只選路由。

  2. 不快取/儲存你的 prompt(託管持久快取為付費能力)。

  3. 不保證免費額度即時可用(額度由廠商控制,可能 429)。

  4. 價格隨廠商變動,請以官方為準。

開發

腳本

用途

node src/server.mjs

stdio MCP 伺服器(零依賴)

node src/http.mjs

HTTP MCP 伺服器(PORT=3000),本地/開發用

node --test tests/

路由核心單元測試

node scripts/mcp-smoke.mjs

stdio MCP 協定冒煙測試

node scripts/http-e2e-test.mjs

HTTP 傳輸端到端(PORT=3100)

node scripts/showcase.mjs

用真實輸出生成 showcase.html

專案結構:

cheapest-llm-router/
├── data/models.json        策展模型註冊表(複用 FACT 已核驗清單)
├── src/core/router.mjs     路由引擎(零依賴,唯一事實源)
├── src/server.mjs          零依賴 stdio MCP 伺服器(預設、已驗證)
├── src/http.mjs            自託管 HTTP MCP 伺服器(Apify Standby)
├── src/handler.mjs         MCP 伺服器工廠(兩個傳輸共用)
├── src/billing.mjs         Apify 按事件計費(僅託管)
├── tests/router.test.mjs   單元測試(node --test)
├── scripts/                冒煙 / HTTP e2e / 展示腳本
└── .actor/                 Apify actor.json + Dockerfile + schema

授權

MIT——可 fork、自部署、自託管、自修改。


Support · 赞助 · 贊助

EN — Cheapest-LLM Router is open source (MIT), ad-free. It is funded by the community, not by ads. If it powers your agents or workflow, please support it:

简体中文 — Cheapest-LLM Router 开源(MIT)、无广告,由社区资助而非广告。若它支撑了你的智能体或工作流,欢迎赞助:点本仓库的 Sponsor 按钮(跳转 Ko-fi)或前往 https://ko-fi.com/panstories

繁體中文 — Cheapest-LLM Router 開源(MIT)、無廣告,由社群資助而非廣告。若它支撐了你的智能體或工作流,歡迎贊助:點本倉庫的 Sponsor 按鈕(導向 Ko-fi)或前往 https://ko-fi.com/panstories

Thank you! · 谢谢 · 謝謝 💙

Available Tools

4 tools
cache_routeB

Same as route, but checks an in-process cache first. Demonstrates the advanced caching layer: repeated identical requests return the cached plan with cached=true. A persistent per-account cache is a hosted paid feature.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
regionNoglobal
priorityNocost
include_paidNo
max_output_tokensNo
required_capabilitiesNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose meaningful traits: in-process (not persistent) caching, cache-hit responses marked cached=true, and that persistent per-account caching is a paid feature. It does not say whether the call has side effects, what auth it needs, or how cache scope/eviction works.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core differentiator (cache-first) front-loaded and no filler. The middle clause about demonstrating the caching layer is marginally self-referential but still informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter tool with no annotations and no output schema, the description covers the caching behavior and one return field (cached=true) but omits any parameter meaning and any note on the interaction between include_paid and the paid persistent cache. Adequate on behavior, thin on inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across six parameters (prompt, region, priority, include_paid, max_output_tokens, required_capabilities), and the description explains none of them. Only 'identical requests' loosely hints that prompt is the cache key, leaving the remaining parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool returns a routing plan and frames it relative to the sibling 'route' ('Same as route, but checks an in-process cache first'), which lets an agent distinguish the two. The gap is that it assumes the agent already knows what 'route' does rather than restating the core verb+resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Repeated identical requests return the cached plan' implies when this tool pays off versus plain route, but there is no explicit when-to-use/when-not statement or named alternative. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cost_compareB

Rank every reachable model by estimated cost for a given prompt size. Produces a cost-comparison report (the monetization "report layer") including max savings vs the most expensive reachable model.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe task prompt to size.
regionNoglobal
max_output_tokensNo
required_capabilitiesNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, yet it usefully discloses that costs are estimates (not live/actual billing), that the result is a report layer, and that it covers all 'reachable' models. It omits whether the call is side-effect-free, whether it hits a live pricing API, or any rate/auth constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded. The parenthetical '(the monetization "report layer")' is internal jargon that does little work for an agent, a minor deduction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly sketches the return (cost report including max savings vs the most expensive reachable model). However, half its parameters are undocumented anywhere, leaving an agent guessing about region and capability filtering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% — only 'prompt' is documented, while region (with its global/CN enum), max_output_tokens, and required_capabilities are bare. The description only echoes 'prompt size' and adds no semantics for the three undocumented parameters, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Rank every reachable model by estimated cost') plus the output artifact (cost-comparison report with max savings). This clearly distinguishes it from the sibling list_models (enumerate only) and cache_route/route (actual routing), though it never names those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a pre-routing comparison use case but never states when to call this versus route, cache_route, or list_models, nor any exclusions or prerequisites. No explicit 'use this when...' guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA

List the curated model registry with optional filters (capability / region / free-only).

ParametersJSON Schema
NameRequiredDescriptionDefault
regionNoFilter by region.
free_onlyNoOnly free-tier models.
capabilityNoFilter by capability, e.g. "vision".

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. 'List' implies a safe read, and 'curated' hints that results are a maintained subset rather than a raw feed, which is a small useful signal. However, it says nothing about result count, ordering, pagination, or whether any filtering requires auth, leaving real behavioral gaps for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler, and the primary action plus the filtering capability are front-loaded. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-optional-parameter listing tool with no output schema, the description covers purpose and filters adequately. It omits what the listing contains (fields returned) and any ordering or pagination behavior, which a caller in a routing workflow would plausibly want to know, so it is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each of the three filters is documented in the schema with enums and defaults, so the schema does the heavy lifting. The description names the same three filters, adding no syntax, format, or combination semantics beyond what the schema already states; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the curated model registry') and enumerates the three filter axes, so the agent knows exactly what the tool returns and how it can be narrowed. It does not explicitly differentiate itself from the routing siblings, but those names (route, cache_route, cost_compare) are operationally distinct enough that confusion is unlikely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'optional filters' implies the tool works with or without arguments, which is useful, but there is no explicit when-to-use guidance or mention of alternatives (e.g. routing through a model vs. enumerating candidates). Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

routeA

Given a prompt, route it to the cheapest reachable free/cheap LLM. Returns the chosen model, estimated cost (USD + CNY), a fallback chain, and reasoning. Reuses Free & Cheap Tokens model channels (Kimi K2.6, Qwen, DeepSeek, Cloudflare Workers AI, Groq, Gemini, etc.).

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe task prompt to route. Token estimate is derived from it.
regionNoCN = mainland-accessible without VPN.global
priorityNoRouting objective.cost
include_paidNoAllow paid models as fallback when no free model fits.
max_output_tokensNoExpected output tokens for cost estimation.
required_capabilitiesNoCapabilities the model must support, e.g. ["chinese","reasoning"].

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It adds real value by disclosing the return payload (chosen model, USD+CNY cost estimate, fallback chain, reasoning) and the underlying channel pool, but it never resolves whether the tool actually dispatches the prompt or only returns a routing recommendation, and says nothing about auth or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and followed by the payoff and source pool. The trailing channel list is slightly decorative but does convey scope; nothing is bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, no output schema, and no annotations, the description does the important extra work of naming the return fields and cost units. The remaining gap is the execute-vs-recommend ambiguity, which an agent needs in order to use the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already explains region, priority, include_paid, max_output_tokens, and required_capabilities with defaults and enums. The description adds no parameter-level detail beyond 'prompt', so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'route [a prompt] to the cheapest reachable free/cheap LLM', and enumerates what comes back (model, cost, fallback chain, reasoning). This is clearly distinguishable from cost_compare and list_models on intent, though it never names a sibling to sharpen the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied rather than stated: an agent can infer 'use me when you want a cheap model selected for a prompt', and the sample model channels suggest scope. There is no explicit when-not condition and no mention of cache_route or cost_compare as alternatives for similar needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcache_route
    • First observedcost_compare
    • First observedlist_models
    • First observedroute

TDQS

A3.5/5.0

Scored across 4 tools

Disambiguation4/5

route and cache_route overlap significantly—cache_route is explicitly described as 'same as route' with caching—so an agent may hesitate between them. cost_compare and list_models are clearly distinct, and cost_compare vs route is mostly clear (ranking vs selection).

Naming Consistency4/5

All tools use snake_case consistently. cost_compare, list_models, and cache_route follow a verb_noun pattern, while route is a single verb, which is a minor deviation.

Tool Count5/5

Four tools is well-scoped for a router server: one for routing, one for cached routing, one for cost comparison, and one for listing models. Each earns its place without bloat.

Completeness4/5

The surface covers routing, cached routing, cost comparison, and model listing, which are the core operations. Minor gaps exist, such as no direct single-model cost lookup or cache management, but core workflows are supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers