Cheapest-LLM Router
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| routeA | Given a prompt, route it to the cheapest reachable free/cheap LLM. Returns the chosen model, estimated cost (USD + CNY), a fallback chain, and reasoning. Reuses Free & Cheap Tokens model channels (Kimi K2.6, Qwen, DeepSeek, Cloudflare Workers AI, Groq, Gemini, etc.). |
| cost_compareB | Rank every reachable model by estimated cost for a given prompt size. Produces a cost-comparison report (the monetization "report layer") including max savings vs the most expensive reachable model. |
| list_modelsA | List the curated model registry with optional filters (capability / region / free-only). |
| cache_routeB | Same as route, but checks an in-process cache first. Demonstrates the advanced caching layer: repeated identical requests return the cached plan with cached=true. A persistent per-account cache is a hosted paid feature. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
route and cache_route overlap significantly—cache_route is explicitly described as 'same as route' with caching—so an agent may hesitate between them. cost_compare and list_models are clearly distinct, and cost_compare vs route is mostly clear (ranking vs selection).
All tools use snake_case consistently. cost_compare, list_models, and cache_route follow a verb_noun pattern, while route is a single verb, which is a minor deviation.
Four tools is well-scoped for a router server: one for routing, one for cached routing, one for cost comparison, and one for listing models. Each earns its place without bloat.
The surface covers routing, cached routing, cost comparison, and model listing, which are the core operations. Minor gaps exist, such as no direct single-model cost lookup or cache management, but core workflows are supported.