Skip to main content
Glama
alialtunar

model-price-radar

by alialtunar

model-price-radar-mcp

What does each LLM cost today, and did it get cheaper? An MCP server over OpenRouter's public model catalog (~450 models) plus monthly price history since 2023, rebuilt from Wayback Machine copies of the catalog. Ask Claude (or any MCP client) for the cheapest model that can do a job, what a workload will cost, or how a model's price moved.

No API keys. One line to install.

You:    Did LLM prices really go down? Show me GPT-4o and DeepSeek Chat.
Claude: [model_price_history ×2, price_changes]
        • GPT-4o: $5 / $15 per 1M tokens until Oct 2024, $2.50 / $10 since Nov 2024.
        • DeepSeek Chat changed price 12 times since mid-2024; input went from
          $0.14 to $0.26 per 1M.
        • Not everything gets cheaper: since Oct 2025 older open models went up
          as cheap hosts dropped them, e.g. Qwen 2.5 Coder 32B input $0.04 → $0.66.

Summarized from real tool output, 1 Oct 2026.

DeepSeek Chat price history by model-price-radar

Why

You want to know

Without it

With model-price-radar

Cheapest model with tools + vision + 200k context

Scroll pricing pages

find_models(needs=["tools","vision"], min_context=200000)

What 100k calls/month will cost on 5 models

Spreadsheet

estimate_cost ranks them with a "× cheapest" column

Did this model's price change?

Nobody keeps history

model_price_history since the model appeared

What moved in LLM pricing this quarter

Twitter

price_changes lists drops, increases, newly free models

Which provider serves Llama cheapest

Compare tabs

compare_providers with quantization and uptime

Related MCP server: tokenomics

How it works

flowchart LR
    C[Claude / MCP client] -->|tool call| S[model-price-radar-mcp]
    S --> L[OpenRouter /api/v1/models: live prices, context, capabilities]
    S --> E[OpenRouter /models/id/endpoints: per-provider prices]
    S --> H[Bundled monthly history 2023→build date]
    S --> W[Wayback Machine: months archived after the build]
    L & H & W --> C

Price history ships inside the package (18 KB, one snapshot per month since July 2023), so history answers are instant; months archived after the release are fetched from the Wayback Machine on demand and cached. Prices are US$ per 1M tokens; "blended" assumes 3 input tokens per output token.

Install

Requires uv.

Claude Code

claude mcp add model-price-radar -- uvx model-price-radar-mcp

Claude Desktop / Cursor (claude_desktop_config.json / .cursor/mcp.json)

{
  "mcpServers": {
    "model-price-radar": {
      "command": "uvx",
      "args": ["model-price-radar-mcp"]
    }
  }
}

Tools

Tool

What it does

find_models

Search the live catalog by name, capabilities (tools, vision, reasoning, structured outputs, audio, free), price ceilings, context

estimate_cost

Cost of a workload (tokens per call × calls) on up to 10 models, cheapest first

model_price_history

Every price change of one model since it appeared, with % changes

price_changes

Since a month: biggest drops and increases, newly free models, models added/removed

compare_providers

Same model across providers: price, context, max output, quantization, uptime

Model names are forgiving: "claude sonnet" resolves to the newest Sonnet, "llama 3.1 70b" to meta-llama/llama-3.1-70b-instruct.

Prompts: cheapest_model_for (a task → recommended model with monthly cost), monthly_price_report.

Try these

  • "Cheapest model with tool calling and vision for 50k support tickets a day? Show monthly cost."

  • "How did Claude and GPT prices change since 2024?"

  • "What got cheaper in LLM pricing since January?"

  • "Which provider serves Llama 4 Maverick cheapest, and at what quantization?"

Limits

  • Prices are OpenRouter's, which usually equal each provider's list price. Batch and enterprise discounts are not included.

  • History has one point per archived month, so a change is dated "between these two months".

  • Models are tracked by OpenRouter id; a renamed id starts a new history.

Part of the keyless MCP series

Open-source MCP servers that answer one market question each, with public data and no API keys.

Server

Question it answers

review-miner-mcp

What do users hate about competitor apps and games? (App Store + Steam reviews)

pricing-time-machine-mcp

How did a SaaS pricing page change over the years? (Wayback Machine)

hn-hiring-trends-mcp

Which skills are tech companies hiring for, and which are rising? (HN Who is hiring)

model-price-radar-mcp (this one)

What does each LLM cost, and did it get cheaper? (OpenRouter + price history)

launch-detector-mcp

What is a company about to launch? (certificate transparency logs)

Development

uv sync --extra dev
uv run pytest                              # offline tests with mocked OpenRouter + Wayback
uv run python scripts/smoke_live.py        # live check
uv run python scripts/build_history.py     # rebuild the bundled history (several minutes)
uv run --with rich python scripts/demo.py deepseek/deepseek-chat   # terminal demo (vhs docs/demo.tape records the GIF)

MIT © Ali Altunar

Available Tools

5 tools
compare_providersA
Read-onlyIdempotent

Same model, different providers: price, context, max output, quantization and recent uptime for every provider that serves it on OpenRouter, cheapest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel id or name, e.g. 'meta-llama/llama-3.3-70b-instruct'.
response_formatNo'markdown' (default) or 'json'.markdown

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, and open-world behavior, so safety is covered. The description adds useful behavioral context beyond that: results are sorted cheapest first and cover every provider on OpenRouter, which helps set expectations for output ordering and completeness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence front-loads the core idea ('Same model, different providers') and then enumerates returned fields with the sorting rule at the end. No sentence is wasted and the structure is easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity, full schema coverage, rich annotations, and the presence of an output schema, the description supplies everything an agent needs: the comparison scope, returned dimensions, and ordering. Nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters (model id/name and response_format enum). The description adds no additional parameter syntax or format details beyond what is already structured, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (compare providers for the same model) and lists the exact dimensions returned (price, context, max output, quantization, uptime), making it easy to distinguish from siblings like find_models or estimate_cost. An agent can tell at a glance that this tool is for provider-level comparison rather than model discovery or cost estimation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Same model, different providers' implies the scenario, but there is no explicit guidance on when to prefer this over siblings such as model_price_history or price_changes. Usage is inferable from the purpose but not stated, so it lands at the minimum-viable level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

estimate_costA
Read-onlyIdempotent

Cost of a workload on several models at today's prices, cheapest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
callsNoNumber of calls, e.g. per day or per month.
modelsYesModel ids or names, e.g. ['openai/gpt-4o-mini', 'claude haiku'].
input_tokensYesInput tokens per call.
output_tokensYesOutput tokens per call.
response_formatNo'markdown' (default) or 'json'.markdown

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive, open-world. The description adds the useful result-ordering trait (cheapest first) and signals current pricing, but says nothing about currency, whether figures are estimates vs billed rates, or how model-name resolution failures behave.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One line, front-loaded with the core action and scope, with the ordering detail appended. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists and schema coverage is total, the description need not explain return values. It conveys the essential framing (today's prices, cheapest first) for a 5-param tool; only edge-case behavior is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including 'calls', token counts, and response_format is already documented. The description adds no syntax or semantics beyond what the schema provides, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (estimate cost) and resource (a workload across several models), plus the ordering (cheapest first). It is distinguishable from find_models and price_changes, though it never names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'at today's prices' implicitly contrasts with sibling tools like model_price_history and price_changes, so usage context is inferable. But there is no explicit when-to-use statement, no exclusions, and no routing to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_modelsA
Read-onlyIdempotent

Search OpenRouter's live catalog by name, capabilities, price ceilings and context size. 'cheapest' ranks by blended price (3 input : 1 output tokens).

ParametersJSON Schema
NameRequiredDescriptionDefault
sortNoRanking.cheapest
limitNo
needsNoRequired capabilities, e.g. ['tools', 'vision'].
queryNoWords in the model id or name, e.g. 'claude', 'llama 70b', 'qwen coder'.
min_contextNoMinimum context window in tokens, e.g. 128000.
max_input_priceNoMax $ per 1M input tokens.
response_formatNo'markdown' (default) or 'json'.markdown
max_output_priceNoMax $ per 1M output tokens.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the description's job is to add operation-specific insight — and it does, by disclosing that 'cheapest' uses a blended price weighting of 3 input : 1 output tokens. That formula is not recoverable from annotations or the schema's bare 'Ranking.' default.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with zero filler: the first states scope and filters, the second qualifies the default sort's semantics. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-format detail is rightly omitted, and the description covers the search filters, the default sort, and the ranking definition. It is complete enough to invoke correctly, only lacking routing guidance relative to sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 88%, giving a baseline of 3, and the description adds meaning on top by explaining how the sort='cheapest' option is computed. It does not, however, expand on any of the other seven parameters beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Search') and resource ('OpenRouter's live catalog') plus the four filter dimensions (name, capabilities, price ceilings, context size), so the agent knows exactly what it returns. It doesn't name any sibling (compare_providers, estimate_cost, price_changes) to draw a boundary, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not guidance and no mention of alternatives among the siblings (estimate_cost, price_history, compare_providers). Usage is only inferable from the stated filters, so the agent must guess whether to reach for this versus a cost-comparison tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

model_price_historyA
Read-onlyIdempotent

Monthly price history of one model since it appeared on OpenRouter (archived copies since 2023) plus today's price. Shows every price change with its % and the overall change.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel id or name, e.g. 'openai/gpt-4o' or 'llama 3.1 70b'.
response_formatNo'markdown' (default) or 'json'.markdown

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds genuine context about data provenance (archived copies since 2023) and what the output includes (per-change % and overall change), but says nothing about rate limits, coverage gaps, or freshness beyond 'today's price'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the resource and time scope, then the returned content. No filler or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the description covers data range and change reporting. The one gap is the absence of routing guidance among the four sibling pricing tools, which an agent would need in order to choose correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (model, response_format) are already documented in the schema, including an example model id and enum values. The description only reinforces 'one model' per call and adds no format or syntax detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (retrieves price history) and resource (one model's pricing) with scope qualifiers: monthly granularity, since the model appeared on OpenRouter, plus today's price. An agent can distinguish this from siblings like price_changes or estimate_cost, which do not cover a single model's full historical timeline.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance is given, and no sibling tool is named as an alternative. The existence of a sibling named price_changes makes the omission notable, since the agent must guess which tool covers change history versus general price history.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

price_changesB
Read-onlyIdempotent

What changed in LLM pricing since a month: biggest price drops and increases, newly free models, and how many models were added or removed.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
sinceNoCompare against the archived month closest to this, 'YYYY-MM' (default: 3 months ago).
response_formatNo'markdown' (default) or 'json'.markdown

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds what the report actually contains (drops, increases, free models, model counts), which is useful content-level context, but it says nothing about aggregation window behavior, freshness, or data source. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that enumerates the report's contents with no filler. It is slightly list-like and could be tightened, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-format explanation is rightly omitted, and the three parameters are simple. However, the description leaves the `limit` parameter unexplained and introduces a 'since a month' claim inconsistent with the documented 3-month default, gaps an agent would notice when setting the window.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: `since` and `response_format` are documented in the schema, while `limit` is not. The description's 'since a month' loosely gestures at the comparison window but actually conflicts with the schema's own default of '3 months ago', so it adds ambiguity rather than meaning for the one parameter it touches. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (LLM pricing) and an explicit scope (changes over a period), enumerating the outputs: biggest drops/increases, newly free models, models added/removed. It does not, however, distinguish itself from close siblings like model_price_history or compare_providers, leaving the agent to infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no stated alternatives, and no exclusions. Given siblings such as model_price_history (per-model history) and compare_providers (cross-provider comparison), the agent needs a routing rule that this description never supplies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedcompare_providers
    • First observedestimate_cost
    • First observedfind_models
    • First observedmodel_price_history
    • First observedprice_changes

TDQS

A3.8/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct question: catalog search, workload cost estimation, per-model price history, market-wide price changes, and provider-level comparison. The descriptions clarify scope, especially between model_price_history (single model over time) and price_changes (market-wide since a month), so misselection is unlikely.

Naming Consistency4/5

All tool names use snake_case and are descriptive. Three follow a verb_noun pattern (find_models, estimate_cost, compare_providers) while two are noun phrases (model_price_history, price_changes), a minor deviation but still readable and predictable overall.

Tool Count5/5

Five tools is well-scoped for an LLM pricing radar: search, cost estimation, historical tracking, market change detection, and provider comparison. Each tool earns its place without redundancy or thin coverage.

Completeness5/5

The surface covers the core lifecycle of price discovery: finding models, estimating workload costs, tracking per-model price history, detecting market-wide changes, and comparing providers. No obvious dead ends exist for the stated domain of OpenRouter price monitoring.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides real-time LLM pricing and availability data as an MCP server, enabling AI agents to make optimal model routing decisions at inference time with cited pricing sources.
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    MCP server that provides OpenRouter model pricing data, enabling price lookups, trending/cheapest lists, and model searches without an API key.
    -