whichapiapi
Provides a speech-to-text transport integration for evaluating and comparing transcription offers under the speech.stt capability.
Pulls model offers and pricing from Hugging Face Inference Providers (HF router) and uses Hugging Face Open ASR Leaderboard benchmarks for model ranking.
Allows auditing Make automation blueprints for AI call overpay, cheaper model alternatives, missing fallbacks, training-data risks, deprecated models, and hardcoded API keys.
Allows auditing n8n workflows to detect overpay, cheaper model alternatives, missing fallbacks, training-data risks, deprecated models, and hardcoded API keys; also converts n8n nodes into eval suites to test suggested model switches.
Provides OpenAI-compatible transcription integration for speech-to-text evaluation and uses OpenAI Whisper normalizers for WER scoring.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@whichapiapirank the best coding models by real cost per task"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Which API API

Pick the right API for your automation — tested on your own tasks, priced for your real load.
Status: v0.3 — evaluator, MCP/REST, offers + benchmarks, non-LLM (search, speech-to-text), automation audit (n8n, Make, code) and a task-aware model picker are done; applying changes to automations is next. License: AGPL-3.0.
Describe a task (text, voice, or an existing n8n workflow). Which API API figures out which capabilities it
needs (llm.chat, search.web, speech.stt, geo.geocode, transit.realtime, …), finds every way to buy
each one — official providers, aggregators, resellers, your own channels — with their real conditions
(batch, cache, off-peak, data-sharing discounts, top-up fees, KYC, geo-blocks), runs the candidates on your
examples, and recommends a primary option plus a fallback from a different provider. Later phases implement
the choice directly in your automation.
Core idea: don't reinvent — data, evals and catalogs come from open sources (models.dev, OpenRouter, LiteLLM price data, new-api resellers, APIs.guru, MCP Registry…). The value is the glue: normalised offers with provenance, cost for your load profile, and evaluation on your tasks.
Quick start
uv sync
uv run whichapiapi offers pull all # ~10k offers: models.dev, OpenRouter, DeepInfra, HF router, Eden AI
uv run whichapiapi offers pull-benchmarks # LMArena, Open ASR, SWE-bench, BFCL (+ Artificial Analysis with a free key)
uv run whichapiapi models rank -p optimal # best model variants: quality, cost per task, speed
uv run whichapiapi audit n8n --docker n8n # read-only audit of your n8n workflows
uv run pytest -qDocker: docker build -t whichapiapi . && docker run -e WHICHAPIAPI_MCP_TOKEN=... -p 8765:8765 -v wa:/data whichapiapi
(MCP at /mcp, read-only REST at /api/offers, /api/rank, /api/models/<model>/benchmarks, /api/audit/n8n).
Pick a model for a task
whichapiapi models rank -p coding (or MCP rank_models, REST /api/rank) ranks every model variant — reasoning
effort included — by a task-specific quality blend of benchmark percentiles, cost per task that counts the
thinking tokens a variant spends before answering, and speed. 14 presets: best, optimal, value, fast, ultrafast,
realtime, coding, math, agents, russian, writing, vision, long_context, open. whichapiapi models page picker.html
builds the interactive picker (sliders for quality / cost / speed). Method and limits: docs/SELECTION.md.
Audit an automation
whichapiapi audit n8n <export.json> | --docker <container>, audit make <blueprint.json>, audit code <dir> (MCP
audit_n8n): every AI call is checked for overpay (same model cheaper elsewhere), cheaper models with an
equal-or-better benchmark score, missing fallback, training-data risk, deprecated/unknown models and API keys
written into the automation. audit n8n-suite <node> turns a node into an eval suite so a suggested switch is
tested on that node's own inputs first.
Use it from an agent (MCP)
claude mcp add whichapiapi -- uv run --project /path/to/which_apiapi whichapiapi mcpTools: rank_models, audit_n8n, find_offers, compare_prices, estimate_cost, learned_prices, pull_offers,
pull_benchmarks, benchmark_boards, model_benchmarks, pull_conditions, conditions_priors, cross_check_price,
eval_plan, run_eval, run_canary (run_eval/run_canary spend money within a hard budget — the agent
is told to ask first). Streamable HTTP: whichapiapi mcp --http (127.0.0.1:8765/mcp).
Evaluate candidates on your own task
Write a suite (promptfoo-compatible YAML) with deterministic checks and an LLM rubric — or let an agent draft one
(MCP create_suite, suite_from_traces); see examples/canary/ for a small one. whichapiapi eval plan|run|learn|reconcile|report.
Real cost is read from the channel's own call log when available (new-api resellers), so reports show
list price vs what you were actually charged.
Canary suite: catch a channel silently substituting the model
uv run whichapiapi eval canary examples/canary/suite.yaml # ~$0.0005/run5 cheap, deterministic cases (no LLM judge) run repeatedly and get diffed against a stored per-provider
baseline; alerts flags a pass-rate/score drop or a latency spike. A clean run refreshes the baseline; a run
WITH alerts does not, so a regression can't quietly become "normal" — re-run with --reset-baseline once
you've reviewed it. It exits with code 2 on alerts; systemd/cron recipes: docs/OPERATIONS.md.
Optional: watch pricing pages with changedetection.io
whichapiapi mcp --http also serves a webhook at /webhooks/changedetection that records "this pricing page
changed" notifications (self-hosted, never spends money, never auto-refetches — it just tells you which
offers pull* to re-run). To wire it up:
Run changedetection.io yourself (
docker run ghcr.io/dgtlmoon/changedetection.io) and add a watch per pricing page you care about.On that watch, set the notification URL to
json://<host>:8765/webhooks/changedetection(orjsons://for TLS) and the notification body to a JSON template, e.g.:{"watch_url": "{{watch_url}}", "watch_uuid": "{{watch_uuid}}", "watch_title": "{{watch_title}}", "diff_added": "{{diff_added}}", "diff_removed": "{{diff_removed}}"}(token names verified against changedetection.io's own notification context, not guessed). The endpoint accepts this JSON directly at the top level, or nested as a string under
body/message— Apprise's exact wrapping forjson:///jsons://wasn't confirmed from its docs, so both forms are handled defensively rather than assumed.whichapiapi offers price-changes(CLI) orprice_change_alerts(MCP tool) lists what came in, newest first, with any of our offers whose source domain matches the changed page.
Non-LLM capabilities
The evaluator and selector are capability-agnostic. Real transports: Tavily (search.web, keyword + LLM-judged
relevance, examples/search/suite.live.yaml), OpenAI-compatible transcription (OpenAI, Groq, new-api resellers) and
Deepgram (speech.stt, examples/speech_stt/suite.live.yaml; WER with OpenAI Whisper's normalizers).
Related MCP server: AI Model Advisor MCP Server
Custom providers
Which API API works with custom providers that have their own pricing — resellers, gateways, self-hosted
OpenAI-compatible endpoints. Describe them once, outside the repository, in ~/.config/whichapiapi/channels.yaml
(or $WHICHAPIAPI_CHANNELS); keys stay in environment variables:
myreseller:
base_url: https://api.example.com/v1
key_env: MYRESELLER_API_KEY
kind: newapi # new-api/one-api: price groups, per-call log → real cost, balance forecast
primary: true # default channel for judges, the balance forecast and the dashboard
judge_model: some-modelThen myreseller:<model> works everywhere (suites, router, integrity checks), offers come from whichapiapi offers pull newapi --url … --channel myreseller, and real per-call cost is measured from its log. Your own keys for
official free tiers go to the encrypted pool: whichapiapi providers add groq.
Documentation
Doc | What |
Modules, data model, flows, extension points | |
How model variants are scored (quality, cost per task, speed) | |
Scheduled refresh and canary (systemd / cron) | |
Trust model: suites are code, MCP/REST auth, secrets | |
Releases | |
Phases and exit criteria | |
What we reuse from GitHub and why | |
Architecture decisions |
Data attribution
Price and catalog data come from third-party sources, each credited in reports: models.dev (MIT), OpenRouter, DeepInfra, Hugging Face Inference Providers, Eden AI, LiteLLM (MIT), InferIndex, LMArena (CC BY 4.0), Artificial Analysis (attribution required), Hugging Face Open ASR Leaderboard, SWE-bench (MIT), BFCL (Apache-2.0). Speech test clips: LibriSpeech and FLEURS (CC BY 4.0), downloaded locally, never committed. WER normalizers vendored from OpenAI Whisper (MIT). Free-tier limits from the FreeLLMAPI catalog (MIT, freellmapi.co); several router mechanisms are adapted from FreeLLMAPI's design. Full notices: THIRD_PARTY_NOTICES.md.
This server cannot be deployed
Maintenance
Related MCP Connectors
Compare LLM API prices, search models and providers, and access reviewed benchmark results.
Live LLM API pricing: token prices, comparisons, cheapest-model lookups. No key required.
Find AI model pricing, estimate token costs and compare offers. No API key required.
Pay-per-use tool marketplace for AI agents. Search, price-check, and call APIs via MCP.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to compare LLM benchmarks, get pricing, and receive model recommendations for tasks and budgets.1MIT
- AlicenseAqualityDmaintenanceEnables AI agents to discover, compare, and select the best AI models across multiple providers based on pricing, performance, and capabilities, with real-time cost estimation and benchmarking.91,988 npm1MIT
- FlicenseNot gradedqualityCmaintenanceEnables real-time access to LLM pricing, benchmarks, deprecation alerts, and cost optimization for over 30 models across 8 providers, allowing AI agents to make cost-effective model selections.-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to calculate, compare, and recommend AI API costs from multiple providers, with support for currency conversion and platform fee analysis.23 PyPI1MIT