models-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@models-mcpfind models under $0.01 per 1M tokens with 128k context"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Models MCP
Search, compare, and inspect AI models by pricing, context window, and capabilities. An MCP server over the models.dev catalog (models.dev/api.json), so your agent always has current model data without you hand-maintaining a list.
models.dev itself doesn't ship an MCP server, just a JSON API and a TypeScript SDK for reading it. This fills that gap.
Runs two ways from the same tool code:
stdio (
src/index.ts) for local MCP clientsCloudflare Worker (
src/worker.ts) as a remote Streamable HTTP endpoint at/mcp
Tools
Tool | What it does |
| Lists every provider (anthropic, openai, google, ...) with model counts |
| Filters models by name, provider, min context window, max input cost, or capability flags (reasoning, tool_call, attachment) |
| Full metadata for one model, by |
| Side-by-side diff of 2-6 models on pricing, context, and capabilities |
| Forces a re-fetch, bypassing the 1-hour cache |
Install
npm install
npm run buildRun standalone over stdio (for testing)
npm startIt speaks MCP over stdio, so you won't see much directly; use the MCP Inspector to poke at it:
npx @modelcontextprotocol/inspector node dist/index.jsHost on Cloudflare Workers
The Worker entry (src/worker.ts) serves the same tools over Streamable HTTP at /mcp, with:
Catalog caching in the Workers Cache API (
caches.default) with a 1-hour TTL, shared across requests and isolates.Per-IP rate limiting via a Workers rate limiting binding: 60 requests/minute per IP, enforced per Cloudflare location. Excess requests get
429withRetry-After: 60.
# local dev at http://localhost:8787/mcp
npm run dev:worker
# deploy
npm run deployAfter deploy, your endpoint is https://models-mcp.<your-subdomain>.workers.dev/mcp.
Point MCP clients at it:
Claude Code:
claude mcp add --transport http models-mcp https://models-mcp.<your-subdomain>.workers.dev/mcpGeneric client config (anything that speaks Streamable HTTP):
{
"mcpServers": {
"models-mcp": {
"url": "https://models-mcp.<your-subdomain>.workers.dev/mcp"
}
}
}For stdio-only clients (Claude Desktop), bridge with mcp-remote:
{
"mcpServers": {
"models-mcp": {
"command": "npx",
"args": ["mcp-remote", "https://models-mcp.<your-subdomain>.workers.dev/mcp"]
}
}
}No API keys required anywhere. All data comes from the public models.dev/api.json endpoint.
Tests
npm testCovers the catalog client (flattening, TTL caching, force refresh, stale-on-failure fallback, id resolution) and all five tools end-to-end through a real MCP client session over an in-memory transport.
Notes on the data
The catalog is cached for 1 hour: in the Workers Cache API when hosted, in process memory over stdio. Call
refresh_catalogto force an update. If a refetch fails, the last good catalog keeps being served.models.dev doesn't publish a versioned schema for consumers, so the types in
src/types.tsare intentionally loose (index signatures preserve any fields not explicitly typed).Model ids follow the
provider/modelconvention used by the AI SDK and OpenCode, e.g.anthropic/claude-sonnet-4-5.get_modelandcompare_modelsalso accept a bare model id if it's unambiguous across providers.
Possible extensions
A
list_facetstool (modalities, tokenizers) similar to what other model-catalog MCPs expose.A
test_modeltool that makes a live call through whichever provider key you have configured, for latency/cost sanity checks.OAuth or Cloudflare Access in front of the Worker, if you want it private.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Search 200+ UnoRouter models (most free), check pricing, and chat through one key
Search, compare, and find alternatives across a curated catalog of 1,100+ AI tools.
@latest documentation and code examples to 9000+ libraries for LLMs and AI code editors in a singl…
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/QAInsights/models-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server