LLM Benchmark MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@LLM Benchmark MCP ServerCompare GPT-4o and Claude 3.5 Sonnet on benchmarks"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
LLM Benchmark MCP Server
MCP server that gives AI agents access to LLM benchmark data, pricing comparisons, and model recommendations.
Features
compare_models — Side-by-side benchmark comparison of LLMs (MMLU, HumanEval, MATH, GPQA, ARC, HellaSwag)
get_model_details — Detailed info about a specific model including strengths/weaknesses
recommend_model — Get the best model recommendation for your task and budget
list_top_models — Top models ranked by category (coding, math, reasoning, chat)
get_pricing — Pricing comparison via OpenRouter API
Related MCP server: Artificial Analysis MCP Server
Supported Models
GPT-4o, GPT-4o-mini, GPT-4 Turbo, o1, o3-mini, Claude 3.5 Sonnet, Claude 3.5 Haiku, Claude 3 Opus, Gemini 2.0 Flash, Gemini 2.0 Pro, Gemini 1.5 Pro, Llama 3.1 (8B/70B/405B), Llama 3.3 70B, Mistral Large, Mistral Small, Mixtral 8x22B, DeepSeek V3, DeepSeek R1, Qwen 2.5 72B
Installation
pip install llm-benchmark-mcp-serverUsage with Claude Desktop
Add to your claude_desktop_config.json:
{
"mcpServers": {
"llm-benchmark": {
"command": "benchmark-server"
}
}
}Or via uvx (no install needed):
{
"mcpServers": {
"llm-benchmark": {
"command": "uvx",
"args": ["llm-benchmark-mcp-server"]
}
}
}Example Queries
"Compare GPT-4o vs Claude 3.5 Sonnet vs Gemini 2.0 Pro"
"Which model is best for coding on a low budget?"
"Show me the top 10 models for math"
"What does GPT-4o cost compared to Claude?"
"Give me details about DeepSeek R1"
Data Sources
Benchmarks: Hardcoded from official papers and public leaderboards (MMLU, HumanEval, MATH, GPQA, ARC-Challenge, HellaSwag)
Pricing: Live data from OpenRouter API
Arena Rankings: Chatbot Arena Leaderboard (when available)
More MCP Servers by AiAgentKarl
Category | Servers |
🔗 Blockchain | |
🌍 Data | Weather · Germany · Agriculture · Space · Aviation · EU Companies |
🔒 Security | |
🤖 Agent Infra | Memory · Directory · Hub · Reputation |
🔬 Research |
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Sourced AI-model pricing and capability data — compare and route to the cheapest capable model.
Live LLM API pricing: token prices, comparisons, cheapest-model lookups. No key required.
Source-backed AI model pricing, rankings, history, and benchmark data.
Compare, estimate, and deploy cloud infrastructure across AWS, GCP, and Azure for AI agents.
Related MCP Servers
- AlicenseAqualityFmaintenanceProvides AI assistants with real-time access to 1000+ AI models including their latest pricing, context windows, capabilities, and specifications. Supports model search, comparison, recommendations, and live testing.326 npm1MIT
- AlicenseAqualityDmaintenanceProvides access to real-time LLM pricing, speed metrics, and performance benchmarks for over 300 models from Artificial Analysis. It enables users to list, filter, and compare models based on costs, tokens per second, and intelligence indices.228 npm9MIT
- AlicenseAqualityDmaintenanceEnables AI agents to discover, compare, and select the best AI models across multiple providers based on pricing, performance, and capabilities, with real-time cost estimation and benchmarking.939 npm1MIT
- FlicenseNot gradedqualityCmaintenanceEnables real-time access to LLM pricing, benchmarks, deprecation alerts, and cost optimization for over 30 models across 8 providers, allowing AI agents to make cost-effective model selections.-