The Aggregate — LLM benchmark aggregate
Server Details
LLM rankings from public benchmarks, model comparisons and source-linked results. Updated daily.
- Status
- Healthy
- Uptime
- 100.0% over 55 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 8 tools
Tools mostly target distinct entities and actions: model search/get/compare, benchmark search/get, aggregate leaderboard, prediction duel, and metadata. There is minor overlap between get_leaderboard and search_models (both surface model rankings), but descriptions distinguish top aggregate ranking from name/provider lookup.
Names follow a mostly consistent verb_noun pattern with get_*, search_*, and compare_* prefixes, all in snake_case. about_the_aggregate is a minor deviation from the verb_noun convention but is a reasonable one-off for metadata.
Eight tools are well-scoped for a read-only benchmark aggregate, covering essential query needs without bloat. Each tool earns its place and the count is comfortably within the ideal 3-15 range.
The surface covers core read-only workflows: discovery (search), detail (get), comparison, aggregate leaderboard, metadata, and prediction standings. Minor gaps include no explicit list-all benchmarks/providers or bulk export, though search tools likely suffice for navigation.
Available Tools
8 toolsabout_the_aggregateAbout The AggregateRead-onlyIdempotentInspect
What this data is: how the IRT fusion works, current coverage counts, update cadence, and how to cite it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
compare_modelsCompare modelsRead-onlyIdempotentInspect
Head-to-head between 2-4 models: aggregate ranks, Elo gap with a significance note based on the standard errors, and notable benchmarks they share.
| Name | Required | Description | Default |
|---|---|---|---|
| models | Yes | Two to four model names or slugs. |
get_benchmarkBenchmark detailRead-onlyIdempotentInspect
One benchmark in depth: what it measures, the original source leaderboard URL, IRT stats (difficulty, noise, model coverage), skill weights, and the current top models on it.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | How many top models to list (1-50, default 10). | |
| benchmark | Yes | Benchmark name or slug, e.g. "Aider polyglot". |
get_leaderboardAggregate leaderboardRead-onlyIdempotentInspect
Top of the cross-benchmark aggregate ranking: every model placed on one Elo scale by an IRT model fit over public benchmark leaderboards (call about_the_aggregate for the current coverage counts). One row per model by default, fused across reasoning-effort settings. Supports paging via limit/offset.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Rows to return (1-100, default 25). | |
| offset | No | Rows to skip from the top (default 0). | |
| include_variants | No | Rank each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false. |
get_modelModel profileRead-onlyIdempotentInspect
One model in depth: aggregate rank, Elo with standard error, provider, what it is, cost per task where known, and its most notable benchmark results (with percentiles).
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Model name or slug, e.g. "Claude Opus 4.5" or "gpt-5-5". |
get_prediction_duelPrediction duel standingsRead-onlyIdempotentInspect
Guesswork — the public prediction duel: every day frontier LLMs and The Aggregate's own IRT model predict newly scraped benchmark scores before seeing them, and the errors are scored. Returns the current monthly standings, wins and losses included.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
search_benchmarksSearch benchmarksRead-onlyIdempotentInspect
Find benchmarks in the aggregate by (partial) name. Returns model coverage, difficulty on the Elo scale, and the benchmark page URL.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (1-25, default 10). | |
| query | Yes | Benchmark name fragment, e.g. "swe-bench" or "arena". |
search_modelsSearch modelsRead-onlyIdempotentInspect
Find ranked models by (partial) name or provider. Returns rank, Elo and the model page URL. One row per model by default, fused across reasoning-effort settings.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (1-25, default 10). | |
| query | Yes | Model or provider name fragment, e.g. "opus" or "deepseek". | |
| include_variants | No | Return each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false. |
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
get_leaderboard1 field changed- added
Input schema / properties / include_variantsAdded value: +{ + "description": "Rank each reasoning-effort variant separately (e.g. \"Claude Opus 4.6 (High)\") instead of one fused row per model. Default false.", + "type": "boolean" +}
- Changed
search_models1 field changed- added
Input schema / properties / include_variantsAdded value: +{ + "description": "Return each reasoning-effort variant separately (e.g. \"Claude Opus 4.6 (High)\") instead of one fused row per model. Default false.", + "type": "boolean" +}
8 tool updates
- First observed
about_the_aggregate - First observed
compare_models - First observed
get_benchmark - First observed
get_leaderboard - First observed
get_model - First observed
get_prediction_duel - First observed
search_benchmarks - First observed
search_models
Related MCP Connectors
Source-backed AI model pricing, rankings, history, and benchmark data.
Pick the right LLM for any task. Ranked shortlist with rationale across 8 evaluators.
Compare LLM API prices, search models and providers, and access reviewed benchmark results.
Live rankings and pricing for 10,000+ AI products and LLMs. Free tier; paid data via x402.
Related MCP Servers
- AlicenseAqualityDmaintenanceGive your AI assistant real-time LLM/VLM knowledge. Pricing, benchmarks, and recommendations — updated every hour, not every training cycle.485 npm2MIT
- AlicenseAqualityAmaintenanceLive LLM API pricing: current token prices, model comparisons, cheapest-model lookups, and The LLM Price Index for 150+ models across 20+ providers, re-verified daily. No API key required.41MIT
- AlicenseAqualityFmaintenanceEnables users to select the optimal LLM for their specific task by aggregating benchmark data and user intent, returning a ranked shortlist with plain-English rationale.5MIT
- AlicenseAqualityCmaintenanceRoutes tasks to the optimal AI model based on task type and benchmark scores across 25+ platforms. Automatically selects the best model for coding, reasoning, writing, and more using public benchmark data.5414 npm1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.