Skip to main content
Glama

The Aggregate — LLM benchmark aggregate

Server Details

LLM rankings from public benchmarks, model comparisons and source-linked results. Updated daily.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Uptime
100.0% over 55 days
Last Tested
Transport
Streamable HTTP · MCP 2025-06-18
URL

TDQS

A3.9/5.0

Scored across 8 tools

Disambiguation4/5

Tools mostly target distinct entities and actions: model search/get/compare, benchmark search/get, aggregate leaderboard, prediction duel, and metadata. There is minor overlap between get_leaderboard and search_models (both surface model rankings), but descriptions distinguish top aggregate ranking from name/provider lookup.

Naming Consistency4/5

Names follow a mostly consistent verb_noun pattern with get_*, search_*, and compare_* prefixes, all in snake_case. about_the_aggregate is a minor deviation from the verb_noun convention but is a reasonable one-off for metadata.

Tool Count5/5

Eight tools are well-scoped for a read-only benchmark aggregate, covering essential query needs without bloat. Each tool earns its place and the count is comfortably within the ideal 3-15 range.

Completeness4/5

The surface covers core read-only workflows: discovery (search), detail (get), comparison, aggregate leaderboard, metadata, and prediction standings. Minor gaps include no explicit list-all benchmarks/providers or bulk export, though search tools likely suffice for navigation.

Available Tools

8 tools
about_the_aggregateAbout The Aggregate
Read-onlyIdempotent
Inspect

What this data is: how the IRT fusion works, current coverage counts, update cadence, and how to cite it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

compare_modelsCompare models
Read-onlyIdempotent
Inspect

Head-to-head between 2-4 models: aggregate ranks, Elo gap with a significance note based on the standard errors, and notable benchmarks they share.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelsYesTwo to four model names or slugs.
get_benchmarkBenchmark detail
Read-onlyIdempotent
Inspect

One benchmark in depth: what it measures, the original source leaderboard URL, IRT stats (difficulty, noise, model coverage), skill weights, and the current top models on it.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoHow many top models to list (1-50, default 10).
benchmarkYesBenchmark name or slug, e.g. "Aider polyglot".
get_leaderboardAggregate leaderboard
Read-onlyIdempotent
Inspect

Top of the cross-benchmark aggregate ranking: every model placed on one Elo scale by an IRT model fit over public benchmark leaderboards (call about_the_aggregate for the current coverage counts). One row per model by default, fused across reasoning-effort settings. Supports paging via limit/offset.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoRows to return (1-100, default 25).
offsetNoRows to skip from the top (default 0).
include_variantsNoRank each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.
get_modelModel profile
Read-onlyIdempotent
Inspect

One model in depth: aggregate rank, Elo with standard error, provider, what it is, cost per task where known, and its most notable benchmark results (with percentiles).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel name or slug, e.g. "Claude Opus 4.5" or "gpt-5-5".
get_prediction_duelPrediction duel standings
Read-onlyIdempotent
Inspect

Guesswork — the public prediction duel: every day frontier LLMs and The Aggregate's own IRT model predict newly scraped benchmark scores before seeing them, and the errors are scored. Returns the current monthly standings, wins and losses included.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

search_benchmarksSearch benchmarks
Read-onlyIdempotent
Inspect

Find benchmarks in the aggregate by (partial) name. Returns model coverage, difficulty on the Elo scale, and the benchmark page URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (1-25, default 10).
queryYesBenchmark name fragment, e.g. "swe-bench" or "arena".
search_modelsSearch models
Read-onlyIdempotent
Inspect

Find ranked models by (partial) name or provider. Returns rank, Elo and the model page URL. One row per model by default, fused across reasoning-effort settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (1-25, default 10).
queryYesModel or provider name fragment, e.g. "opus" or "deepseek".
include_variantsNoReturn each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updates
    • Changedget_leaderboard1 field changed
      • addedInput schema / properties / include_variants
        Added value: +{
        +  "description": "Rank each reasoning-effort variant separately (e.g. \"Claude Opus 4.6 (High)\") instead of one fused row per model. Default false.",
        +  "type": "boolean"
        +}
    • Changedsearch_models1 field changed
      • addedInput schema / properties / include_variants
        Added value: +{
        +  "description": "Return each reasoning-effort variant separately (e.g. \"Claude Opus 4.6 (High)\") instead of one fused row per model. Default false.",
        +  "type": "boolean"
        +}
  2. 8 tool updates
    • First observedabout_the_aggregate
    • First observedcompare_models
    • First observedget_benchmark
    • First observedget_leaderboard
    • First observedget_model
    • First observedget_prediction_duel
    • First observedsearch_benchmarks
    • First observedsearch_models

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Live LLM API pricing: current token prices, model comparisons, cheapest-model lookups, and The LLM Price Index for 150+ models across 20+ providers, re-verified daily. No API key required.
    4
    1
    MIT
  • A
    license
    A
    quality
    F
    maintenance
    Enables users to select the optimal LLM for their specific task by aggregating benchmark data and user intent, returning a ranked shortlist with plain-English rationale.
    5
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Routes tasks to the optimal AI model based on task type and benchmark scores across 25+ platforms. Automatically selects the best model for coding, reasoning, writing, and more using public benchmark data.
    5
    414 npm
    1
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources