Skip to main content
Glama

The Aggregate — LLM benchmark aggregate

Aggregate leaderboard

get_leaderboard
Read-onlyIdempotent

Top of the cross-benchmark aggregate ranking: every model placed on one Elo scale by an IRT model fit over public benchmark leaderboards (call about_the_aggregate for the current coverage counts). One row per model by default, fused across reasoning-effort settings. Supports paging via limit/offset.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoRows to return (1-100, default 25).
offsetNoRows to skip from the top (default 0).
include_variantsNoRank each reasoning-effort variant separately (e.g. "Claude Opus 4.6 (High)") instead of one fused row per model. Default false.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / include_variants
      Added value: +{
      +  "description": "Rank each reasoning-effort variant separately (e.g. \"Claude Opus 4.6 (High)\") instead of one fused row per model. Default false.",
      +  "type": "boolean"
      +}
  2. First observed

TDQS

Score is being calculated.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources