Benchmark detail
get_benchmarkOne benchmark in depth: what it measures, the original source leaderboard URL, IRT stats (difficulty, noise, model coverage), skill weights, and the current top models on it.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | How many top models to list (1-50, default 10). | |
| benchmark | Yes | Benchmark name or slug, e.g. "Aider polyglot". |