Skip to main content
Glama

Get benchmark results

benchmark_results
Read-onlyIdempotent

Get reviewed scores for an exact benchmark_id from benchmark_list. Join result.modelId to catalog_search model.id. Compare scores only within the same benchmark edition and configuration. Includes source provenance and reported coding-agent usage when published. Follow pagination.hasMore.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
benchmark_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataYes
metaYes
paginationYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changed
    • addedInput schema / $schema
      Added value: +"https://json-schema.org/draft/2020-12/schema"
    • removedInput schema / properties / benchmark_id / enum
      Removed value: -[
      -  "livebench",
      -  "aa-intelligence",
      -  "aa-coding",
      -  "aa-agentic",
      -  "swe-bench-verified-epoch-2-0-2",
      -  "terminal-bench-2-1-terminus-2",
      -  "terminal-bench-2-1-claude-code",
      -  "terminal-bench-2-1-codex",
      -  "gpqa-diamond",
      -  "otis-mock-aime-2024-2025",
      -  "livebench-coding",
      -  "livebench-agentic-coding",
      -  "livebench-reasoning",
      -  "livebench-math"
      -]
    • addedInput schema / properties / benchmark_id / maxLength
      Added value: +120
    • addedInput schema / properties / benchmark_id / minLength
      Added value: +1
    • addedInput schema / properties / benchmark_id / pattern
      Added value: +"^[a-z0-9-]+$"
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "properties": {
      +    "data": {
      +      "properties": {
      +        "benchmark": {
      +          "properties": {
      +            "cohort": {
      +              "enum": [
      +                "frontier",
      +                "measured"
      +              ],
      +              "type": "string"
      +            },
      +            "cohortLabel": {
      +              "type": "string"
      +            },
      +            "configurationNote": {
      +              "type": "string"
      +            },
      +            "date": {
      +              "description": "Source date, or an empty string when the upstream source supplies none.",
      +              "type": "string"
      +            },
      +            "description": {
      +              "type": "string"
      +            },
      +            "evaluator": {
      +              "type": "string"
      +            },
      +            "id": {
      +              "type": "string"
      +            },
      +            "label": {
      +              "type": "string"
      +            },
      +            "provenance": {
      +              "properties": {
      +                "sourceLabel": {
      +                  "type": "string"
      +                },
      +                "sourceUrl": {
      +                  "format": "uri",
      +                  "type": "string"
      +                }
      +              },
      +              "required": [
      +                "sourceUrl",
      +                "sourceLabel"
      +              ],
      +              "type": "object"
      +            },
      +            "scoreUnit": {
      +              "enum": [
      +                "points",
      +                "percent"
      +              ],
      +              "type": "string"
      +            },
      +            "version": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "id",
      +            "label",
      +            "description",
      +            "version",
      +            "date",
      +            "scoreUnit",
      +            "evaluator",
      +            "configurationNote",
      +            "cohort",
      +            "cohortLabel",
      +            "provenance"
      +          ],
      +          "type": "object"
      +        },
      +        "results": {
      +          "items": {
      +            "properties": {
      +              "configuration": {
      +                "type": "string"
      +              },
      +              "date": {
      +                "format": "date",
      +                "type": "string"
      +              },
      +              "dateKind": {
      +                "enum": [
      +                  "source-update",
      +                  "run"
      +                ],
      +                "type": "string"
      +              },
      +              "evaluator": {
      +                "type": "string"
      +              },
      +              "harness": {
      +                "type": "string"
      +              },
      +              "modelId": {
      +                "type": "string"
      +              },
      +              "provenance": {
      +                "properties": {
      +                  "sourceDataUrl": {
      +                    "format": "uri",
      +                    "type": "string"
      +                  },
      +                  "sourceRecordId": {
      +                    "type": "string"
      +                  },
      +                  "sourceUrl": {
      +                    "format": "uri",
      +                    "type": "string"
      +                  }
      +                },
      +                "required": [
      +                  "sourceUrl",
      +                  "sourceDataUrl",
      +                  "sourceRecordId"
      +                ],
      +                "type": "object"
      +              },
      +              "reportedAgentUsage": {
      +                "anyOf": [
      +                  {
      +                    "properties": {
      +                      "agent": {
      +                        "type": "string"
      +                      },
      +                      "agentVersion": {
      +                        "type": "string"
      +                      },
      +                      "averageTrialDurationSeconds": {
      +                        "minimum": 0,
      +                        "type": "number"
      +                      },
      +                      "derived": {
      +                        "description": "Values derived from the published run totals and rounded benchmark score.",
      +                        "properties": {
      +                          "costPerSuccessfulTrialUsd": {
      +                            "minimum": 0,
      +                            "type": "number"
      +                          },
      +                          "costPerTrialUsd": {
      +                            "minimum": 0,
      +                            "type": "number"
      +                          },
      +                          "inputCacheShare": {
      +                            "maximum": 1,
      +                            "minimum": 0,
      +                            "type": "number"
      +                          },
      +                          "tokensPerSuccessfulTrial": {
      +                            "minimum": 0,
      +                            "type": "number"
      +                          }
      +                        },
      +                        "required": [
      +                          "costPerTrialUsd",
      +                          "costPerSuccessfulTrialUsd",
      +                          "inputCacheShare",
      +                          "tokensPerSuccessfulTrial"
      +                        ],
      +                        "type": "object"
      +                      },
      +                      "reportedCostUsd": {
      +                        "minimum": 0,
      +                        "type": "number"
      +                      },
      +                      "tokens": {
      +                        "properties": {
      +                          "cachedInput": {
      +                            "minimum": 0,
      +                            "type": "integer"
      +                          },
      +                          "output": {
      +                            "minimum": 0,
      +                            "type": "integer"
      +                          },
      +                          "total": {
      +                            "minimum": 0,
      +                            "type": "integer"
      +                          },
      +                          "uncachedInput": {
      +                            "minimum": 0,
      +                            "type": "integer"
      +                          }
      +                        },
      +                        "required": [
      +                          "uncachedInput",
      +                          "cachedInput",
      +                          "output",
      +                          "total"
      +                        ],
      +                        "type": "object"
      +                      },
      +                      "trialCount": {
      +                        "minimum": 1,
      +                        "type": "integer"
      +                      }
      +                    },
      +                    "required": [
      +                      "agent",
      +                      "agentVersion",
      +                      "trialCount",
      +                      "tokens",
      +                      "reportedCostUsd",
      +                      "averageTrialDurationSeconds",
      +                      "derived"
      +                    ],
      +                    "type": "object"
      +                  },
      +                  {
      +                    "type": "null"
      +                  }
      +                ]
      +              },
      +              "runStartedAt": {
      +                "anyOf": [
      +                  {
      +                    "format": "date-time",
      +                    "type": "string"
      +                  },
      +                  {
      +                    "type": "null"
      +                  }
      +                ]
      +              },
      +              "sampleCount": {
      +                "anyOf": [
      +                  {
      +                    "minimum": 1,
      +                    "type": "integer"
      +                  },
      +                  {
      +                    "type": "null"
      +                  }
      +                ]
      +              },
      +              "score": {
      +                "maximum": 100,
      +                "minimum": 0,
      +                "type": "number"
      +              },
      +              "scoreUnit": {
      +                "enum": [
      +                  "points",
      +                  "percent"
      +                ],
      +                "type": "string"
      +              },
      +              "sourceModel": {
      +                "type": "string"
      +              },
      +              "standardError": {
      +                "anyOf": [
      +                  {
      +                    "minimum": 0,
      +                    "type": "number"
      +                  },
      +                  {
      +                    "type": "null"
      +                  }
      +                ]
      +              },
      +              "version": {
      +                "type": "string"
      +              }
      +            },
      +            "required": [
      +              "modelId",
      +              "sourceModel",
      +              "configuration",
      +              "score",
      +              "scoreUnit",
      +              "evaluator",
      +              "harness",
      +              "version",
      +              "date",
      +              "dateKind",
      +              "runStartedAt",
      +              "sampleCount",
      +              "standardError",
      +              "reportedAgentUsage",
      +              "provenance"
      +            ],
      +            "type": "object"
      +          },
      +          "type": "array"
      +        }
      +      },
      +      "required": [
      +        "benchmark",
      +        "results"
      +      ],
      +      "type": "object"
      +    },
      +    "meta": {
      +      "properties": {
      +        "apiVersion": {
      +          "const": "v1"
      +        },
      +        "currency": {
      +          "const": "USD"
      +        },
      +        "tokenPriceUnit": {
      +          "const": "per_million_tokens"
      +        }
      +      },
      +      "required": [
      +        "apiVersion",
      +        "currency",
      +        "tokenPriceUnit"
      +      ],
      +      "type": "object"
      +    },
      +    "pagination": {
      +      "properties": {
      +        "hasMore": {
      +          "type": "boolean"
      +        },
      +        "limit": {
      +          "minimum": 1,
      +          "type": "integer"
      +        },
      +        "offset": {
      +          "minimum": 0,
      +          "type": "integer"
      +        },
      +        "page": {
      +          "minimum": 1,
      +          "type": "integer"
      +        },
      +        "pageCount": {
      +          "minimum": 0,
      +          "type": "integer"
      +        },
      +        "pageSize": {
      +          "minimum": 1,
      +          "type": "integer"
      +        },
      +        "total": {
      +          "minimum": 0,
      +          "type": "integer"
      +        }
      +      },
      +      "type": "object"
      +    }
      +  },
      +  "required": [
      +    "meta",
      +    "data",
      +    "pagination"
      +  ],
      +  "type": "object"
      +}
  2. First observed

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and closed-world behavior, so the bar is lower. The description adds real context beyond them: provenance and coding-agent usage are included only when published, and pagination must be driven by pagination.hasMore.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and identity constraint, then layered caveats. The prose is telegraphic but every clause (join, comparison scope, provenance, pagination) carries distinct information; nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required, yet the description still flags provenance/usage fields and pagination handling. Combined with the join and comparison caveats, an agent has enough to call it correctly; only the limit/offset range semantics are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It clarifies that benchmark_id must be an exact value from benchmark_list, which is genuinely useful, but limit and offset are never explained beyond the indirect pagination.hasMore hint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Get reviewed scores") and pins the scope to an exact benchmark_id sourced from benchmark_list. It explicitly differentiates itself from siblings benchmark_list and catalog_search by naming them and describing the join.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies an ordering (obtain the id from benchmark_list) and gives a comparison rule ("only within the same benchmark edition and configuration"), plus a downstream routing hint to catalog_search. It stops short of explicit when-not-to-use exclusions, so it lands just below the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources