Skip to main content
Glama

Get ranked benchmark scenario results

get_benchmark_scenario_results
Read-only

Filter, rank and limit a scenario's per-model benchmark results. No LLM call. overall blends quality, speed and cost using organization task weights and is null if a component is missing. Status tags are independent: success does not exclude stale or stale_score results. Missing sort metrics come last in either direction. See enricher://docs/model-benchmark for interpretation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoKeep only the top N after sorting (None = all).
statusNoKeep results carrying ANY of these tags: success | failed | stale (config_hash changed since this run, re-run it) | stale_score (reference/scoring config changed since scored, rescore it) | unscored (ran fine, never scored). Empty/None = every status.
sort_byNoMetric to sort by.overall
providersNoKeep only these provider names (empty/None = every provider).
model_keysNoKeep only these model composite keys (empty/None = every model).
sort_orderNoSort direction.desc
scenario_idYesUUID of the scenario.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/destructiveHint annotations, the description discloses key behavioral nuances: 'overall' is a weighted blend that becomes null if any component is missing, status tags are independent (success does not exclude stale/stale_score), and missing sort metrics always sort last. These are critical for correctly interpreting results and are not visible in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no wasted words. The opening verb phrase immediately states the action, followed by essential behavioral caveats and a pointer to docs. Each sentence contributes distinct information, and the most important constraints (status independence, null handling) are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema means return values need not be described. The description covers filtering, ranking, limiting, and the tricky semantics of nulls and status tags. It also provides a documentation link for deeper interpretation. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the baseline is 3. The description adds value by explaining the meaning of 'overall' (weighted blend) and the independence of status tags, which are not fully captured in the parameter descriptions. This goes beyond the schema's per-parameter docs, though it doesn't exhaustively cover every parameter nuance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (filter, rank, limit), a specific resource (a scenario's per-model benchmark results), and the scope (per-model). It distinguishes itself from siblings by focusing on retrieving and ranking results rather than creating or running benchmarks. The 'No LLM call' clarification reinforces its read-only purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this is the tool for retrieving and ranking existing benchmark results, with explicit filtering and sorting semantics. It does not explicitly name alternatives or when-not-to-use conditions, but the context (read-only, results-focused) makes the use case unambiguous. A brief reference to the docs for interpretation adds some guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.