Skip to main content
Glama

Get benchmark scenario

get_benchmark_scenario
Read-only

Read one benchmark scenario with per-model quality, cost and speed results. include_reference=true adds the reference, the fixed entity_data and the schema-generation entity_samples. Config changes can make existing results stale. No LLM call. For a ranked subset use get_benchmark_scenario_results. Read after run_benchmark completes.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
scenario_idYesUUID of the scenario.
include_referenceNoInclude reference_output + entity_data (can be large).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=falsehistorically. The description adds meaningful behavioral context beyond this: 'No LLM call' clarifies cost/latency characteristics, while 'Config changes can make existing results stale' warns about data freshness. It also explains the side effects of include_reference=true, which is a behavioral trait not evident from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four short sentences, each contributing essential information: what the tool returns, the parameter effect, a staleness warning, and a routing hint. It is front-loaded with the core purpose and has zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple read with an output schema present, so return values need no elaboration. The description covers the main use case (after run_benchmark), the caveat (config changes), the parameter semantics, and the alternative tool. Nothing an agent needs to safely and effectively invoke it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so a baseline of 3 is expected. However, the description adds value by explaining that include_reference=true adds not only reference_output and entity_data (already in the schema) but also 'the schema-generation entity_samples', which is not mentioned in the schema. This clarifies the parameter's full impact beyond the structured definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read one benchmark scenario') and the resource with its content ('per-model quality, cost and speed results'). It also explicitly distinguishes from the sibling get_benchmark_scenario_results by noting that the sibling returns a 'ranked subset', which is a clear differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage guidance: 'Read after run_benchmark completes' and 'For a ranked subset use get_benchmark_scenario_results'. It also warns about staleness with 'Config changes can make existing results stale', helping agents decide when to refresh or trust the data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.