Skip to main content
Glama

Revert benchmark reference updates

revert_benchmark_reference_updates
Idempotent

Undo automatic edits a scoring pass made to a scenario's reference. Every scoring pass folds what the scored models showed the reference should be (a candidate the judge found better, a rule the samples prove, a value the reference lacked) and writes it into the reference — an edit a pass already made is only replaced by stronger evidence (samples, a wrong verdict, more agreeing models), never by one more model's better verdict; the log is reference_meta.auto_applied on get_benchmark_scenario, each entry with its inverse patch. Pass the revision ids to undo: the inverse is applied, the entry is marked reverted and its (path, attribute) is pinned so no later pass re-applies it (a manual set_benchmark_reference lifts every pin). Scores never go stale from this. Requires owner and a plan with benchmarks. No LLM call.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
scenario_idYesUUID of the scenario.
revision_idsYesIds of the reference_meta.auto_applied entries to undo (1..200).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by explaining exactly what happens on revert: the inverse patch is applied, the entry is marked reverted, the path/attribute is pinned to prevent re-application, and manual set_benchmark_reference lifts pins. It also states that scores never go stale and that no LLM call is made, which is valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence contributes: purpose, background on how edits are made, the revert mechanism, pinning behavior, prerequisites, and side effects. It is front-loaded with the core purpose and avoids redundant restatement of the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers what the tool does, where to find revision ids, what happens after reversion, how pins interact with future passes and manual sets, prerequisites, and the fact that scores remain valid. An output schema exists, so return-value details are not required here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning by explaining that revision_ids are entries from reference_meta.auto_applied on get_benchmark_scenario and that passing them triggers the inverse patch. This enriches the schema's bare parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Undo automatic edits a scoring pass made to a scenario's reference.' It also distinguishes this from manual reference changes by referencing set_benchmark_reference and automatic scoring-pass edits, making the tool's unique role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use the tool: when you need to undo automatic scoring-pass edits to a scenario reference. It points to get_benchmark_scenario for finding revision ids and mentions set_benchmark_reference as the manual counterpart, but it does not explicitly state exclusions or a direct 'use X instead' rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.