Skip to main content
Glama

Create benchmark scenario

create_benchmark_scenario

Create a reusable benchmark with a mandatory scoring judge. scenario_type='enrichment' needs schema_id and entity_data (the entity to enrich, as enrich_entity takes it — refused when it carries no value; never put it in description); 'sample_generation' needs sample_request; 'schema_generation' needs entity_samples (1..20 samples of one entity type). Enrichment and schema generation need a verified gold reference via set_benchmark_reference before running. Sample generation is rubric-scored and takes no reference. Requires owner and a plan with benchmarks; creating the scenario does not run the models. Returns the scenario and link. Next call run_benchmark when its reference requirements are satisfied. See enricher://docs/model-benchmark.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameYes
languageNosample_generation: output language for names + values.en
strategyNoEnrichment: pinned strategy (no 'auto'): single_pass | expert_domains | multi_expertisesingle_pass
languagesNoEnrichment: defaults to ['en'].
schema_idNoUUID of the saved schema to enrich against (enrichment only, required there).
descriptionNoFree-text note shown in the Benchmarks tab; no model ever reads it. The entity to enrich goes in entity_data, never here.
entity_dataNoEnrichment (required there): the fixed entity input every model enriches — the same JSON enrich_entity takes. Read get_schema.input_contract first: identifying fields, preserve paths and keys for supplied array items. Refused when it carries no value.
repetitionsNoRun each model N times per run; keeps mean + consistency spread.
scenario_typeNoenrichment | sample_generation | schema_generation (immutable).enrichment
attachment_idsNoAttachment UUIDs included in every run.
entity_samplesNoschema_generation: the 1..20 fixed input samples (JSON objects of one entity type) every model converts to a schema. Several samples let the scoring read evidence the reference cannot state alone (nullable, types, identity).
sample_requestNosample_generation: the free-text sample request every model answers — the kind of entity, what the sample must contain, any budget or structural preference (same contract as generate_sample's request).
typical_objectNosample_generation: a specific instance to model (e.g. 'Serena Williams').
enable_web_searchNosample_generation: ground values with the model's web search.
naming_conventionNosample_generation: auto | snake_case | camelCase.auto
generate_semantic_idsNoschema_generation: add semantic_id properties to keyed objects.
scoring_judge_model_keyYesLLM judge composite key used to score results (required).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint=false annotation, the description discloses that the tool only creates and does not run models, requires an owner and plan, returns the scenario and link, and refuses entity_data that carries no value. It also surfaces the preconditions for later execution, giving a clear behavioral contract.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause carries routing or behavioral information; the mandatory judge and scenario-type dispatch are front-loaded, and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with three scenario modes, the description provides type-specific requirements, reference preconditions, owner/plan prerequisites, a link to docs, and a follow-up call, while the output schema covers return values. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already high (94%), but the description adds cross-tool semantics: entity_data is 'the same JSON enrich_entity takes', sample_request follows 'same contract as generate_sample's request', entity_samples are '1..20 samples of one entity type', and the entity must 'never' go in description. This materially improves parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb-resource pair: 'Create a reusable benchmark with a mandatory scoring judge.' It then enumerates the three scenario_type branches and references related tools, making it clearly distinct from execution tools like run_benchmark and update/delete siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives direct when-to-use rules per scenario_type ('needs schema_id and entity_data', 'needs sample_request', 'needs entity_samples'), states the gold-reference prerequisite via set_benchmark_reference, and names the next action (run_benchmark) once reference requirements are met. This is explicit routing with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.