Create benchmark scenario
create_benchmark_scenarioCreate a reusable benchmark with a mandatory scoring judge. scenario_type='enrichment' needs schema_id and entity_data (the entity to enrich, as enrich_entity takes it — refused when it carries no value; never put it in description); 'sample_generation' needs sample_request; 'schema_generation' needs entity_samples (1..20 samples of one entity type). Enrichment and schema generation need a verified gold reference via set_benchmark_reference before running. Sample generation is rubric-scored and takes no reference. Requires owner and a plan with benchmarks; creating the scenario does not run the models. Returns the scenario and link. Next call run_benchmark when its reference requirements are satisfied. See enricher://docs/model-benchmark.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| language | No | sample_generation: output language for names + values. | en |
| strategy | No | Enrichment: pinned strategy (no 'auto'): single_pass | expert_domains | multi_expertise | single_pass |
| languages | No | Enrichment: defaults to ['en']. | |
| schema_id | No | UUID of the saved schema to enrich against (enrichment only, required there). | |
| description | No | Free-text note shown in the Benchmarks tab; no model ever reads it. The entity to enrich goes in entity_data, never here. | |
| entity_data | No | Enrichment (required there): the fixed entity input every model enriches — the same JSON enrich_entity takes. Read get_schema.input_contract first: identifying fields, preserve paths and keys for supplied array items. Refused when it carries no value. | |
| repetitions | No | Run each model N times per run; keeps mean + consistency spread. | |
| scenario_type | No | enrichment | sample_generation | schema_generation (immutable). | enrichment |
| attachment_ids | No | Attachment UUIDs included in every run. | |
| entity_samples | No | schema_generation: the 1..20 fixed input samples (JSON objects of one entity type) every model converts to a schema. Several samples let the scoring read evidence the reference cannot state alone (nullable, types, identity). | |
| sample_request | No | sample_generation: the free-text sample request every model answers — the kind of entity, what the sample must contain, any budget or structural preference (same contract as generate_sample's request). | |
| typical_object | No | sample_generation: a specific instance to model (e.g. 'Serena Williams'). | |
| enable_web_search | No | sample_generation: ground values with the model's web search. | |
| naming_convention | No | sample_generation: auto | snake_case | camelCase. | auto |
| generate_semantic_ids | No | schema_generation: add semantic_id properties to keyed objects. | |
| scoring_judge_model_key | Yes | LLM judge composite key used to score results (required). |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||