Run benchmark
run_benchmarkStart billed asynchronous execution and scoring of a benchmark. Requires owner, a benchmark-enabled plan and a judge; a verified reference is also required except for sample generation. Supply model_keys or providers; omitting both runs all active models with usable provider keys. Repetitions and judging increase cost. Returns job_id and total_models: poll get_job_status, then read get_benchmark_scenario_results. Re-running replaces each selected model's previous result. An organization's benchmark runs and scoring passes execute one at a time: a launch while another is in flight is queued (queue_position, status 'pending'), and a launch on a scenario whose run is still queued folds its models into that run (merged=true, job_id names the queued run). Reference setup and score interpretation: enricher://docs/model-benchmark.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| providers | No | Provider names (e.g. ['anthropic', 'mistral']) — runs every active model of those providers that has a valid key. | |
| model_keys | No | Explicit model composite keys. Overrides `providers`. | |
| scenario_id | Yes | UUID of the scenario to run. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||