Get ranked benchmark scenario results
get_benchmark_scenario_resultsFilter, rank and limit a scenario's per-model benchmark results. No LLM call. overall blends quality, speed and cost using organization task weights and is null if a component is missing. Status tags are independent: success does not exclude stale or stale_score results. Missing sort metrics come last in either direction. See enricher://docs/model-benchmark for interpretation.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Keep only the top N after sorting (None = all). | |
| status | No | Keep results carrying ANY of these tags: success | failed | stale (config_hash changed since this run, re-run it) | stale_score (reference/scoring config changed since scored, rescore it) | unscored (ran fine, never scored). Empty/None = every status. | |
| sort_by | No | Metric to sort by. | overall |
| providers | No | Keep only these provider names (empty/None = every provider). | |
| model_keys | No | Keep only these model composite keys (empty/None = every model). | |
| sort_order | No | Sort direction. | desc |
| scenario_id | Yes | UUID of the scenario. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||