Read one eval run (per-item, per-criterion scores)
caliper_evals_runs_getOne run: status, overall score, and by default the ten lowest-scoring items with their per-criterion scores and the judge's reasoning (items: all for every item, none for totals only). THE ANSWER to 'why did the score drop' — read this, then name the criterion that slipped and quote the reasoning on the lowest items. Works for runs Caliper ran and for runs submitted from CI (triggeredBy 'ci', usually labeled with a pull request number). Still no polling loops; read it when the user asks.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| items | No | Which items to include. `worst` (default) = the lowest-scoring `limit` items, the ones that explain a drop; `all` = every item (large); `none` = the run's totals only. | |
| limit | No | How many items for `worst`, default 10. | |
| runId | Yes | Run id, from caliper_evals_runs_list or caliper_evals_run. | |
| evalId | Yes | Eval the run belongs to. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |