Run a Caliper eval
caliper_evals_runQueues a new run of an eval — inference over the dataset, then LLM-judge scoring. THE VERIFY STEP of the improvement loop: after an approved workbench_flows_edit_text, run the eval again and report the score delta vs the previous run. Runs take a while — but you're brought back into THIS conversation automatically with the scores the moment it finishes, so tell the user it's queued and that you'll follow up here; never poll or ask them to check back. Costs workspace LLM budget, so it sits behind the approval gate: may return needs_confirmation.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Short label for the run, e.g. 'after refund-policy fix'. | |
| evalId | Yes | Eval to run. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |