Eval performance for a Workbench flow
caliper_flow_performanceHow a Workbench flow is ACTUALLY doing, with receipts: every Caliper eval targeting the flow, recent runs with scores, the latest run decomposed into per-criterion averages, the score delta vs the previous run, and the worst-scoring items WITH the judge's reasoning. Use this BEFORE claiming a flow works or proposing changes — and cite the runId + scores when you do. The worst items are diagnostic: failures clustered around missing company facts suggest a knowledge gap (consider proposing a Compass interview with the workflow owner) rather than a prompt problem.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Workbench flow id to report on. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |