Create a Caliper eval
caliper_evals_createBinds a dataset + rubric to something under test as a repeatable eval — THE LAST SETUP STEP before scoring. Two targets: a Workbench flow (pass flowId; Caliper runs inference itself, then caliper_evals_run scores it), or 'external' (target: 'external'; the user's own model runs elsewhere and their script submits outputs through the public API, usually from CI — see the docs guide 'Run evals in CI'). For an external eval, hand back the eval id and tell the user to create a workspace API key with the ci_evals preset in their workspace settings; keys can't be minted from here. flowStage 'draft' evals the live draft (pre-publish); 'published' (default) evals the latest published revision at run time, or one pinned with flowRevisionId. Reuse an existing rubric from caliper_rubrics_list when one already scores this job.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Eval name (2-120 chars). | |
| flowId | No | Workbench flow id to evaluate. Required unless target is 'external'. | |
| target | No | 'workbench_flow' (default; needs flowId) or 'external' (the user's own model; outputs arrive through the public API). | |
| rubricId | Yes | Rubric the judge scores with. | |
| datasetId | Yes | Dataset of test items. | |
| flowStage | No | Which stage to run against. Use 'draft' while iterating pre-publish. Default 'published'. | |
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | ||
| externalLabel | No | External evals only: what produces the outputs, e.g. 'CI' or 'prod pipeline'. Free text, shown on the eval. | |
| flowRevisionId | No | Published-stage evals only: pin to one published revision (id from workbench_flows_revisions_list). Omit to always eval the latest published version at run time. |