Create a Caliper dataset
caliper_datasets_createCreates a dataset of test items — THE FIRST STEP of setting up evaluation for a flow. Three item shapes: Q&A (input + optional expectedOutput, the golden answer); SEQUENCE (turns: 2-20 scripted user messages the model answers one at a time with its own earlier replies in front of it, + expectedResponse for the final reply, optional expectedBehavior for the whole conversation); SIMULATED (goal + optional persona/disposition/strategy/maxTurns + expectedBehavior; a platform flow plays a person adaptively, Caliper-run evals only; disposition is a preset id like genuine, pressure, confused, impatient, or free text). Use sequences and simulated items for the slow attacks and for real customers with real needs: a model that holds on message one often folds on message ten. Write good inputs from real usage: the Compass pages the flow was built from are the best source of realistic scenarios.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Test items (1-100). | |
| title | Yes | Dataset name (2-120 chars). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | What this dataset covers. |