Generate a Caliper dataset from a flow (spends credit)
caliper_datasets_generateCreates a NEW dataset of model-written items — Q&A by default, or scripted sequences / simulated people via shape — pass flowId and the generator reads the flow's prompt, mode, and schema to write realistic cases for THAT flow; description adds guidance (or stands alone when there's no flow). Use this when the user wants test cases fast and has none; prefer caliper_datasets_create with hand-written items when real scenarios are already in hand (Compass pages, a transcript). This SPENDS workspace inference credit (one generator call), so it sits behind the approval gate: say so and expect needs_confirmation. Returns the dataset summary; read the items with caliper_datasets_get and tell the user to review them before trusting an eval built on them.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Items to generate, 1-50. Default 10. | |
| shape | No | What each generated item is: `qa` (one input, the default), `sequence` (2-6 scripted user turns the model answers one at a time, with expectedResponse + expectedBehavior), or `simulated` (a person Caliper plays adaptively: goal, persona, disposition, expectedBehavior). outputsMode applies to `qa` only. | |
| title | Yes | Dataset name. | |
| flowId | No | Workbench flow to generate cases for (from workbench_flows_list). Required unless description is given. | |
| presetId | No | Framing preset. Default `blank`; the others bias every item toward that attack class. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | What the target system does / what to cover. Required (≥10 chars) when flowId is omitted; optional guidance otherwise. | |
| disposition | No | With shape `simulated`: how every generated person behaves — a preset id (genuine, pressure, confused, impatient, vague, non-native) or free text. Omit to let the generator vary it from person to person. | |
| outputsMode | No | What each generated item carries beyond the input: `expected` (golden answers — eval-ready, the default), `none` (inputs only — for a spec others fill in), `captured` (sample answers to rate in a review). | |
| expectedStyle | No | With outputsMode `expected`: `verbatim` literal reference answers (default) or `conditions` — what a correct answer must do, when there's no single right wording. | |
| datasetDescription | No | Description stored on the dataset. |