Skip to main content
Glama

ZeroWidth

Generate a Caliper dataset from a flow (spends credit)

caliper_datasets_generate

Creates a NEW dataset of model-written items — Q&A by default, or scripted sequences / simulated people via shape — pass flowId and the generator reads the flow's prompt, mode, and schema to write realistic cases for THAT flow; description adds guidance (or stands alone when there's no flow). Use this when the user wants test cases fast and has none; prefer caliper_datasets_create with hand-written items when real scenarios are already in hand (Compass pages, a transcript). This SPENDS workspace inference credit (one generator call), so it sits behind the approval gate: say so and expect needs_confirmation. Returns the dataset summary; read the items with caliper_datasets_get and tell the user to review them before trusting an eval built on them.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
countNoItems to generate, 1-50. Default 10.
shapeNoWhat each generated item is: `qa` (one input, the default), `sequence` (2-6 scripted user turns the model answers one at a time, with expectedResponse + expectedBehavior), or `simulated` (a person Caliper plays adaptively: goal, persona, disposition, expectedBehavior). outputsMode applies to `qa` only.
titleYesDataset name.
flowIdNoWorkbench flow to generate cases for (from workbench_flows_list). Required unless description is given.
presetIdNoFraming preset. Default `blank`; the others bias every item toward that attack class.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
visibilityNoWho can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE.
descriptionNoWhat the target system does / what to cover. Required (≥10 chars) when flowId is omitted; optional guidance otherwise.
dispositionNoWith shape `simulated`: how every generated person behaves — a preset id (genuine, pressure, confused, impatient, vague, non-native) or free text. Omit to let the generator vary it from person to person.
outputsModeNoWhat each generated item carries beyond the input: `expected` (golden answers — eval-ready, the default), `none` (inputs only — for a spec others fill in), `captured` (sample answers to rate in a review).
expectedStyleNoWith outputsMode `expected`: `verbatim` literal reference answers (default) or `conditions` — what a correct answer must do, when there's no single right wording.
datasetDescriptionNoDescription stored on the dataset.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: discloses that it SPENDS workspace inference credit (one generator call), sits behind an approval gate producing needs_confirmation, returns only a dataset summary, and requires reading items via caliper_datasets_get before trusting an eval.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded and every clause carries information (routing, credit cost, confirmation flow, follow-up). Sentence length runs long, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-param, no-output-schema mutation tool, the description supplies the missing pieces: cost, approval lifecycle, return shape, and the recommended follow-up call. An agent can invoke this correctly without further inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real meaning by explaining that flowId causes the generator to read the flow's prompt/mode/schema, and how description can substitute for or augment flowId.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Creates a NEW dataset of model-written items') and enumerates the shapes it can produce. It clearly distinguishes itself from the hand-written sibling caliper_datasets_create.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('user wants test cases fast and has none') and when-not-to-use ('prefer caliper_datasets_create with hand-written items when real scenarios are already in hand'). It also flags the approval-gate flow the agent should expect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources