Skip to main content
Glama

ZeroWidth

Generate more items for a Caliper dataset (spends credit)

caliper_datasets_generate_items

Appends model-written items to an EXISTING dataset, in the style of what's already there (existing items are the few-shot examples; the output shape matches theirs unless overridden). Use to widen coverage when the user says 'more like these' or 'add edge cases'; write them by hand with caliper_datasets_add_items when the scenarios are known. SPENDS workspace inference credit, so it sits behind the approval gate — say so. Existing items and their ratings are untouched.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
countNoItems to add, 1-20. Default 5.
shapeNoWhat each generated item is: `qa` (one input, the default), `sequence` (2-6 scripted user turns the model answers one at a time, with expectedResponse + expectedBehavior), or `simulated` (a person Caliper plays adaptively: goal, persona, disposition, expectedBehavior). outputsMode applies to `qa` only.
datasetIdYesDataset to extend, from caliper_datasets_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
descriptionNoWhat to bias toward, e.g. 'angry customers', 'ambiguous refund questions'.
dispositionNoWith shape `simulated`: how every generated person behaves — a preset id (genuine, pressure, confused, impatient, vague, non-native) or free text. Omit to let the generator vary it from person to person.
outputsModeNoWhat each generated item carries beyond the input: `expected` (golden answers — eval-ready, the default), `none` (inputs only — for a spec others fill in), `captured` (sample answers to rate in a review).
anchorItemIdNoAn item id (from caliper_datasets_get) the new items should resemble most.
expectedStyleNoWith outputsMode `expected`: `verbatim` literal reference answers (default) or `conditions` — what a correct answer must do, when there's no single right wording.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover safety hints (readOnlyHint false, destructiveHint false), while the description adds the economically material facts: it SPENDS workspace inference credit, sits behind an approval gate, and leaves existing items and ratings untouched. These are beyond anything structured fields convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what the tool does, then routing, then the credit/approval warning, then the non-destructive guarantee. Dense but every sentence earns its place; slight compression of the style-inheritance clause could improve readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters, full schema coverage, and no output schema, the description supplies the missing behavioral framing an agent needs: cost, approval flow, how generation derives from existing data, and that prior content is preserved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so per-parameter docs already exist. The description still adds meaning the schema doesn't state: existing items act as few-shot examples and the output shape matches theirs unless overridden. It does not explain count defaults or workspace/approvalId semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('appends model-written items to an EXISTING dataset') and immediately qualifies how the generated content is derived from existing items. It reads clearly as distinct from both caliper_datasets_add_items (manual) and caliper_datasets_generate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the trigger phrasing ('more like these', 'add edge cases') and the alternative (caliper_datasets_add_items) with the condition that selects it ('when the scenarios are known'). It also flags the approval-gate workflow, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources