ZeroWidth Caliper
Server Details
Build datasets and rubrics, then run evals that score your AI in Caliper.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 45 tools
Most tools target clearly distinct resources and actions, and descriptions explicitly distinguish similar operations like dataset generation versus appending items. A few boundaries are still porous—such as datasets_create/add_items/generate/generate_items and evals_create/run/runs_get—but the descriptions make these readable and actionable.
The caliper_* tools follow a mostly consistent domain_resource_action snake_case pattern, but the set also contains comments_*, entity_tags_*, get_doc, list_docs, search_docs, and search_workspace. The mixed prefixing and verb-first naming in the docs/search tools break full consistency, though names remain readable.
With 45 tools, the server is substantially overloaded for a single MCP surface. While the breadth reflects a large platform, the count forces agents to navigate many near-adjacent operations and increases selection cost.
The toolset covers a broad lifecycle: datasets, evals, runs, rubrics, sources, traces, source feeds, starters, comments, entity tags, documentation, and cross-tool search. Some lifecycle edges are absent—such as direct review/spec creation or flow editing—but the core evaluation and observability surface is well covered.
Available Tools
45 toolscaliper_datasets_add_itemsAdd items to a Caliper datasetAInspect
Appends items to an existing dataset — use this to grow coverage (new edge cases, scenarios from a completed interview) instead of creating a parallel dataset. Same three shapes as caliper_datasets_create (Q&A, sequence, simulated). Existing items and their ratings are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Items to append (1-100). | |
| datasetId | Yes | Id of the dataset to extend. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so the safety profile is known. The description adds real value by disclosing that existing items and their ratings are untouched, i.e. the append is purely additive. It does not mention the needs_confirmation/approvalId flow, which is only visible in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the alternative/create contrast immediately after. Every clause earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an append tool with no output schema and full schema coverage, the description supplies the needed framing: additive semantics, shape parity with create, and non-destruction of existing data. It omits return/response behavior and the approval pathway, which are minor but leave small gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by mapping the items payload to 'the same three shapes as caliper_datasets_create (Q&A, sequence, simulated)', giving useful semantic orientation for how the item fields combine. It adds no per-field detail, so it stops short of 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Appends items to an existing dataset') and explicitly frames the scope contrast with creating a parallel dataset, which separates it from caliper_datasets_create. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use condition ('grow coverage — new edge cases, scenarios from a completed interview') and names the alternative behavior it replaces ('instead of creating a parallel dataset'). The routing decision versus the create sibling is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_createCreate a Caliper datasetAInspect
Creates a dataset of test items — THE FIRST STEP of setting up evaluation for a flow. Three item shapes: Q&A (input + optional expectedOutput, the golden answer); SEQUENCE (turns: 2-20 scripted user messages the model answers one at a time with its own earlier replies in front of it, + expectedResponse for the final reply, optional expectedBehavior for the whole conversation); SIMULATED (goal + optional persona/disposition/strategy/maxTurns + expectedBehavior; a platform flow plays a person adaptively, Caliper-run evals only; disposition is a preset id like genuine, pressure, confused, impatient, or free text). Use sequences and simulated items for the slow attacks and for real customers with real needs: a model that holds on message one often folds on message ten. Write good inputs from real usage: the Compass pages the flow was built from are the best source of realistic scenarios.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Test items (1-100). | |
| title | Yes | Dataset name (2-120 chars). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | What this dataset covers. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral constraint — simulated items are 'Caliper-run evals only' — but says nothing about the notable approvalId/needs_confirmation flow this mutation triggers, leaving a real gap for a write tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and 'first step' positioning, then structured into the three item shapes. It is long, and phrases like the message-ten anecdote and the Compass-source advice are more persuasive than operational, but most content earns its place for a complex nested schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with a nested items array and no output schema, the description covers the item taxonomy thoroughly and explains how each shape is used in evals. The main omission is any mention of the approval/confirmation handshake that the approvalId parameter implies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes beyond it by organizing the nested item parameters into three named shapes (Q&A / SEQUENCE / SIMULATED) and clarifying the role of fields like disposition, expectedResponse, and expectedBehavior. This adds conceptual meaning over the flat per-field schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a dataset of test items') and explicitly positions it as 'THE FIRST STEP of setting up evaluation for a flow', which cleanly separates it from siblings like caliper_datasets_add_items and caliper_datasets_update. An agent knows exactly what this does and where it sits in the workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context ('first step' of eval setup) and prescriptive guidance on when to use sequence/simulated items versus Q&A, plus where to source realistic inputs (Compass pages). It does not explicitly name the alternative sibling for appending items to an existing dataset, so routing is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_deleteDelete a Caliper datasetADestructiveInspect
Deletes a dataset. Evals bound to it stop being runnable (their binding shows datasetOk: false), and reviews/specs over its items lose their source — check caliper_evals_list for evals that reference it and say so before proposing. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| datasetId | Yes | Dataset id, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While annotations already flag destructiveHint=true, the description goes well beyond them by disclosing the downstream blast radius: bound evals stop being runnable (datasetOk: false) and reviews/specs lose their source. It also surfaces the two-step `needs_confirmation` / approvalId flow, which is critical operational context for a two-phase destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Deletes a dataset') and then spends the remaining words on consequences, the cross-tool check, and the confirmation flow. No filler sentences; every clause carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output-schema tool with full schema coverage, this is exactly the right content: consequences, a related-tool verification step, and the confirmation handshake. Nothing an agent needs to invoke it safely and correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three params are already documented and the baseline is 3. The description earns a point above baseline by explaining that the call 'May return `needs_confirmation`,' which clarifies the purpose of the approvalId parameter and how the two-call confirmation sequence works.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Deletes) and resource (a dataset), making it trivially distinguishable from siblings like caliper_datasets_update, caliper_datasets_remove_item, and caliper_datasets_get. The cascade consequences further pin the scope to whole-dataset deletion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit pre-flight step: 'check caliper_evals_list for evals that reference it and say so before proposing.' This names a concrete alternative tool and the condition that invokes it. It does not, however, state when not to delete or the exact criteria for bailing out, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_generateGenerate a Caliper dataset from a flow (spends credit)AInspect
Creates a NEW dataset of model-written items — Q&A by default, or scripted sequences / simulated people via shape — pass flowId and the generator reads the flow's prompt, mode, and schema to write realistic cases for THAT flow; description adds guidance (or stands alone when there's no flow). Use this when the user wants test cases fast and has none; prefer caliper_datasets_create with hand-written items when real scenarios are already in hand (Compass pages, a transcript). This SPENDS workspace inference credit (one generator call), so it sits behind the approval gate: say so and expect needs_confirmation. Returns the dataset summary; read the items with caliper_datasets_get and tell the user to review them before trusting an eval built on them.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Items to generate, 1-50. Default 10. | |
| shape | No | What each generated item is: `qa` (one input, the default), `sequence` (2-6 scripted user turns the model answers one at a time, with expectedResponse + expectedBehavior), or `simulated` (a person Caliper plays adaptively: goal, persona, disposition, expectedBehavior). outputsMode applies to `qa` only. | |
| title | Yes | Dataset name. | |
| flowId | No | Workbench flow to generate cases for (from workbench_flows_list). Required unless description is given. | |
| presetId | No | Framing preset. Default `blank`; the others bias every item toward that attack class. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | What the target system does / what to cover. Required (≥10 chars) when flowId is omitted; optional guidance otherwise. | |
| disposition | No | With shape `simulated`: how every generated person behaves — a preset id (genuine, pressure, confused, impatient, vague, non-native) or free text. Omit to let the generator vary it from person to person. | |
| outputsMode | No | What each generated item carries beyond the input: `expected` (golden answers — eval-ready, the default), `none` (inputs only — for a spec others fill in), `captured` (sample answers to rate in a review). | |
| expectedStyle | No | With outputsMode `expected`: `verbatim` literal reference answers (default) or `conditions` — what a correct answer must do, when there's no single right wording. | |
| datasetDescription | No | Description stored on the dataset. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: discloses that it SPENDS workspace inference credit (one generator call), sits behind an approval gate producing needs_confirmation, returns only a dataset summary, and requires reading items via caliper_datasets_get before trusting an eval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded and every clause carries information (routing, credit cost, confirmation flow, follow-up). Sentence length runs long, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-param, no-output-schema mutation tool, the description supplies the missing pieces: cost, approval lifecycle, return shape, and the recommended follow-up call. An agent can invoke this correctly without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real meaning by explaining that flowId causes the generator to read the flow's prompt/mode/schema, and how description can substitute for or augment flowId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a NEW dataset of model-written items') and enumerates the shapes it can produce. It clearly distinguishes itself from the hand-written sibling caliper_datasets_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('user wants test cases fast and has none') and when-not-to-use ('prefer caliper_datasets_create with hand-written items when real scenarios are already in hand'). It also flags the approval-gate flow the agent should expect.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_generate_itemsGenerate more items for a Caliper dataset (spends credit)AInspect
Appends model-written items to an EXISTING dataset, in the style of what's already there (existing items are the few-shot examples; the output shape matches theirs unless overridden). Use to widen coverage when the user says 'more like these' or 'add edge cases'; write them by hand with caliper_datasets_add_items when the scenarios are known. SPENDS workspace inference credit, so it sits behind the approval gate — say so. Existing items and their ratings are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Items to add, 1-20. Default 5. | |
| shape | No | What each generated item is: `qa` (one input, the default), `sequence` (2-6 scripted user turns the model answers one at a time, with expectedResponse + expectedBehavior), or `simulated` (a person Caliper plays adaptively: goal, persona, disposition, expectedBehavior). outputsMode applies to `qa` only. | |
| datasetId | Yes | Dataset to extend, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| description | No | What to bias toward, e.g. 'angry customers', 'ambiguous refund questions'. | |
| disposition | No | With shape `simulated`: how every generated person behaves — a preset id (genuine, pressure, confused, impatient, vague, non-native) or free text. Omit to let the generator vary it from person to person. | |
| outputsMode | No | What each generated item carries beyond the input: `expected` (golden answers — eval-ready, the default), `none` (inputs only — for a spec others fill in), `captured` (sample answers to rate in a review). | |
| anchorItemId | No | An item id (from caliper_datasets_get) the new items should resemble most. | |
| expectedStyle | No | With outputsMode `expected`: `verbatim` literal reference answers (default) or `conditions` — what a correct answer must do, when there's no single right wording. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover safety hints (readOnlyHint false, destructiveHint false), while the description adds the economically material facts: it SPENDS workspace inference credit, sits behind an approval gate, and leaves existing items and ratings untouched. These are beyond anything structured fields convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool does, then routing, then the credit/approval warning, then the non-destructive guarantee. Dense but every sentence earns its place; slight compression of the style-inheritance clause could improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, full schema coverage, and no output schema, the description supplies the missing behavioral framing an agent needs: cost, approval flow, how generation derives from existing data, and that prior content is preserved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so per-parameter docs already exist. The description still adds meaning the schema doesn't state: existing items act as few-shot examples and the output shape matches theirs unless overridden. It does not explain count defaults or workspace/approvalId semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('appends model-written items to an EXISTING dataset') and immediately qualifies how the generated content is derived from existing items. It reads clearly as distinct from both caliper_datasets_add_items (manual) and caliper_datasets_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger phrasing ('more like these', 'add edge cases') and the alternative (caliper_datasets_add_items) with the condition that selects it ('when the scenarios are known'). It also flags the approval-gate workflow, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_getRead a Caliper dataset (items paged)ARead-onlyInspect
One dataset with a page of its items — read this BEFORE editing items (caliper_datasets_update_item needs the item id) and before extending coverage, so you don't add cases that already exist. Items come limit at a time (default 25) from offset; total is the full count. Long fields are cut at ~800 characters with a truncation marker. Get the datasetId from caliper_datasets_list.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Items per page, 1-100. Default 25. | |
| offset | No | Items to skip. Default 0. | |
| datasetId | Yes | Dataset id, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description still adds real behavioral context beyond them: page size default (25) via limit/offset, that total is the full count not the page, and that long fields are truncated at ~800 chars with a marker.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core operation and the read-before-write guidance; every clause carries distinct information (dependency, paging, truncation, id source) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must cover return shape, and it does: one dataset plus a paged item list, total count, and truncation behavior. Combined with annotations and full schema coverage, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents all four params. The description goes a bit further by framing limit/offset as the paging mechanism and stating the source of datasetId (caliper_datasets_list), tying parameters to workflow rather than restating types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (a Caliper dataset) and adds the scope detail 'with a page of its items', which separates it from caliper_datasets_list (no items) and the item-mutating siblings. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to read this BEFORE editing items and BEFORE extending coverage, and names the dependency (caliper_datasets_update_item needs the item id) plus a preventive reason (avoid adding duplicate cases). When-to-use and the alternative are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_listList Caliper datasetsARead-onlyInspect
Lists datasets in the active workspace. Returns summaries (id, name, item count); items are not included.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so safety is covered. The description adds genuinely useful behavioral context beyond that: the exact return shape (id, name, item count) and the negative constraint that items are not included, which prevents an agent from expecting embedded item data. It omits pagination, ordering, and result-size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the scope ('active workspace') and the return contract front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on the burden of describing what is returned, and it does so briefly and accurately. It leaves pagination and ordering unspecified, which is a minor gap for a list tool with a single optional filter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one optional parameter and schema description coverage is 100%; the schema fully documents the workspace slug, including the personal-token vs API-key behavior. The description adds no parameter-level detail beyond the phrase 'active workspace', so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists datasets') with a scope qualifier ('in the active workspace'), so the operation is unambiguous. It stops short of naming the sibling it contrasts with (caliper_datasets_get for single-dataset detail), relying on 'items are not included' to imply the distinction rather than stating it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Use is implied by 'Lists datasets in the active workspace' and by the caveat that items are excluded, which hints an agent should call caliper_datasets_get for detail. No explicit when-to-use, when-not-to-use, or named alternative is given, so the agent must infer the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_remove_itemRemove one item from a Caliper datasetADestructiveInspect
Drops one item. Ratings and run results keyed to it are orphaned (kept in history, gone from the dataset). A dataset must keep at least one item. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | Item id, from caliper_datasets_get. | |
| datasetId | Yes | Dataset id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description goes well beyond them: it discloses the exact consequence of the mutation (ratings and run results are orphaned — kept in history, removed from the dataset), an invariant (minimum of one item), and a non-trivial response behavior (may return `needs_confirmation`). This is precisely the added context destructive annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then consequences, then constraints/response behavior. Every sentence carries distinct information and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no output schema, the description covers the destructive effect, the precondition, and the confirmation round-trip. Minor gaps remain (no mention of permissions or what the return payload looks like beyond needs_confirmation), but the schema covers the remaining parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: it explains the trigger for `needs_confirmation` and therefore the existence/purpose of the `approvalId` parameter, which the schema only describes circularly. It does not clarify `workspace` scoping, so it is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb+resource ("Drops one item"), which is clear and scoped. However, it largely restates the tool title ("Remove one item from a Caliper dataset") rather than distinguishing itself from siblings like caliper_datasets_delete (whole dataset) or caliper_datasets_update_item; the differentiation is inferred, not stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives operational constraints ("A dataset must keep at least one item") and a confirmation flow, but never says when to reach for this tool versus caliper_datasets_update_item, caliper_datasets_delete, or other siblings. No when-not guidance either.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_updateRename / re-describe / re-scope a Caliper datasetADestructiveInspect
Changes a dataset's title, description, or visibility. Pass only what changes; items are untouched (use caliper_datasets_add_items / caliper_datasets_update_item / caliper_datasets_remove_item for those). May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New name. | |
| datasetId | Yes | Dataset id, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description; null clears it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so safety is partly covered, but the description adds non-obvious behavior: this is a PATCH-style partial update where unspecified fields survive, and the call may return a `needs_confirmation` state requiring a follow-up with an approvalId. It does not explain what is actually destructive here (e.g. a visibility downgrade revoking access), which is the one gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences/fragments, front-loaded with the primary action and scope, then the exclusion with alternatives, then the confirmation caveat. No filler and nothing rhetorical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still flags the one non-obvious return path (`needs_confirmation`) and the schema documents the approvalId continuation. For a six-parameter mutation with destructive semantics, it is complete enough, though the lack of any note on which changes are destructive or reversible leaves a minor hole.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including a fully documented visibility enum with a natural-language mapping and the approvalId/workspace auth notes, so the schema does the heavy lifting. The description only echoes the three top-level fields and adds the patch-semantics cue, which does not meaningfully exceed what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (changes) plus the exact mutable fields (title, description, visibility) of a named resource (dataset), which matches the title. It also explicitly excludes dataset item content and names the siblings that own those operations, so the agent can separate this from add_items/update_item/remove_item without reading schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Supplies an explicit when-not rule ('items are untouched') and names three alternative tools for that case, which is exactly the when/when-not/alternatives pattern. The 'Pass only what changes' instruction also tells the agent how to call it (partial update) versus a full replacement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_update_itemReplace one item in a Caliper datasetADestructiveInspect
Rewrites one item's content, keeping its id (ratings and eval results keyed to it stay attached). The item's KIND can't change — a Q&A item stays Q&A, a sequence stays a sequence — so pass the same shape you read from caliper_datasets_get. Chat and raw items can't be edited from here. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| item | Yes | The full replacement item, same three shapes as caliper_datasets_create. | |
| itemId | Yes | Item id, from caliper_datasets_get. | |
| datasetId | Yes | Dataset id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only tell the agent this is a destructive non-read-only write; the description adds the important fallout semantics — the id and its attached ratings/eval results survive, the item KIND is immutable, and the call can return `needs_confirmation`. That is materially more than the annotation set provides and does not contradict it (the destructiveHint refers to overwriting the item's content in place).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler: the core action and its id-preservation consequence come first, the kind constraint and shape instruction second, the exclusions and confirmation caveat last. Each sentence carries information an agent needs before invoking.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the essential gotchas (immutable kind, preserved id/ratings, non-editable item types, possible confirmation response). It would be complete if it explained how the returned `needs_confirmation` maps onto the `approvalId` parameter, since that round-trip is otherwise left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description still adds real meaning by constraining the `item` payload to the same three shapes returned by caliper_datasets_get and warning that the shape/kind cannot change. The `approvalId`/`needs_confirmation` connection is only hinted at rather than explained, keeping this short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ("Rewrites one item's content") and pins the scope to a single item within a dataset, which separates it from the dataset-level caliper_datasets_update and from caliper_datasets_add_items. It further clarifies the item retains its id, so an agent knows this is an in-place replacement rather than an add/remove.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear prerequisites ("pass the same shape you read from caliper_datasets_get") and explicit exclusions ("Chat and raw items can't be edited from here"), plus the kind-preservation constraint. It stops short of naming the alternative tool for those excluded shapes or for replacing an item by delete+add, so it is clear context without full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_createCreate a Caliper evalAInspect
Binds a dataset + rubric to something under test as a repeatable eval — THE LAST SETUP STEP before scoring. Two targets: a Workbench flow (pass flowId; Caliper runs inference itself, then caliper_evals_run scores it), or 'external' (target: 'external'; the user's own model runs elsewhere and their script submits outputs through the public API, usually from CI — see the docs guide 'Run evals in CI'). For an external eval, hand back the eval id and tell the user to create a workspace API key with the ci_evals preset in their workspace settings; keys can't be minted from here. flowStage 'draft' evals the live draft (pre-publish); 'published' (default) evals the latest published revision at run time, or one pinned with flowRevisionId. Reuse an existing rubric from caliper_rubrics_list when one already scores this job.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Eval name (2-120 chars). | |
| flowId | No | Workbench flow id to evaluate. Required unless target is 'external'. | |
| target | No | 'workbench_flow' (default; needs flowId) or 'external' (the user's own model; outputs arrive through the public API). | |
| rubricId | Yes | Rubric the judge scores with. | |
| datasetId | Yes | Dataset of test items. | |
| flowStage | No | Which stage to run against. Use 'draft' while iterating pre-publish. Default 'published'. | |
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | ||
| externalLabel | No | External evals only: what produces the outputs, e.g. 'CI' or 'prod pipeline'. Free text, shown on the eval. | |
| flowRevisionId | No | Published-stage evals only: pin to one published revision (id from workbench_flows_revisions_list). Omit to always eval the latest published version at run time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare this is a non-destructive, non-open-world write; the description adds real operational context beyond that: external evals require a workspace API key with the ci_evals preset that 'can't be minted from here,' and it explains the draft-vs-published execution behavior at run time. These are meaningful prerequisites an agent would otherwise not know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The lead sentence front-loads the core purpose and lifecycle position, and the remaining text is information-dense rather than filler. It runs long and a couple of clauses are heavily packed, but given 12 parameters and two targets, nearly every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description steps in to state what to do with the result ('hand back the eval id and tell the user to create a workspace API key'), and it covers the branching setup, prerequisites, and next-step tool. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 92%, so the schema does most of the work (baseline 3). The description still adds relational meaning the schema can't express — flowId is only needed for the workbench target, flowRevisionId applies only to the published stage, and the draft/published distinction is spelled out — lifting it above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — it 'binds a dataset + rubric to something under test as a repeatable eval' — and positions itself in the lifecycle as 'THE LAST SETUP STEP before scoring.' It also distinguishes its job from caliper_evals_run, which does the scoring, so an agent can separate it from siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes the two invocation paths (workbench flow via flowId vs target:'external'), when to pick each, when to use flowStage 'draft' vs 'published', and routes to caliper_rubrics_list when a rubric already exists. It even points at the 'Run evals in CI' docs guide for the external path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_deleteDelete a Caliper evalADestructiveInspect
Deletes an eval and stops its schedule. Run history is kept but no longer reachable from the eval. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| evalId | Yes | Eval id, from caliper_evals_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known, but the description adds real value beyond them: the schedule is stopped, run history becomes unreachable, and the call may return needs_confirmation. It stops short of stating auth/permission requirements or whether deletion is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying distinct information (what is deleted, what happens to history, what a confirmation response implies), with the primary action front-loaded. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter destructive tool with no output schema, the description covers the salient outcomes an agent needs: the eval is gone, its schedule stops, history is orphaned, and a confirmation round-trip may be required. Only the permission/auth prerequisites are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description earns above that by explaining the needs_confirmation/approvalId retry cycle, which gives the approvalId parameter meaning the schema alone only partially conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Deletes an eval') and immediately names the collateral effect ('stops its schedule'). It reads clearly against caliper_evals_update/run/get, though it does not explicitly differentiate itself from the closer sibling caliper_evals_runs_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no statement of prerequisites, and no pointer to alternatives (e.g., cancel a run vs. delete the eval). The only usage-adjacent signal is the approval flow implied by the closing sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_getRead one Caliper evalARead-onlyInspect
One eval: what it targets (Workbench flow + stage, or external), its dataset and rubric ids, the rubric criteria it scores with (the snapshot taken at creation), schedule + runOnPublish + regression settings, run counts, latest score, and binding health (datasetOk / flowOk false = a run would fail at resolution — say so before proposing caliper_evals_run). Get the evalId from caliper_evals_list or caliper_flow_performance.
| Name | Required | Description | Default |
|---|---|---|---|
| evalId | Yes | Eval id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this as a safe read (readOnlyHint=true, destructiveHint=false), so the safety profile is covered. The description adds genuine behavioral context beyond that: the rubric criteria are a creation-time snapshot, and datasetOk/flowOk=false means a run would fail at resolution, which is decision-relevant context the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is front-loaded and every clause describes a distinct returned field, so nothing is filler. It is a single dense run-on sentence, which slightly hurts scanability for an agent parsing it quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of describing the return payload, and it does so comprehensively: it lists every meaningful field, explains the binding-health semantics, and routes the agent to a follow-up action. Nothing an agent needs to call or interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both evalId and workspace are already fully documented by the schema, setting the baseline at 3. The description adds only a sourcing hint for evalId (get it from caliper_evals_list / caliper_flow_performance) rather than any new format or constraint meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates precisely what a single eval record contains (target, dataset/rubric ids, rubric snapshot, schedule, run counts, latest score, binding health), so the agent knows exactly what this returns. The verb itself is only implied via the name/title, and sibling differentiation against caliper_evals_list is present but indirect (it only says where to source the evalId).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete task-oriented guidance: check binding health and 'say so before proposing caliper_evals_run,' which is a real precondition an agent should act on. It names caliper_evals_list and caliper_flow_performance as where to obtain the id, but states no explicit when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_listList Caliper evalsARead-onlyInspect
Lists evals in the active workspace with their target config (which Workbench flow, which dataset/rubric). Use to find the evalId for caliper_evals_run.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds that results are scoped to the active workspace and include target config, but says nothing about pagination, result ordering, or volume.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The resource/scope statement comes first and the routing hint second, so the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing what each listed eval contains (target config with flow and dataset/rubric). For a simple read-only list tool with full schema coverage on its one parameter, this is nearly sufficient; only pagination/ordering behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single optional workspace parameter with 100% schema description coverage, so the schema carries the semantics (default workspace, override behavior, API-key case). The description's 'active workspace' phrasing aligns with that. Zero-param-adjacent baseline applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) plus resource (evals) and narrows the scope to the active workspace, then names the returned content (target config: which Workbench flow, dataset/rubric). This distinguishes it from caliper_evals_get (single eval) and caliper_evals_run, which it routes the agent toward.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: to find the evalId needed by caliper_evals_run. That is a clear downstream purpose, but it names no alternative for other discovery needs (e.g. caliper_evals_get for a known id) and gives no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runRun a Caliper evalAInspect
Queues a new run of an eval — inference over the dataset, then LLM-judge scoring. THE VERIFY STEP of the improvement loop: after an approved workbench_flows_edit_text, run the eval again and report the score delta vs the previous run. Runs take a while — but you're brought back into THIS conversation automatically with the scores the moment it finishes, so tell the user it's queued and that you'll follow up here; never poll or ask them to check back. Costs workspace LLM budget, so it sits behind the approval gate: may return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Short label for the run, e.g. 'after refund-policy fix'. | |
| evalId | Yes | Eval to run. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly=false, destructive=false, openWorld=true); the description adds substantial context beyond them — the run is long-running and re-enters THIS conversation automatically on completion, it consumes workspace LLM budget, and it may return a `needs_confirmation` envelope behind an approval gate. That is exactly the behavioral detail an agent needs to avoid polling or misreporting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and well organized, but the mid-section sentence about auto-return, budget, and approval is dense and runs long. Nearly every sentence earns its place, so it stays efficient despite the density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining that the scores arrive via an automatic follow-up in this conversation, plus the approval-gate response shape. For a cost-bearing async mutation tool, nothing essential for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (evalId, label, workspace, approvalId) are already documented in the schema; the description's mention of the approval gate loosely motivates approvalId but adds no syntax or format detail. Baseline 3 applies when the schema carries the parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Queues a new run of an eval' — and immediately defines the operation as inference over the dataset followed by LLM-judge scoring. It is clearly distinguishable from siblings like caliper_evals_runs_get, caliper_evals_runs_list, and caliper_evals_runs_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions the tool as 'THE VERIFY STEP of the improvement loop' and gives the trigger condition: after an approved workbench_flows_edit_text, run the eval again and report the score delta. It also tells the agent not to poll and what to tell the user, which fully closes the usage loop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_cancelStop a running Caliper eval runADestructiveInspect
Stops a run that is still PENDING / RUNNING / SCORING — the brake on a run that's spending more than expected or was started by mistake. The queue stops at once; an item already handed to the flow finishes on its own timeout. The run settles as FAILED with a 'cancelled' reason and keeps the items it completed. A run that already finished returns run_not_live. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | Run id, from caliper_evals_run or caliper_evals_runs_list. | |
| evalId | Yes | Eval the run belongs to. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint=true annotation by disclosing exactly how destruction is scoped: the queue stops immediately, an in-flight item finishes on its own timeout, the run settles as FAILED with a 'cancelled' reason, and completed items are retained. It also surfaces the run_not_live failure mode and a possible needs_confirmation response that requires a follow-up call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the operation and its eligible states, then layers consequence, error, and confirmation behavior in short sentences with no filler. Every sentence carries information an agent needs before calling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by covering the terminal state of the run, what is preserved, and both the error and confirmation paths. Combined with destructiveHint annotations and fully documented params, nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so runId, evalId, workspace, and approvalId are all documented in the schema; the description adds only the indirect hint that needs_confirmation implies a second call carrying an approvalId. Baseline 3 is appropriate given the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Stops') and resource ('a run') with an explicit eligible-state set (PENDING / RUNNING / SCORING), which distinguishes it from read-only siblings like caliper_evals_runs_get and caliper_evals_runs_list. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete motivating contexts — a run spending more than expected, or started by mistake — and an implicit exclusion via 'A run that already finished returns run_not_live.' It does not name an alternative tool for the already-finished case, but no sibling offers a competing cancellation path, so little is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_getRead one eval run (per-item, per-criterion scores)ARead-onlyInspect
One run: status, overall score, and by default the ten lowest-scoring items with their per-criterion scores and the judge's reasoning (items: all for every item, none for totals only). THE ANSWER to 'why did the score drop' — read this, then name the criterion that slipped and quote the reasoning on the lowest items. Works for runs Caliper ran and for runs submitted from CI (triggeredBy 'ci', usually labeled with a pull request number). Still no polling loops; read it when the user asks.
| Name | Required | Description | Default |
|---|---|---|---|
| items | No | Which items to include. `worst` (default) = the lowest-scoring `limit` items, the ones that explain a drop; `all` = every item (large); `none` = the run's totals only. | |
| limit | No | How many items for `worst`, default 10. | |
| runId | Yes | Run id, from caliper_evals_runs_list or caliper_evals_run. | |
| evalId | Yes | Eval the run belongs to. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a safe, read-only, non-destructive, closed operation, so the safety profile is covered. The description adds genuine output-shape context: by default returns the ten lowest-scoring items, 'all' is large, 'none' yields totals only, and it works for CI-submitted runs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the return content ('One run: status, overall score...') and stays compact, though the capitalized 'THE ANSWER' aside is slightly editorial. Every sentence carries information about output or usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey return values, and it does: status, overall score, per-item per-criterion scores, judge reasoning, and the items/limit default behavior. An agent has enough to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates the items behavior (all/none) that the schema already documents and adds only the 'explain a drop' framing, adding little beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: reads one eval run, returning status, overall score, and by default the lowest-scoring items with per-criterion scores and judge reasoning. It clarifies the run types it covers (Caliper-run and CI-submitted), though it never names sibling tools like caliper_evals_runs_list to differentiate by task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: 'THE ANSWER to why did the score drop — read this, then name the criterion that slipped,' and an explicit exclusion, 'Still no polling loops; read it when the user asks.' It does not name an alternative tool to use instead when the user just wants a run list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_item_executionRead the flow trace behind one scored itemARead-onlyInspect
What the flow actually DID on one run item: every step in order (what it said, which tools it called with what arguments, what came back) and the final answer, plus status, duration, and cost. Read this when a low score needs explaining beyond the judge's reasoning — a wrong tool call or an empty tool result is usually the cause, and the fix is different from a prompt fix. Pass the run item's id (NOT itemId) from caliper_evals_runs_get. Items whose outputs were supplied from outside (external evals) have no trace and return not_found. Steps are capped for transport; the run page in Caliper has the full record.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | Run id. | |
| evalId | Yes | Eval the run belongs to. | |
| runItemId | Yes | The run item's `id` from caliper_evals_runs_get (not its dataset itemId). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-destructive read (readOnlyHint=true, destructiveHint=false, openWorldHint=false), and the description adds real behavioral context beyond that: not_found for external-eval items, transport-level step capping, and a pointer to the full record on the Caliper run page. It does not discuss auth or pagination semantics for the capped step list, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all load-bearing: what it returns, when to reach for it, and the two failure/limitation caveats. It is front-loaded with the payload description. Minor redundancy between the schema's runItemId note and the description's id-vs-itemId warning costs it a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of describing returns — and it does, enumerating ordered steps (utterances, tool calls with arguments, tool results), final answer, status, duration and cost. Combined with the not_found and truncation caveats, an agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including the note that runItemId is the item's `id` rather than its dataset itemId. The description's "pass the run item's `id` (NOT `itemId`) from caliper_evals_runs_get" largely restates that schema text, adding only reinforcement and the sibling that produces the value. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns what the flow actually DID on one run item — ordered steps, tool calls with arguments, tool results, final answer, plus status, duration and cost. This is clearly distinguishable from caliper_evals_runs_get (which yields scores/metadata) and caliper_traces_get, so an agent can pick it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ("Read this when a low score needs explaining beyond the judge's reasoning") and explains the diagnostic value (a wrong tool call or empty result is usually the cause, and the fix differs from a prompt fix). It also names a when-not case: externally supplied outputs have no trace and return not_found.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_listList an eval's runs (status + scores)ARead-onlyInspect
Recent runs for one eval, newest first: status (PENDING/RUNNING/SCORING/DONE/FAILED), overall score once DONE, label, and timestamps. THE CHECK-BACK for caliper_evals_run: when the user asks how the run went, read this — cite the run id and score, and compare against the PREVIOUS run's score for the delta. Still don't poll in a loop; check when the user asks or when reporting.
| Name | Required | Description | Default |
|---|---|---|---|
| evalId | Yes | Eval whose runs to list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safe-read profile (readOnlyHint, destructiveHint=false). The description goes well beyond them: newest-first ordering, the PENDING/RUNNING/SCORING/DONE/FAILED lifecycle, that a score exists only once DONE, and an anti-polling behavioral constraint. That is substantial behavioral context layered on top of the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource and return fields, then the routing guidance. Dense but every clause carries information; the em-dash-heavy final sentence runs slightly long but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly enumerates the returned fields (status, score once DONE, label, timestamps, ordering) and pairs that with the delta-comparison workflow. Nothing an agent needs to call and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so evalId and workspace are already fully documented. The description adds only the single-eval scoping ('for one eval') and does not elaborate on the workspace/permission nuances beyond what the schema states. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Recent runs for one eval') plus scope ('newest first') and the exact fields returned (status, score, label, timestamps). An agent can immediately distinguish this from caliper_evals_get and caliper_evals_runs_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames itself as 'THE CHECK-BACK for caliper_evals_run', names the triggering condition ('when the user asks how the run went'), and adds a when-not rule ('don't poll in a loop; check when the user asks or when reporting'). This is exactly the when/when-not/related-tool guidance the dimension asks for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_updateUpdate a Caliper eval (metadata, schedule, alerts)ADestructiveInspect
Changes an eval's title, description, visibility, schedule (interval + time anchor), runOnPublish (queue a run whenever the flow publishes), or regressionThreshold (score drop vs the previous run that triggers an alert; null = off). Pass only what changes. The dataset, rubric snapshot, and target are fixed at creation — create a new eval to change those. Scheduled and on-publish runs spend credit on their own, so state that plainly when turning them on. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New name. | |
| evalId | Yes | Eval id, from caliper_evals_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description; null clears it. | |
| runOnPublish | No | Run whenever the targeted flow publishes a version. | |
| scheduleTime | No | Daily/weekly anchor, HH:MM 24-hour in scheduleTimezone. Null = one interval from now. | |
| scheduleInterval | No | Run cadence; null clears the schedule. Each scheduled run spends credit — say so. | |
| scheduleTimezone | No | IANA timezone for scheduleTime, e.g. America/Chicago. | |
| scheduleDayOfWeek | No | Weekly only: 0 = Sunday … 6 = Saturday. | |
| regressionThreshold | No | Score drop vs the previous run that raises an alert; null turns alerts off. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and readOnlyHint=false, so the safety profile is already partly covered; the description still adds real value by disclosing the approval flow ('May return needs_confirmation'), the ongoing credit cost of scheduled/on-publish runs, and which fields are permanently fixed after creation. It never explains what 'destructive' means here (e.g., that null clears description/schedule), leaving one gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the mutable fields in the first sentence, then immutability, then confirmation/credit notes — a logical order with little waste. The parenthetical glosses on each field make it slightly long, but every clause carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-param mutation tool with no output schema, the description covers mutation scope, immutability of non-editable fields, the confirmation/approval round-trip, and the cost side effects of enabling schedules — everything an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter (including 'null turns alerts off' and the timezone/HH:MM formats). The description restates a few of these (threshold semantics, runOnPublish trigger) without adding new syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Changes/Update) plus the resource (a Caliper eval) and enumerates exactly which attributes are mutable (title, description, visibility, schedule, runOnPublish, regressionThreshold). The immutable-field callout sharpens the boundary versus caliper_evals_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use rules ('Pass only what changes') and an explicit alternative with condition ('The dataset, rubric snapshot, and target are fixed at creation — create a new eval to change those'). It also instructs the agent to surface credit spend when enabling scheduled/on-publish runs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_flow_performanceEval performance for a Workbench flowARead-onlyInspect
How a Workbench flow is ACTUALLY doing, with receipts: every Caliper eval targeting the flow, recent runs with scores, the latest run decomposed into per-criterion averages, the score delta vs the previous run, and the worst-scoring items WITH the judge's reasoning. Use this BEFORE claiming a flow works or proposing changes — and cite the runId + scores when you do. The worst items are diagnostic: failures clustered around missing company facts suggest a knowledge gap (consider proposing a Compass interview with the workflow owner) rather than a prompt problem.
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Workbench flow id to report on. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description adds real behavioral value beyond that: it discloses the composite/aggregated nature of the output (deltas, per-criterion decomposition, judge reasoning) and gives interpretive guidance on failure clustering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the 'with receipts' framing, then progresses through outputs, usage trigger, and diagnostics in a logical order. Dense but each clause carries information; slightly verbose with the long colon-separated inventory, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry the return-value burden, and it does so thoroughly (eval coverage, run history, per-criterion averages, deltas, judge reasoning). Combined with the 100% schema coverage on inputs, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so flowId and workspace are already documented in the schema. The description adds no additional semantics about the parameters themselves (its 'runId' mention is about output citation, not an input). Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('how a Workbench flow is actually doing') and enumerates the exact contents of the report: eval coverage, recent runs with scores, per-criterion averages, delta vs previous run, and worst items with judge reasoning. This distinguishes it clearly from raw siblings like caliper_evals_runs_list or caliper_evals_runs_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use this BEFORE claiming a flow works or proposing changes') and a follow-through requirement (cite runId + scores). It also explains how to interpret worst items diagnostically. It does not name alternative sibling tools (e.g., evals_runs_list for raw run data), so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_createCreate a Caliper rubricAInspect
Creates the scoring rubric an eval's LLM judge uses — 1-10 criteria, each scored on a numeric scale (default 1-5), judged pass/fail (kind pass_fail), or checked in code with no judge (kind check: contains, not_contains, matches a regex, valid_json, max_chars, equals_expected). Write criteria about the FLOW'S JOB (accuracy to source material, tone, refusal behavior), not generic 'quality'.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Rubric name (2-120 chars). | |
| criteria | Yes | ||
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the create-and-safe profile is covered. The description adds the default scale (1-5) and the semantics of each criteria kind, but does not mention the approval/confirmation flow implied by approvalId or any workspace scoping behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the guiding heuristic is well placed, but the middle sentence is a dash-chained dump of every check type that reproduces the schema enum verbatim. It is dense without being fully wasteful, but those enumerated values do not earn their place in the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description plus the fairly rich input schema together cover the required title/criteria and their structure well. Remaining gaps (approvalId, visibility, workspace) are already documented in the schema, so the description is nearly complete for this create operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema documents most fields. The description restates the kind values (scale, pass_fail, check) and enumerates the check types, largely duplicating schema enum descriptions rather than adding syntax or constraints. Baseline 3 is appropriate for mid-level coverage where the description re-explains rather than extends.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (creates the scoring rubric) and pins down its domain role ('an eval's LLM judge uses'), plus the concrete shape (1-10 criteria, three kinds). It is clearly distinguishable from rubrics_get/list/update/delete siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers real guidance on how to author criteria ('about the FLOW'S JOB ... not generic quality'), which is valuable. However, it never says when to reach for this tool versus caliper_rubrics_update or how it relates to evals_create — usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_deleteDelete a Caliper rubricADestructiveInspect
Deletes a rubric. Existing evals keep their snapshot of it and keep running; nothing new can bind to it. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| rubricId | Yes | Rubric id, from caliper_rubrics_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds real value beyond them: existing evals retain a snapshot and keep running, nothing new can bind to the rubric, and a `needs_confirmation` response may occur (linking to the approvalId param). This is exactly the consequence-level context a destructive tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no waste; the core action leads and the consequences and confirmation behavior follow compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output-schema tool, the description covers the key behavioral facts (retention, binding block, confirmation flow). Minor omissions remain, such as explicit irreversibility wording, but nothing an agent critically needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented (including the approvalId/needs_confirmation linkage). The description adds nothing beyond what the schema provides for parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Deletes a rubric'), which is unambiguous against the sibling set (rubrics_create/get/list/update). It does not explicitly differentiate itself from alternative lifecycle operations, but the operation is clear from the phrasing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains consequences of deleting but gives no guidance on when to delete versus alternatives (e.g., update vs delete), nor any prerequisites for invoking it. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_getRead a Caliper rubric (criteria + scales)ARead-onlyInspect
One rubric with its criteria — each criterion's name, what the judge looks for, and the score scale. Read this to explain a score (which criterion slipped and what it asks for) or before caliper_rubrics_update. Get the rubricId from caliper_rubrics_list or an eval's rubricId.
| Name | Required | Description | Default |
|---|---|---|---|
| rubricId | Yes | Rubric id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and a closed world, so the safety profile is covered. The description adds real value by describing the returned structure (criteria, judge focus, scale), which matters because there is no output schema, though it says nothing about pagination or size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; contents are front-loaded, followed by usage conditions and the id source. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only fetch with no output schema, the description covers return shape, usage context, and prerequisite sourcing of the required id. An agent has everything needed to select and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by telling the agent where to source the required rubricId (rubrics_list or an eval's rubricId) and notes the read-before-update relationship. That provenance guidance is genuine added meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('One rubric with its criteria') and enumerates the payload contents (criterion name, what the judge looks for, score scale). It is cleanly distinguishable from caliper_rubrics_list and caliper_rubrics_update in the sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives two when-to-use conditions ('to explain a score' and 'before caliper_rubrics_update') and names the alternatives for obtaining the required id ('caliper_rubrics_list or an eval's rubricId'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_listList Caliper rubricsARead-onlyInspect
Lists the rubrics in the active workspace (id, name, description, criterion count). Check here BEFORE caliper_rubrics_create — reuse an existing rubric's id in caliper_evals_create when one already scores the same job. Read a rubric's criteria with caliper_rubrics_get.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful context beyond that: the listing is scoped to the active workspace and returns id, name, description, and criterion count, which helps an agent interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: what it returns, when to call it relative to create, and where to read criteria. The workflow-critical instruction is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by naming the returned fields, and the single optional parameter is fully documented. An agent has everything needed to call it and act on the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional parameter, so the schema already carries the full load. The description adds no syntax or format detail about the workspace parameter beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (rubrics), scopes it to the active workspace, and enumerates the returned fields. It is clearly distinguishable from sibling tools like caliper_rubrics_get and caliper_rubrics_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to check here BEFORE caliper_rubrics_create and gives the reason (reuse an existing rubric id in caliper_evals_create), and routes deep reading to caliper_rubrics_get. When-to-use and alternative tool are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_updateUpdate a Caliper rubricADestructiveInspect
Changes a rubric's title, description, visibility, or replaces its criteria wholesale (pass the FULL list — criteria get fresh ids). Evals snapshot the rubric when they're created, so an existing eval keeps scoring with the criteria it started with; say so, and offer to create a new eval when the criteria change materially. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New name. | |
| criteria | No | Replacement criteria (1-10) — the whole list, not a diff. | |
| rubricId | Yes | Rubric id, from caliper_rubrics_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description; null clears it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond destructiveHint=true by disclosing that criteria are replaced wholesale with fresh ids, that existing evals keep scoring against their snapshotted criteria, and that a needs_confirmation response may come back. This is meaningful behavioral context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The what-it-does clause is front-loaded and the criteria-replacement warning follows immediately. The embedded instruction to 'say so, and offer to create a new eval' is slightly meta but earns its place as routing guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description covers the mutation semantics, the confirmation flow, and the eval-snapshot consequence, while parameters are fully documented in the schema. Only the mechanics of the needs_confirmation response are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics the schema doesn't carry: that criteria ids are regenerated on replacement and that the eval snapshot behavior affects whether the change is safe to make.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+field set: changes a rubric's title, description, visibility, or replaces its criteria. An agent can immediately distinguish this from caliper_rubrics_create/delete/get/list without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when the operation applies (changing mutable rubric fields, wholesale criteria replacement) and workflow guidance around eval snapshots, but never states explicit exclusions versus the create/delete siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_source_feeds_listList a source's live datasetsARead-onlyInspect
Datasets this source keeps adding to as conversations arrive: which dataset, for what (review, eval, spec), the pick it matches, sampling, how many have landed, and whether it's still running or why it stopped.
| Name | Required | Description | Default |
|---|---|---|---|
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and destructiveHint=false, so safety is covered. With no output schema, the description usefully discloses the return shape (dataset, purpose, matching pick, sampling, count landed, running/stopped status), which is genuine behavioral context beyond the annotations. It omits pagination/rate-limit behavior, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a compact colon list of returned fields. Dense but every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description compensates by enumerating the returned fields, and the annotations cover the safety profile while the schema covers both parameters. Nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with sourceId and workspace fully documented in the schema, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource: datasets a source keeps appending to as conversations arrive, i.e. the source's live feeds. This is specific and distinguishable from generic dataset listing, but it never names or contrasts with siblings like caliper_source_feeds_stop or caliper_datasets_list, so the boundary must be inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no mention of the alternative caliper_source_feeds_stop or how this differs from caliper_datasets_list. The agent must infer context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_source_feeds_stopStop a live datasetADestructiveInspect
Stops a source from adding new conversations to a dataset. What it already added stays. Get the feedId from caliper_source_feeds_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| feedId | Yes | Live dataset id, from caliper_source_feeds_list. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds real behavioral value beyond them: it clarifies the destruction is scoped to future additions only ('What it already added stays') and warns that the call 'May return needs_confirmation', which maps to the approvalId two-step flow. It stops short of explaining the confirmation response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the effect of the operation front-loaded before the parameter hint and the confirmation caveat. Every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and an approval flow, the description covers effect, scope of data retention, prerequisite id, and the possible confirmation response. It could say slightly more about how to proceed after needs_confirmation, but the approvalId parameter fills that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all four parameters including workspace and approvalId are already documented in the schema. The description only reinforces the feedId provenance, adding no format or constraint detail beyond what the schema states. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with precise scope: 'Stops a source from adding new conversations to a dataset.' The follow-up 'What it already added stays' distinguishes it from removal-style siblings like caliper_datasets_remove_item, so an agent can differentiate without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies a precondition for invoking the tool ('Get the feedId from caliper_source_feeds_list'), which is useful routing context. However, it never states when to stop a feed versus deleting the dataset or removing items, and gives no exclusions, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_sources_listList Caliper sourcesARead-onlyInspect
Apps sending their agent's traces to Caliper — from their own code, OpenTelemetry, or a published Workbench flow — busiest first: name, how it sends, traces in the last 14 days (per day), and the tools its agent calls most with the share of traces using each. Start here when the user asks how their agent behaves on real traffic; then caliper_sources_over_time for one source.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, openWorldHint=false), so the description only needs to add context — which it does: sort order ('busiest first'), the 14-day window for trace counts, and the shape of the payload. It omits pagination/limits and auth expectations, keeping it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource definition and packed with useful detail in two sentences. Density is high and the nested em-dash clauses ask for careful reading, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the burden of describing returns — and it does, enumerating source name, ingestion method, per-day trace counts, and top tools with shares. It stops short of mentioning result limits or paging, a minor gap for a list endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage on a single parameter, the schema already documents the workspace slug, defaults, and API-key behavior fully; the description adds nothing about parameters. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list Caliper sources) and immediately defines what a 'source' is — an app sending traces via its own code, OpenTelemetry, or a published Workbench flow. This clearly distinguishes it from siblings like caliper_sources_over_time, caliper_source_feeds_list, and caliper_sources_to_dataset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to reach for it ('Start here when the user asks how their agent behaves on real traffic') and names the follow-on tool with its condition ('then caliper_sources_over_time for one source'). This is textbook routing guidance with an alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_sources_over_timeRead a Caliper source over timeARead-onlyInspect
One source day by day: traces, failures, tool calls, lookups that found nothing, tokens, cost, and speed (meanMs; p50/p95 as 'answered within' bucket edges); each tool with its calls and the share of each day's traces that used it; the documents lookups landed on most; models used. Use it to answer 'how often is it calling web search and is that changing' — compare the first and last weeks and name the day it moved. Follow with caliper_traces_list on that day or tool to show the conversations behind it.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | How many days back, default 30. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so safety is covered. The description usefully adds the shape of the aggregation and the meaning of latency values (p50/p95 as "answered within" bucket edges), which the annotations cannot convey. It omits any note on cost of computing 365-day windows, caching, or auth constraints beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource, then the metrics, then the recommended workflow. The metric enumeration is dense and the parenthetical on p95 bucket edges is slightly awkward, but each clause carries real information about the return shape rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns and does so thoroughly — metrics, per-tool breakdowns, share-of-traces, top documents, models — and tells the agent what to do next. Nothing essential for calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so sourceId, days (default 30, max 365) and workspace are already fully documented, putting the baseline at 3. The description adds no syntax or default information beyond the schema, though it does frame what the returned window represents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — reading one Caliper source day by day — and enumerates exactly what is computed (traces, failures, tool calls, tokens, cost, latency percentiles, documents, models). This distinguishes it from listing tools like caliper_sources_list and from trace-level siblings such as caliper_traces_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ("how often is it calling web search and is that changing") and a concrete analytic procedure (compare first/last weeks, name the day it moved), plus a follow-up route to caliper_traces_list. It does not state when NOT to use it or contrast it with a potentially overlapping analytics sibling like caliper_flow_performance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_sources_to_datasetSave a source's conversations to a datasetAInspect
Saves the conversations behind a pick (a day, a tool, a document, failures — the same narrowing as caliper_traces_list) into a new dataset (newDatasetName) or an existing one (datasetId), newest first up to limit. purpose decides what each item keeps: review = the agent's reply, tool calls included, for people to rate; eval = the reply becomes the expected answer; spec = only the questions, for people to answer. Set keepAdding to make it live: new matching conversations keep arriving (every one, or 1 in 10 / 1 in 100), up to 1,000. Then offer the next step — caliper_evals_create, or a review in Caliper. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | One UTC day, YYYY-MM-DD. | |
| tool | No | Only traces that called this tool. | |
| limit | No | Most recent N, default 100. | |
| purpose | Yes | review | eval | spec. | |
| document | No | Only traces whose lookups hit this document. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| datasetId | No | Add to this dataset… | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| errorsOnly | No | Only failed traces. | |
| keepAdding | No | Keep adding new matches as they arrive, taking one in this many (1 = every one). | |
| newDatasetName | No | …or start one with this name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-destructive, closed-world write (readOnlyHint=false, destructiveHint=false, openWorldHint=false), and the description adds genuinely non-derivable behavior: keepAdding turns the dataset into a live feed capped at 1,000 items, one-in-10/one-in-100 sampling, and that the call may return `needs_confirmation` with an approvalId follow-up. It omits any note about what happens to pre-existing dataset contents when adding, which is the one remaining gap for a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, then layers purpose semantics, live-mode behavior, and the next step in a tight sequence with no filler. It is on the denser side for one paragraph, but every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters, a 100% documented schema, and no output schema, the description covers the non-obvious flow (new vs existing dataset, live mode caps, confirmation handshake, follow-on tools) without needing to restate return shape. Minor gaps remain around workspace/token handling and interaction with existing dataset contents, both of which are partially covered in the schema itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description goes further by explaining what each `purpose` value actually retains (review = reply plus tool calls for rating; eval = reply becomes the expected answer; spec = questions only) and the meaning of datasetId vs newDatasetName and the sampling factor. This is real semantic value the bare enum cannot convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — saving the conversations behind a pick into a new or existing dataset — and immediately scopes it against the sibling narrowing tool caliper_traces_list, which it shares semantics with. An agent can distinguish this from caliper_datasets_add_items and caliper_datasets_create without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the decision-relevant context: which purpose (review/eval/spec) fits which downstream intent, when to use a new dataset vs an existing datasetId, and explicitly names the follow-on step (caliper_evals_create or a Caliper review). It stops short of stating when not to use it, so it lands just under an explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_starters_installInstall a Caliper starter (dataset + rubric)AInspect
Copies one starter into the workspace as a new dataset and a new rubric the user owns and can edit. Returns both ids; the next step is caliper_evals_create binding them to the flow (or target 'external'). Needs caliper:datasets:write and caliper:evals:write.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Starter slug from caliper_starters_list. | |
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this non-read-only and non-destructive, and the description adds real substance beyond them: what gets created (user-owned, editable dataset + rubric), what is returned (both ids), and the exact permissions required (caliper:datasets:write, caliper:evals:write). Does not mention idempotency or what happens if the slug already exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the action and result come first, then the next step and prerequisites. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the mutation, its outputs (both ids), the follow-up tool, and permission requirements, which is sufficient for a 3-param install tool with no output schema. Missing only edge-case behavior such as re-installing an existing slug.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the slug enum documented and sourced to caliper_starters_list, so the schema carries the burden. The description's mention of 'returns both ids' adds nothing about the workspace override or approvalId flow, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (copies/installs) plus the concrete resources produced (a new dataset and a new rubric the user owns). An agent can distinguish this from caliper_starters_list (read-only listing) and caliper_datasets_create (blank dataset) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the workflow: install a starter, then bind the returned ids via caliper_evals_create. It does not explicitly say when to choose this over caliper_datasets_create or caliper_rubrics_create, but the context is unambiguous for a starter-based workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_starters_listList Caliper starters (ready-made dataset + rubric pairs)ARead-onlyInspect
The starter sets Caliper ships: prompt injection, system-prompt leaking, over-refusal, personal-data handling, bias under ambiguity. Each is a small original dataset (single messages AND multi-turn build-ups: false memory, fabricated earlier turns, slow escalation) paired with an anchored rubric. Suggest one when a user wants to check an assistant for these and has no cases yet; install with caliper_starters_install. Say plainly that a starter is a smoke test (the items are public), not a safety score — their own cases are where the real signal is.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely new framing: the items are public, the starter is a 'smoke test' and not a safety score, which materially shapes how the agent should present results. It does not discuss return format or ordering, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The inventory of starters is front-loaded and the guidance follows in a logical order. It runs a bit long and embeds an instruction about how to phrase the answer ('Say plainly...'), but every sentence conveys distinct, useful information rather than restating the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description does tell the agent what it will get back (the named starter sets with their dataset and rubric composition) and how to act on it. For a zero-parameter list tool backed by read-only annotations, this is essentially complete, with only return ordering/pagination left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema carries nothing for the description to complement. Baseline 4 applies; there is no parameter ambiguity to resolve.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description concretely enumerates what the tool surfaces (five named starter sets, each an original dataset paired with an anchored rubric) and distinguishes itself from the install sibling. It never uses an explicit verb like 'list', relying on the title for that, but the content is specific enough that an agent can tell it apart from caliper_datasets_list or caliper_rubrics_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the trigger ('when a user wants to check an assistant for these and has no cases yet'), names the follow-up action and its tool ('install with caliper_starters_install'), and implies the exclusion via 'has no cases yet' plus 'their own cases are where the real signal is.' That is explicit when, when-not, and alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_traces_getRead one traceARead-onlyInspect
One trace: what came in, what went out, and every step in order — model calls (model, tokens, cost, reasoning), tool calls (arguments and result), lookups (query and documents with scores), anything else — each with timing and status. Long values are cut at 2,000 characters. Use it to explain why one conversation went the way it did.
| Name | Required | Description | Default |
|---|---|---|---|
| traceId | Yes | Trace id, from caliper_traces_list. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, non-destructive, and closed-world, so the safety profile is covered. The description adds genuinely useful behavior the annotations cannot: the enumeration of step types (model calls with tokens/cost, tool calls, lookups) and, most importantly, the 2,000-character truncation on long values. It omits any note on pagination, result size limits, or what happens with a stale/invalid traceId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with 'One trace' and then the content inventory; the second and third sentences add the truncation caveat and the use case without padding. The middle enumeration is long but each item earns its place by telling the agent what to expect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return contents, and it does so thoroughly (steps, timing, status, truncation threshold). It is slightly incomplete on error/edge behavior for invalid ids and on whether very large traces are paginated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the workspace override rules and where traceId/sourceId come from, so the schema does the heavy lifting. The description contributes no additional parameter meaning, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('One trace: what came in, what went out') and enumerates exactly what the retrieval returns, so an agent can immediately distinguish it from the list-style sibling. It stops short of naming caliper_traces_list as the way to discover trace ids, so sibling differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use case — 'explain why one conversation went the way it did' — which tells the agent what question this tool answers. However, it offers no when-not guidance, no prerequisite chain (list first, then get), and never names an alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_traces_listList a source's conversationsARead-onlyInspect
The traces behind a point on a source's chart, newest first, 50 a page: the question and answer (cut short), tools called, time, tokens, cost, status. Narrow by day, tool, document, or failures; page with before (the nextBefore from the last page). Open one with caliper_traces_get.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | One UTC day, YYYY-MM-DD. | |
| tool | No | Only traces that called this tool. | |
| before | No | Paging: the nextBefore from the previous page. | |
| document | No | Only traces whose lookups hit this document. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| errorsOnly | No | Only failed traces. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint=false, so safety is settled. The description adds real behavioral context beyond that: newest-first ordering, 50-per-page batching, truncated Q/A text, and the fields returned (tools, time, tokens, cost, status). It doesn't state a total count or rate limits, keeping it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is listed and how it is ordered, followed by filtering, paging, and the sibling handoff. Dense but every clause carries information; only the return-field enumeration is slightly list-like.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description compensates by naming the returned fields and pagination contract, and it covers the filter and paging parameters. Nothing needed to call or interpret this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters, including the workspace/token caveat and the YYYY-MM-DD pattern. The description echoes the filter set (day, tool, document, failures) without adding syntax beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('traces behind a point on a source's chart') plus ordering and page size, which distinguishes it from caliper_traces_get, the single-trace sibling. An agent knows exactly what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete narrowing dimensions (day, tool, document, failures), the paging protocol (`before` = nextBefore from last page), and explicitly routes to caliper_traces_get for a single trace. Both when-to-use and the alternative are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_createComment on an entityAInspect
Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| rootId | No | Reply into this thread; omit to start a new one. | |
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| entityKind | Yes | What the thread hangs on. | |
| mentionedUserIds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the write/safety profile is covered. The description adds value beyond that by disclosing a side effect annotations cannot express: mentioning users rings their notification bell, with an accompanying social caution. It omits permission/auth requirements for posting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and its two modes, then layers usage and mention etiquette compactly. Two sentences, minimal waste, though the trailing 'never mention someone who didn't ask' is advisory padding rather than invocation-critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with 43% schema coverage and no output schema, the description covers the social/mention dimension well but leaves the entity-targeting parameters and the undocumented approvalId unexplained, which an agent would need to call this reliably in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, so the description must compensate, and it partially does: rootId (reply target), mentionedUserIds (workspace member ids, notification behavior), and body are implied. However entityKind's 17-value enum, entityId, and especially approvalId are unexplained in both schema and description, leaving real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Posts a comment on a workspace entity') and distinguishes the two modes of operation: new thread vs. reply when rootId is given. It never names its closest siblings (comments_list, comments_resolve), so it stops just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage context ('leave findings where the discussion already lives') with two illustrative scenarios (eval result on a debated flow, summary on a long thread). It does not state when not to use it or point to a sibling alternative, so it lacks the explicit routing of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_listRead an entity's comment threadsARead-onlyInspect
Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.
| Name | Required | Description | Default |
|---|---|---|---|
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| entityKind | Yes | What the thread hangs on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavior: threads are returned open-before-resolved and include authors and timestamps, which shapes how an agent interprets output. It stops short of pagination or volume limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the key facts front-loaded and no redundancy. The second sentence is motivational framing that carries mild value but is slightly softer than a hard routing rule.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema and a mostly documented schema, the description covers ordering, content, and the entity scoping needed to call it. Pagination/result-size behavior is the only notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the schema itself handles the workspace slug nuances and the enum list. The description only says 'one workspace entity', adding little beyond the structured fields, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (lists) and resource (comment threads on one workspace entity), plus the ordering rule (open first, then resolved) and payload (authors, timestamps). This clearly separates it from comments_create and comments_resolve, which mutate rather than read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives a soft situational cue ('read this before weighing in on contested work'), which implies when the tool is useful. However it names no alternatives and gives no explicit when-not or prerequisite guidance, so usage remains inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_resolveResolve or reopen a threadADestructiveInspect
Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| rootId | Yes | ||
| resolved | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=false, so the mutation/safety profile is covered. The description adds real behavioral context beyond that: the precondition (a reply explaining what settled the thread) and the reopen semantics. It doesn't clarify reversibility or how the state change affects existing replies, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the action and the primary precondition are front-loaded before the reopening clause. Every clause carries decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small toggle tool with no output schema, the description covers the action, the target, and the conditions well. It leaves approvalId and workspace behavior unexplained, which matters for a destructive write, but overall an agent has enough to act correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25%, and the description compensates for rootId only (root comment id). 'resolved' is implied by resolve/reopen framing, but 'workspace' is documented only in the schema and 'approvalId' is explained nowhere — a notable gap for a mutation tool with an approval parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (sets resolved state) plus resource (comment thread), and covers both directions — resolve and reopen. The parenthetical 'rootId = the thread's root comment id' disambiguates the target, making it clearly distinct from comments_create/comments_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gates the action: resolve ONLY when the human asked or the question is demonstrably settled, and reply first explaining what settled it; reopen is for new evidence. This is genuine when/when-not guidance with a required prior step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_browseBrowse the workspace's tagsARead-onlyInspect
Without a tag: every tag in use across the workspace with how many entities carry it, most-used first — the vocabulary the team already organizes by. With a tag: everything filed under it across every tool, each with its kind, id, title, and path. Use it to reuse existing labels instead of inventing near-duplicates, and to answer 'show me everything about X' when X is a label.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | A tag to expand into its items. Omit to list tags. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive/openWorld, so safety is covered, and the description adds real behavioral detail: result ordering (most-used first), entity counts per tag, and the per-item fields returned in tag mode (kind, id, title, path). No pagination or size-limit behavior is mentioned, keeping this from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and cleanly parallel: 'Without a tag:' then 'With a tag:' then the usage clause. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return-value burden for both modes, and annotations cover the safety profile while the schema covers both parameters. Nothing an agent needs to select or call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by characterizing what each mode of the tag parameter actually returns, turning a bare 'omit to list tags' into the two distinct result shapes. The workspace parameter semantics remain entirely schema-borne.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific dual operation (list every tag in the workspace, or expand one tag into all entities filed under it) with the exact resource and scope. It is immediately distinguishable from the write-oriented siblings entity_tags_get and entity_tags_set because the browse mode and cross-tool aggregation are spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use guidance: reuse existing labels rather than inventing near-duplicates, and answer 'show me everything about X' when X is a label. It stops short of naming sibling alternatives or stating when-not-to-use (e.g. versus search_workspace), so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_getRead the tags on entitiesARead-onlyInspect
Returns the tags on a batch of entities of one kind — the labels galleries organize by. Ids come from the kind's list/get tool or from search_workspace. Use it before entity_tags_set so you replace the full set knowingly, and to answer 'what is this filed under'. Entities the user can't see are omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| entityKind | Yes | Which kind the ids belong to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds the non-obvious behavioral fact that 'Entities the user can't see are omitted,' i.e. results are permission-filtered rather than erroring. It does not mention limits (e.g. the 100-id cap or ordering), but the value-add beyond annotations is genuine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all front-loaded: what it returns first, then usage, then the permission caveat. No filler and each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does state what comes back (tags) and the omission behavior. It leaves minor gaps — tag value format, ordering, and whether unknown ids are silently dropped or error — but nothing that blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema documents most params, but the description adds real provenance for the hardest parameter: ids 'come from the kind's list/get tool or from search_workspace.' That tells an agent where to obtain valid ids, which the bare array schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Returns the tags on a batch of entities of one kind.' The parenthetical 'the labels galleries organize by' distinguishes tags from other entity metadata, and the named siblings (entity_tags_set, search_workspace) make the boundary clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit usage contexts: 'Use it before entity_tags_set so you replace the full set knowingly, and to answer what is this filed under.' That is a real when-to-use plus a stated alternative. It does not mention entity_tags_browse, the other obvious read-side sibling, so it falls short of fully disambiguating the read alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_setSet an entity's tagsADestructiveInspect
Replaces the FULL tag set on one entity (an empty list clears it). Read the current tags with entity_tags_get first and pass the merged list — this is not additive. Tags are lowercase letters, numbers, spaces, and hyphens; prefer labels already in use (entity_tags_browse) so the workspace's vocabulary stays small. The id comes from the kind's list/get tool or search_workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | ||
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| entityKind | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, and the description reinforces this with the replace-not-merge semantics and the empty-list-clears behavior. It does not mention the needs_confirmation/approvalId retry flow that the schema implies, which is the one notable behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, all front-loaded with the destructive replace semantics first, followed by workflow and vocabulary guidance. Every sentence earns its place; the parentheticals make it slightly busy but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description covers destructive semantics, prerequisites, id sourcing, and tag format well enough to invoke the tool correctly. The confirmation/approval retry path is left to the schema field description rather than the tool description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, and the description compensates by documenting the tag character set (lowercase letters, numbers, spaces, hyphens), the merge requirement for the tags array, and the origin of entityId. The workspace and approvalId parameters are only explained in the schema, so coverage is good but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Replaces the FULL tag set on one entity') and immediately clarifies the destructive scope with '(an empty list clears it)'. This cleanly separates it from the sibling readers entity_tags_get and entity_tags_browse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit workflow: read current tags with entity_tags_get first, pass the merged list, and warns 'this is not additive'. Also routes to entity_tags_browse for vocabulary reuse and names where the entityId comes from (kind's list/get tool or search_workspace), so the agent knows both prerequisites and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_docFetch ZeroWidth doc by slugARead-onlyInspect
Fetch the full Markdown body of a specific docs page by its slug. Use this after search_docs when the user needs the complete content of a page. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Page slug. Accepts 'compass/api', '/compass/api', or 'docs/compass/api'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, covering the safety profile. The description adds valuable context by stating 'No authentication required' and specifying the return body as full Markdown, which agents need to know beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose ('Fetch the full Markdown body...'), followed by usage routing and a salient behavioral note. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with full schema coverage and annotations covering safety, the description is complete: it states the return format, the required input type, usage context, and authentication status. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single `slug` parameter is fully documented in the schema with accepted formats. The description adds no syntax or format details beyond what the schema provides, so it meets the baseline rather than exceeding it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch'), resource ('full Markdown body of a specific docs page'), and retrieval key ('by its slug'). It distinguishes itself from the search-oriented sibling by positioning as the step after `search_docs` for complete page content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this after `search_docs` when the user needs the complete content of a page, which names the alternative and the condition that selects it. It does not state when not to use it (e.g., for listing or metadata), but the context is clear enough for a simple read tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_docsList ZeroWidth docs pagesARead-onlyInspect
Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling get_doc. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| product | No | Optional product slug filter (e.g. 'compass', 'legal', 'overview'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds context beyond them with 'No authentication required,' a genuinely useful operational fact for callers, though it says nothing about pagination or result size for a full enumeration.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with purpose, routing guidance, and the auth fact front-loaded in order of importance. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-optional-param listing tool whose annotations cover safety, the description is nearly sufficient. Without an output schema it could note the return shape (e.g. that results are slugs/pages), but the 'slugs' reference largely covers that, so only a small gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the single 'product' parameter. The description's example values ('compass', 'legal', 'overview') duplicate the schema description verbatim, adding no meaning beyond it. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Enumerate') and resource ('all available docs pages') with scope, plus the optional product filter. It names the sibling get_doc and frames itself as the discovery step before retrieval, letting an agent distinguish it from single-doc fetch without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it: 'to discover what slugs exist before calling get_doc,' giving a clear dependency flow and naming the alternative. It does not address when NOT to use it or whether search_docs is a better discovery path, leaving a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsSearch ZeroWidth docsARead-onlyInspect
Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max number of results. Defaults to 10. | |
| query | Yes | Search query — keywords or natural-language phrase. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful non-annotation context: 'No authentication required — the docs corpus is public' and the shape of returned results (title, slug, public URL, snippet). It does not mention pagination or ordering, but it goes beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four short sentences, front-loaded with the core action, followed by return information, usage trigger, and auth note. Every sentence earns its place, and there is no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only search tool, the description covers purpose, return format, auth requirements, and usage triggers. It does not explain how this differs from search_workspace or list_docs/get_doc, and it lacks result-ordering or empty-result behavior, but it is otherwise sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters are documented in the schema itself. The description says the query returns a 'query-relevant snippet' but adds no syntax, format, or constraint details beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Search ZeroWidth product documentation.' It also names the return fields and the product scope with concrete examples (Compass, Workbench, Caliper, etc.), which lets an agent distinguish it from siblings like list_docs, get_doc, and search_workspace without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: 'Use this when the user asks about a ZeroWidth product ..., an API behavior, or a policy.' However, it does not name alternative tools (search_workspace, list_docs, get_doc) or state when not to use this tool, so it stops short of explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_workspaceSearch the whole workspaceARead-onlyInspect
Finds entities across every tool by name in one call — Workbench flows, Compass pages, Caliper datasets, evals, rubrics, reviews, specs and sources (apps sending agent traces), Ledger entries, Napkin sketches and decks. Use it FIRST when the user names something without saying where it lives ('the onboarding flow', 'that invoice page'); reach for a tool's own list only when you already know the tool. Each hit carries its id, kind, and workspace-relative path, so the id feeds the matching *_get tool and the path makes a link. Results only include what the user can see, and only kinds this token may read.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Case-insensitive substring matched against names/titles. | |
| kinds | No | Restrict to these kinds (flow, page, dataset, eval, entry, board). Omit to search everything. | |
| limit | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context beyond that: results are filtered to what the user can see and to kinds the token may read, and it discloses the hit shape (id, kind, workspace-relative path).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences with no filler: purpose and coverage first, routing second, return/scope semantics last. Every clause carries information the agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden and does so by describing the id/kind/path tuple and how the id feeds *_get tools. Combined with the permission scoping note, an agent has everything needed to call and use this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most parameters are already documented structurally. The description reinforces name/title matching but adds no new syntax or format detail for kinds, limit, or workspace, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb (Finds) and resource (entities across every tool by name), then enumerates the concrete kinds covered (flows, pages, datasets, evals, rubrics, etc.). An agent can immediately distinguish this cross-tool search from the many per-tool list siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('Use it FIRST when the user names something without saying where it lives'), gives concrete examples ('the onboarding flow'), and names the alternative plus its selection condition ('reach for a tool's own list only when you already know the tool'). This is textbook when/when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
45 tool updates
- First observed
caliper_datasets_add_items - First observed
caliper_datasets_create - First observed
caliper_datasets_delete - First observed
caliper_datasets_generate - First observed
caliper_datasets_generate_items - First observed
caliper_datasets_get - First observed
caliper_datasets_list - First observed
caliper_datasets_remove_item - First observed
caliper_datasets_update - First observed
caliper_datasets_update_item - First observed
caliper_evals_create - First observed
caliper_evals_delete - First observed
caliper_evals_get - First observed
caliper_evals_list - First observed
caliper_evals_run - First observed
caliper_evals_runs_cancel - First observed
caliper_evals_runs_get - First observed
caliper_evals_runs_item_execution - First observed
caliper_evals_runs_list - First observed
caliper_evals_update - First observed
caliper_flow_performance - First observed
caliper_rubrics_create - First observed
caliper_rubrics_delete - First observed
caliper_rubrics_get - First observed
caliper_rubrics_list - First observed
caliper_rubrics_update - First observed
caliper_source_feeds_list - First observed
caliper_source_feeds_stop - First observed
caliper_sources_list - First observed
caliper_sources_over_time - First observed
caliper_sources_to_dataset - First observed
caliper_starters_install - First observed
caliper_starters_list - First observed
caliper_traces_get - First observed
caliper_traces_list - First observed
comments_create - First observed
comments_list - First observed
comments_resolve - First observed
entity_tags_browse - First observed
entity_tags_get - First observed
entity_tags_set - First observed
get_doc - First observed
list_docs - First observed
search_docs - First observed
search_workspace
Related MCP Connectors
Run AI agent evaluations on your own model keys, and publish runs others can review and re-run.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Read and score Opik traces, experiments, datasets and prompts from your AI assistant.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.1MIT

io.github.HZYAI/ragscoreofficial
AlicenseNot gradedqualityBmaintenanceEnables generating QA datasets from documents and evaluating RAG systems with detailed metrics, failure diagnosis, and visual reports through natural language.41 PyPI15Apache 2.0
Trustwise MCP Serverofficial
AlicenseNot gradedqualityCmaintenanceProvides advanced evaluation tools for assessing AI safety, alignment, and performance of LLM outputs. Enables programmatic evaluation of quality, safety metrics like toxicity and PII detection, and operational metrics including carbon footprint and cost estimation.4Apache 2.0
Coval MCP Serverofficial
AlicenseAqualityBmaintenanceEnables AI assistants to interact with Coval's evaluation platform for launching and monitoring evaluation runs, managing agents and test sets, and retrieving evaluation metrics.1816 npm2MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.