ZeroWidth
Server Details
Search, write and run work across Compass, Workbench, Caliper, Ledger, Prism and Napkin.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 248 tools
Tools are heavily namespaced by product and entity (caliper_datasets_*, compass_pages_*, ledger_entries_*, napkin_*), and descriptions go out of their way to disambiguate neighbors (e.g. napkin_docs_create vs napkin_boards_create vs napkin_deck_write, routines_create vs workbench_flows_schedule). Real overlap clusters remain — six 'create a Napkin thing' tools, compass_interviews_create vs prism_interviews_create, three metric-adoption paths — but descriptions explain the boundaries well. Only the sheer density of 248 names creates residual misselection risk.
Overwhelmingly a consistent snake_case namespace_entity_verb pattern (…_create / _get / _list / _update / _delete) across Caliper, Compass, Ledger, Napkin, Prism, and Workbench. A few deviations exist (get_doc, search_workspace, comments_*, entity_tags_*, caliper_flow_performance, workbench_flow_authoring_guide) but they are readable and follow the same product-prefix convention.
248 tools is far beyond what an agent can hold or select among reliably, dwarfing even the 25+ 'too many' bucket. Each tool does map to a real feature of a genuinely multi-product platform, so it is proportional to scope rather than arbitrary, but the total is impractical in one surface.
Coverage is exceptional: near-full CRUD/lifecycle sets for datasets, evals, rubrics, Compass pages/links/opportunities/gaps/interviews, Ledger entries/metrics/feeds/plan, Napkin boards/docs/decks/sheets/diagrams/interfaces/brand, and Workbench flows/kb/tasks. Minor gaps remain — no routine, series, or shim delete/update, no comment edit/delete — but they are small and workaroundable.
Available Tools
248 toolscaliper_datasets_add_itemsAdd items to a Caliper datasetAInspect
Appends items to an existing dataset — use this to grow coverage (new edge cases, scenarios from a completed interview) instead of creating a parallel dataset. Same three shapes as caliper_datasets_create (Q&A, sequence, simulated). Existing items and their ratings are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Items to append (1-100). | |
| datasetId | Yes | Id of the dataset to extend. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so the safety profile is known. The description adds real value by disclosing that existing items and their ratings are untouched, i.e. the append is purely additive. It does not mention the needs_confirmation/approvalId flow, which is only visible in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the alternative/create contrast immediately after. Every clause earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an append tool with no output schema and full schema coverage, the description supplies the needed framing: additive semantics, shape parity with create, and non-destruction of existing data. It omits return/response behavior and the approval pathway, which are minor but leave small gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by mapping the items payload to 'the same three shapes as caliper_datasets_create (Q&A, sequence, simulated)', giving useful semantic orientation for how the item fields combine. It adds no per-field detail, so it stops short of 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Appends items to an existing dataset') and explicitly frames the scope contrast with creating a parallel dataset, which separates it from caliper_datasets_create. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use condition ('grow coverage — new edge cases, scenarios from a completed interview') and names the alternative behavior it replaces ('instead of creating a parallel dataset'). The routing decision versus the create sibling is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_createCreate a Caliper datasetAInspect
Creates a dataset of test items — THE FIRST STEP of setting up evaluation for a flow. Three item shapes: Q&A (input + optional expectedOutput, the golden answer); SEQUENCE (turns: 2-20 scripted user messages the model answers one at a time with its own earlier replies in front of it, + expectedResponse for the final reply, optional expectedBehavior for the whole conversation); SIMULATED (goal + optional persona/disposition/strategy/maxTurns + expectedBehavior; a platform flow plays a person adaptively, Caliper-run evals only; disposition is a preset id like genuine, pressure, confused, impatient, or free text). Use sequences and simulated items for the slow attacks and for real customers with real needs: a model that holds on message one often folds on message ten. Write good inputs from real usage: the Compass pages the flow was built from are the best source of realistic scenarios.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Test items (1-100). | |
| title | Yes | Dataset name (2-120 chars). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | What this dataset covers. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral constraint — simulated items are 'Caliper-run evals only' — but says nothing about the notable approvalId/needs_confirmation flow this mutation triggers, leaving a real gap for a write tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and 'first step' positioning, then structured into the three item shapes. It is long, and phrases like the message-ten anecdote and the Compass-source advice are more persuasive than operational, but most content earns its place for a complex nested schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with a nested items array and no output schema, the description covers the item taxonomy thoroughly and explains how each shape is used in evals. The main omission is any mention of the approval/confirmation handshake that the approvalId parameter implies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes beyond it by organizing the nested item parameters into three named shapes (Q&A / SEQUENCE / SIMULATED) and clarifying the role of fields like disposition, expectedResponse, and expectedBehavior. This adds conceptual meaning over the flat per-field schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a dataset of test items') and explicitly positions it as 'THE FIRST STEP of setting up evaluation for a flow', which cleanly separates it from siblings like caliper_datasets_add_items and caliper_datasets_update. An agent knows exactly what this does and where it sits in the workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context ('first step' of eval setup) and prescriptive guidance on when to use sequence/simulated items versus Q&A, plus where to source realistic inputs (Compass pages). It does not explicitly name the alternative sibling for appending items to an existing dataset, so routing is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_deleteDelete a Caliper datasetADestructiveInspect
Deletes a dataset. Evals bound to it stop being runnable (their binding shows datasetOk: false), and reviews/specs over its items lose their source — check caliper_evals_list for evals that reference it and say so before proposing. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| datasetId | Yes | Dataset id, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While annotations already flag destructiveHint=true, the description goes well beyond them by disclosing the downstream blast radius: bound evals stop being runnable (datasetOk: false) and reviews/specs lose their source. It also surfaces the two-step `needs_confirmation` / approvalId flow, which is critical operational context for a two-phase destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Deletes a dataset') and then spends the remaining words on consequences, the cross-tool check, and the confirmation flow. No filler sentences; every clause carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output-schema tool with full schema coverage, this is exactly the right content: consequences, a related-tool verification step, and the confirmation handshake. Nothing an agent needs to invoke it safely and correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three params are already documented and the baseline is 3. The description earns a point above baseline by explaining that the call 'May return `needs_confirmation`,' which clarifies the purpose of the approvalId parameter and how the two-call confirmation sequence works.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Deletes) and resource (a dataset), making it trivially distinguishable from siblings like caliper_datasets_update, caliper_datasets_remove_item, and caliper_datasets_get. The cascade consequences further pin the scope to whole-dataset deletion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit pre-flight step: 'check caliper_evals_list for evals that reference it and say so before proposing.' This names a concrete alternative tool and the condition that invokes it. It does not, however, state when not to delete or the exact criteria for bailing out, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_generateGenerate a Caliper dataset from a flow (spends credit)AInspect
Creates a NEW dataset of model-written items — Q&A by default, or scripted sequences / simulated people via shape — pass flowId and the generator reads the flow's prompt, mode, and schema to write realistic cases for THAT flow; description adds guidance (or stands alone when there's no flow). Use this when the user wants test cases fast and has none; prefer caliper_datasets_create with hand-written items when real scenarios are already in hand (Compass pages, a transcript). This SPENDS workspace inference credit (one generator call), so it sits behind the approval gate: say so and expect needs_confirmation. Returns the dataset summary; read the items with caliper_datasets_get and tell the user to review them before trusting an eval built on them.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Items to generate, 1-50. Default 10. | |
| shape | No | What each generated item is: `qa` (one input, the default), `sequence` (2-6 scripted user turns the model answers one at a time, with expectedResponse + expectedBehavior), or `simulated` (a person Caliper plays adaptively: goal, persona, disposition, expectedBehavior). outputsMode applies to `qa` only. | |
| title | Yes | Dataset name. | |
| flowId | No | Workbench flow to generate cases for (from workbench_flows_list). Required unless description is given. | |
| presetId | No | Framing preset. Default `blank`; the others bias every item toward that attack class. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | What the target system does / what to cover. Required (≥10 chars) when flowId is omitted; optional guidance otherwise. | |
| disposition | No | With shape `simulated`: how every generated person behaves — a preset id (genuine, pressure, confused, impatient, vague, non-native) or free text. Omit to let the generator vary it from person to person. | |
| outputsMode | No | What each generated item carries beyond the input: `expected` (golden answers — eval-ready, the default), `none` (inputs only — for a spec others fill in), `captured` (sample answers to rate in a review). | |
| expectedStyle | No | With outputsMode `expected`: `verbatim` literal reference answers (default) or `conditions` — what a correct answer must do, when there's no single right wording. | |
| datasetDescription | No | Description stored on the dataset. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: discloses that it SPENDS workspace inference credit (one generator call), sits behind an approval gate producing needs_confirmation, returns only a dataset summary, and requires reading items via caliper_datasets_get before trusting an eval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded and every clause carries information (routing, credit cost, confirmation flow, follow-up). Sentence length runs long, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-param, no-output-schema mutation tool, the description supplies the missing pieces: cost, approval lifecycle, return shape, and the recommended follow-up call. An agent can invoke this correctly without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real meaning by explaining that flowId causes the generator to read the flow's prompt/mode/schema, and how description can substitute for or augment flowId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a NEW dataset of model-written items') and enumerates the shapes it can produce. It clearly distinguishes itself from the hand-written sibling caliper_datasets_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('user wants test cases fast and has none') and when-not-to-use ('prefer caliper_datasets_create with hand-written items when real scenarios are already in hand'). It also flags the approval-gate flow the agent should expect.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_generate_itemsGenerate more items for a Caliper dataset (spends credit)AInspect
Appends model-written items to an EXISTING dataset, in the style of what's already there (existing items are the few-shot examples; the output shape matches theirs unless overridden). Use to widen coverage when the user says 'more like these' or 'add edge cases'; write them by hand with caliper_datasets_add_items when the scenarios are known. SPENDS workspace inference credit, so it sits behind the approval gate — say so. Existing items and their ratings are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Items to add, 1-20. Default 5. | |
| shape | No | What each generated item is: `qa` (one input, the default), `sequence` (2-6 scripted user turns the model answers one at a time, with expectedResponse + expectedBehavior), or `simulated` (a person Caliper plays adaptively: goal, persona, disposition, expectedBehavior). outputsMode applies to `qa` only. | |
| datasetId | Yes | Dataset to extend, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| description | No | What to bias toward, e.g. 'angry customers', 'ambiguous refund questions'. | |
| disposition | No | With shape `simulated`: how every generated person behaves — a preset id (genuine, pressure, confused, impatient, vague, non-native) or free text. Omit to let the generator vary it from person to person. | |
| outputsMode | No | What each generated item carries beyond the input: `expected` (golden answers — eval-ready, the default), `none` (inputs only — for a spec others fill in), `captured` (sample answers to rate in a review). | |
| anchorItemId | No | An item id (from caliper_datasets_get) the new items should resemble most. | |
| expectedStyle | No | With outputsMode `expected`: `verbatim` literal reference answers (default) or `conditions` — what a correct answer must do, when there's no single right wording. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover safety hints (readOnlyHint false, destructiveHint false), while the description adds the economically material facts: it SPENDS workspace inference credit, sits behind an approval gate, and leaves existing items and ratings untouched. These are beyond anything structured fields convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool does, then routing, then the credit/approval warning, then the non-destructive guarantee. Dense but every sentence earns its place; slight compression of the style-inheritance clause could improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, full schema coverage, and no output schema, the description supplies the missing behavioral framing an agent needs: cost, approval flow, how generation derives from existing data, and that prior content is preserved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so per-parameter docs already exist. The description still adds meaning the schema doesn't state: existing items act as few-shot examples and the output shape matches theirs unless overridden. It does not explain count defaults or workspace/approvalId semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('appends model-written items to an EXISTING dataset') and immediately qualifies how the generated content is derived from existing items. It reads clearly as distinct from both caliper_datasets_add_items (manual) and caliper_datasets_generate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger phrasing ('more like these', 'add edge cases') and the alternative (caliper_datasets_add_items) with the condition that selects it ('when the scenarios are known'). It also flags the approval-gate workflow, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_getRead a Caliper dataset (items paged)ARead-onlyInspect
One dataset with a page of its items — read this BEFORE editing items (caliper_datasets_update_item needs the item id) and before extending coverage, so you don't add cases that already exist. Items come limit at a time (default 25) from offset; total is the full count. Long fields are cut at ~800 characters with a truncation marker. Get the datasetId from caliper_datasets_list.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Items per page, 1-100. Default 25. | |
| offset | No | Items to skip. Default 0. | |
| datasetId | Yes | Dataset id, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description still adds real behavioral context beyond them: page size default (25) via limit/offset, that total is the full count not the page, and that long fields are truncated at ~800 chars with a marker.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core operation and the read-before-write guidance; every clause carries distinct information (dependency, paging, truncation, id source) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must cover return shape, and it does: one dataset plus a paged item list, total count, and truncation behavior. Combined with annotations and full schema coverage, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents all four params. The description goes a bit further by framing limit/offset as the paging mechanism and stating the source of datasetId (caliper_datasets_list), tying parameters to workflow rather than restating types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (a Caliper dataset) and adds the scope detail 'with a page of its items', which separates it from caliper_datasets_list (no items) and the item-mutating siblings. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to read this BEFORE editing items and BEFORE extending coverage, and names the dependency (caliper_datasets_update_item needs the item id) plus a preventive reason (avoid adding duplicate cases). When-to-use and the alternative are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_listList Caliper datasetsARead-onlyInspect
Lists datasets in the active workspace. Returns summaries (id, name, item count); items are not included.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so safety is covered. The description adds genuinely useful behavioral context beyond that: the exact return shape (id, name, item count) and the negative constraint that items are not included, which prevents an agent from expecting embedded item data. It omits pagination, ordering, and result-size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the scope ('active workspace') and the return contract front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on the burden of describing what is returned, and it does so briefly and accurately. It leaves pagination and ordering unspecified, which is a minor gap for a list tool with a single optional filter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one optional parameter and schema description coverage is 100%; the schema fully documents the workspace slug, including the personal-token vs API-key behavior. The description adds no parameter-level detail beyond the phrase 'active workspace', so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists datasets') with a scope qualifier ('in the active workspace'), so the operation is unambiguous. It stops short of naming the sibling it contrasts with (caliper_datasets_get for single-dataset detail), relying on 'items are not included' to imply the distinction rather than stating it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Use is implied by 'Lists datasets in the active workspace' and by the caveat that items are excluded, which hints an agent should call caliper_datasets_get for detail. No explicit when-to-use, when-not-to-use, or named alternative is given, so the agent must infer the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_remove_itemRemove one item from a Caliper datasetADestructiveInspect
Drops one item. Ratings and run results keyed to it are orphaned (kept in history, gone from the dataset). A dataset must keep at least one item. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | Item id, from caliper_datasets_get. | |
| datasetId | Yes | Dataset id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description goes well beyond them: it discloses the exact consequence of the mutation (ratings and run results are orphaned — kept in history, removed from the dataset), an invariant (minimum of one item), and a non-trivial response behavior (may return `needs_confirmation`). This is precisely the added context destructive annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then consequences, then constraints/response behavior. Every sentence carries distinct information and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no output schema, the description covers the destructive effect, the precondition, and the confirmation round-trip. Minor gaps remain (no mention of permissions or what the return payload looks like beyond needs_confirmation), but the schema covers the remaining parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: it explains the trigger for `needs_confirmation` and therefore the existence/purpose of the `approvalId` parameter, which the schema only describes circularly. It does not clarify `workspace` scoping, so it is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb+resource ("Drops one item"), which is clear and scoped. However, it largely restates the tool title ("Remove one item from a Caliper dataset") rather than distinguishing itself from siblings like caliper_datasets_delete (whole dataset) or caliper_datasets_update_item; the differentiation is inferred, not stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives operational constraints ("A dataset must keep at least one item") and a confirmation flow, but never says when to reach for this tool versus caliper_datasets_update_item, caliper_datasets_delete, or other siblings. No when-not guidance either.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_updateRename / re-describe / re-scope a Caliper datasetADestructiveInspect
Changes a dataset's title, description, or visibility. Pass only what changes; items are untouched (use caliper_datasets_add_items / caliper_datasets_update_item / caliper_datasets_remove_item for those). May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New name. | |
| datasetId | Yes | Dataset id, from caliper_datasets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description; null clears it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so safety is partly covered, but the description adds non-obvious behavior: this is a PATCH-style partial update where unspecified fields survive, and the call may return a `needs_confirmation` state requiring a follow-up with an approvalId. It does not explain what is actually destructive here (e.g. a visibility downgrade revoking access), which is the one gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences/fragments, front-loaded with the primary action and scope, then the exclusion with alternatives, then the confirmation caveat. No filler and nothing rhetorical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still flags the one non-obvious return path (`needs_confirmation`) and the schema documents the approvalId continuation. For a six-parameter mutation with destructive semantics, it is complete enough, though the lack of any note on which changes are destructive or reversible leaves a minor hole.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including a fully documented visibility enum with a natural-language mapping and the approvalId/workspace auth notes, so the schema does the heavy lifting. The description only echoes the three top-level fields and adds the patch-semantics cue, which does not meaningfully exceed what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (changes) plus the exact mutable fields (title, description, visibility) of a named resource (dataset), which matches the title. It also explicitly excludes dataset item content and names the siblings that own those operations, so the agent can separate this from add_items/update_item/remove_item without reading schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Supplies an explicit when-not rule ('items are untouched') and names three alternative tools for that case, which is exactly the when/when-not/alternatives pattern. The 'Pass only what changes' instruction also tells the agent how to call it (partial update) versus a full replacement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_datasets_update_itemReplace one item in a Caliper datasetADestructiveInspect
Rewrites one item's content, keeping its id (ratings and eval results keyed to it stay attached). The item's KIND can't change — a Q&A item stays Q&A, a sequence stays a sequence — so pass the same shape you read from caliper_datasets_get. Chat and raw items can't be edited from here. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| item | Yes | The full replacement item, same three shapes as caliper_datasets_create. | |
| itemId | Yes | Item id, from caliper_datasets_get. | |
| datasetId | Yes | Dataset id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only tell the agent this is a destructive non-read-only write; the description adds the important fallout semantics — the id and its attached ratings/eval results survive, the item KIND is immutable, and the call can return `needs_confirmation`. That is materially more than the annotation set provides and does not contradict it (the destructiveHint refers to overwriting the item's content in place).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler: the core action and its id-preservation consequence come first, the kind constraint and shape instruction second, the exclusions and confirmation caveat last. Each sentence carries information an agent needs before invoking.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the essential gotchas (immutable kind, preserved id/ratings, non-editable item types, possible confirmation response). It would be complete if it explained how the returned `needs_confirmation` maps onto the `approvalId` parameter, since that round-trip is otherwise left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description still adds real meaning by constraining the `item` payload to the same three shapes returned by caliper_datasets_get and warning that the shape/kind cannot change. The `approvalId`/`needs_confirmation` connection is only hinted at rather than explained, keeping this short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ("Rewrites one item's content") and pins the scope to a single item within a dataset, which separates it from the dataset-level caliper_datasets_update and from caliper_datasets_add_items. It further clarifies the item retains its id, so an agent knows this is an in-place replacement rather than an add/remove.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear prerequisites ("pass the same shape you read from caliper_datasets_get") and explicit exclusions ("Chat and raw items can't be edited from here"), plus the kind-preservation constraint. It stops short of naming the alternative tool for those excluded shapes or for replacing an item by delete+add, so it is clear context without full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_createCreate a Caliper evalAInspect
Binds a dataset + rubric to something under test as a repeatable eval — THE LAST SETUP STEP before scoring. Two targets: a Workbench flow (pass flowId; Caliper runs inference itself, then caliper_evals_run scores it), or 'external' (target: 'external'; the user's own model runs elsewhere and their script submits outputs through the public API, usually from CI — see the docs guide 'Run evals in CI'). For an external eval, hand back the eval id and tell the user to create a workspace API key with the ci_evals preset in their workspace settings; keys can't be minted from here. flowStage 'draft' evals the live draft (pre-publish); 'published' (default) evals the latest published revision at run time, or one pinned with flowRevisionId. Reuse an existing rubric from caliper_rubrics_list when one already scores this job.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Eval name (2-120 chars). | |
| flowId | No | Workbench flow id to evaluate. Required unless target is 'external'. | |
| target | No | 'workbench_flow' (default; needs flowId) or 'external' (the user's own model; outputs arrive through the public API). | |
| rubricId | Yes | Rubric the judge scores with. | |
| datasetId | Yes | Dataset of test items. | |
| flowStage | No | Which stage to run against. Use 'draft' while iterating pre-publish. Default 'published'. | |
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | ||
| externalLabel | No | External evals only: what produces the outputs, e.g. 'CI' or 'prod pipeline'. Free text, shown on the eval. | |
| flowRevisionId | No | Published-stage evals only: pin to one published revision (id from workbench_flows_revisions_list). Omit to always eval the latest published version at run time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare this is a non-destructive, non-open-world write; the description adds real operational context beyond that: external evals require a workspace API key with the ci_evals preset that 'can't be minted from here,' and it explains the draft-vs-published execution behavior at run time. These are meaningful prerequisites an agent would otherwise not know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The lead sentence front-loads the core purpose and lifecycle position, and the remaining text is information-dense rather than filler. It runs long and a couple of clauses are heavily packed, but given 12 parameters and two targets, nearly every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description steps in to state what to do with the result ('hand back the eval id and tell the user to create a workspace API key'), and it covers the branching setup, prerequisites, and next-step tool. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 92%, so the schema does most of the work (baseline 3). The description still adds relational meaning the schema can't express — flowId is only needed for the workbench target, flowRevisionId applies only to the published stage, and the draft/published distinction is spelled out — lifting it above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — it 'binds a dataset + rubric to something under test as a repeatable eval' — and positions itself in the lifecycle as 'THE LAST SETUP STEP before scoring.' It also distinguishes its job from caliper_evals_run, which does the scoring, so an agent can separate it from siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes the two invocation paths (workbench flow via flowId vs target:'external'), when to pick each, when to use flowStage 'draft' vs 'published', and routes to caliper_rubrics_list when a rubric already exists. It even points at the 'Run evals in CI' docs guide for the external path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_deleteDelete a Caliper evalADestructiveInspect
Deletes an eval and stops its schedule. Run history is kept but no longer reachable from the eval. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| evalId | Yes | Eval id, from caliper_evals_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known, but the description adds real value beyond them: the schedule is stopped, run history becomes unreachable, and the call may return needs_confirmation. It stops short of stating auth/permission requirements or whether deletion is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying distinct information (what is deleted, what happens to history, what a confirmation response implies), with the primary action front-loaded. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter destructive tool with no output schema, the description covers the salient outcomes an agent needs: the eval is gone, its schedule stops, history is orphaned, and a confirmation round-trip may be required. Only the permission/auth prerequisites are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description earns above that by explaining the needs_confirmation/approvalId retry cycle, which gives the approvalId parameter meaning the schema alone only partially conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Deletes an eval') and immediately names the collateral effect ('stops its schedule'). It reads clearly against caliper_evals_update/run/get, though it does not explicitly differentiate itself from the closer sibling caliper_evals_runs_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no statement of prerequisites, and no pointer to alternatives (e.g., cancel a run vs. delete the eval). The only usage-adjacent signal is the approval flow implied by the closing sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_getRead one Caliper evalARead-onlyInspect
One eval: what it targets (Workbench flow + stage, or external), its dataset and rubric ids, the rubric criteria it scores with (the snapshot taken at creation), schedule + runOnPublish + regression settings, run counts, latest score, and binding health (datasetOk / flowOk false = a run would fail at resolution — say so before proposing caliper_evals_run). Get the evalId from caliper_evals_list or caliper_flow_performance.
| Name | Required | Description | Default |
|---|---|---|---|
| evalId | Yes | Eval id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this as a safe read (readOnlyHint=true, destructiveHint=false), so the safety profile is covered. The description adds genuine behavioral context beyond that: the rubric criteria are a creation-time snapshot, and datasetOk/flowOk=false means a run would fail at resolution, which is decision-relevant context the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is front-loaded and every clause describes a distinct returned field, so nothing is filler. It is a single dense run-on sentence, which slightly hurts scanability for an agent parsing it quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of describing the return payload, and it does so comprehensively: it lists every meaningful field, explains the binding-health semantics, and routes the agent to a follow-up action. Nothing an agent needs to call or interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both evalId and workspace are already fully documented by the schema, setting the baseline at 3. The description adds only a sourcing hint for evalId (get it from caliper_evals_list / caliper_flow_performance) rather than any new format or constraint meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates precisely what a single eval record contains (target, dataset/rubric ids, rubric snapshot, schedule, run counts, latest score, binding health), so the agent knows exactly what this returns. The verb itself is only implied via the name/title, and sibling differentiation against caliper_evals_list is present but indirect (it only says where to source the evalId).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete task-oriented guidance: check binding health and 'say so before proposing caliper_evals_run,' which is a real precondition an agent should act on. It names caliper_evals_list and caliper_flow_performance as where to obtain the id, but states no explicit when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_listList Caliper evalsARead-onlyInspect
Lists evals in the active workspace with their target config (which Workbench flow, which dataset/rubric). Use to find the evalId for caliper_evals_run.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds that results are scoped to the active workspace and include target config, but says nothing about pagination, result ordering, or volume.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The resource/scope statement comes first and the routing hint second, so the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing what each listed eval contains (target config with flow and dataset/rubric). For a simple read-only list tool with full schema coverage on its one parameter, this is nearly sufficient; only pagination/ordering behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single optional workspace parameter with 100% schema description coverage, so the schema carries the semantics (default workspace, override behavior, API-key case). The description's 'active workspace' phrasing aligns with that. Zero-param-adjacent baseline applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) plus resource (evals) and narrows the scope to the active workspace, then names the returned content (target config: which Workbench flow, dataset/rubric). This distinguishes it from caliper_evals_get (single eval) and caliper_evals_run, which it routes the agent toward.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: to find the evalId needed by caliper_evals_run. That is a clear downstream purpose, but it names no alternative for other discovery needs (e.g. caliper_evals_get for a known id) and gives no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runRun a Caliper evalAInspect
Queues a new run of an eval — inference over the dataset, then LLM-judge scoring. THE VERIFY STEP of the improvement loop: after an approved workbench_flows_edit_text, run the eval again and report the score delta vs the previous run. Runs take a while — but you're brought back into THIS conversation automatically with the scores the moment it finishes, so tell the user it's queued and that you'll follow up here; never poll or ask them to check back. Costs workspace LLM budget, so it sits behind the approval gate: may return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Short label for the run, e.g. 'after refund-policy fix'. | |
| evalId | Yes | Eval to run. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly=false, destructive=false, openWorld=true); the description adds substantial context beyond them — the run is long-running and re-enters THIS conversation automatically on completion, it consumes workspace LLM budget, and it may return a `needs_confirmation` envelope behind an approval gate. That is exactly the behavioral detail an agent needs to avoid polling or misreporting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and well organized, but the mid-section sentence about auto-return, budget, and approval is dense and runs long. Nearly every sentence earns its place, so it stays efficient despite the density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining that the scores arrive via an automatic follow-up in this conversation, plus the approval-gate response shape. For a cost-bearing async mutation tool, nothing essential for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (evalId, label, workspace, approvalId) are already documented in the schema; the description's mention of the approval gate loosely motivates approvalId but adds no syntax or format detail. Baseline 3 applies when the schema carries the parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Queues a new run of an eval' — and immediately defines the operation as inference over the dataset followed by LLM-judge scoring. It is clearly distinguishable from siblings like caliper_evals_runs_get, caliper_evals_runs_list, and caliper_evals_runs_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions the tool as 'THE VERIFY STEP of the improvement loop' and gives the trigger condition: after an approved workbench_flows_edit_text, run the eval again and report the score delta. It also tells the agent not to poll and what to tell the user, which fully closes the usage loop.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_cancelStop a running Caliper eval runADestructiveInspect
Stops a run that is still PENDING / RUNNING / SCORING — the brake on a run that's spending more than expected or was started by mistake. The queue stops at once; an item already handed to the flow finishes on its own timeout. The run settles as FAILED with a 'cancelled' reason and keeps the items it completed. A run that already finished returns run_not_live. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | Run id, from caliper_evals_run or caliper_evals_runs_list. | |
| evalId | Yes | Eval the run belongs to. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint=true annotation by disclosing exactly how destruction is scoped: the queue stops immediately, an in-flight item finishes on its own timeout, the run settles as FAILED with a 'cancelled' reason, and completed items are retained. It also surfaces the run_not_live failure mode and a possible needs_confirmation response that requires a follow-up call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the operation and its eligible states, then layers consequence, error, and confirmation behavior in short sentences with no filler. Every sentence carries information an agent needs before calling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by covering the terminal state of the run, what is preserved, and both the error and confirmation paths. Combined with destructiveHint annotations and fully documented params, nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so runId, evalId, workspace, and approvalId are all documented in the schema; the description adds only the indirect hint that needs_confirmation implies a second call carrying an approvalId. Baseline 3 is appropriate given the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Stops') and resource ('a run') with an explicit eligible-state set (PENDING / RUNNING / SCORING), which distinguishes it from read-only siblings like caliper_evals_runs_get and caliper_evals_runs_list. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete motivating contexts — a run spending more than expected, or started by mistake — and an implicit exclusion via 'A run that already finished returns run_not_live.' It does not name an alternative tool for the already-finished case, but no sibling offers a competing cancellation path, so little is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_getRead one eval run (per-item, per-criterion scores)ARead-onlyInspect
One run: status, overall score, and by default the ten lowest-scoring items with their per-criterion scores and the judge's reasoning (items: all for every item, none for totals only). THE ANSWER to 'why did the score drop' — read this, then name the criterion that slipped and quote the reasoning on the lowest items. Works for runs Caliper ran and for runs submitted from CI (triggeredBy 'ci', usually labeled with a pull request number). Still no polling loops; read it when the user asks.
| Name | Required | Description | Default |
|---|---|---|---|
| items | No | Which items to include. `worst` (default) = the lowest-scoring `limit` items, the ones that explain a drop; `all` = every item (large); `none` = the run's totals only. | |
| limit | No | How many items for `worst`, default 10. | |
| runId | Yes | Run id, from caliper_evals_runs_list or caliper_evals_run. | |
| evalId | Yes | Eval the run belongs to. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a safe, read-only, non-destructive, closed operation, so the safety profile is covered. The description adds genuine output-shape context: by default returns the ten lowest-scoring items, 'all' is large, 'none' yields totals only, and it works for CI-submitted runs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the return content ('One run: status, overall score...') and stays compact, though the capitalized 'THE ANSWER' aside is slightly editorial. Every sentence carries information about output or usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey return values, and it does: status, overall score, per-item per-criterion scores, judge reasoning, and the items/limit default behavior. An agent has enough to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates the items behavior (all/none) that the schema already documents and adds only the 'explain a drop' framing, adding little beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: reads one eval run, returning status, overall score, and by default the lowest-scoring items with per-criterion scores and judge reasoning. It clarifies the run types it covers (Caliper-run and CI-submitted), though it never names sibling tools like caliper_evals_runs_list to differentiate by task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: 'THE ANSWER to why did the score drop — read this, then name the criterion that slipped,' and an explicit exclusion, 'Still no polling loops; read it when the user asks.' It does not name an alternative tool to use instead when the user just wants a run list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_item_executionRead the flow trace behind one scored itemARead-onlyInspect
What the flow actually DID on one run item: every step in order (what it said, which tools it called with what arguments, what came back) and the final answer, plus status, duration, and cost. Read this when a low score needs explaining beyond the judge's reasoning — a wrong tool call or an empty tool result is usually the cause, and the fix is different from a prompt fix. Pass the run item's id (NOT itemId) from caliper_evals_runs_get. Items whose outputs were supplied from outside (external evals) have no trace and return not_found. Steps are capped for transport; the run page in Caliper has the full record.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | Run id. | |
| evalId | Yes | Eval the run belongs to. | |
| runItemId | Yes | The run item's `id` from caliper_evals_runs_get (not its dataset itemId). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-destructive read (readOnlyHint=true, destructiveHint=false, openWorldHint=false), and the description adds real behavioral context beyond that: not_found for external-eval items, transport-level step capping, and a pointer to the full record on the Caliper run page. It does not discuss auth or pagination semantics for the capped step list, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all load-bearing: what it returns, when to reach for it, and the two failure/limitation caveats. It is front-loaded with the payload description. Minor redundancy between the schema's runItemId note and the description's id-vs-itemId warning costs it a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of describing returns — and it does, enumerating ordered steps (utterances, tool calls with arguments, tool results), final answer, status, duration and cost. Combined with the not_found and truncation caveats, an agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including the note that runItemId is the item's `id` rather than its dataset itemId. The description's "pass the run item's `id` (NOT `itemId`) from caliper_evals_runs_get" largely restates that schema text, adding only reinforcement and the sibling that produces the value. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns what the flow actually DID on one run item — ordered steps, tool calls with arguments, tool results, final answer, plus status, duration and cost. This is clearly distinguishable from caliper_evals_runs_get (which yields scores/metadata) and caliper_traces_get, so an agent can pick it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ("Read this when a low score needs explaining beyond the judge's reasoning") and explains the diagnostic value (a wrong tool call or empty result is usually the cause, and the fix differs from a prompt fix). It also names a when-not case: externally supplied outputs have no trace and return not_found.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_runs_listList an eval's runs (status + scores)ARead-onlyInspect
Recent runs for one eval, newest first: status (PENDING/RUNNING/SCORING/DONE/FAILED), overall score once DONE, label, and timestamps. THE CHECK-BACK for caliper_evals_run: when the user asks how the run went, read this — cite the run id and score, and compare against the PREVIOUS run's score for the delta. Still don't poll in a loop; check when the user asks or when reporting.
| Name | Required | Description | Default |
|---|---|---|---|
| evalId | Yes | Eval whose runs to list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safe-read profile (readOnlyHint, destructiveHint=false). The description goes well beyond them: newest-first ordering, the PENDING/RUNNING/SCORING/DONE/FAILED lifecycle, that a score exists only once DONE, and an anti-polling behavioral constraint. That is substantial behavioral context layered on top of the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource and return fields, then the routing guidance. Dense but every clause carries information; the em-dash-heavy final sentence runs slightly long but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly enumerates the returned fields (status, score once DONE, label, timestamps, ordering) and pairs that with the delta-comparison workflow. Nothing an agent needs to call and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so evalId and workspace are already fully documented. The description adds only the single-eval scoping ('for one eval') and does not elaborate on the workspace/permission nuances beyond what the schema states. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Recent runs for one eval') plus scope ('newest first') and the exact fields returned (status, score, label, timestamps). An agent can immediately distinguish this from caliper_evals_get and caliper_evals_runs_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames itself as 'THE CHECK-BACK for caliper_evals_run', names the triggering condition ('when the user asks how the run went'), and adds a when-not rule ('don't poll in a loop; check when the user asks or when reporting'). This is exactly the when/when-not/related-tool guidance the dimension asks for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_evals_updateUpdate a Caliper eval (metadata, schedule, alerts)ADestructiveInspect
Changes an eval's title, description, visibility, schedule (interval + time anchor), runOnPublish (queue a run whenever the flow publishes), or regressionThreshold (score drop vs the previous run that triggers an alert; null = off). Pass only what changes. The dataset, rubric snapshot, and target are fixed at creation — create a new eval to change those. Scheduled and on-publish runs spend credit on their own, so state that plainly when turning them on. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New name. | |
| evalId | Yes | Eval id, from caliper_evals_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description; null clears it. | |
| runOnPublish | No | Run whenever the targeted flow publishes a version. | |
| scheduleTime | No | Daily/weekly anchor, HH:MM 24-hour in scheduleTimezone. Null = one interval from now. | |
| scheduleInterval | No | Run cadence; null clears the schedule. Each scheduled run spends credit — say so. | |
| scheduleTimezone | No | IANA timezone for scheduleTime, e.g. America/Chicago. | |
| scheduleDayOfWeek | No | Weekly only: 0 = Sunday … 6 = Saturday. | |
| regressionThreshold | No | Score drop vs the previous run that raises an alert; null turns alerts off. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and readOnlyHint=false, so the safety profile is already partly covered; the description still adds real value by disclosing the approval flow ('May return needs_confirmation'), the ongoing credit cost of scheduled/on-publish runs, and which fields are permanently fixed after creation. It never explains what 'destructive' means here (e.g., that null clears description/schedule), leaving one gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the mutable fields in the first sentence, then immutability, then confirmation/credit notes — a logical order with little waste. The parenthetical glosses on each field make it slightly long, but every clause carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-param mutation tool with no output schema, the description covers mutation scope, immutability of non-editable fields, the confirmation/approval round-trip, and the cost side effects of enabling schedules — everything an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter (including 'null turns alerts off' and the timezone/HH:MM formats). The description restates a few of these (threshold semantics, runOnPublish trigger) without adding new syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Changes/Update) plus the resource (a Caliper eval) and enumerates exactly which attributes are mutable (title, description, visibility, schedule, runOnPublish, regressionThreshold). The immutable-field callout sharpens the boundary versus caliper_evals_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use rules ('Pass only what changes') and an explicit alternative with condition ('The dataset, rubric snapshot, and target are fixed at creation — create a new eval to change those'). It also instructs the agent to surface credit spend when enabling scheduled/on-publish runs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_flow_performanceEval performance for a Workbench flowARead-onlyInspect
How a Workbench flow is ACTUALLY doing, with receipts: every Caliper eval targeting the flow, recent runs with scores, the latest run decomposed into per-criterion averages, the score delta vs the previous run, and the worst-scoring items WITH the judge's reasoning. Use this BEFORE claiming a flow works or proposing changes — and cite the runId + scores when you do. The worst items are diagnostic: failures clustered around missing company facts suggest a knowledge gap (consider proposing a Compass interview with the workflow owner) rather than a prompt problem.
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Workbench flow id to report on. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description adds real behavioral value beyond that: it discloses the composite/aggregated nature of the output (deltas, per-criterion decomposition, judge reasoning) and gives interpretive guidance on failure clustering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the 'with receipts' framing, then progresses through outputs, usage trigger, and diagnostics in a logical order. Dense but each clause carries information; slightly verbose with the long colon-separated inventory, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry the return-value burden, and it does so thoroughly (eval coverage, run history, per-criterion averages, deltas, judge reasoning). Combined with the 100% schema coverage on inputs, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so flowId and workspace are already documented in the schema. The description adds no additional semantics about the parameters themselves (its 'runId' mention is about output citation, not an input). Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('how a Workbench flow is actually doing') and enumerates the exact contents of the report: eval coverage, recent runs with scores, per-criterion averages, delta vs previous run, and worst items with judge reasoning. This distinguishes it clearly from raw siblings like caliper_evals_runs_list or caliper_evals_runs_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use this BEFORE claiming a flow works or proposing changes') and a follow-through requirement (cite runId + scores). It also explains how to interpret worst items diagnostically. It does not name alternative sibling tools (e.g., evals_runs_list for raw run data), so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_createCreate a Caliper rubricAInspect
Creates the scoring rubric an eval's LLM judge uses — 1-10 criteria, each scored on a numeric scale (default 1-5), judged pass/fail (kind pass_fail), or checked in code with no judge (kind check: contains, not_contains, matches a regex, valid_json, max_chars, equals_expected). Write criteria about the FLOW'S JOB (accuracy to source material, tone, refusal behavior), not generic 'quality'.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | Rubric name (2-120 chars). | |
| criteria | Yes | ||
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the create-and-safe profile is covered. The description adds the default scale (1-5) and the semantics of each criteria kind, but does not mention the approval/confirmation flow implied by approvalId or any workspace scoping behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the guiding heuristic is well placed, but the middle sentence is a dash-chained dump of every check type that reproduces the schema enum verbatim. It is dense without being fully wasteful, but those enumerated values do not earn their place in the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description plus the fairly rich input schema together cover the required title/criteria and their structure well. Remaining gaps (approvalId, visibility, workspace) are already documented in the schema, so the description is nearly complete for this create operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema documents most fields. The description restates the kind values (scale, pass_fail, check) and enumerates the check types, largely duplicating schema enum descriptions rather than adding syntax or constraints. Baseline 3 is appropriate for mid-level coverage where the description re-explains rather than extends.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (creates the scoring rubric) and pins down its domain role ('an eval's LLM judge uses'), plus the concrete shape (1-10 criteria, three kinds). It is clearly distinguishable from rubrics_get/list/update/delete siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers real guidance on how to author criteria ('about the FLOW'S JOB ... not generic quality'), which is valuable. However, it never says when to reach for this tool versus caliper_rubrics_update or how it relates to evals_create — usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_deleteDelete a Caliper rubricADestructiveInspect
Deletes a rubric. Existing evals keep their snapshot of it and keep running; nothing new can bind to it. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| rubricId | Yes | Rubric id, from caliper_rubrics_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds real value beyond them: existing evals retain a snapshot and keep running, nothing new can bind to the rubric, and a `needs_confirmation` response may occur (linking to the approvalId param). This is exactly the consequence-level context a destructive tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no waste; the core action leads and the consequences and confirmation behavior follow compactly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output-schema tool, the description covers the key behavioral facts (retention, binding block, confirmation flow). Minor omissions remain, such as explicit irreversibility wording, but nothing an agent critically needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented (including the approvalId/needs_confirmation linkage). The description adds nothing beyond what the schema provides for parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Deletes a rubric'), which is unambiguous against the sibling set (rubrics_create/get/list/update). It does not explicitly differentiate itself from alternative lifecycle operations, but the operation is clear from the phrasing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains consequences of deleting but gives no guidance on when to delete versus alternatives (e.g., update vs delete), nor any prerequisites for invoking it. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_getRead a Caliper rubric (criteria + scales)ARead-onlyInspect
One rubric with its criteria — each criterion's name, what the judge looks for, and the score scale. Read this to explain a score (which criterion slipped and what it asks for) or before caliper_rubrics_update. Get the rubricId from caliper_rubrics_list or an eval's rubricId.
| Name | Required | Description | Default |
|---|---|---|---|
| rubricId | Yes | Rubric id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and a closed world, so the safety profile is covered. The description adds real value by describing the returned structure (criteria, judge focus, scale), which matters because there is no output schema, though it says nothing about pagination or size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; contents are front-loaded, followed by usage conditions and the id source. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only fetch with no output schema, the description covers return shape, usage context, and prerequisite sourcing of the required id. An agent has everything needed to select and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by telling the agent where to source the required rubricId (rubrics_list or an eval's rubricId) and notes the read-before-update relationship. That provenance guidance is genuine added meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('One rubric with its criteria') and enumerates the payload contents (criterion name, what the judge looks for, score scale). It is cleanly distinguishable from caliper_rubrics_list and caliper_rubrics_update in the sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives two when-to-use conditions ('to explain a score' and 'before caliper_rubrics_update') and names the alternatives for obtaining the required id ('caliper_rubrics_list or an eval's rubricId'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_listList Caliper rubricsARead-onlyInspect
Lists the rubrics in the active workspace (id, name, description, criterion count). Check here BEFORE caliper_rubrics_create — reuse an existing rubric's id in caliper_evals_create when one already scores the same job. Read a rubric's criteria with caliper_rubrics_get.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful context beyond that: the listing is scoped to the active workspace and returns id, name, description, and criterion count, which helps an agent interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: what it returns, when to call it relative to create, and where to read criteria. The workflow-critical instruction is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by naming the returned fields, and the single optional parameter is fully documented. An agent has everything needed to call it and act on the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional parameter, so the schema already carries the full load. The description adds no syntax or format detail about the workspace parameter beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (rubrics), scopes it to the active workspace, and enumerates the returned fields. It is clearly distinguishable from sibling tools like caliper_rubrics_get and caliper_rubrics_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to check here BEFORE caliper_rubrics_create and gives the reason (reuse an existing rubric id in caliper_evals_create), and routes deep reading to caliper_rubrics_get. When-to-use and alternative tool are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_rubrics_updateUpdate a Caliper rubricADestructiveInspect
Changes a rubric's title, description, visibility, or replaces its criteria wholesale (pass the FULL list — criteria get fresh ids). Evals snapshot the rubric when they're created, so an existing eval keeps scoring with the criteria it started with; say so, and offer to create a new eval when the criteria change materially. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | New name. | |
| criteria | No | Replacement criteria (1-10) — the whole list, not a diff. | |
| rubricId | Yes | Rubric id, from caliper_rubrics_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description; null clears it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond destructiveHint=true by disclosing that criteria are replaced wholesale with fresh ids, that existing evals keep scoring against their snapshotted criteria, and that a needs_confirmation response may come back. This is meaningful behavioral context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The what-it-does clause is front-loaded and the criteria-replacement warning follows immediately. The embedded instruction to 'say so, and offer to create a new eval' is slightly meta but earns its place as routing guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description covers the mutation semantics, the confirmation flow, and the eval-snapshot consequence, while parameters are fully documented in the schema. Only the mechanics of the needs_confirmation response are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics the schema doesn't carry: that criteria ids are regenerated on replacement and that the eval snapshot behavior affects whether the change is safe to make.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+field set: changes a rubric's title, description, visibility, or replaces its criteria. An agent can immediately distinguish this from caliper_rubrics_create/delete/get/list without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when the operation applies (changing mutable rubric fields, wholesale criteria replacement) and workflow guidance around eval snapshots, but never states explicit exclusions versus the create/delete siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_source_feeds_listList a source's live datasetsARead-onlyInspect
Datasets this source keeps adding to as conversations arrive: which dataset, for what (review, eval, spec), the pick it matches, sampling, how many have landed, and whether it's still running or why it stopped.
| Name | Required | Description | Default |
|---|---|---|---|
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and destructiveHint=false, so safety is covered. With no output schema, the description usefully discloses the return shape (dataset, purpose, matching pick, sampling, count landed, running/stopped status), which is genuine behavioral context beyond the annotations. It omits pagination/rate-limit behavior, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a compact colon list of returned fields. Dense but every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description compensates by enumerating the returned fields, and the annotations cover the safety profile while the schema covers both parameters. Nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with sourceId and workspace fully documented in the schema, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource: datasets a source keeps appending to as conversations arrive, i.e. the source's live feeds. This is specific and distinguishable from generic dataset listing, but it never names or contrasts with siblings like caliper_source_feeds_stop or caliper_datasets_list, so the boundary must be inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no mention of the alternative caliper_source_feeds_stop or how this differs from caliper_datasets_list. The agent must infer context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_source_feeds_stopStop a live datasetADestructiveInspect
Stops a source from adding new conversations to a dataset. What it already added stays. Get the feedId from caliper_source_feeds_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| feedId | Yes | Live dataset id, from caliper_source_feeds_list. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds real behavioral value beyond them: it clarifies the destruction is scoped to future additions only ('What it already added stays') and warns that the call 'May return needs_confirmation', which maps to the approvalId two-step flow. It stops short of explaining the confirmation response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the effect of the operation front-loaded before the parameter hint and the confirmation caveat. Every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and an approval flow, the description covers effect, scope of data retention, prerequisite id, and the possible confirmation response. It could say slightly more about how to proceed after needs_confirmation, but the approvalId parameter fills that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all four parameters including workspace and approvalId are already documented in the schema. The description only reinforces the feedId provenance, adding no format or constraint detail beyond what the schema states. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with precise scope: 'Stops a source from adding new conversations to a dataset.' The follow-up 'What it already added stays' distinguishes it from removal-style siblings like caliper_datasets_remove_item, so an agent can differentiate without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies a precondition for invoking the tool ('Get the feedId from caliper_source_feeds_list'), which is useful routing context. However, it never states when to stop a feed versus deleting the dataset or removing items, and gives no exclusions, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_sources_listList Caliper sourcesARead-onlyInspect
Apps sending their agent's traces to Caliper — from their own code, OpenTelemetry, or a published Workbench flow — busiest first: name, how it sends, traces in the last 14 days (per day), and the tools its agent calls most with the share of traces using each. Start here when the user asks how their agent behaves on real traffic; then caliper_sources_over_time for one source.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, openWorldHint=false), so the description only needs to add context — which it does: sort order ('busiest first'), the 14-day window for trace counts, and the shape of the payload. It omits pagination/limits and auth expectations, keeping it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource definition and packed with useful detail in two sentences. Density is high and the nested em-dash clauses ask for careful reading, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the burden of describing returns — and it does, enumerating source name, ingestion method, per-day trace counts, and top tools with shares. It stops short of mentioning result limits or paging, a minor gap for a list endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage on a single parameter, the schema already documents the workspace slug, defaults, and API-key behavior fully; the description adds nothing about parameters. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list Caliper sources) and immediately defines what a 'source' is — an app sending traces via its own code, OpenTelemetry, or a published Workbench flow. This clearly distinguishes it from siblings like caliper_sources_over_time, caliper_source_feeds_list, and caliper_sources_to_dataset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to reach for it ('Start here when the user asks how their agent behaves on real traffic') and names the follow-on tool with its condition ('then caliper_sources_over_time for one source'). This is textbook routing guidance with an alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_sources_over_timeRead a Caliper source over timeARead-onlyInspect
One source day by day: traces, failures, tool calls, lookups that found nothing, tokens, cost, and speed (meanMs; p50/p95 as 'answered within' bucket edges); each tool with its calls and the share of each day's traces that used it; the documents lookups landed on most; models used. Use it to answer 'how often is it calling web search and is that changing' — compare the first and last weeks and name the day it moved. Follow with caliper_traces_list on that day or tool to show the conversations behind it.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | How many days back, default 30. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so safety is covered. The description usefully adds the shape of the aggregation and the meaning of latency values (p50/p95 as "answered within" bucket edges), which the annotations cannot convey. It omits any note on cost of computing 365-day windows, caching, or auth constraints beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource, then the metrics, then the recommended workflow. The metric enumeration is dense and the parenthetical on p95 bucket edges is slightly awkward, but each clause carries real information about the return shape rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns and does so thoroughly — metrics, per-tool breakdowns, share-of-traces, top documents, models — and tells the agent what to do next. Nothing essential for calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so sourceId, days (default 30, max 365) and workspace are already fully documented, putting the baseline at 3. The description adds no syntax or default information beyond the schema, though it does frame what the returned window represents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — reading one Caliper source day by day — and enumerates exactly what is computed (traces, failures, tool calls, tokens, cost, latency percentiles, documents, models). This distinguishes it from listing tools like caliper_sources_list and from trace-level siblings such as caliper_traces_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ("how often is it calling web search and is that changing") and a concrete analytic procedure (compare first/last weeks, name the day it moved), plus a follow-up route to caliper_traces_list. It does not state when NOT to use it or contrast it with a potentially overlapping analytics sibling like caliper_flow_performance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_sources_to_datasetSave a source's conversations to a datasetAInspect
Saves the conversations behind a pick (a day, a tool, a document, failures — the same narrowing as caliper_traces_list) into a new dataset (newDatasetName) or an existing one (datasetId), newest first up to limit. purpose decides what each item keeps: review = the agent's reply, tool calls included, for people to rate; eval = the reply becomes the expected answer; spec = only the questions, for people to answer. Set keepAdding to make it live: new matching conversations keep arriving (every one, or 1 in 10 / 1 in 100), up to 1,000. Then offer the next step — caliper_evals_create, or a review in Caliper. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | One UTC day, YYYY-MM-DD. | |
| tool | No | Only traces that called this tool. | |
| limit | No | Most recent N, default 100. | |
| purpose | Yes | review | eval | spec. | |
| document | No | Only traces whose lookups hit this document. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| datasetId | No | Add to this dataset… | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| errorsOnly | No | Only failed traces. | |
| keepAdding | No | Keep adding new matches as they arrive, taking one in this many (1 = every one). | |
| newDatasetName | No | …or start one with this name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-destructive, closed-world write (readOnlyHint=false, destructiveHint=false, openWorldHint=false), and the description adds genuinely non-derivable behavior: keepAdding turns the dataset into a live feed capped at 1,000 items, one-in-10/one-in-100 sampling, and that the call may return `needs_confirmation` with an approvalId follow-up. It omits any note about what happens to pre-existing dataset contents when adding, which is the one remaining gap for a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, then layers purpose semantics, live-mode behavior, and the next step in a tight sequence with no filler. It is on the denser side for one paragraph, but every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters, a 100% documented schema, and no output schema, the description covers the non-obvious flow (new vs existing dataset, live mode caps, confirmation handshake, follow-on tools) without needing to restate return shape. Minor gaps remain around workspace/token handling and interaction with existing dataset contents, both of which are partially covered in the schema itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description goes further by explaining what each `purpose` value actually retains (review = reply plus tool calls for rating; eval = reply becomes the expected answer; spec = questions only) and the meaning of datasetId vs newDatasetName and the sampling factor. This is real semantic value the bare enum cannot convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — saving the conversations behind a pick into a new or existing dataset — and immediately scopes it against the sibling narrowing tool caliper_traces_list, which it shares semantics with. An agent can distinguish this from caliper_datasets_add_items and caliper_datasets_create without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the decision-relevant context: which purpose (review/eval/spec) fits which downstream intent, when to use a new dataset vs an existing datasetId, and explicitly names the follow-on step (caliper_evals_create or a Caliper review). It stops short of stating when not to use it, so it lands just under an explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_starters_installInstall a Caliper starter (dataset + rubric)AInspect
Copies one starter into the workspace as a new dataset and a new rubric the user owns and can edit. Returns both ids; the next step is caliper_evals_create binding them to the flow (or target 'external'). Needs caliper:datasets:write and caliper:evals:write.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Starter slug from caliper_starters_list. | |
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this non-read-only and non-destructive, and the description adds real substance beyond them: what gets created (user-owned, editable dataset + rubric), what is returned (both ids), and the exact permissions required (caliper:datasets:write, caliper:evals:write). Does not mention idempotency or what happens if the slug already exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the action and result come first, then the next step and prerequisites. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the mutation, its outputs (both ids), the follow-up tool, and permission requirements, which is sufficient for a 3-param install tool with no output schema. Missing only edge-case behavior such as re-installing an existing slug.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the slug enum documented and sourced to caliper_starters_list, so the schema carries the burden. The description's mention of 'returns both ids' adds nothing about the workspace override or approvalId flow, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (copies/installs) plus the concrete resources produced (a new dataset and a new rubric the user owns). An agent can distinguish this from caliper_starters_list (read-only listing) and caliper_datasets_create (blank dataset) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the workflow: install a starter, then bind the returned ids via caliper_evals_create. It does not explicitly say when to choose this over caliper_datasets_create or caliper_rubrics_create, but the context is unambiguous for a starter-based workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_starters_listList Caliper starters (ready-made dataset + rubric pairs)ARead-onlyInspect
The starter sets Caliper ships: prompt injection, system-prompt leaking, over-refusal, personal-data handling, bias under ambiguity. Each is a small original dataset (single messages AND multi-turn build-ups: false memory, fabricated earlier turns, slow escalation) paired with an anchored rubric. Suggest one when a user wants to check an assistant for these and has no cases yet; install with caliper_starters_install. Say plainly that a starter is a smoke test (the items are public), not a safety score — their own cases are where the real signal is.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely new framing: the items are public, the starter is a 'smoke test' and not a safety score, which materially shapes how the agent should present results. It does not discuss return format or ordering, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The inventory of starters is front-loaded and the guidance follows in a logical order. It runs a bit long and embeds an instruction about how to phrase the answer ('Say plainly...'), but every sentence conveys distinct, useful information rather than restating the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description does tell the agent what it will get back (the named starter sets with their dataset and rubric composition) and how to act on it. For a zero-parameter list tool backed by read-only annotations, this is essentially complete, with only return ordering/pagination left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema carries nothing for the description to complement. Baseline 4 applies; there is no parameter ambiguity to resolve.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description concretely enumerates what the tool surfaces (five named starter sets, each an original dataset paired with an anchored rubric) and distinguishes itself from the install sibling. It never uses an explicit verb like 'list', relying on the title for that, but the content is specific enough that an agent can tell it apart from caliper_datasets_list or caliper_rubrics_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the trigger ('when a user wants to check an assistant for these and has no cases yet'), names the follow-up action and its tool ('install with caliper_starters_install'), and implies the exclusion via 'has no cases yet' plus 'their own cases are where the real signal is.' That is explicit when, when-not, and alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_traces_getRead one traceARead-onlyInspect
One trace: what came in, what went out, and every step in order — model calls (model, tokens, cost, reasoning), tool calls (arguments and result), lookups (query and documents with scores), anything else — each with timing and status. Long values are cut at 2,000 characters. Use it to explain why one conversation went the way it did.
| Name | Required | Description | Default |
|---|---|---|---|
| traceId | Yes | Trace id, from caliper_traces_list. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, non-destructive, and closed-world, so the safety profile is covered. The description adds genuinely useful behavior the annotations cannot: the enumeration of step types (model calls with tokens/cost, tool calls, lookups) and, most importantly, the 2,000-character truncation on long values. It omits any note on pagination, result size limits, or what happens with a stale/invalid traceId.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with 'One trace' and then the content inventory; the second and third sentences add the truncation caveat and the use case without padding. The middle enumeration is long but each item earns its place by telling the agent what to expect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return contents, and it does so thoroughly (steps, timing, status, truncation threshold). It is slightly incomplete on error/edge behavior for invalid ids and on whether very large traces are paginated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the workspace override rules and where traceId/sourceId come from, so the schema does the heavy lifting. The description contributes no additional parameter meaning, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('One trace: what came in, what went out') and enumerates exactly what the retrieval returns, so an agent can immediately distinguish it from the list-style sibling. It stops short of naming caliper_traces_list as the way to discover trace ids, so sibling differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use case — 'explain why one conversation went the way it did' — which tells the agent what question this tool answers. However, it offers no when-not guidance, no prerequisite chain (list first, then get), and never names an alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caliper_traces_listList a source's conversationsARead-onlyInspect
The traces behind a point on a source's chart, newest first, 50 a page: the question and answer (cut short), tools called, time, tokens, cost, status. Narrow by day, tool, document, or failures; page with before (the nextBefore from the last page). Open one with caliper_traces_get.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | One UTC day, YYYY-MM-DD. | |
| tool | No | Only traces that called this tool. | |
| before | No | Paging: the nextBefore from the previous page. | |
| document | No | Only traces whose lookups hit this document. | |
| sourceId | Yes | Source id, from caliper_sources_list or search_workspace. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| errorsOnly | No | Only failed traces. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint=false, so safety is settled. The description adds real behavioral context beyond that: newest-first ordering, 50-per-page batching, truncated Q/A text, and the fields returned (tools, time, tokens, cost, status). It doesn't state a total count or rate limits, keeping it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is listed and how it is ordered, followed by filtering, paging, and the sibling handoff. Dense but every clause carries information; only the return-field enumeration is slightly list-like.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description compensates by naming the returned fields and pagination contract, and it covers the filter and paging parameters. Nothing needed to call or interpret this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters, including the workspace/token caveat and the YYYY-MM-DD pattern. The description echoes the filter set (day, tool, document, failures) without adding syntax beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('traces behind a point on a source's chart') plus ordering and page size, which distinguishes it from caliper_traces_get, the single-trace sibling. An agent knows exactly what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete narrowing dimensions (day, tool, document, failures), the paging protocol (`before` = nextBefore from last page), and explicitly routes to caliper_traces_get for a single trace. Both when-to-use and the alternative are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_createComment on an entityAInspect
Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| rootId | No | Reply into this thread; omit to start a new one. | |
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| entityKind | Yes | What the thread hangs on. | |
| mentionedUserIds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the write/safety profile is covered. The description adds value beyond that by disclosing a side effect annotations cannot express: mentioning users rings their notification bell, with an accompanying social caution. It omits permission/auth requirements for posting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and its two modes, then layers usage and mention etiquette compactly. Two sentences, minimal waste, though the trailing 'never mention someone who didn't ask' is advisory padding rather than invocation-critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with 43% schema coverage and no output schema, the description covers the social/mention dimension well but leaves the entity-targeting parameters and the undocumented approvalId unexplained, which an agent would need to call this reliably in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, so the description must compensate, and it partially does: rootId (reply target), mentionedUserIds (workspace member ids, notification behavior), and body are implied. However entityKind's 17-value enum, entityId, and especially approvalId are unexplained in both schema and description, leaving real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Posts a comment on a workspace entity') and distinguishes the two modes of operation: new thread vs. reply when rootId is given. It never names its closest siblings (comments_list, comments_resolve), so it stops just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage context ('leave findings where the discussion already lives') with two illustrative scenarios (eval result on a debated flow, summary on a long thread). It does not state when not to use it or point to a sibling alternative, so it lacks the explicit routing of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_listRead an entity's comment threadsARead-onlyInspect
Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.
| Name | Required | Description | Default |
|---|---|---|---|
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| entityKind | Yes | What the thread hangs on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavior: threads are returned open-before-resolved and include authors and timestamps, which shapes how an agent interprets output. It stops short of pagination or volume limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the key facts front-loaded and no redundancy. The second sentence is motivational framing that carries mild value but is slightly softer than a hard routing rule.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema and a mostly documented schema, the description covers ordering, content, and the entity scoping needed to call it. Pagination/result-size behavior is the only notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the schema itself handles the workspace slug nuances and the enum list. The description only says 'one workspace entity', adding little beyond the structured fields, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (lists) and resource (comment threads on one workspace entity), plus the ordering rule (open first, then resolved) and payload (authors, timestamps). This clearly separates it from comments_create and comments_resolve, which mutate rather than read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives a soft situational cue ('read this before weighing in on contested work'), which implies when the tool is useful. However it names no alternatives and gives no explicit when-not or prerequisite guidance, so usage remains inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comments_resolveResolve or reopen a threadADestructiveInspect
Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| rootId | Yes | ||
| resolved | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=false, so the mutation/safety profile is covered. The description adds real behavioral context beyond that: the precondition (a reply explaining what settled the thread) and the reopen semantics. It doesn't clarify reversibility or how the state change affects existing replies, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the action and the primary precondition are front-loaded before the reopening clause. Every clause carries decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small toggle tool with no output schema, the description covers the action, the target, and the conditions well. It leaves approvalId and workspace behavior unexplained, which matters for a destructive write, but overall an agent has enough to act correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25%, and the description compensates for rootId only (root comment id). 'resolved' is implied by resolve/reopen framing, but 'workspace' is documented only in the schema and 'approvalId' is explained nowhere — a notable gap for a mutation tool with an approval parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (sets resolved state) plus resource (comment thread), and covers both directions — resolve and reopen. The parenthetical 'rootId = the thread's root comment id' disambiguates the target, making it clearly distinct from comments_create/comments_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gates the action: resolve ONLY when the human asked or the question is demonstrably settled, and reply first explaining what settled it; reopen is for new evidence. This is genuine when/when-not guidance with a required prior step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_changes_upsertBring changes in from another trackerADestructiveInspect
Create or update up to 100 changes in one call, each matched by url — the GitHub issue, Linear or Jira ticket, PR or doc it mirrors. A url already linked to a change updates that change (label, description, status, tags, experiment, workflow — omitted fields are left alone); any other url creates a change with that link. Run it again with the same urls to keep Compass in step; nothing duplicates. Statuses are by name (compass_statuses_list). Returns one result per item — created / updated with the change's key, or failed with why; a failed item doesn't stop the rest. Summarize what you're about to bring in before calling. May return needs_confirmation — one approval covers the whole batch.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (which only flag the write + destructive + closed-world profile): it discloses partial-update semantics ('omitted fields are left alone'), idempotency ('nothing duplicates'), per-item isolation ('a failed item doesn't stop the rest'), the 100-item cap, and the needs_confirmation approval flow where one approval covers the batch. This is unusually rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded — the core upsert action leads, followed by matching semantics, side effects, and confirmation behavior. Parenthetical and em-dash clauses pack multiple facts per sentence, which is efficient though slightly heavy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains the return shape (one result per item: created/updated with key, or failed with reason) and the confirmation path. For a batched mutation tool this covers everything an agent needs to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and the description closes gaps: `url` is the match key mirroring a GitHub/Linear/Jira ticket, PR or doc; omitted fields are left untouched; statuses are referenced by name via compass_statuses_list. It does not explain `workspace` or `approvalId`, but those already carry schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (create or update) and resource (changes), plus scope: up to 100 per call, each matched by `url`. The upsert-by-url framing is unambiguous and no sibling tool operates on changes, so an agent can select this without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: run again with the same urls to keep Compass in step, summarize before calling, and use compass_statuses_list for status names. It does not name a when-not-to-use case or an alternative tool, but no overlapping sibling exists, so the routing guidance is effectively complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_createRaise a Compass open question (gap)AInspect
Records something you don't yet know — the gap registry is your working memory, so use it liberally while mapping or interviewing. Give one clear question, WHY it matters (what answering it unblocks), and how you'd resolve it (ask the user / interview a specific person / connect a source). No approval needed — noting your own uncertainty isn't acting on the user's behalf. Attach it to a subject (subjectType/subjectId) when it's ABOUT a specific page or person.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | The open question, plainly ('Who approves refunds over $500?'). | |
| rationale | No | Why it matters / what answering it unblocks. | |
| subjectId | No | Id of the subject. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| subjectType | No | What it's about, e.g. compass_page. | |
| suggestedResolution | No | How to resolve it: 'ask the user', 'interview Dana', 'connect Drive'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag a non-destructive write (readOnlyHint=false, destructiveHint=false), and the description adds meaningful behavioral context beyond them: no approval is required and why ('noting your own uncertainty isn't acting on the user's behalf'). It doesn't describe the returned record or any limits, but the auth/approval disclosure is genuinely additive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences that move from what to record, to why no gating is needed, to when to attach a subject. Dense but each sentence carries guidance; no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-record write with no output schema and full schema coverage, the description supplies the missing 'why' and 'approval' context plus subject-attachment guidance. It leaves workspace selection entirely to the schema but otherwise covers what an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, and the description earns an increment by tying the fields to intent — the question, the WHY (rationale), and the resolution path (suggestedResolution) — plus the conditional semantics of subjectType/subjectId ('when it's ABOUT a specific page or person').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — recording an open question into the gap registry — and immediately distinguishes itself from siblings by describing the registry's role as working memory. An agent can separate it from compass_gaps_list/resolve/update without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it ('while mapping or interviewing', 'use it liberally') and specifies the condition for attaching a subject ('when it's ABOUT a specific page or person'). It does not explicitly name the sibling tools it supersedes (gaps_update/resolve), so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_listList Compass open questions (gaps)ARead-onlyInspect
Lists the workspace's open questions — the gap registry: what the map doesn't know yet. Defaults to OPEN gaps. This is where you keep your head — check it before asking the user something you may already have flagged, and compose interview briefs from a subject's open questions.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter by lifecycle stage. Default OPEN. | |
| subjectId | No | Only gaps ABOUT this subject id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| subjectType | No | Only gaps ABOUT this subject type (e.g. compass_page). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds the default-OPEN behavior and the pre-ask check habit, but says nothing about result volume, pagination, or ordering — modest added value beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core statement of what is listed and the default, followed by usage. The colloquial aside ('This is where you keep your head') costs a few words but does carry usage guidance, so little is wasted overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with four fully documented, all-optional parameters and no output schema, the description covers purpose, default filtering, and workflow integration adequately. Only pagination/return-shape behavior is left unstated, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (status enum, subjectId, subjectType, workspace) are already documented in the schema. The description only restates the OPEN default and the 'subject's open questions' notion, adding no syntax or semantics beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the workspace's open questions — the gap registry') and characterizes what a gap is ('what the map doesn't know yet'). This clearly distinguishes it from the sibling mutations compass_gaps_create/resolve/update without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete workflow triggers: check this before asking the user something already flagged, and use it to compose interview briefs from a subject's open questions. That is strong when-to-use guidance, though it names no explicit alternative or exclusion (e.g. vs. compass_inbox_list).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_resolveResolve a Compass open question (gap)ADestructiveInspect
Closes a gap once you've learned the answer (ANSWERED, with the resolution) or decided it doesn't matter (DISMISSED). Keep the registry honest — resolve gaps as their answers land (from an interview, a doc, or the user) so it always reflects what's still unknown. Gaps you raised close without a card; a question a person wrote needs their approval.
| Name | Required | Description | Default |
|---|---|---|---|
| gapId | Yes | Id of the gap (from compass_gaps_list). | |
| status | Yes | ANSWERED (you learned it) or DISMISSED (irrelevant). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Re-call with the approvalId from a needs_confirmation response after the user approves. | |
| resolution | No | The answer / note, when marking ANSWERED. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the terminal-state safety profile is covered. The description adds genuinely new behavioral context: gaps you raised close without a card, while a person-written question requires their approval — an authorization workflow not encoded in the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and its two outcomes. The middle 'Keep the registry honest' sentence is mildly exhortative but carries the when-to-use guidance, so it earns its place; nothing dangles.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers purpose, trigger conditions, status semantics, and the approval prerequisite. It does not describe the success response or reversibility, but the annotations plus 100% schema coverage leave the agent adequately equipped.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so gapId, status, workspace, approvalId, and resolution are all documented in the schema. The description reinforces the ANSWERED-with-resolution relationship and the approvalId flow but adds little beyond what the schema already states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Closes/resolve) and resource (a gap = open question), and enumerates the two terminal outcomes ANSWERED and DISMISSED with their meanings. It clearly conveys the terminal-state nature, but never names the adjacent sibling compass_gaps_update, so an agent must infer the resolve-vs-update boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says when to act — resolve gaps as their answers land, explicitly listing sources (interview, doc, or the user) and the DISMISSED case when it proves irrelevant. It lacks explicit 'instead of' routing to siblings, but the trigger conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_gaps_updateReword a Compass open question (gap)ADestructiveInspect
Edits an open question's wording, rationale, or suggested resolution — for sharpening a vague question or fixing one you phrased badly. Closing a gap is compass_gaps_resolve, not this. Gap ids come from compass_gaps_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| gapId | Yes | Id of the gap (from compass_gaps_list). | |
| question | No | New wording of the question. | |
| rationale | No | Why it matters / what answering it unblocks. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| suggestedResolution | No | How to resolve it: 'ask the user', 'interview Dana', 'connect Drive'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so the safety profile is carried structurally. The description adds genuine context by disclosing the 'needs_confirmation' return path, but it never explains why an edit is destructive (e.g. what prior wording is overwritten) nor the approval requirement behind the destructive hint, which is the notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, then exclusions, then id provenance and the confirmation behavior. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation with full schema coverage and annotations covering safety, the description covers purpose, alternative, id source, and the confirmation flow. It stops short of explaining the destructive nature of the edit or the workspace/token rules that the schema mentions, and with no output schema it only names 'needs_confirmation' rather than describing returns, but it is nearly sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so gapId, question, rationale, workspace, approvalId, and suggestedResolution are all documented in the schema. The description only reinforces the field set and the provenance of gapId, adding nothing on syntax or format beyond what the schema already provides, which is the expected baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Edits) and resource (an open question's wording, rationale, or suggested resolution), enumerating exactly which fields change. It names the sibling it is not (compass_gaps_resolve) so an agent can distinguish editing from closing without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('for sharpening a vague question or fixing one you phrased badly') and an explicit when-not, routing closing to compass_gaps_resolve. It also tells the agent where the required gapId comes from (compass_gaps_list), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_inbox_listWhat's waiting for a decision in CompassARead-onlyInspect
The Compass Inbox in one read: finished interviews awaiting review (transcript in, not yet turned into pages), AI-proposed opportunities awaiting accept / dismiss, and the caller's pending approval cards for Compass writes. THE place to answer 'what needs me?' for the map. Next moves: compass_interviews_get to read a transcript, then propose pages from it with compass_pages_create; compass_opportunities_accept or compass_opportunities_delete for a proposal.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuine context beyond that: the aggregate nature of the result, the state filter ('transcript in, not yet turned into pages'), and that approval cards are scoped to the caller. Pagination/limits are not mentioned, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the three inbox categories before the routing sentence, so the agent gets the core scope first. The colon-list and semicolon-chained next-moves sentence are dense but every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe what comes back, and it does: the three item categories and their pending states. Combined with annotations covering safety and the schema covering the sole parameter, an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional workspace parameter, so the schema fully documents it. The description adds nothing about the workspace parameter (e.g., how it affects which inbox items appear), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and a precisely scoped resource (the Compass Inbox) and enumerates the three categories it aggregates: finished interviews awaiting review, AI-proposed opportunities awaiting accept/dismiss, and the caller's pending approval cards. This distinguishes it from the narrower siblings compass_interviews_list and compass_opportunities_list without the agent needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions itself as 'THE place to answer what needs me?' and gives concrete next moves that route to compass_interviews_get/compass_pages_create and compass_opportunities_accept/delete. It lacks an explicit when-not (e.g., pointing to compass_opportunities_list for a full unfiltered list), so it stops short of full alternative coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interview_inviteEmail an interview inviteAInspect
Emails the guest link for an existing interview, framed as coming from the REQUESTING USER (their name signs it; replies go to them). Author message in their voice — short, human, says why THEIR knowledge matters and that it takes ~15 minutes, no account needed. The approval card shows the exact subject + message before anything sends. Limit: the invite plus ONE reminder; a third ask is the user's conversation to have. Completion arrives as a notification with draft-page counts — don't poll.
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | The interviewee's email address. | ||
| message | Yes | The body, written in the requesting user's voice. Greeting + link + signature are added automatically — write only the middle. | |
| subject | Yes | Email subject, e.g. '15 minutes on how refunds actually work?' | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| interviewId | Yes | Interview id from compass_interviews_create. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: discloses the approval-card preview of subject+message before sending, the email is sent as the requesting user (replies route to them), a hard limit of one reminder, and that completion arrives as a notification rather than a polled response. This is exactly the behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then packs voice, approval, limit, and completion into tight clauses with no filler. It is dense with several distinct concerns in one paragraph, but each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining how the result surfaces (notification with draft-page counts, don't poll) plus the confirmation flow and send limits. An agent has everything needed to invoke and follow through correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the field descriptions already carry the load (baseline 3). The description adds authoring guidance beyond the schema — 'short, human, says why THEIR knowledge matters and that it takes ~15 minutes' — which shapes how `message` and `subject` should be written.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Emails the guest link for an existing interview.' The word 'existing' distinguishes it from compass_interviews_create, and the framing detail (from the requesting user) further narrows its purpose. An agent can place it without opening any sibling schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Says to use it for existing interviews and gives clear operational rules — invite plus ONE reminder, the third ask is the user's conversation, and don't poll. It implies but never explicitly names the sibling alternatives (create vs invite vs revoke), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_createInvite someone to a Compass interviewAInspect
Mints a stakeholder-interview invite: a no-account guest link where the person talks to an interviewer agent briefed by your focusPrompt, and the transcript flows back into Compass as reviewable draft pages. THE KNOWLEDGE-GAP MOVE: when caliper_flow_performance shows failures clustered on missing company facts (the judge says the flow invented a policy, missed a rule, didn't know who owns something), the fix is usually not a prompt edit — it's asking the human who actually knows. Write a focusPrompt that names the SPECIFIC gaps (cite the eval run id), pick the owner of the relevant workflow as interviewee when the map knows one, and hand the user the invite link to forward. May return needs_confirmation — tell the user who you want to interview and why, then wait.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| focusPrompt | Yes | What the interviewer should dig into — specific, grounded in the gap you found (≤2000 chars). Never shown verbatim to the guest. | |
| contextPageIds | No | Compass page ids the interviewer gets as briefing context (e.g. the workflow page whose flow underperformed). | |
| intervieweeName | No | Who this invite is for, when known. | |
| intervieweeEmail | No | Their email, when known — enables compass_interview_invite to send the link directly (with approval). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare openWorldHint=true and readOnlyHint=false; the description adds real context on top of that — the invite is a no-account guest link, the transcript returns as reviewable draft pages, and the call may return needs_confirmation requiring a user pause. It does not cover auth or rate-limit behavior, but the confirmation and side-effect profile is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition, then the usage rule, then the closing procedure. Dense but nearly every clause carries information; the final sentence is a slightly compressed run-on covering needs_confirmation and the user hand-off, but it is not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param tool with no output schema, the description covers the important return behavior (needs_confirmation pause) and the downstream outcome (draft pages), and pairs naturally with approvalId in the schema. Enough for an agent to invoke and handle the response correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100% (baseline 3), the description adds guidance the schema lacks: focusPrompt should name SPECIFIC gaps and cite the eval run id, and interviewee should be the workflow owner when the map knows one. This meaningfully shapes how an agent fills the parameters rather than restating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('mints a stakeholder-interview invite') and describes the concrete artifact produced: a no-account guest link, an interviewer agent briefed by focusPrompt, and a transcript flowing back as draft pages. It is clearly distinguishable from sibling compass_interview_invite, which is referenced as the send step rather than the mint step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names when to use this over the obvious alternative: when caliper_flow_performance shows failures clustered on missing company facts, 'the fix is usually not a prompt edit — it's asking the human who actually knows.' It also prescribes the operational sequence (name gaps with eval run id, pick the workflow owner, hand over the link) and the needs_confirmation handling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_getRead one Compass interview, transcript includedARead-onlyInspect
One interview in full — metadata plus the transcript of what the guest actually said. Read this before summarizing an interview or drafting pages from it: the transcript is the source you work from, and your compass skill has the rules for typing and splitting what's in it. Ids come from compass_interviews_list or compass_inbox_list.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| interviewId | Yes | Interview id (from compass_interviews_list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and closed-world scope, so the safety profile is covered. The description adds genuinely useful behavioral context with no output schema: what the payload contains (metadata plus the actual transcript) and that downstream typing/splitting rules live in the compass skill.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what the tool returns before the usage advice. The clause about the compass skill's typing/splitting rules is slightly tangential but still earns its place by explaining why the transcript matters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-resource read with full schema coverage, covered annotations and no output schema, the description supplies what an agent needs: the return content, the id source, and the workflow context. It does not address failure modes (unknown id, missing workspace) but little more is required here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline would be 3. The description goes slightly beyond the schema by naming compass_inbox_list as an additional source of interviewId, which the schema's own description does not mention. It says nothing about the workspace parameter, so this is a modest lift, not a full one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource: fetch one interview in full, including metadata and transcript. It is immediately distinguishable from compass_interviews_list (plural enumeration) and from the create/revoke siblings that mutate interview state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the trigger condition explicitly ('read this before summarizing an interview or drafting pages from it') and tells the agent where the required id comes from (compass_interviews_list or compass_inbox_list). That is concrete routing guidance rather than implied usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_listList Compass interviewsARead-onlyInspect
Every stakeholder interview in the workspace, newest first: who was invited, status (INVITED / IN_PROGRESS / COMPLETED), focus, message count, whether the transcript was already reviewed (documentPageId set), and the invite URL while the link is still live. Check this before inviting someone again — interview fatigue is real — and to answer 'who have we already asked?'. Read one in full with compass_interviews_get.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: ordering (newest first), the status lifecycle values, and the fact that the invite URL is only present while the link is live. It does not mention pagination or result-size limits, which is the one remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with scope and ordering, then the returned-field inventory, then the usage trigger and the sibling pointer. The informal aside ('interview fatigue is real') is short and justifies the routing guidance rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by enumerating the returned fields and their semantics. Combined with annotations covering the read-only/non-destructive nature and the explicit pointer to the detail tool, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single workspace parameter is fully documented in the schema, so the baseline of 3 applies. The description adds nothing about the workspace scoping or how default-workspace tokens behave, leaving all parameter meaning to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Every stakeholder interview in the workspace, newest first') and enumerates exactly what each record contains — invitee, status enum, focus, message count, documentPageId, invite URL. It is immediately distinguishable from compass_interviews_get, which it explicitly routes to for full detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('Check this before inviting someone again') plus a concrete user question it answers ('who have we already asked?'), and names the alternative tool (compass_interviews_get) with the condition that selects it. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interviews_revokeRevoke an interview invite linkADestructiveInspect
Kills an interview's invite link — the guest's next visit sees that the link is no longer active. Use when an invite went to the wrong person, the user changed their mind, or the link leaked. Idempotent. Ids come from compass_interviews_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| interviewId | Yes | Interview id (from compass_interviews_list). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint=true/readOnlyHint=false annotations, the description discloses idempotency, the exact user-visible consequence of revocation, and the `needs_confirmation` approval flow. These are meaningful behavioral traits an agent could not infer from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and effect before usage triggers and edge behaviors. No filler; every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive single-object mutation with no output schema, the description covers effect, idempotency, id provenance, and the confirmation follow-up path. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so interviewId, workspace, and approvalId are already documented. The description only restates the interviewId source already present in the schema, so with the schema carrying the load a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Kills an interview's invite link') plus the concrete effect ('the guest's next visit sees that the link is no longer active'). This clearly separates it from the sibling compass_interview_invite, which creates invites rather than revoking them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions: wrong recipient, changed mind, or leaked link. It covers when to use but does not name a contrasting alternative or state when-not to use (e.g., a different tool for regenerating vs. fully revoking).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_interview_targetsWho should be interviewed about a pageARead-onlyInspect
Ranks the PEOPLE the map says know about a page (workflow, system, pain point): owns links first, then involved_in, then weaker edges. Each target carries contact email + role from their PERSON page and their latest interview (skip someone who just gave one — interview fatigue is real). Empty result = the map doesn't know an owner: ASK THE USER who runs this, create the PERSON page + owns link from the answer, and the map gets smarter. Use before compass_interviews_create to pick the interviewee.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The page the knowledge gap is about. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read (readOnlyHint=true, destructiveHint=false), and the description adds substantial extra behavior: the ranking order (owns > involved_in > weaker edges), fatigue-avoidance skip logic for recent interviewees, and the empty-result recovery path. This is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then ranking order, then empty-result handling, then the sibling pointer. Dense but each sentence carries actionable information; minor length from parentheticals keeps it just short of ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains what each target carries (contact email, role, latest interview), so an agent knows what it gets back. It also covers the failure mode, making it complete for a 2-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents pageId and workspace semantics in detail. The description adds no syntax or format detail beyond 'a page', so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (ranks) and resource (PEOPLE who know about a page), with explicit scope and an ordered edge taxonomy. It is clearly distinguishable from siblings like compass_pages_get or compass_interviews_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use before compass_interviews_create to pick the interviewee', naming the sibling and the sequencing. It also prescribes what to do on an empty result (ask the user, create PERSON + owns link), covering the when and the when-not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_links_createConnect two Compass pagesAInspect
Creates a typed, directed edge between two pages — the knowledge graph's connective tissue. Canonical directions: PERSON owns WORKFLOW/SYSTEM, PERSON involved_in WORKFLOW, WORKFLOW uses SYSTEM, SYSTEM uses SYSTEM, PAIN_POINT affects WORKFLOW/SYSTEM/PERSON, DOCUMENT documents anything, relates_to as fallback. Idempotent on (from, to, kind) — re-creating an existing edge returns it. May return needs_confirmation; tell the user what you're proposing, wait for their approval, then re-call with the approvalId.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Edge type, in canonical direction. | |
| note | No | ≤500-char qualifier when the kind alone undersells it. | |
| source | No | Provenance label. Defaults AI_ACCEPTED (agent writing under live human direction); pass USER for human-driven scripts. | |
| toPageId | Yes | Edge target page id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| fromPageId | Yes | Edge source page id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the basic safety profile (not read-only, not destructive, closed-world). The description adds real behavioral traits beyond that: idempotency keyed on (from, to, kind), the returned-existing-edge behavior, and the needs_confirmation → user approval → re-call with approvalId loop.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in sentence one, followed by direction semantics, then idempotency, then the confirmation protocol — a deliberate and dense ordering with no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with no output schema, the description covers the create semantics, idempotency, and the approval round-trip. It omits error/failure behavior and permission requirements, but the schema already carries workspace/auth and provenance details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes further by enumerating canonical direction pairs for the `kind` parameter — semantics the schema's circular 'Edge type, in canonical direction' text does not supply. The approvalId and workspace behavior are already well documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
First sentence states a specific verb and resource — creates a typed, directed edge between two pages — plus a concrete metaphor for its role in the knowledge graph. It is clearly distinguishable from siblings like compass_links_update and compass_links_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides substantive when-to-use guidance via the canonical direction list (PERSON owns WORKFLOW, relates_to as fallback), which tells the agent how to pick the right invocation. It does not explicitly name when to prefer links_update over create, but the idempotency note covers the main overlap case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_links_deleteRemove a connection between two pagesADestructiveInspect
Deletes one edge from the graph (the pages stay). Use when a connection is simply wrong — the system isn't used by that workflow, the person left the team. Link ids come from compass_page_links_list. Re-creating the same edge later revives it. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| linkId | Yes | Id of the link (from compass_page_links_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds genuinely non-obvious behavior: deleting is reversible ('Re-creating the same edge later revives it') and the call may return needs_confirmation, implying a gated approval flow. That is real context beyond the structured fields; it stops short of stating permission requirements or exactly what the confirmation response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences: the destructive scope first, the usage trigger second, the reversibility/confirmation caveat last. No filler, and the most decision-relevant fact (edges only, pages stay) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does flag the needs_confirmation outcome and its approval follow-up. Coverage is good for a single-edge delete; only details like idempotency on a stale linkId are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — linkId sourcing, workspace rules and the approvalId flow are all documented in the schema. The description only echoes the linkId provenance already stated in the schema, so this is the baseline case where the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Deletes one edge from the graph') and immediately scopes it against the obvious confusion — the pages themselves survive, which separates it from compass_pages_delete. An agent can distinguish it from compass_links_create/update and the page tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rationale with concrete triggers ('the system isn't used by that workflow, the person left the team') and points to compass_page_links_list as the source of ids. It does not name the sibling alternative (e.g., compass_links_update) or state when not to delete, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_links_updateChange a connection's reading or noteADestructiveInspect
Edits an existing edge between two pages: its kind (how the two relate) and/or its note. Use to CORRECT a connection typed wrongly — 'Dana doesn't own billing, she's involved in it'. Link ids come from compass_page_links_list. Fails with conflict when the new reading already exists between the same two pages (delete this one instead). May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | New edge type, in canonical direction. | |
| note | No | New qualifier note (≤500 chars); empty string clears it. | |
| linkId | Yes | Id of the link (from compass_page_links_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the mutation/destructive profile; the description adds real behavioral content beyond that — the conflict failure mode when an identical reading already exists, the routing advice to delete instead, and the possibility of a 'needs_confirmation' response. It stops short of stating reversibility or permission requirements for a destructive edit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then usage, then failure behavior in roughly four tight sentences with no filler. The ALL-CAPS emphasis and quoted example are slightly chatty but earn their space by disambiguating the correction use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns and does mention the 'needs_confirmation' response and conflict failure, which are the two non-obvious outcomes. It is complete enough to call correctly, though it omits the ordinary success return shape and any permission prerequisites for this destructive edit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters including the kind enum and the 500-char note limit are already documented in the schema. The description's gloss of `kind` as 'how the two relate' adds only marginal meaning beyond the schema's 'New edge type, in canonical direction.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edits an existing edge between two pages') and enumerates the two mutable fields, kind and note. This cleanly separates it from compass_links_create and compass_links_delete without the agent needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit triggering scenario ('Use to CORRECT a connection typed wrongly') with a concrete example, names the prerequisite source for link ids (compass_page_links_list), and routes the agent to an alternative ('delete this one instead') when the new reading already exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_map_viewView the Compass map (rendered image)ARead-onlyInspect
Renders the workspace's visual map to an image and returns it so you can SEE it the way the user does: an isometric drawing where every page is a structure whose shape is its type — people are figures, audiences are crowds, workflows are gears lying on the ground, systems are database drums, offerings are price tags, pain points are warning signs, documents are standing sheets of paper, values are shields, priorities are flags — sized by how many other pages connect to them, placed near what they link to, with pages connected to nothing parked to one side, and routes drawn between linked pages. A coloured ring on the ground around a page shows the changes touching it by stage (blue undecided, amber not started, green in progress, violet done; thicker = more), with a key in the top-left corner. Use this when the user asks about the shape of their map, where the problems are, what connects to what, or anything spatial. The text part counts the pages of each type, the links, and how many pages are unconnected, plus mapUrl. The image is for you to look at; do NOT put it in your reply as a markdown image (the chat can't show it). When the user should see the map, link mapUrl — e.g. [Open your map](mapUrl) — and describe what you saw.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only establish that this is a safe read; the description goes well beyond them by disclosing what the non-image payload contains (page-type counts, link count, unconnected-page count, mapUrl) and, critically, an anti-pattern instruction: do NOT embed the image in a reply because the chat cannot render it. That is exactly the kind of behavioral guidance structured fields cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The usage trigger is front-loaded and the closing instructions are tight, but the long enumeration of every page type and its symbol, plus the full colour-ring legend, is bulky and would be needed only when interpreting a specific image. It earns most of its length but not all of it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a single trivial parameter, the description carries the full burden and does so: it explains what the image depicts, what the text block returns, and how the agent should present the result. Nothing needed to call or act on this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single optional 'workspace' parameter, so the schema already documents the token/default-workspace semantics. The description adds nothing about the parameter, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('renders the workspace's visual map to an image and returns it') and the vivid type-to-shape mapping makes clear exactly what the artifact contains. No sibling tool renders a map, so it is trivially distinguishable from compass_pages_list, compass_page_links_list, and the rest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use trigger ('when the user asks about the shape of their map, where the problems are, what connects to what, or anything spatial') and pairs it with presentation guidance ('link mapUrl... and describe what you saw'). The routing condition is unambiguous rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_acceptAccept an AI-proposed opportunityADestructiveInspect
The review verb for the Inbox: promotes an AI-PROPOSED opportunity (from compass_opportunities_propose or the opportunity scout) into the workspace's own pipeline. compass_opportunities_set_status does NOT do this — a proposal stays in the Inbox until accepted. Optional status decides its lane in the same step (BACKLOG to park it, EXPERIMENTING to start it); omitted, it lands in NEW. Fails with not_proposed when the row isn't an AI proposal. Ids come from compass_inbox_list / compass_opportunities_list. May return needs_confirmation — name the proposal and wait.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Lane to place it in on accept (default NEW). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| opportunityId | Yes | The AI-proposed opportunity's id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnly=false, so mutation semantics are covered. The description adds genuinely useful behavior beyond that: the not_proposed failure condition, the needs_confirmation round-trip with approvalId, and the status defaulting to NEW. It stops short of explaining what 'promote' does to the source proposal or any irreversibility, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core action and sibling distinction lead, followed by parameter behavior, failure mode, and the confirmation caveat. Every sentence carries information, though the several clauses make it slightly heavy for a single-paragraph description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description covers the return-signal territory an agent needs (not_proposed failure, needs_confirmation handling, where ids originate). For a mutation tool with full annotation coverage, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description nonetheless adds real meaning: status chooses the lane in the same step with a stated default, and approvalId is tied to the two-step needs_confirmation flow (omit on first call). This goes beyond the schema's field-level text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (promotes an AI-proposed opportunity into the workspace pipeline) and explicitly differentiates from compass_opportunities_set_status, which 'does NOT do this.' An agent can distinguish it from all siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the source of proposals (compass_opportunities_propose / opportunity scout), the alternative that won't work (set_status), and where ids come from (compass_inbox_list / compass_opportunities_list). When-to-use is fully covered with the counter-case named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_createCapture an opportunityAInspect
Record a change — a fix, a chore, a new step, or an experiment worth trying. Anchor it to a workflow page when one fits (compass_pages_list); leave unanchored otherwise. It lands in the board's first status unless you name another (compass_statuses_list). Set experiment only when the user wants it measured against the Ledger. May return needs_confirmation — summarize and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Tags to file it under. | |
| label | Yes | ≤200 chars, the change in one line. | |
| links | No | Where the work also lives — issue, ticket, PR or doc URLs. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| experiment | No | Measure it against the Ledger. Default false. | |
| statusName | No | Start in this status (a name from compass_statuses_list). | |
| description | No | What the change is, and why it's worth doing. | |
| workflowPageId | No | Workflow page to anchor to, when one fits. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already signaling a non-destructive, closed-world write, the description adds genuinely new behavior: the default that it 'lands in the board's first status' and the two-phase approval flow ('May return `needs_confirmation` — summarize and wait for approval'). These are non-obvious traits an agent must handle, not restatements of annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then each sentence covers one decision (anchoring, default status, experiment, approval). No filler; every clause carries an actionable instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-param create tool with no output schema, the description covers anchoring, status defaulting, experiment gating, and the needs_confirmation exceptional return. It does not describe the success return shape (e.g., the created opportunity's id/link), which is the only meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with rich per-parameter descriptions, so the baseline would be 3, but the prose adds decision logic the schema lacks: when to anchor workflowPageId, when to set statusName, when experiment applies, and how approvalId follows a needs_confirmation response. It reinforces and contextualizes rather than introducing new syntax.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Record a change') and enumerates concrete examples (fix, chore, new step, experiment), so the agent knows what lands in the board. It does not, however, differentiate itself from close siblings like compass_opportunities_propose or compass_changes_upsert. Clear on what it does, silent on how it differs from neighbors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditions for the key decisions: anchor to a workflow page 'when one fits' (with compass_pages_list), leave unanchored otherwise, and set `experiment` 'only when the user wants it measured against the Ledger.' It also names compass_statuses_list for status selection. No exclusion against the sibling `propose` tool, so it stops short of full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_deleteDelete an opportunityADestructiveInspect
Removes an opportunity from the pipeline (soft delete). For an AI proposal the user doesn't want, this is the dismiss verb; for a captured experiment that was a duplicate or a mistake, the remove verb. To conclude a real experiment without evidence, prefer compass_opportunities_set_status REJECTED — that keeps the record. Ids come from compass_opportunities_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description adds crucial nuance beyond them: this is a soft delete, and the call may return `needs_confirmation` requiring an approvalId on a follow-up call. That disclosure of the confirmation flow and the soft-delete semantics is exactly the behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and its soft-delete nature, followed by the routing guidance and the confirmation caveat. Every sentence earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive-flagged mutation with no output schema, the description covers the operation semantics, when to prefer an alternative, id provenance, and the needs_confirmation return path. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters, making 3 the baseline. The description earns a bump by tying the `needs_confirmation` return to the approvalId parameter and stating that ids come from compass_opportunities_list, adding source-of-value context the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Removes an opportunity from the pipeline') and immediately qualifies the operation as a soft delete. It further disambiguates from sibling operations by naming the two distinct intents — dismiss for AI proposals, remove for duplicates/mistakes — so an agent can tell it apart from compass_opportunities_accept or set_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: names the alternative tool (compass_opportunities_set_status REJECTED) and the condition that selects it ('to conclude a real experiment without evidence ... that keeps the record'). It also states where ids originate, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_listList Compass opportunitiesBRead-onlyInspect
The workspace's changes — fixes, chores, new steps, experiments. Each carries statusName (the workspace's own word for where it is) and statusCategory (TRIAGE undecided, BACKLOG not started, ACTIVE in progress, DONE, CANCELED dropped), plus the older lane key in status. experiment: true marks a change measured against the Ledger: ledgerEntryId null means no expectation registered; ledger.verdict carries confirmed/missed after settlement. Scores (value/feasibility/risk) and a workflow anchor are optional. Filter by lane key in status or by workflowPageId; compass_statuses_list has the status names.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter to one lane. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| workflowPageId | No | Only opportunities anchored to this workflow page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, non-destructive, closed-world safety profile, so the bar is lower. The description adds real context beyond that: it decodes statusCategory values (TRIAGE/BACKLOG/ACTIVE/DONE/CANCELED), explains the experiment/ledger semantics (ledgerEntryId null = no expectation, verdict = confirmed/missed after settlement), and notes scores/anchor are optional. It stops short of listing behavior such as pagination or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads a data-model definition rather than the tool's action, and is written in heavy shorthand ('the older lane key in `status`'). Almost every clause carries information, but the structure assumes prior domain knowledge and is not scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields (statusName, statusCategory, status, ledger fields, scores, workflow anchor). But for a list tool with 0 required params it omits the default scope, ordering, result limits, and what filtering by lane key actually returns, leaving gaps in how to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters, making 3 the baseline. The description adds only marginal meaning, restating that filtering is by lane key in `status` or workflowPageId and pointing to compass_statuses_list; it does not clarify the enum values (NEW/QUALIFYING/...) mapped to the `status` param.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource ('the workspace's changes — fixes, chores, new steps, experiments') but never states the verb; the listing action is only implied by the trailing 'Filter by lane key...'. It does not distinguish this tool from siblings like compass_changes_upsert or the other compass_*_list tools, so an agent must infer the read/list role from the title alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete usage hint — filter by lane key in `status` or by `workflowPageId` — and routes the agent to compass_statuses_list for status names. However, it never states when to use this list versus create/update/set_status siblings, nor what happens when no filter is supplied (all opportunities? default scope?).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_mark_implementedRecord that the change landedAInspect
The measurement window's boundary: when the user says the experiment's change actually shipped / went live / rolled out, record the landing date. Readings before it are baseline; after it, evidence of effect. Recorded once — it cannot move afterward, so confirm the date. Attaches to the linked Ledger entry as evidence. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | ISO date the change landed — omit for today. | |
| note | No | What shipped, if worth recording. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=false) by disclosing the irreversible, write-once nature ('cannot move afterward'), the side effect of attaching to the linked Ledger entry as evidence, and the possible `needs_confirmation` response with its approvalId loop.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five dense sentences that are front-loaded with the trigger and the boundary meaning. The opening label 'The measurement window's boundary:' is slightly abstract, but every sentence carries distinct information and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers the mutation's effect, its permanence, the confirmation/approval flow, and the evidence attachment, which is everything an agent needs to invoke it correctly against a 5-param, 100%-covered schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds conceptual meaning to `at` (baseline before, effect evidence after) and ties `approvalId` to the needs_confirmation flow that is not spelled out in the annotated description alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: recording the date the experiment's change shipped, framed as the measurement window boundary. This is clearly distinguishable from sibling mutations like compass_opportunities_set_status, update, or register_expectation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when the user says the experiment's change actually shipped / went live / rolled out') and a confirmation precondition ('Recorded once — it cannot move afterward, so confirm the date'). It does not name or exclude specific sibling tools, but the when-to-use condition is unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_proposePropose an AI opportunityAInspect
File an AI-PROPOSED opportunity into the workspace's review queue — the scouting verb (the opportunity-scout routine's main move). Unlike compass_opportunities_create this needs NO approval: the proposal itself is the human gate — it lands in the Compass Inbox and the cockpit's Needs-you for accept/dismiss. Check compass_opportunities_list first so you never duplicate an idea. Anchor to a workflow page when one fits; score value/feasibility/risk 1–5.
| Name | Required | Description | Default |
|---|---|---|---|
| label | Yes | ≤200 chars, the opportunity in one line. | |
| reasoning | No | One or two sentences on why you're proposing this. | |
| riskScore | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| valueScore | No | ||
| description | Yes | What it is and why it's worth trying, grounded in the map. | |
| workflowPageId | No | Workflow page to anchor to, when one fits. | |
| feasibilityScore | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the bar is lower. The description still adds real workflow context beyond them: no approval gate, the item lands in the Compass Inbox and the cockpit's Needs-you for accept/dismiss, and the inherent duplicate risk. It omits auth/permission specifics, which is the only real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: the core action and its routing distinction come first, followed by the guard-rail and the scoring hint. Slightly packed with parentheticals but every clause carries routing or behavioral value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 8 parameters at 63% coverage, the description fills the important gaps: what happens after the call (review queue, accept/dismiss) and the duplicate-check prerequisite. It is complete enough for correct invocation, though workspace/auth edge cases are left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 63%, so the schema does much of the work. The description adds the scoring scale for value/feasibility/risk (1-5) and the anchoring intent for workflowPageId, but says nothing about reasoning or the workspace parameter's token semantics beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('File an AI-PROPOSED opportunity into the workspace's review queue') and frames it as 'the scouting verb'. It explicitly distinguishes itself from the sibling compass_opportunities_create, so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use routing ('Unlike compass_opportunities_create this needs NO approval'), a prerequisite ('Check compass_opportunities_list first so you never duplicate an idea'), and a conditional recommendation ('Anchor to a workflow page when one fits'). This is exactly the when/alternatives guidance expected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_register_expectationPre-register an experiment's expectationAInspect
The honesty mechanism: write what the experiment is expected to change BEFORE evidence exists. Creates a Ledger decision entry and links it to the opportunity — never backfill an expectation to match an outcome. Bind a metric (ledger_metrics_list) + comparator + target when the expectation is measurable; readings then land on the entry as evidence automatically. One expectation per opportunity — revise by superseding in Ledger. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Target value, in the metric's unit. | |
| deadline | No | ISO date the expectation is due by. | |
| metricId | No | Ledger metric to bind (requires target). From ledger_metrics_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| comparator | No | ||
| expectation | Yes | One sentence — what we expect this to change. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as a non-destructive write (readOnlyHint=false, destructiveHint=false), and the description adds substantial context beyond them: it creates a Ledger entry linked to the opportunity, forbids backfilling, auto-attaches readings as evidence, enforces a single expectation, and discloses the possible `needs_confirmation` return plus the supersede-to-revise workflow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core purpose and the honesty rationale, then workflow and constraints. Every clause carries information, though the density leaves little breathing room and some cross-references (ledger_metrics_list) are embedded mid-sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with 88% schema coverage and no output schema, the description covers the return signal (needs_confirmation), the evidence-linking behavior, and the one-per-opportunity rule. It leaves workspace-scoping and the comparator enum semantics to the schema, which is reasonable given the high coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 88% schema coverage the schema already documents parameters, but the description adds genuine meaning: metricId must be bound with comparator + target for the expectation to be measurable, and readings then flow in as evidence. It also implicitly ties the approvalId flow to the needs_confirmation response.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (register/pre-register an experiment's expectation) and frames it as 'the honesty mechanism' with the concrete action 'Creates a Ledger decision entry and links it to the opportunity.' It distinguishes itself from siblings by the one-per-opportunity constraint and the superseding revision path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use ('write what the experiment is expected to change BEFORE evidence exists', 'never backfill'), a hard constraint ('One expectation per opportunity — revise by superseding in Ledger'), and a conditional path for measurable expectations. It references ledger_metrics_list but does not explicitly contrast with siblings like compass_opportunities_update for non-expectation edits, so it falls just short of explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_set_statusMove a change to another statusADestructiveInspect
Move a change to one of the workspace's statuses, by name (compass_statuses_list has them — e.g. "In review"). The old lane keys (NEW / QUALIFYING / BACKLOG / EXPERIMENTING / SETTLED / REJECTED) still work and land on that lane's default status. Only changes marked as experiments follow the Ledger: when an experiment enters an In progress status with no registered expectation, offer compass_opportunities_register_expectation; when it enters Done with its Ledger entry still open, offer ledger_entries_settle. Other changes just move. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | A lane key. Use this or `statusName`. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| statusName | No | A status name from compass_statuses_list. Use this or `status`. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description complements this by disclosing the legacy lane-key fallback behavior, the experiment-only Ledger coupling, and that the call "May return needs_confirmation" (tying to the approvalId param). It doesn't explain reversibility or what a status change does to dependent data, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the name-based addressing before the conditional Ledger guidance, and every sentence carries operative information. It is dense-to-long, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the addressing options, the experiment-conditional side effects, and the confirmation return shape. It leaves the mutation's side effects on the change itself unstated, but the essential call-time context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents every parameter, including the enum values and the status/statusName alternative. The description adds only marginal value by naming example lane keys and default-status landing behavior, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Move") and resource ("a change") with the target (one of the workspace's statuses), and clarifies the two accepted addressing modes. An agent can distinguish it from compass_opportunities_update and compass_opportunities_mark_implemented without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use compass_statuses_list for names, offer compass_opportunities_register_expectation when an experiment enters In progress without an expectation, and ledger_entries_settle when it enters Done with an open Ledger entry. It also states the negative case ("Other changes just move").
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_opportunities_updateEdit an opportunityADestructiveInspect
Corrects an opportunity's content: label, description, priority, what done means, the 1–5 value / feasibility / risk scores, the owner (a PERSON page), notes, tags, workflow anchor, external link, visibility. Fields you omit are untouched; the lane is NOT here — move it with compass_opportunities_set_status. Opportunity ids come from compass_opportunities_list. May return needs_confirmation — say what you're changing and wait.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | New tag set — replaces existing tags. | |
| label | No | New one-line label (2–200 chars). | |
| assignee | No | The workspace member doing the work, by email or user id; null unassigns. They're notified. | |
| doneWhen | No | What's true when it's done, that someone else could check (≤1,000 chars). | |
| priority | No | How soon it matters; null clears it, and the map's suggestion shows again. | |
| riskScore | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| experiment | No | Measure it against the Ledger (expectation when it starts, verdict when it's done). | |
| valueScore | No | ||
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description (≤5,000 chars). | |
| externalUrl | No | Where the doing happens off-platform (URL); null clears. | |
| ownerPageId | No | PERSON page id who owns the work; null clears. | |
| opportunityId | Yes | The change's key (e.g. ACME-12) or id, from compass_opportunities_list. | |
| workflowPageId | No | Workflow page to anchor to; null un-anchors. | |
| feasibilityScore | No | ||
| qualitativeNotes | No | Free-form notes (≤20,000 chars) — replaces the current notes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply destructiveHint=true and readOnlyHint=false, so the safety profile is covered; the description adds genuinely non-obvious behavior: partial-update semantics, the needs_confirmation/approvalId loop, and that an assignee triggers a notification. It does not explicitly warn that tags/notes replace rather than merge, which the destructive annotation implies but the prose only hints at.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences front-loaded with the core action and the mutable-field list, followed by the constraint/alternative and the confirmation flow. No filler; every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 18-parameter mutation with no output schema, the description covers id provenance, the confirmation protocol, and sibling routing, which is most of what an agent needs. It stops short of documenting return values or the exact consequences of null-clearing fields, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so most parameters are already self-documented. The description restates a handful of fields (owner is a PERSON page, 1–5 scores) that the schema largely covers, adding little syntax or format detail beyond it. Baseline 3 is appropriate when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb+resource ('Corrects an opportunity's content') and enumerates exactly which fields are editable, then explicitly carves out what is NOT editable ('the lane is NOT here — move it with compass_opportunities_set_status'). An agent can distinguish this from sibling updaters without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the partial-update contract ('Fields you omit are untouched'), routes the lane change to the correct sibling, tells where opportunity ids come from, and explains the needs_confirmation handshake including the instruction to pause and confirm. Multiple when/when-not conditions, nothing left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_page_links_listList a page's connectionsARead-onlyInspect
All typed edges touching one page, both directions, each hydrated with the other endpoint's page summary. Use this to understand a subject's neighborhood before adding to it.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The page id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe read-only profile, so the description's real contribution is disclosing the result shape: bidirectional edges each hydrated with the other endpoint's page summary. That return-format detail is genuinely useful and absent from annotations, though nothing is said about volume limits or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the scope of what is returned front-loaded before the usage hint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with fully documented schema and no output schema, the description covers what it returns and when to reach for it. The only gap is any mention of result size or pagination for a page with many edges.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (pageId, workspace, including the workspace default/override semantics) are documented there. The description adds no syntax or format meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) plus the exact resource and scope: 'all typed edges touching one page, both directions.' Clearly separates it from the write siblings (compass_links_create/update/delete) without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this to understand a subject's neighborhood before adding to it' supplies a clear context and implicitly routes the agent to the link-mutation siblings for the 'adding' step, but it names no alternative tool explicitly and gives no when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_createCreate a Compass pageAInspect
Creates a typed page in the workspace's Compass compendium. Check compass_pages_list first — don't create a page for a subject the map already has; link to it instead. May return needs_confirmation — if so, tell the user what you're proposing, wait for their approval, then re-call with the approvalId. Only record what the human has actually told you.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Markdown body shaped to the type (steps for WORKFLOW, role/contact sections for PERSON, etc.). | |
| tags | No | Optional tags. | |
| type | Yes | Page type for this subject. VALUE = something the business won't sacrifice; PRIORITY = an outcome it's pushing toward; AUDIENCE = who the work is for; OFFERING = what the business delivers. | |
| title | Yes | ≤80 chars, sentence case, no filler verbs. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare it is a non-destructive write. The description adds real behavioral context beyond them: a possible `needs_confirmation` return, the requirement to surface the proposal to the user and re-call with an `approvalId`, and a grounding constraint ('Only record what the human has actually told you').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences, front-loaded with purpose then routing, confirmation handling, and a grounding rule. No filler and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with no output schema, the description covers the creation semantics, the dedup prerequisite, the confirmation handshake, and the data-grounding constraint. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters including type values, workspace rules, visibility enum, and approvalId. The description reinforces the approvalId round-trip but adds little parameter detail beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a typed page in the workspace's Compass compendium') and distinguishes itself from siblings by naming `compass_pages_list` and the alternative action (link instead of create). An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-not guidance ('don't create a page for a subject the map already has; link to it instead') plus a named prerequisite (`compass_pages_list`). It also spells out the confirmation flow and what to do when `needs_confirmation` is returned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_deleteDelete a Compass pageADestructiveInspect
Moves a page to the trash (soft delete — its connections and anchored opportunities go with it, and compass_pages_restore brings the whole set back). Use this to REMOVE a page you or an extraction created wrongly, or one the user says no longer belongs on the map; to fix a wrong title or body, use compass_pages_update instead. Page ids come from compass_pages_list / compass_pages_search. May return needs_confirmation — name the page and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Id of the page to delete. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true but say nothing about scope or recovery; the description fills both gaps by stating the cascade (connections and anchored opportunities go with it) and the reversal path via compass_pages_restore. It also discloses the interactive approval behavior (may return `needs_confirmation`, name the page and wait), which is non-obvious runtime behavior not captured anywhere in structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the operation and its soft-delete semantics before the usage routing. Every sentence does distinct work: mechanism, cascade, reversal, use-case, counter-case, id provenance, confirmation flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, confirmation-gated tool with no output schema, the description covers mutation scope, reversibility, the required interaction loop, and prerequisite id lookup. An agent has everything needed to call it correctly, including knowing it may have to pause for user approval.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the schema already carries per-parameter text. The description still adds meaning the schema does not: where pageId values originate, and the approval protocol that gives approvalId its purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Moves a page to the trash') and immediately qualifies the mechanism as a soft delete, so the agent knows this is recoverable rather than permanent. It explicitly distinguishes itself from compass_pages_update and points at compass_pages_restore, so no sibling ambiguity remains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('a page you or an extraction created wrongly, or one the user says no longer belongs') and an explicit when-not ('to fix a wrong title or body, use compass_pages_update instead'). It also routes the agent to the id source (compass_pages_list / compass_pages_search) before the call is made.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_getGet a Compass pageARead-onlyInspect
Fetches one Compass page by id, including its full Markdown body, header fields, and any attached Napkin sketches — view an attached sketch's actual drawing with napkin_boards_view (its boardId) before discussing it.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The page id (from compass_pages_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds value beyond that by disclosing the return contents (Markdown body, header fields, sketches) and a downstream workflow dependency, though it says nothing about pagination, size limits, or missing-page behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence that names the resource and then the payload, with the napkin hint appended as a useful aside. Slightly dense but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by listing the returned contents, and annotations cover the safety profile. Remaining gaps (error/not-found behavior, pagination of attached sketches) are minor for a single-resource fetch.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema itself documents both pageId (with its source, compass_pages_list) and the workspace slug rules. The description only restates 'by id', so it adds essentially nothing over the structured fields — baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetches one Compass page by id') and enumerates the payload ('full Markdown body, header fields, and any attached Napkin sketches'), clearly distinguishing it from the sibling list/search/update tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit cross-tool routing: view an attached sketch's actual drawing with napkin_boards_view (its boardId) before discussing it. However, it never states when to prefer this over compass_pages_list or compass_pages_search, so guidance is contextual rather than complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_listList Compass pagesARead-onlyInspect
Lists pages in the active workspace's Compass compendium. Optional type filter narrows to one node type (WORKFLOW, PERSON, SYSTEM, PAIN_POINT, DOCUMENT, VALUE, PRIORITY, AUDIENCE, OFFERING). Returns summaries, 50 at a time (limit / offset, total and nextOffset in the result) — fetch one with compass_pages_get for the full body. When you know what you're looking for, compass_pages_search is the better first call.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Narrow to a specific page type (WORKFLOW, PERSON, SYSTEM, PAIN_POINT, DOCUMENT, VALUE, PRIORITY, AUDIENCE, OFFERING). | |
| limit | No | Page size, default 50. Prefer compass_pages_search when you know what you're looking for. | |
| offset | No | Skip this many, for the next page. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: page size of 50, `limit`/`offset` paging, and the presence of `total` and `nextOffset` in the result. It stops short of describing auth or rate-limit behavior, which the schema carries for `workspace`.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what it lists and the workspace scope, then layers filtering, pagination, and sibling routing in three tight sentences. Nothing is padding, and the most decision-relevant routing advice is easy to find.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description discloses the return shape (summaries, total, nextOffset) and points to `compass_pages_get` for full bodies, which is everything an agent needs to page and follow up correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, establishing a baseline of 3. The description reinforces the `type` enum and pagination parameters but adds no syntax or format meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (lists pages in the active workspace's Compass compendium) plus scope. It explicitly routes to sibling tools (`compass_pages_get` for full bodies, `compass_pages_search` as a better first call), so an agent can differentiate it without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit alternatives with conditions: use `compass_pages_get` when you have a page and want the full body, and use `compass_pages_search` when you already know what you're looking for. This is concrete when-to-use guidance, not inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_restoreRestore a deleted Compass pageAInspect
Puts a trashed page back on the map, together with exactly the connections and opportunities its delete took. Use when the user wants a deleted page back (the pageId from the earlier compass_pages_delete, or from the Compass trash). Fails with not_found when the page is already live or never existed. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Id of the deleted page. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false); the description goes well beyond them by disclosing that the restore is not just the page but also its connections and opportunities, that a not_found failure mode exists for already-live/nonexistent pages, and that a needs_confirmation response is possible. That last point is critical behavioral context an agent must plan for.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and scope before the trigger, parameter sourcing, and failure modes. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and covers the failure (not_found) and confirmation (needs_confirmation) paths, which is the important part. It does not describe the success payload beyond the restored-page implication, a minor gap for a restore operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description exceeds it by explaining the workflow semantics behind approvalId ('from a prior needs_confirmation response, after the user has approved') and pageId's origin, tying the parameters into the confirmation loop rather than just restating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Puts a trashed page back on the map') and adds scope detail that distinguishes it from compass_pages_delete: it restores the page *plus* the connections and opportunities the delete removed. An agent can tell this apart from compass_pages_get, compass_pages_list, or compass_pages_delete without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger ('Use when the user wants a deleted page back') and it even tells the agent where the pageId comes from (a prior compass_pages_delete, or the Compass trash). The failure condition ('not_found when the page is already live or never existed') effectively defines the when-not-to-call case, though no alternative tool is named for that situation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_searchSearch Compass pagesARead-onlyInspect
Keyword search over page titles, bodies, and tags (case-insensitive substring match), paginated. The fast way to check whether a subject already has a page before creating or linking. Returns summaries with a 300-char body snippet — fetch full text with compass_pages_get.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Search term. Use the most distinctive word or phrase for the subject ('Zendesk', 'invoice approval') — not full sentences. | |
| type | No | Narrow to one page type. | |
| limit | No | Results per page, 1-25. Default 10. | |
| offset | No | Pagination offset. Default 0. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/non-destructive/openWorld=false, so the safety profile is covered; the description adds real value by disclosing match semantics (case-insensitive substring), that results are paginated summaries with a 300-char body snippet, and the handoff to compass_pages_get. It does not state result ordering or how snippet truncation affects matching, which is the only remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, all load-bearing: match semantics first, then the routing use case, then the return shape and follow-up tool. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by specifying the return shape (summaries with 300-char snippets), the pagination contract, and the fetch-full-text alternative. For a 5-parameter read tool with full schema coverage and clean annotations, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the useful 'use the most distinctive word, not full sentences' guidance for q, the limit range, and the workspace-auth caveat, so the schema does the heavy lifting. The description only restates the keyword/pagination concept, adding nothing beyond it; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (keyword search) and resource (Compass pages), plus the exact match scope: titles, bodies, tags, case-insensitive substring. An agent can distinguish it from compass_pages_list (unfiltered listing) and compass_pages_get (full text retrieval) without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context — 'the fast way to check whether a subject already has a page before creating or linking' — and names compass_pages_get as the follow-up for full text. It does not explicitly say when to prefer it over compass_pages_list or search_docs/search_workspace, so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_pages_updateUpdate a Compass pageADestructiveInspect
Edits an existing page's title, description, body, tags, header fields, or visibility — use this to FIX what you (or an extraction) got wrong instead of creating a duplicate. Fetch the current page with compass_pages_get first and preserve what the human wrote; title / body / tags replace the field wholesale, while headerFields MERGES over the current ones (send only the keys you're changing — 'set the owner' leaves status alone). Page ids come from compass_pages_list. May return needs_confirmation — tell the user what you're changing, wait for approval, then re-call with the approvalId.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | New Markdown body — replaces the whole body. | |
| tags | No | New tag set — replaces existing tags. | |
| title | No | New title (≤200 chars). | |
| pageId | Yes | Id of the page to edit. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response, after the user has approved. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New one-line gallery summary (≤2,000 chars). | |
| headerFields | No | Partial header fields for the page's type (PERSON: role, email, …; WORKFLOW: owner, status, …) — merged over the current values; the server validates the merged result against the type. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true / readOnly=false, and the description meaningfully adds to that: it explains the replace-wholesale vs headerFields-merge semantics, the need to preserve human-authored content, and the needs_confirmation → approvalId round-trip. Some of the replace language overlaps the schema, but the merge/confirmation behavior is extra context the annotations don't carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and target, then packed with routing, prerequisite, and merge/confirmation guidance in which every sentence earns its place — no filler or restated schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter destructive edit with a nested object and no output schema, the description covers the risky behaviors an agent needs (field replacement vs merge, approval gating, id provenance, preservation of human content), leaving nothing critical to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description goes further by contrasting replace semantics (title/body/tags) against merge semantics for headerFields with a concrete example ('set the owner' leaves status alone) and by explaining where pageId values originate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Edits) and resource (an existing page), enumerates the editable fields, and explicitly distinguishes itself from the create path ('instead of creating a duplicate'). An agent can separate it from compass_pages_create and compass_pages_get without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (fixing mistakes/extra value from extractions rather than duplicating), prerequisites (fetch with compass_pages_get first, page ids from compass_pages_list), and the confirmation workflow to follow when needs_confirmation is returned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compass_statuses_listList the workspace's change statusesARead-onlyInspect
The statuses this workspace's changes move through, in board order — the names people use. Each sits in a group: Undecided (TRIAGE), Not started (BACKLOG), In progress (ACTIVE), Done (DONE), Dropped (CANCELED). Groups are stable across workspaces; names aren't, so use the names when talking to people and the groups when reasoning about progress. Move a change with compass_opportunities_set_status.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuine domain context beyond structured data: the five group taxonomy, the fact that groups are stable while names are not, and the ordering guarantee ('in board order'). It does not discuss return format or pagination, which is acceptable for a small list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the identity of the resource, then layers in the taxonomy and the stable-vs-unstable distinction. Four sentences, all informative, though the group enumeration and cross-reference sentence make it slightly denser than strictly necessary for a one-parameter list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of conveying what comes back; it does this well by explaining names, groups, and ordering. Lacking only explicit return-shape/pagination detail, which is minor for a simple status lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single workspace parameter is fully documented in the schema (default-workspace and API-key behavior). The description adds nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — listing the workspace's change statuses — and uniquely characterizes them as the values changes 'move through, in board order.' It also names the sibling that mutates them (compass_opportunities_set_status), so an agent can place it precisely among the many compass_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives practical guidance for how to consume the result (use names with people, groups for reasoning) and points to compass_opportunities_set_status for the corresponding mutation. It stops short of an explicit 'call this when you need X' trigger or any when-not, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_browseBrowse the workspace's tagsARead-onlyInspect
Without a tag: every tag in use across the workspace with how many entities carry it, most-used first — the vocabulary the team already organizes by. With a tag: everything filed under it across every tool, each with its kind, id, title, and path. Use it to reuse existing labels instead of inventing near-duplicates, and to answer 'show me everything about X' when X is a label.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | A tag to expand into its items. Omit to list tags. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive/openWorld, so safety is covered, and the description adds real behavioral detail: result ordering (most-used first), entity counts per tag, and the per-item fields returned in tag mode (kind, id, title, path). No pagination or size-limit behavior is mentioned, keeping this from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and cleanly parallel: 'Without a tag:' then 'With a tag:' then the usage clause. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return-value burden for both modes, and annotations cover the safety profile while the schema covers both parameters. Nothing an agent needs to select or call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by characterizing what each mode of the tag parameter actually returns, turning a bare 'omit to list tags' into the two distinct result shapes. The workspace parameter semantics remain entirely schema-borne.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific dual operation (list every tag in the workspace, or expand one tag into all entities filed under it) with the exact resource and scope. It is immediately distinguishable from the write-oriented siblings entity_tags_get and entity_tags_set because the browse mode and cross-tool aggregation are spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use guidance: reuse existing labels rather than inventing near-duplicates, and answer 'show me everything about X' when X is a label. It stops short of naming sibling alternatives or stating when-not-to-use (e.g. versus search_workspace), so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_getRead the tags on entitiesARead-onlyInspect
Returns the tags on a batch of entities of one kind — the labels galleries organize by. Ids come from the kind's list/get tool or from search_workspace. Use it before entity_tags_set so you replace the full set knowingly, and to answer 'what is this filed under'. Entities the user can't see are omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| entityKind | Yes | Which kind the ids belong to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds the non-obvious behavioral fact that 'Entities the user can't see are omitted,' i.e. results are permission-filtered rather than erroring. It does not mention limits (e.g. the 100-id cap or ordering), but the value-add beyond annotations is genuine.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all front-loaded: what it returns first, then usage, then the permission caveat. No filler and each sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does state what comes back (tags) and the omission behavior. It leaves minor gaps — tag value format, ordering, and whether unknown ids are silently dropped or error — but nothing that blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema documents most params, but the description adds real provenance for the hardest parameter: ids 'come from the kind's list/get tool or from search_workspace.' That tells an agent where to obtain valid ids, which the bare array schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Returns the tags on a batch of entities of one kind.' The parenthetical 'the labels galleries organize by' distinguishes tags from other entity metadata, and the named siblings (entity_tags_set, search_workspace) make the boundary clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit usage contexts: 'Use it before entity_tags_set so you replace the full set knowingly, and to answer what is this filed under.' That is a real when-to-use plus a stated alternative. It does not mention entity_tags_browse, the other obvious read-side sibling, so it falls short of fully disambiguating the read alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_tags_setSet an entity's tagsADestructiveInspect
Replaces the FULL tag set on one entity (an empty list clears it). Read the current tags with entity_tags_get first and pass the merged list — this is not additive. Tags are lowercase letters, numbers, spaces, and hyphens; prefer labels already in use (entity_tags_browse) so the workspace's vocabulary stays small. The id comes from the kind's list/get tool or search_workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | ||
| entityId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| entityKind | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, and the description reinforces this with the replace-not-merge semantics and the empty-list-clears behavior. It does not mention the needs_confirmation/approvalId retry flow that the schema implies, which is the one notable behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, all front-loaded with the destructive replace semantics first, followed by workflow and vocabulary guidance. Every sentence earns its place; the parentheticals make it slightly busy but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description covers destructive semantics, prerequisites, id sourcing, and tag format well enough to invoke the tool correctly. The confirmation/approval retry path is left to the schema field description rather than the tool description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, and the description compensates by documenting the tag character set (lowercase letters, numbers, spaces, hyphens), the merge requirement for the tags array, and the origin of entityId. The workspace and approvalId parameters are only explained in the schema, so coverage is good but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Replaces the FULL tag set on one entity') and immediately clarifies the destructive scope with '(an empty list clears it)'. This cleanly separates it from the sibling readers entity_tags_get and entity_tags_browse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit workflow: read current tags with entity_tags_get first, pass the merged list, and warns 'this is not additive'. Also routes to entity_tags_browse for vocabulary reuse and names where the entityId comes from (kind's list/get tool or search_workspace), so the agent knows both prerequisites and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_docFetch ZeroWidth doc by slugARead-onlyInspect
Fetch the full Markdown body of a specific docs page by its slug. Use this after search_docs when the user needs the complete content of a page. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Page slug. Accepts 'compass/api', '/compass/api', or 'docs/compass/api'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, covering the safety profile. The description adds valuable context by stating 'No authentication required' and specifying the return body as full Markdown, which agents need to know beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose ('Fetch the full Markdown body...'), followed by usage routing and a salient behavioral note. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with full schema coverage and annotations covering safety, the description is complete: it states the return format, the required input type, usage context, and authentication status. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single `slug` parameter is fully documented in the schema with accepted formats. The description adds no syntax or format details beyond what the schema provides, so it meets the baseline rather than exceeding it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch'), resource ('full Markdown body of a specific docs page'), and retrieval key ('by its slug'). It distinguishes itself from the search-oriented sibling by positioning as the step after `search_docs` for complete page content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this after `search_docs` when the user needs the complete content of a page, which names the alternative and the condition that selects it. It does not state when not to use it (e.g., for listing or metadata), but the context is clear enough for a simple read tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_postFetch a Distributed Cognition essayARead-onlyInspect
Fetch the full Markdown body of one Distributed Cognition essay by slug (from list_posts). Link the public URL when citing it to the user. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Essay slug, e.g. from `list_posts`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavior context beyond that: the returned artifact is the full Markdown body, and no authentication is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and resource, with zero filler. Every sentence carries actionable information (what is returned, where the slug comes from, how to cite it, auth requirement).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description covers the essentials: return format (Markdown body), key source, auth requirement, and citation behavior. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 1 parameter and 100% schema description coverage, the baseline is 3. The description's note that the slug comes 'from `list_posts`' duplicates the schema's own description ('Essay slug, e.g. from `list_posts`'), so it adds no meaning beyond the structured field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Fetch') plus precise resource ('full Markdown body of one Distributed Cognition essay by slug'), including the retrieval key and format. It clearly distinguishes itself from the sibling `list_posts`, which is identified as the source of slugs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes the workflow prerequisite by pointing to `list_posts` for slugs and gives an output-handling instruction (link the public URL when citing). It does not state explicit when-not-to-use conditions, but the context is clear for a single-purpose fetch tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_add_evidenceAttach evidence to a Ledger entryAInspect
Records what happened against an open decision — a manual observation the user reports ('the pilot team says triage feels faster'), an implementation note, or a reference to a Caliper run. Evidence is what settlement later reads, so attach it as it arrives and cite it in the lesson. Metric readings attach themselves via ledger_metrics_record_reading — don't duplicate them here. Fails with conflict on superseded entries. Ids come from ledger_entries_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | Structured detail worth keeping with the summary. | |
| kind | No | Usually `manual`; `caliper_eval_run` with refId for a run; `implementation` when the change shipped. | manual |
| refId | No | Source record id (a Caliper run id, …) when there is one. | |
| entryId | Yes | The entry the evidence is about. | |
| summary | Yes | What happened, in one or two sentences. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-destructive, non-open-world write, and the description adds real behavioral context beyond them: failure with conflict on superseded entries, a possible `needs_confirmation` envelope, and the overall importance of evidence for later settlement. It doesn't describe the success return shape, but with an envelope caveat and failure mode disclosed this is stronger than average.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and follows with routing, failure, and id-sourcing facts. Some narrative (the quoted user anecdote, 'cite it in the lesson') is looser than needed, but every sentence carries usable signal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, it covers routing, prerequisites, duplicate-avoidance, failure modes, and the confirmation envelope, which are the pieces an agent needs. It omits nothing critical, though it could state whether the appended entry becomes immutable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters including the `kind` enum and `workspace` token rules. The description adds usage flavor (evidence kinds, refId for a Caliper run) but no syntax the schema lacks, so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (records/attaches evidence against an open decision) and enumerates concrete kinds of evidence. It explicitly distinguishes itself from the sibling ledger_metrics_record_reading, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use ('attach it as it arrives and cite it in the lesson'), a when-not ('Metric readings attach themselves via ledger_metrics_record_reading — don't duplicate them here'), and prerequisite sourcing ('Ids come from ledger_entries_list'). Alternatives and conditions are named explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_createRecord a decision, lesson, or observation in the LedgerAInspect
Records a memory entry. Every entry is a TITLE (summary: one short plain sentence) over a BODY (rationale: the detail, markdown welcome) — never put the detail in the title. Three kinds: decision — a change being made now; PRE-REGISTRATION IS THE POINT, so prediction (what we expect) must be written NOW, before any evidence exists, and never backfilled to match an outcome. lesson — a distilled belief the team already holds (lesson text required, no prediction). observation — a durable fact worth remembering (no prediction): something already true, never something planned. Ideas, pitches, backlog items, and upcoming work are NOT entries — a dated piece of work that carries out a decision is a plan item (ledger_plan_add on that decision), and a running list you keep across runs belongs in a Napkin doc or sheet. When a user states something durable about their business in conversation, offer to capture it as a lesson or observation. When a user shares MEETING NOTES, propose the decisions you find with ledger_entries_draft so they keep or drop each one in Ledger; use this tool for an entry they've asked you to record. Link the Compass pages the entry touches so future consult-before-acting finds it. May return needs_confirmation — summarize the entry and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | decision = a change with a pre-registered expectation; lesson = a belief arriving already settled; observation = a fact that is already true (not a plan, idea, or pitch). | decision |
| lesson | No | What we learned — required for kind=lesson, optional for kind=observation, forbidden on decisions (their lesson is written at settlement). | |
| summary | Yes | The entry's TITLE: one short plain sentence naming what we did (decision) or the fact itself (lesson, observation). Plain text, no markdown, at most 280 characters — e.g. "Newsletter moves to a biweekly cadence". Everything longer goes in rationale. | |
| rationale | No | The entry's BODY: the detail under the title — why, context, lists, steps. Markdown is fine here and renders as formatted text. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| prediction | No | Decisions ONLY. What we expect: a list of `claims` plus one shared `deadline`, or `freeform` for room decisions nothing can settle. A claim either REACHES a value (comparator + target, e.g. CTA clicks >= 400) or HOLDS one (`hold`, e.g. newsletter reads no more than 5% below the 28 days before this). Name every number the change is expected to move AND every number it shouldn't cost — a decision that claims only what it hopes will rise gets to pick its own evidence. Set `watch: true` on a number worth following that the decision isn't committing to. When the change aims at part of an event metric ("new users in Germany"), narrow the claim with `slice` ({ country: ["DE"] }) using label values from ledger_metrics_get — a claim on an event metric settles on the total counted since landing (reach) or events per day (hold). | |
| compassPageIds | No | Where — Compass page ids this decision touches. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=false, so the write/safety profile is largely covered. The description adds genuinely non-annotation behavior: it may return `needs_confirmation`, in which case the agent should summarize and wait for approval, plus the pre-registration invariant (prediction written before evidence, never backfilled) that constrains correct use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the title/body model and the three kinds, and each subsequent sentence carries a rule or a routing decision rather than filler. It is dense and uses heavy caps for emphasis, which is slightly noisy, but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with nested prediction objects and no output schema, the description covers kinds, the pre-registration constraint, the approval flow, and the Compass linking purpose. Return-shape detail is the only meaningful omission, and the needs_confirmation envelope is at least flagged.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds framing the schema does not: the title/body split (summary as TITLE, rationale as BODY, never detail in the title) and the rule that a decision's lesson is written at settlement. The prediction block is well covered by its own schema description, so the added value is modest rather than transformative.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (records a memory entry in the Ledger) and immediately breaks the resource into its three kinds (decision/lesson/observation) with a one-line definition for each. It further distinguishes itself from siblings by name (ledger_plan_add, ledger_entries_draft), so an agent can route without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use routing: use ledger_entries_draft for meeting notes so the user keeps/drops each proposal, use ledger_plan_add for dated work that carries out a decision, use this tool for an entry the user asked to record, and put running lists in a Napkin doc. It also names what is NOT an entry (ideas, pitches, backlog items). This is about as complete as usage guidance gets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_draftPropose Ledger entries for reviewAInspect
Writes up to 50 entries as DRAFTS: proposals that stay out of the record until a person confirms each one in Ledger's Drafts view. Use it for entries you found rather than were told — decisions in meeting notes, or a team's past changes read from its tracker, pull requests or launch posts (set fromHistory: true). For history: take each claim from what the source said AT THE TIME, set landedAt to when it shipped, attach the source URL, and leave rollout as full unless the source says otherwise. These read as low confidence because they were written down after the fact; say so plainly rather than overstating them. Tell the person how many drafts are waiting and that they review them under Decisions → Drafts. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| drafts | Yes | ||
| workspace | No | Workspace slug. Ignored for workspace API keys. | |
| approvalId | No | ||
| fromHistory | No | True when the drafts come from a team's past work, not the conversation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations state readOnlyHint=false, destructiveHint=false, openWorldHint=false, but the description adds the crucial behavioral fact that these are non-committing proposals reviewed in Decisions → Drafts, that confidence is low for historical claims, and that the call may return `needs_confirmation`. Those are traits the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and safety are front-loaded in the first sentence, and the remaining sentences are operational instructions rather than filler. It is dense and fairly long, with some stylistic guidance ('say so plainly rather than overstating them') that is useful but could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write-of-drafts tool with no output schema, no nested-object flag, and 50% schema coverage, the description supplies the workflow, the history-specific field rules, the confidence caveat, the user-facing next step, and the possible `needs_confirmation` return. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 50% schema coverage, the description compensates well: it explains fromHistory semantics, prescribes landedAt ('when it shipped'), rollout defaults ('full unless the source says otherwise'), source URL attachment, and that prediction is decisions-only. It doesn't touch workspace or approvalId, but the high-value parameters are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb, resource, and scope: 'Writes up to 50 entries as DRAFTS: proposals that stay out of the record until a person confirms each one in Ledger's Drafts view.' The 'draft' framing distinguishes it cleanly from ledger_entries_create, which commits records, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use it ('entries you found rather than were told') and gives concrete trigger examples — decisions in meeting notes, a team's past changes from its tracker, PRs, or launch posts — plus the fromHistory condition. This is explicit when-to-use guidance with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_getRead one Ledger entry in fullARead-onlyInspect
One entry's full anatomy — the six fields, attached evidence (Caliper runs, measurements, observations), and the supersede chain (what replaced it, or what it replaced). Cite entry ids when telling the user about prior related decisions.
| Name | Required | Description | Default |
|---|---|---|---|
| entryId | Yes | Entry id (from ledger_entries_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds real value beyond that by enumerating the returned content (six fields, Caliper runs, measurements, observations, supersede chain), which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no filler. The most useful content (what the entry contains) is front-loaded, and the second sentence is a targeted usage tip rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description responsibly summarizes the return payload, and annotations cover the safety profile while the schema fully covers parameters. Only the relationship to sibling tools (list, update, retract) is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both entryId and workspace are fully documented in the schema, including the personal-token workspace requirement. The description adds no parameter syntax or meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Read one Ledger entry in full') and then details exactly what the entry contains: the six fields, attached evidence, and the supersede chain. An agent can tell it apart from ledger_entries_list, though the description never explicitly contrasts with that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied. The closing sentence ('Cite entry ids when telling the user about prior related decisions') advises how to use the output, not when to select this tool over ledger_entries_list or ledger_entries_update. No exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_listList Ledger decision entriesARead-onlyInspect
The workspace's memory: decisions (changes with pre-registered expectations), lessons (distilled beliefs), and observations (captured facts). CONSULT BEFORE ACTING — before proposing a flow change, prompt edit, or process decision, filter by the Compass page it touches (compassPageId) and check whether prior attempts exist and how they settled. Filter status=open for unsettled expectations awaiting evidence (each carries a derived lapsed flag — true when its deadline has passed; surface lapsed ones when the user asks what needs attention); kind=lesson for what the team already believes.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Free-text search across an entry's summary, rationale, and lesson (case-insensitive). Use it to find the entry about a topic before acting. | |
| kind | No | Filter by species: decision (pre-registered expectations), lesson (distilled beliefs), observation (captured facts). | |
| cursor | No | Pagination cursor from a prior page. | |
| origin | No | Filter by who wrote it (manual, workbench, …). | |
| status | No | Filter by lifecycle state (open = awaiting evidence). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| compassPageId | No | Compass page id — every decision touching that workflow / system / person. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds genuine behavioral context beyond them: that status=open entries carry a derived `lapsed` flag and that lapsed entries should be surfaced when the user asks what needs attention — meaningful with no output schema to document it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the ledger's purpose before the usage directives, and each clause maps to an actionable filter. The semicolon-heavy second half packs three filter recipes into one dense sentence, which slightly reduces scannability but doesn't waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param, zero-required list tool with annotations covering safety and no output schema, the description supplies the missing behavioral detail (the derived lapsed flag) and the consult-before-acting workflow. Return shape and pagination semantics remain undocumented, which keeps it short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema coverage at 100%, the baseline is 3, but the description goes further by explaining how filters compose (use compassPageId to find prior attempts and how they settled; status=open means awaiting evidence) rather than restating field descriptions. It clarifies the intent behind combining parameters, which the schema alone does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (the ledger: decisions, lessons, observations) and the operation by naming it a filtered list, but the listing verb is more implied than stated. It is distinguishable from siblings like ledger_entries_get or ledger_entries_settle through the filtering framing, though it never explicitly contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('CONSULT BEFORE ACTING — before proposing a flow change, prompt edit, or process decision') and concrete selection recipes for parameters (compassPageId to check prior attempts, status=open for unsettled expectations, kind=lesson for team beliefs). This is a strong routing instruction rather than a vague hint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_retractRetract a Ledger entryADestructiveInspect
Takes an entry out of the curated ledger — THE remedy when you recorded something wrong (a decision that wasn't made, a duplicate, a fact the user corrects). Reversible from Ledger and fully audited, so it is safe to offer as soon as the user says 'that's not right'. Not for overturning a settled claim the team once believed — a human supersedes that. Fails with conflict if already retracted. Ids come from ledger_entries_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| entryId | Yes | The entry to retract. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true/readOnly=false, but the description adds substantial context beyond them: reversibility from Ledger, full audit trail, the conflict failure mode when already retracted, and a possible `needs_confirmation` envelope. Reversibility does not contradict destructiveHint — it refines it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and largely earn-its-place, with the core remedy statement first and exclusions after. A couple of clauses (the quoted 'that's not right' reassurance, the 'THE remedy' emphasis) are slightly redundant with earlier sentences, adding density without new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers what an agent needs: the confirmation envelope it may receive, the conflict behavior, and the safety/reversibility profile. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value by stating where entryId comes from ('Ids come from ledger_entries_list') and by tying `needs_confirmation` to the approvalId flow. It stops short of describing the workspace parameter's semantics, which the schema already covers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Takes an entry out of the curated ledger') and immediately frames the purpose in operationally meaningful terms (correcting a mistaken record). It distinguishes itself from siblings like ledger_entries_settle and ledger_entries_update by scoping to retroactive removal of a wrongly recorded entry.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('THE remedy when you recorded something wrong — a decision that wasn't made, a duplicate, a fact the user corrects') with concrete triggers, plus an explicit when-NOT-to-use ('Not for overturning a settled claim the team once believed — a human supersedes that'). The alternative and the condition selecting it are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_settleSettle an open Ledger decisionADestructiveInspect
The ritual moment: evidence has landed, the expectation closes, the lesson is written. Only propose settlement when attached evidence actually answers the prediction — check ledger_entries_get first and cite the evidence in the lesson. lesson (what we now believe) is required and permanent; settlement happens exactly once. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| lesson | Yes | What we now believe — the distilled, citable lesson. | |
| entryId | Yes | The open entry to settle. | |
| verdict | No | Did the pre-registered expectation hold? `confirmed` every claim came out as hoped, `missed` none did, `mixed` some did and some didn't. Mixed is not a softer miss — "it worked and it cost us something" is usually the most informative result a change can produce, so use it rather than rounding to either side. Only when the entry carries an expectation and the evidence gives a clear answer. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false; the description meaningfully extends this by stating the lesson is 'permanent' and that 'settlement happens exactly once', which makes the irreversibility concrete. It also discloses a `needs_confirmation` response path despite there being no output schema — useful granularity beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loaded, but the opening clause ('The ritual moment: evidence has landed, the expectation closes, the lesson is written') is atmospheric framing that consumes prime space without adding operational content. Trimming it would leave a purely actionable description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no output schema, the description covers the important ground: prerequisites, irreversibility, one-shot settlement, and the confirmation path. It omits guidance on workspace/approvalId handling, which the schema partially covers, but nothing critical to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema already documents lesson, entryId, verdict and workspace. The description adds only that `lesson` is permanent and should cite evidence; verdict, workspace and approvalId semantics are left entirely to the schema. Baseline 3 is appropriate when the schema does most of the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys the operation — closing an entry once evidence lands and writing a lesson — and distinguishes it from siblings like ledger_entries_add_evidence by requiring that evidence 'actually answers the prediction'. The opening metaphor ('the ritual moment') is evocative rather than literal, so the exact verb+resource must be inferred, but it is recoverable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition for use ('Only propose settlement when attached evidence actually answers the prediction') and routes the agent to `ledger_entries_get` to verify first. It stops short of naming the correct alternative when the precondition fails (e.g., add_evidence), so the exclusion is implied rather than fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_entries_updateEdit a Ledger entryADestructiveInspect
Corrects an entry in place — its title (summary), body (rationale), the pre-registered prediction of an OPEN decision (null clears it), or its facet (filing category; editable even after settlement). Works on open decisions, and on lessons, observations, and beliefs (they carry no expectation, so fixing their wording is fine any time). Use to fix a typo, swapped fields, or a misrecorded detail. Never edits kind or lesson. Fails with conflict on a settled decision — its claim is corrected by a human superseding it in Ledger — and on superseded or retracted entries. Entry ids come from ledger_entries_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| facet | No | Filing facet; empty string = unassigned. | |
| entryId | Yes | The entry to edit: an open decision, or a lesson / observation / belief. | |
| summary | No | New title: one short plain sentence, no markdown, at most 280 characters. | |
| rationale | No | New body: the detail under the title; markdown is fine. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| prediction | No | Replacement expectation (claims + deadline, or freeform), or null to clear it. Correcting a misrecorded expectation is fine; rewriting one to match what happened is not. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructive/not-read-only; the description adds substantial context beyond them: the conflict failure mode on settled decisions, the null-clears-prediction semantics, that facet stays editable post-settlement, that kind/lesson are immutable, and that a needs_confirmation envelope may be returned. This is the behavioral detail an agent needs for a destructive in-place mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what is corrected, then constraints, then failure modes, then the id source — a logical order of decreasing immediate relevance. Every clause carries a distinct rule (editable set, invariants, failure conditions, confirmation flow); no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param destructive mutation with a nested prediction object and no output schema, the description covers the editable surface, immutability rules, failure/conflict behavior, id provenance, and the confirmation envelope. An agent has everything needed to invoke it correctly and to anticipate refusals.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents each field. The description still adds meaning beyond it by mapping summary→title, rationale→body, explaining null clears the prediction, and noting facet is editable even after settlement. Modest added value over an already-complete schema, hence 4 rather than 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Corrects') plus resource ('an entry') and enumerates exactly which fields are editable (summary, rationale, prediction, facet). It also draws hard boundaries against siblings by declaring what it never edits (kind, lesson) and what it will not touch (settled, superseded, retracted entries). An agent can distinguish it from ledger_entries_settle/retract/create without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('fix a typo, swapped fields, or a misrecorded detail') and explicit when-not (settled decisions, superseded/retracted entries), with the reason for the refusal. It also routes the agent to ledger_entries_list for ids, removing the guesswork about where entryId comes from.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feeds_backfillFill in a metric's history from its feedAInspect
Run the same frozen instruction over a past range, so a metric has history instead of starting the day it was set up. Offer this whenever you create a feed — a chart with a year behind it is worth far more than one that begins today, and an expectation can be judged against what normal looked like. Works when the source returns a series; a source that only ever reports 'right now' will write one point. Safe to repeat: overlapping ranges dedupe. Keep ranges to a few hundred points; a run that fills up says so. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Feed id, from ledger_feeds_list. | |
| to | No | End of the range. Defaults to now. | |
| from | Yes | Start of the range, ISO date or instant. | |
| workspace | No | Workspace slug. Omit to use the pinned workspace. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare mutation-safety (readOnly=false, destructive=false, openWorld=true). The description adds genuinely useful behavioral detail beyond them: idempotency via dedupe on overlapping ranges, size-limit feedback ('a run that fills up says so'), the single-point outcome for non-series sources, and a `needs_confirmation` return that ties to the approvalId param. This is well above what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then layers usage, constraints, and re-run safety. Mostly efficient, though the 'worth far more than one that begins today' pitch is mildly promotional filler that could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by disclosing the re-run/dedupe behavior, the range-size limitation, and the `needs_confirmation` outcome. Combined with the annotations, an agent has enough to call and interpret it, though nothing confirms what a successful backfill returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema already documents id/from/to/workspace. The description implies the from-range concept and hints that approval/confirmation is involved (linked to approvalId), but adds no format or semantics beyond the schema. Baseline 3 is appropriate at this coverage level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action: run the frozen feed instruction over a past range to populate history. The verb+resource ('run' + 'frozen instruction'/'range') is concrete. It does not name the closely related sibling ledger_feeds_run, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition for invoking: 'Offer this whenever you create a feed,' with a rationale (a chart with history is more valuable, expectations can be judged). It does not explicitly distinguish when to use this versus ledger_feeds_run or ledger_feeds_preview, so no exclusions are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feeds_createSet up how a metric's numbers arriveAInspect
Freeze a tool call as a metric's standing source: it runs on the cadence you give and records what it finds, with no model involved. metric takes an id or a snake_case slug; an unknown slug starts tracking that metric. Confirm the mapping with ledger_feeds_preview first. Tell the user they can backfill history afterwards with ledger_feeds_backfill. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Short name for the feed, e.g. "Weekly signups". | |
| metric | Yes | Metric id, or a snake_case slug to start tracking. | |
| mapping | Yes | ||
| schedule | Yes | ||
| toolArgs | No | ||
| toolName | Yes | The tool's own name on that server. | |
| workspace | No | Workspace slug. Omit to use the pinned workspace. | |
| approvalId | No | ||
| windowHours | No | How far back each run looks. Sensible per-cadence default. | |
| integrationId | Yes | From ledger_feed_sources_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-destructive, closed-world mutation. The description adds real behavioral context beyond that: the feed runs autonomously with 'no model involved', an unknown slug has the side effect of starting tracking for a new metric, and the call may return `needs_confirmation`. It does not, however, explain scheduling failure modes or what happens on a later conflicting feed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the tool's purpose, then three compact functional sentences covering slug semantics, the preview prerequisite, and the backfill follow-up. Every sentence carries information, though the 'Tell the user they can backfill' clause is more agent-behavior guidance than tool description and slightly dilutes the focus.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation with nested mapping/schedule objects and no output schema, the description covers the important operational surface: cadence, the mapping-confirmation prerequisite, the slug side effect, and the `needs_confirmation` return signal. It leaves the nested mapping and schedule fields entirely to the schema, which is a modest gap given the nesting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 60%, so the schema carries most of the burden. The description's only parameter statement ('metric takes an id or a snake_case slug; an unknown slug starts tracking that metric') largely restates the schema's own description for `metric`, adding the side-effect note but nothing about mapping, schedule, toolArgs, or windowHours. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Freeze a tool call as a metric's standing source') and immediately characterizes the mechanism (runs on cadence, records findings, no model involved). It is immediately distinguishable from siblings like ledger_feeds_preview and ledger_feeds_run, which are named in the same breath.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing instructions: confirm the mapping with ledger_feeds_preview first, then tell the user about ledger_feeds_backfill afterwards. The prerequisite and follow-up alternatives are named with the condition that selects each, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feeds_deleteStop a metric feedADestructiveInspect
Remove a feed. The readings it already wrote stay — the series is the record and outlives the instruction. Prefer pausing with ledger_feeds_update when the user might want it back. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Feed id, from ledger_feeds_list. | |
| workspace | No | Workspace slug. Omit to use the pinned workspace. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the deletion semantics are covered structurally. The description adds real context beyond them: readings already written are retained ('the series is the record and outlives the instruction') and that a needs_confirmation response is possible. It stops short of describing the confirmation/approval flow or permission requirements, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action, then consequence, then the preferred alternative. Every sentence carries distinct information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive no-output-schema tool, it covers the critical agent-facing questions: irreversibility of the instruction vs. retention of data, the safer alternative, and a possible confirmation response. It could still say what confirmation requires or what happens to dependent feed sources, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; the schema already explains 'id' (from ledger_feeds_list) and 'workspace'. The description adds no parameter-level meaning and leaves 'approvalId' undocumented in both places, though the needs_confirmation mention hints at that flow without explaining it. Baseline 3 is appropriate given the schema does most of the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove a feed') and immediately distinguishes itself from the sibling update/pause path. An agent can separate it from ledger_feeds_update, ledger_metrics_archive, and ledger_plan_delete without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('Prefer pausing with ledger_feeds_update') and the condition that selects it ('when the user might want it back'). This is a genuine when/when-not routing rule, not implied usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feeds_listList metric feedsARead-onlyInspect
The standing instructions for how metrics' numbers arrive — which tool each one calls, how often, and how the last run went. Check before proposing a feed so you don't duplicate one, and consult when a user asks why a metric is stale: a feed with a failing last run is usually the answer.
| Name | Required | Description | Default |
|---|---|---|---|
| metricId | No | Narrow to one metric (id or slug). | |
| workspace | No | Workspace slug. Omit to use the pinned workspace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is fully covered. The description adds the useful fact that each feed carries a last-run status, which explains its diagnostic value, but says nothing about ordering, pagination, or result size. With annotations carrying the safety burden, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the resource definition front-loaded ahead of the usage guidance. The em-dash clause packing three attributes is dense but informative. Slightly indirect opening (a noun phrase rather than a verb phrase) costs it the top score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey what a feed record contains, and it does name the key fields (target tool, cadence, last-run outcome). With only two optional, fully documented parameters and complete safety annotations, an agent has what it needs to call this correctly; ordering and pagination behavior are the only real omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (metricId as id-or-slug narrowing, workspace with pinned-workspace default) are already fully documented in the schema. The description adds no parameter detail at all, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description defines the resource concretely and specifically: 'the standing instructions for how metrics' numbers arrive — which tool each one calls, how often, and how the last run went.' That is enough for an agent to distinguish metric feeds from sibling families like ledger_entries_*, ledger_feeds_create/run, and ledger_metrics_list. It never states the verb 'list' explicitly, relying on the title for that, which keeps it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives two concrete triggers: check before proposing a feed to avoid duplication, and consult when a user asks why a metric is stale because a failing last run is usually the cause. That is unusually actionable guidance for a list tool. It does not name competing tools (e.g. ledger_feed_sources_list or ledger_feeds_preview) as alternatives, so no explicit when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feed_sources_listList servers a metric feed could pull fromARead-onlyInspect
The connected servers this workspace exposes to you, with the id a feed needs. Their tools appear to you namespaced as ext__<name>__<tool> — call one directly to see what it returns before proposing a feed. A server the workspace has switched off for you is not listed and cannot be fed from here.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Omit to use the pinned workspace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint/openWorldHint/destructiveHint already declaring a safe, closed-world read, the description still adds real behavioral context: the `ext__<name>__<tool>` namespacing, that server tools can be called directly to inspect returns, and that disabled servers are excluded and cannot be fed from. That goes beyond the annotations, stopping short of describing the return shape or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is returned and the key id, followed by practical usage detail. Dense but every clause carries information; minor polish could tighten the namespacing sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return explanation and does so adequately — servers plus the feed-usable id. It is complete enough to invoke correctly, though it leaves return ordering/shape unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single optional `workspace` parameter is fully documented in the schema (omit to use the pinned workspace). The description alludes to workspace scope but adds no syntax or format beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it lists the connected servers exposed to the workspace, together with the id a feed needs. The feed linkage is clear from the content, though it does not explicitly name the sibling feed tools it precedes (ledger_feeds_create/propose), so sibling differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It offers contextual guidance ('call one directly to see what it returns before proposing a feed') and an exclusion (switched-off servers are not listed and cannot be fed from), which implies when this list is useful. However, it never explicitly says when to use this vs. other ledger_feeds_* tools or states prerequisites, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feeds_previewTry a feed's mapping before creating itAInspect
Call a connected server's tool once and see what a mapping would pull out of the response — nothing is written and no feed is created. Send no mapping for a first look: you get a sample of the response plus the paths that hold numbers. Then send a mapping to confirm it finds the readings you expect. Always do this before ledger_feeds_create; proposing a feed whose mapping you haven't seen work is how a metric fills up with the wrong number. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| mapping | No | How to read the number. Omit on the first look. | |
| toolArgs | No | Arguments, exactly as the tool wants them. Use {{from}} / {{to}} (ISO instants) or {{date}} (YYYY-MM-DD) where a time range goes — those are substituted per run, and are what let one feed also backfill. | |
| toolName | Yes | The tool's own name on that server — NOT the ext__ namespaced form you call it by. | |
| workspace | No | Workspace slug. Omit to use the pinned workspace. | |
| approvalId | No | ||
| windowHours | No | How far back the preview's window reaches. Default 24. | |
| integrationId | Yes | From ledger_feed_sources_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavioral context beyond the annotations: it discloses the possible `needs_confirmation` return, the two-pass sampling behavior, and that the call is non-persistent ('nothing is written and no feed is created'). readOnlyHint=false sits in mild tension with that claim, since it invokes an external, open-world tool, but the description is explicit about what the ledger side will not do.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the operation and its non-persistence, then the workflow, then the failure mode. No sentence is filler; the warning about wrong metrics justifies its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, nested, mutating-adjacent call with no output schema, the description covers both the invocation workflow and the shape of what comes back (a response sample plus the paths holding numbers). The only thin spot is the confirmation flow's mechanics after `needs_confirmation`, which is named but not explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the baseline is 3, but the description adds genuine meaning by explaining the role of `mapping` across the two passes ('Send no mapping for a first look... then send a mapping to confirm'). The toolArgs templating semantics are already in the schema, so this is reinforcement rather than new semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource: call a connected server's tool once to see what a mapping would extract, with the explicit scope 'nothing is written and no feed is created'. It is immediately distinguishable from ledger_feeds_create, which it names by contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an unambiguous when-to-use ('Always do this before ledger_feeds_create'), a two-phase protocol (send no mapping first, then send a mapping to confirm), and the rationale for the exclusion (an unseen mapping is how a metric fills with the wrong number).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feeds_runRun a metric feed nowAInspect
Run a feed immediately over its own window and record what it finds. Use it right after creating one to prove the mapping works — a run that comes back empty names the path that missed, which is what you fix with ledger_feeds_update. Does not move the feed's schedule. Re-running is safe: readings dedupe per time bucket. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Feed id, from ledger_feeds_list. | |
| workspace | No | Workspace slug. Omit to use the pinned workspace. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=false and destructiveHint=false already declared, the description adds real behavioral context beyond annotations: it does not move the schedule, re-running is idempotent ('readings dedupe per time bucket'), and it may return `needs_confirmation`. That is precisely the mutation/idempotency/approval information an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, action front-loaded, and every clause adds distinct information (trigger, failure handling, schedule safety, idempotency, confirmation state). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries return-value burden; it names the two meaningful statuses (`empty`, `needs_confirmation`) but does not describe the normal success payload. Combined with the unexplained `approvalId`, one gap remains, but the core behavior is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; `id` and `workspace` are documented in the schema, while `approvalId` is undocumented. The description's mention of a possible `needs_confirmation` return hints at the approval flow tied to `approvalId`, but never explains the parameter, so it only marginally compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Run a feed immediately over its own window and record what it finds.' The 'over its own window' framing implicitly separates it from ledger_feeds_backfill, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger: 'Use it right after creating one to prove the mapping works.' It also routes the failure case to ledger_feeds_update. No explicit exclusions (e.g. when to prefer preview or backfill), but the context is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_feeds_updateFix or pause a metric feedADestructiveInspect
Change a feed's arguments, mapping, cadence, or pause it. Reach for this when a feed's last run reports empty (the response shape moved, so the mapping needs a new path) or error. Changing the mapping changes what the number means, so say what you're changing and why. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Feed id, from ledger_feeds_list. | |
| label | No | ||
| enabled | No | false pauses it; the definition and history are kept. | |
| mapping | No | ||
| schedule | No | ||
| toolArgs | No | ||
| toolName | No | ||
| workspace | No | Workspace slug. Omit to use the pinned workspace. | |
| approvalId | No | ||
| windowHours | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnly=false, destructive=true, openWorld=false. The description adds value beyond them: it warns that changing the mapping alters what the recorded number means, and discloses a `needs_confirmation` return path. It does not detail reversibility, approval semantics, or permission requirements, so it is not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the action and its scope, then trigger conditions and caveats. The 'say what you're changing and why' line is a soft instruction rather than voidless, but overall it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, nested-schema mutation with 10 params and no output schema, the description covers the trigger and the confirmation return, which is genuinely useful. It nonetheless leaves most parameter meanings and the full behavioral contract (reversibility, approval flow) underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 30% across 10 params, so the description must compensate. It maps concepts to a few params (arguments→toolArgs, mapping→mapping, cadence→schedule, pause→enabled), but leaves label, workspace, approvalId, toolName, and windowHours unexplained. Partial compensation justifies the baseline 3, not more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (change) and narrows the resource to a feed's arguments, mapping, cadence, or pause state. This cleanly separates it from siblings like ledger_feeds_create, ledger_feeds_run, ledger_feeds_delete, and ledger_feeds_list without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger — reach for this when a feed's last run reports `empty` or `error` — and even diagnoses the `empty` cause (response shape moved). It stops short of naming alternative siblings (e.g. when to preview vs. run vs. recreate), so it is strong context without full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metric_presets_adoptStart tracking metrics from the starting catalogAInspect
Creates the named presets as real metrics the workspace owns, tagged from the catalog. Safe to repeat: a preset the workspace already has comes back already_present rather than creating a second series, and a workspace's own renames and targets are never overwritten. Adopt only what the user agreed to — a metric nobody reads is noise on the Metrics tab, and eight thoughtful ones beat forty. Tell them the metrics start empty and the next step is a feed (ledger_feeds_create) or a first reading. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| slugs | Yes | Preset slugs from ledger_metric_presets_list. | |
| workspace | No | Workspace slug. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavior beyond the annotations: idempotent adoption returning `already_present`, protection of existing renames and targets from being overwritten, and a possible `needs_confirmation` result. These are exactly the side-effect and repeat-safety traits an agent needs and that readOnlyHint/destructiveHint alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded with the core action, then behavior, then guidance. Slightly chatty with the 'eight thoughtful ones beat forty' line, but it is short and each sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries return-value explanation well (`already_present`, `needs_confirmation`), plus next-step and repeat-safety context. It omits the workspace parameter's meaning and the approval flow that `approvalId` implies, which is a modest gap for a 3-param write tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema documents slugs (with maxItems 40) and workspace. The description clarifies that slugs come from the preset catalog and that metrics are tagged from it, but says nothing about the `workspace` or `approvalId` parameters, so it does not fully compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Creates the named presets as real metrics the workspace owns, tagged from the catalog.' An agent can distinguish this from ledger_metric_presets_list (read) and ledger_metrics_create (raw creation) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear adoption guidance ('Adopt only what the user agreed to') and names the next step (ledger_feeds_create) plus the source list implied by the slugs. It does not explicitly contrast with the sibling ledger_metric_starters_adopt, leaving the presets-vs-starters choice to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metric_presets_listBrowse the starting catalog of metricsARead-onlyInspect
Ready-made metrics a business can start tracking, each with a unit, a cadence, which direction is good, a business surface (facet), and cross-cutting tags. Reach for this when a workspace has few or no metrics, when someone asks what they should be measuring, or when a decision needs a number to settle against and none exists. Filter by facet (where it lives in the business) or tag (what kind of number it is). Suggest a SMALL set — three to six that fit what you know about this business — and say in one sentence why each one, rather than listing the catalog. Adopt with ledger_metric_presets_adopt. These are starting points: a workspace renames and retargets them freely afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Cross-cutting bucket, e.g. retention, cost, speed. | |
| facet | No | Business surface, matching the Ledger facet taxonomy. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld=false/destructive=false, so safety is covered. The description adds real context beyond them: items are starting points that a workspace may rename and retarget, and adoption happens through a separate tool. It stops short of pagination or size limits, but for a small read-only catalog that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool returns, then when to use it, then filtering, then delivery guidance. Every sentence carries information, though the block is dense enough that it spans several distinct concerns; still efficient for the amount of guidance an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by enumerating the fields each preset contains and clarifying these are editable starting points. Combined with annotations covering the safety profile, an agent has everything needed to call and use this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both enums documented, so baseline is 3. The description goes further by explaining the semantic distinction between the two filters — facet = where it lives in the business, tag = what kind of number it is — which helps an agent choose the right filter rather than just read the enum values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb/resource ('ready-made metrics a business can start tracking') and enumerates what each item carries (unit, cadence, good direction, facet, tags). It also implicitly separates itself from ledger_metrics_list by scoping to workspaces with few or no metrics, so an agent can tell what it returns without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three explicit triggering conditions ('few or no metrics', 'someone asks what they should be measuring', 'a decision needs a number to settle against and none exists') and names the downstream tool (ledger_metric_presets_adopt). It even steers the output behavior — suggest 3-6, not the whole catalog.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metrics_archiveArchive or unarchive a metricADestructiveInspect
Takes a metric out of the gallery (archived: true) or puts it back (false). Readings stay, and entries that settled against it still read correctly — this is the cleanup for a metric minted once and abandoned, or one the workspace stopped watching. metricId accepts the id or slug from ledger_metrics_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| archived | Yes | true to archive, false to restore. | |
| metricId | Yes | Metric id or slug (from ledger_metrics_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations flag destructiveHint=true, and the description adds genuinely useful context beyond them: readings are preserved and settled entries still read correctly, plus the possibility of a `needs_confirmation` envelope. This clarifies the real blast radius of the 'destructive' flag, though it doesn't detail permission or rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then supporting consequences and closure conditions. Three sentences, essentially all earning their place, with only mild density in the middle sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description covers the mutation semantics, data-retention behavior, and the confirmation flow. Combined with annotations covering the safety profile, an agent has enough to invoke it correctly; permissions and default workspace behavior are left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description repeats that `metricId` accepts an id or slug from ledger_metrics_list, which the schema already states, adding no new syntax or constraint detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: toggles a metric's gallery state via `archived` true/false. Clearly distinct from siblings like ledger_metrics_update or ledger_metrics_create, and the bidirectional semantics (archive vs. unarchive) are stated up front.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use guidance — 'cleanup for a metric minted once and abandoned, or one the workspace stopped watching' — which frames the intended scenario well. It stops short of naming an explicit alternative tool (e.g., why not ledger_metrics_update), so it doesn't reach the top bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metrics_createStart tracking a metricAInspect
Create a metric — a number the workspace watches (triage time, weekly signups, cost per run). Check ledger_metrics_list first; names are unique per workspace. When a user says they want to track or measure something, offer this. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| icon | No | The thing the number counts, e.g. phone for calls, landmark for profit. | |
| name | Yes | e.g. "Triage time". | |
| unit | No | Display unit — "min", "%", "$", "tickets/day". | |
| level | No | outcome = what the business is judged on; driver = moves an outcome; activity = daily work. | |
| drives | No | Ids of existing metrics this one moves. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| description | No | What the number means and where it comes from. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, so the mutation profile is known. The description adds two pieces of behavioral context the annotations cannot: the per-workspace uniqueness constraint on names and the fact that the call 'May return `needs_confirmation`', which prepares the agent for an unusual response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the verb and resource, then precondition, then trigger, then response caveat. Every sentence adds a distinct, actionable fact with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers purpose, prerequisite, trigger, and one return-value caveat, which is nearly everything needed. It omits permission requirements and what the created metric's identifier looks like, a minor gap given the annotations already cover the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 88% schema description coverage the schema already documents most parameters, so 3 is the baseline. The description earns above that by defining the `name` concept ('a number the workspace watches') and stating its uniqueness scope, information the schema itself does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a metric') and grounds it with concrete examples (triage time, weekly signups, cost per run), so the agent knows exactly what is produced. It does not, however, differentiate from sibling creators like ledger_metric_presets_adopt or ledger_metric_starters_adopt, so a one-line distinction is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Check ledger_metrics_list first; names are unique per workspace' gives a concrete precondition, and 'When a user says they want to track or measure something, offer this' supplies an explicit trigger condition. It stops short of naming when NOT to use it versus preset/starter adoption tools, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metrics_getRead one Ledger metric with its readingsARead-onlyInspect
One metric's definition (unit, kind, direction, target, cadence) plus its readings newest first and the open decisions bound to it. Use before answering 'how is X trending?' or before recording a reading against it. metricId accepts the id or the snake_case slug from ledger_metrics_list. For an event metric whose readings carry labels, labels lists each label and its values, largest total first. Pass slice to get the series for part of it ("new users in DE"), or by to split the series by one label ("new users by country"); either returns series, summed per bucket.
| Name | Required | Description | Default |
|---|---|---|---|
| by | No | Event metrics only. A label name to split the series by. | |
| slice | No | Event metrics only. Per label, the values to keep: { country: ["DE", "AT"] } keeps readings from DE or AT; several labels must all match. Names and values come from `labels`. | |
| bucket | No | Series bucket when `slice` or `by` is set. Default week. | |
| metricId | Yes | Metric id or slug (from ledger_metrics_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, destructiveHint, openWorldHint), so the description is free to add behavior it uniquely knows: readings are ordered newest first, labels are ordered by largest total, and series results are summed per bucket with a week default. It does not mention result-size limits or pagination on the readings array.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The what-it-returns clause is front-loaded, followed by usage triggers, then parameter mechanics; every sentence carries distinct information with no filler. Density is high but appropriate for a tool with five parameters and no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly describes the return payload (definition, readings, open decisions, labels, series) and its ordering. It omits how many readings are returned or whether they are capped, which is a minor remaining gap for a read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description earns more by explaining that metricId accepts either an id or a snake_case slug, and by giving concrete slice/by examples ('new users in DE', 'new users by country') plus the interaction with bucket. That adds interpretive value beyond the schema's field-level text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (read one metric) and enumerates exactly what is returned: definition fields, readings newest-first, and open decisions bound to it. It is clearly distinguishable from ledger_metrics_list, since it also explains that metricId can come from that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions: 'Use before answering how is X trending?' or before recording a reading against it, which routes the agent away from list/create siblings. It stops short of stating when-not-to-use it or naming the recording tool outright, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metrics_listList the workspace's Ledger metricsARead-onlyInspect
The numbers the workspace watches — each with unit, latest reading, and how many open decisions are bound to it. Consult when a user mentions a number that sounds like a tracked metric, and before recording a reading.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds value beyond that by describing the shape of the returned data (unit, latest reading, bound-decision count). It omits pagination/result-size behavior, hence 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first delivers the content/return summary, the second delivers the usage trigger. Zero filler and front-loaded with the most useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description appropriately compensates by naming the returned fields. With one optional, fully documented parameter and covered annotations, the definition is essentially complete; only pagination/ordering behavior is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'workspace' parameter is thoroughly documented in the schema (default workspace, personal tokens, API-key behavior). The description adds no parameter-level detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description frames the tool as listing 'the numbers the workspace watches,' which together with the title ('List the workspace's Ledger metrics') gives a clear verb+resource. It also enumerates the returned fields (unit, latest reading, open decision count). It does not explicitly distinguish itself from ledger_metrics_get, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete triggers – 'when a user mentions a number that sounds like a tracked metric' and 'before recording a reading' – which effectively routes the agent toward ledger_metrics_record_reading as the follow-up action. No explicit exclusions or named alternatives, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metrics_record_readingRecord a metric readingAInspect
Record one observation of a tracked metric. metricId accepts a metric id OR its snake_case slug from ledger_metrics_list; an unknown one is a not_found error — check ledger_metrics_list, create it with ledger_metrics_create, or pass createIfMissing: true to mint a "measure" metric at that slug in the same call (an "event" metric when the reading carries labels). The reading automatically lands as evidence on every open decision whose prediction is bound to this metric — so when a user reports a number ("triage is down to 12 minutes"), offer to record it. For event-kind metrics, omit value to count one occurrence. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | ISO timestamp the reading is for. Defaults to now; set it when backfilling an earlier reading. | |
| key | No | Idempotency key (e.g. "2026-w32") — a repeat write with the same key returns the original reading instead of doubling the series. Use for scheduled/recurring recordings. The key is per set of labels, so one key per day can cover every country. | |
| note | No | Where the number came from, if worth recording. | |
| value | No | The observed value, in the metric's unit. Omit for event-kind metrics to record one occurrence. | |
| labels | No | Event metrics only. What this count is broken down by, e.g. { country: "DE", plan: "pro" } — snake_case names, string values. Lets the metric be read and claimed by slice later. Refused on a measure. | |
| metricId | Yes | Metric id or slug (from ledger_metrics_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| createIfMissing | No | Mint the metric when the slug is unknown. Default false — an unknown slug is an error, so a typo can't quietly start a second series. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond annotations by disclosing the not_found failure mode, the createIfMissing side effect (minting a measure vs event metric based on labels), and the cross-cutting effect that readings land as evidence on open decisions bound to the metric. Also flags a possible needs_confirmation return. The only untold part is idempotent behavior details, which the schema already covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and then packs error handling, creation flow, and side effects into dense parentheticals. Every clause carries information, though the single paragraph is heavy and would benefit from light separation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 params, 89% schema coverage, nested labels, an annotation set, and no output schema, this description covers the required workflow (resolve-or-create), the event/measure distinction, and a non-obvious cross-tool side effect. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is high (89%), so baseline is 3, but the description adds real meaning: metricId accepts an id or slug and what happens on unknown values, and labels imply event-kind semantics vs a measure. This clarifies parameter interplay the schema states only in fragments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record one observation of a tracked metric'), which cleanly separates it from sibling read/create/list/archive metric tools. The scope is precise and an agent can identify its role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: check ledger_metrics_list for the id/slug, use ledger_metrics_create to make one, or pass createIfMissing to mint inline. It also gives a triggering condition ('when a user reports a number... offer to record it') and the event-metric invocation pattern (omit value).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metric_starters_adoptStart from a starter metric treeAInspect
Creates a starter's metrics at their levels and links them. Pass slugs to keep only some of its metrics; links are made only where both ends exist. Safe to repeat and safe after a different starter: metrics the workspace already has are left as they are and get linked into the tree. Adopt only what the user agreed to. Tell them the metrics start empty and the next step is connecting a source (ledger_feeds_create) or a first reading. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| slugs | No | The starter's metrics to keep. Omit for all of them. | |
| starter | Yes | Starter slug from ledger_metric_starters_list. | |
| workspace | No | Workspace slug. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations: idempotency ('safe to repeat and safe after a different starter'), non-destructive merge semantics ('metrics the workspace already has are left as they are'), and a possible `needs_confirmation` return. This is exactly the beyond-annotation disclosure the dimension rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then scope, then idempotency, then next-step guidance. Dense but every clause carries information; only slightly crowded by the interleaved user-facing instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, so the mention of `needs_confirmation` is doing useful work, and the idempotency/empty-start caveats cover the main surprises. Missing coverage of `approvalId` and workspace scoping keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so the schema already documents `slugs` and `starter`. The description restates the slug-subset behavior but adds nothing on `approvalId` or `workspace`, both of which stay undocumented in the description. Baseline 3 fits when the schema carries most of the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a starter's metrics at their levels and links them') with the linking behavior included, which distinguishes it from a plain metric-create tool. An agent can tell what it does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: 'Adopt only what the user agreed to', `slugs` to subset, and follow-up steps (connect a source or a first reading). It stops short of explicitly naming ledger_metric_presets_adopt as the alternative adopt path, so routing between the two adopt tools is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metric_starters_listBrowse starter metric treesARead-onlyInspect
Starter trees by shape of business or team (subscription software, services firm, online store, sales team, service operation). Each lists its metrics with a level (outcome / driver / activity) and the links between them (which number moves which). Reach for this before ledger_metric_presets_list when a workspace has no metrics yet or asks how its numbers fit together: pick the starter that matches what you know about the business, describe its tree in a sentence or two, and offer to adopt it with ledger_metric_starters_adopt, dropping any metric that doesn't fit.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real value beyond this by disclosing the shape of the returned content (metrics tagged by level plus inter-metric links) and the intended follow-up adoption flow. It stops short of listing pagination or count limits, hence a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource and its contents, then the routing condition, then the workflow. Every sentence earns its place, though the closing adoption advice ('describe its tree in a sentence or two... dropping any metric that doesn't fit') edges toward prescriptive verbosity for a read-only list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description carries the burden of describing returns and does so (metrics, levels, links between them). Combined with the explicit when-to-use condition and the named follow-up tool, an agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-parameter tool is 4. The description's enumeration of business shapes usefully hints at the conceptual categories the caller will see, but these are not input arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific resource (starter metric trees) and states exactly what they contain: business/team shapes, each metric's level (outcome/driver/activity), and the links between numbers. It distinguishes itself from the sibling ledger_metric_presets_list by naming it directly, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to reach for this tool ('before ledger_metric_presets_list when a workspace has no metrics yet or asks how its numbers fit together') and names the alternative. It also spells out the downstream workflow — pick a matching starter, describe its tree, then offer ledger_metric_starters_adopt — leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_metrics_updateEdit a metric's definitionADestructiveInspect
Corrects a metric's name, unit, description, kind (measure / event), direction (which way is good), target, cadence, icon, level, or the metrics it drives (its place in the metric tree). Readings are untouched. Only what you pass changes. The slug is not editable here — external writers address metrics by slug. Fails with conflict when a rename collides with a live metric. metricId is the id from ledger_metrics_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| icon | No | The mark the metric wears in every tool: the thing it counts (phone for calls, landmark for profit). "" goes back to the catalog's icon. | |
| kind | No | measure = a value each reading; event = a count of occurrences. | |
| name | No | ||
| unit | No | "min", "%", "$", "tickets/day". | |
| level | No | outcome = what the business is judged on; driver = a number that moves an outcome; activity = daily work. "" unplaces it. | |
| drives | No | Ids of the metrics this one moves. Replaces the current set; pass the full list. Loops are refused. | |
| target | No | Goal value in the metric's unit; null clears. | |
| cadence | No | How often a reading is expected. | |
| metricId | Yes | Metric id (from ledger_metrics_list). | |
| direction | No | up = higher is better, down = lower is better, none. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint=true annotation: it discloses partial-update behavior, that readings are unaffected, that the slug is immutable, the conflict-on-rename failure mode, and that a needs_confirmation envelope may be returned. These are exactly the behavioral traits an agent needs before invoking a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the verb and editable fields, then layers constraints (immutability, conflict, confirmation) in tight declarative sentences. Every clause carries information; nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter mutation tool with no output schema, it covers the essentials: scope of change, immutable fields, failure mode, and the confirmation flow. It could tie approvalId more explicitly to the needs_confirmation envelope, but the schema covers that parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 85%, so the schema already documents most parameters (including kind, direction, drives, cadence). The description restates the field list and adds only the 'metric tree' framing for drives, so it adds marginal value beyond structured fields, which is the baseline-3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (corrects/edits) and resource (a metric's definition) and enumerates the editable fields. It explicitly distinguishes itself from reading tools with 'Readings are untouched,' so an agent can separate it from ledger_metrics_record_reading without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: partial-update semantics ('Only what you pass changes'), the source for metricId ('from ledger_metrics_list'), and a hard exclusion (slug not editable). It does not explicitly say when to prefer this over ledger_metrics_create or ledger_metrics_archive, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_plan_addAdd dated work to a Ledger decisionAInspect
Adds a plan item to an OPEN decision (the work that carries it out: a post, a launch step, an email). Omit entryId for a standalone date (a holiday, an event you're only watching). Items are all-day unless you pass startTime with timeZone (and optionally endTime), for a webinar, a scheduled post, or a launch at noon. If the date costs money or time and comes with an expectation, record it as a decision with ledger_entries_create instead. Use repeatWeeklyUntil for a weekly cadence: it writes one row per week (max 60), each movable on its own. Fails with conflict once the decision is settled. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Where the work lives (draft, deck, published post). | |
| note | No | ||
| tags | No | What kind of work it is: blog, video, social, event… Reuse the workspace's existing tags (see ledger_plan_list). | |
| dueOn | Yes | A calendar day, YYYY-MM-DD. | |
| title | Yes | ||
| endTime | No | Same-day end, after startTime. | |
| entryId | No | The open decision this carries out. | |
| timeZone | No | IANA zone the time is in (America/Chicago). Required with startTime; use the user's own zone. | |
| startTime | No | Omit for an all-day item. 24-hour HH:MM in timeZone. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| ownerUserId | No | A workspace member's user id. | |
| repeatWeeklyUntil | No | Also add a copy every 7 days through this day. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate it's not read-only and not destructive. The description adds crucial behavioral details: fails with conflict once decision is settled, may return needs_confirmation, and repeatWeeklyUntil writes one row per week (max 60), each movable. This goes beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and immediately addresses entryId. The description is dense but well-structured, covering multiple scenarios efficiently. It could be slightly more concise, but every sentence adds necessary context for a tool with 13 parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 13 parameters, no output schema, and annotations that only cover safety hints, the description does a good job explaining key behaviors, failure modes, and parameter usage. It doesn't cover all parameters (e.g., workspace, tags), but those are documented in the schema. It's nearly complete for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 85%, so baseline is 3. The description adds value by explaining entryId's purpose ('the work that carries it out'), when to omit it ('standalone date'), the all-day vs timed distinction (startTime/timeZone/endTime), and repeatWeeklyUntil's weekly row behavior (max 60). Some parameters like workspace are not covered, but the description enhances understanding of key parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Adds a plan item to an OPEN decision'. It also clarifies the conceptual model (plan item = work that carries out a decision) and distinguishes from the sibling ledger_entries_create by describing when to use which. An agent can immediately understand this adds a dated work item, not a decision or evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use entryId vs omit it (standalone dates), when to use all-day vs timed, and when to redirect to ledger_entries_create instead. However, it doesn't mention ledger_plan_update or ledger_plan_delete as alternatives for editing/removing items, though those siblings exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_plan_deleteRemove a Ledger plan itemADestructiveInspect
Removes one plan item while its decision is open, for an item added by mistake or work that's no longer planned. If the work was planned and then dropped, prefer ledger_plan_update with status 'skipped' so settlement can see it. Fails with conflict once the decision is settled. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds real context beyond that: the operation is only valid while the decision is open, it fails with a conflict after settlement, and it may return a needs_confirmation envelope. It stops short of spelling out the confirmation/approval flow explicitly, keeping it at a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences: purpose, the preferred alternative, then the failure/confirmation behavior. No filler and the destructive scope is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema, the description covers the important return states (conflict, needs_confirmation) and the mutation constraint. The only gap is the expected format/identity semantics of itemId, which is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with workspace and approvalId documented but itemId bare. The description compensates by linking the needs_confirmation return to the approvalId parameter, which is meaning the schema alone does not supply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (removes) and resource (plan item) and adds the key scoping condition 'while its decision is open', which an agent can use to tell it apart from ledger_plan_update and ledger_plan_add without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative and the condition that selects it: 'If the work was planned and then dropped, prefer ledger_plan_update with status \'skipped\''. It also states the failure condition (conflict once settled), so both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_plan_listRead the Ledger calendar or one decision's planARead-onlyInspect
Pass entryId for one decision's plan items, or from + to (YYYY-MM-DD, at most about a year apart) for the calendar: decision spans (recorded day → expectation deadline) plus every dated item inside the window, including standalone dates with no decision. Filter the calendar by ownerUserId to answer 'what's mine this week' or to find tomorrow's items to draft. Never use it to compare or rank people's output.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | A calendar day, YYYY-MM-DD. | |
| tag | No | Only items with this tag (blog, video, event…). | |
| from | No | A calendar day, YYYY-MM-DD. | |
| entryId | No | One decision's plan. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| ownerUserId | No | Only items owned by this member. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safe read-only, non-destructive, closed-world profile, and the description adds real behavioral detail beyond them: what the calendar actually returns (decision spans from recorded day to expectation deadline, plus standalone dated items) and the roughly one-year window limit. Pagination and result volume are not addressed, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the entryId-vs-from/to decision, then the return semantics, then the filtering use case and the exclusion. Dense but nearly every clause earns its place; the single long paragraph is slightly harder to scan than bulleted alternatives.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with a fully documented schema and annotations covering safety, the description supplies the missing pieces: the two call modes, the window constraint, the returned item types, and the anti-use case. No output schema exists, so return-shape detail here is a bonus rather than a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: entryId scopes to one decision's plan, from/to must be paired and kept to about a year, and ownerUserId is framed by its use case. Only `tag` goes unexplained beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific read operation on the ledger plan/calendar and defines its two modes (per-decision via entryId vs. calendar window via from/to) in the first sentence. It is clearly distinguishable from siblings ledger_plan_add/update/delete, which mutate the same resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the caller: pass entryId for one decision, from+to for the calendar, and use ownerUserId for 'what's mine this week' or 'tomorrow's items to draft'. It also states a when-not — never use it to compare or rank people's output — so both selection and exclusion are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_plan_updateMove, rename, re-own, or mark a Ledger plan itemADestructiveInspect
Edits one plan item while its decision is open: move it (dueOn), rename it, change the owner (null clears), attach the link, or set status (planned | done | skipped). Mark done only when the user says it shipped; a link alone doesn't mean done. Fails with conflict once the decision is settled. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| note | No | ||
| tags | No | Replaces the item's tags. | |
| dueOn | No | A calendar day, YYYY-MM-DD. | |
| title | No | ||
| itemId | Yes | ||
| status | No | ||
| endTime | No | ||
| timeZone | No | Required when setting startTime. | |
| startTime | No | null makes it all-day again. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| ownerUserId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the description's added value is the conflict behavior ('Fails with conflict once the decision is settled') and the possibility of a `needs_confirmation` return, which the agent must handle. It does not spell out that field overwrites are irreversible, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each carrying distinct information (capabilities, guardrail, failure modes), with the core capability front-loaded. No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description covers the two return-path facts an agent needs (conflict on settled decisions, `needs_confirmation` envelope, which pairs with the approvalId parameter). It omits the workspace-token requirement, but that is already documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 13 parameters and only 46% schema description coverage, the description partly compensates: it maps dueOn to move, title to rename, url to link, status to the three enum values, and adds the non-obvious 'null clears' rule for ownerUserId. Only `note` and `itemId` are left undocumented in both places, so the compensation is incomplete but substantial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Edits one plan item') and enumerates the exact mutations supported (move via dueOn, rename, re-own, attach link, set status). It is trivially distinguishable from siblings ledger_plan_add, ledger_plan_delete, and ledger_plan_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a real precondition ('while its decision is open') and a substantive when-to-use rule ('Mark done only when the user says it shipped; a link alone doesn't mean done'). It never names a sibling tool as the alternative, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_docsList ZeroWidth docs pagesARead-onlyInspect
Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling get_doc. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| product | No | Optional product slug filter (e.g. 'compass', 'legal', 'overview'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds context beyond them with 'No authentication required,' a genuinely useful operational fact for callers, though it says nothing about pagination or result size for a full enumeration.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with purpose, routing guidance, and the auth fact front-loaded in order of importance. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-optional-param listing tool whose annotations cover safety, the description is nearly sufficient. Without an output schema it could note the return shape (e.g. that results are slugs/pages), but the 'slugs' reference largely covers that, so only a small gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the single 'product' parameter. The description's example values ('compass', 'legal', 'overview') duplicate the schema description verbatim, adding no meaning beyond it. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Enumerate') and resource ('all available docs pages') with scope, plus the optional product filter. It names the sibling get_doc and frames itself as the discovery step before retrieval, letting an agent distinguish it from single-doc fetch without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it: 'to discover what slugs exist before calling get_doc,' giving a clear dependency flow and naming the alternative. It does not address when NOT to use it or whether search_docs is a better discovery path, leaving a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_postsList Distributed Cognition essaysARead-onlyInspect
Lists ZeroWidth's published Distributed Cognition essays (newest first): title, slug, excerpt, byline, public URL. Use this when the user asks what's been published, or when a workspace question might already have an essay behind it — then fetch the full text with get_post. No authentication required.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max essays to return. Defaults to 10. | |
| cursor | No | Pagination cursor from a prior call's nextCursor. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so safety is covered. The description still adds useful context the annotations don't carry: 'newest first' ordering and 'No authentication required'. It doesn't discuss pagination or empty-result behavior, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The resource and returned fields come first, then usage, then the sibling pointer and the auth note — front-loaded exactly as an agent needs it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity list tool with only two optional params, the description covers identity, return shape, ordering, usage triggers, the follow-up tool, and an auth note. With no output schema, enumerating the returned fields compensates adequately; nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with `limit` (default 10, max 20) and `cursor` fully documented in the schema. The description adds no parameter detail — notably it doesn't mention that results are paginated or bound to 20 — so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Lists) plus a clearly scoped resource (ZeroWidth's published Distributed Cognition essays) and an explicit enumeration of the returned fields. It also names `get_post` as the sibling for full text, so an agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states two concrete trigger conditions — 'user asks what's been published' and 'a workspace question might already have an essay behind it' — and explicitly routes to `get_post` for full text. When-to-use and the alternative are both given, nothing inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_boards_createCreate a Napkin boardAInspect
Creates a blank Napkin board — a sketch (free canvas) or a deck (slides). NOT for written documents: 'draft/write a doc, note, memo' is napkin_docs_create. Use when the user asks for one, or proactively when a sketch/deck would carry the conversation better than words; hand it over with [[board:ID]] alone on its own line. For a deck with content, prefer napkin_deck_write (one call, whole deck).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Defaults to "sketch". | |
| name | Yes | Board name. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a write (readOnlyHint=false) but non-destructive and not open-world, so the safety profile is covered. The description adds real context beyond that: the created board is blank, the handover syntax is `[[board:ID]]` on its own line, and it warns that napkin_deck_write is preferable when content already exists. It does not discuss permissions or workspace-token behavior, which the schema covers instead.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying a distinct load: what it makes, what it is not for, and how/when to use it. The exclusion and the primary trigger are front-loaded, and the handover format is given inline rather than left to a later sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity create tool with full schema coverage and annotations covering the safety profile, the description supplies everything still missing: the emotional/practical trigger for proactive use and the exact syntax for surfacing the new board back to the user. No output schema exists, so return-value detail is not required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description earns above baseline by explaining what the `kind` enum values actually mean (sketch = free canvas, deck = slides) and by implying the deck case has a better sibling tool, adding semantic meaning beyond the raw enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Creates a blank Napkin board") and immediately scopes it with the two valid kinds ("sketch (free canvas) or a deck (slides)"). It also names the sibling it is NOT (napkin_docs_create), so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-not guidance ("NOT for written documents: 'draft/write a doc, note, memo' is napkin_docs_create"), a positive trigger ("Use when the user asks for one"), a proactive condition, and a routing rule to a better alternative for the deck-with-content case ("prefer napkin_deck_write"). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_boards_getRead a Napkin board's shapesARead-onlyInspect
Returns a board's shapes as data — each with its id, kind, position, size, and text — so you can TARGET one to change or remove with napkin_draw (updates / deletes take these ids and absolute board coordinates). Pen strokes come back as a point count, not the points. For what the board LOOKS like, use napkin_boards_view instead. Board ids come from napkin_boards_list.
| Name | Required | Description | Default |
|---|---|---|---|
| boardId | Yes | Napkin board id (from napkin_boards_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/non-destructive/closed-world, so the safety profile is covered. The description adds meaningful context beyond structured data: pen strokes return as a point count rather than raw points, and the returned ids/coordinates are the exact inputs napkin_draw expects. Doesn't cover pagination or size limits, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the primary return behavior, then constraints, then alternatives. Every sentence carries distinct information with no redundancy, and emphasis (TARGET, LOOKS) aids scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only board-fetch tool with annotations covering safety and no output schema, the description supplies the return shape, the downstream integration path, the pen-stroke caveat, the id source, and the sibling alternative. An agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so boardId and workspace are already documented by the schema (including sourcing for boardId). The description reinforces that board ids come from napkin_boards_list but adds no format or syntax detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns a board's shapes as data') and enumerates the returned fields (id, kind, position, size, text). It explicitly distinguishes itself from napkin_boards_view ('what the board LOOKS like') and connects to napkin_draw, so an agent can tell it apart from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the use case (to TARGET a shape to change or remove via napkin_draw updates/deletes) and names the alternative tool with its distinct condition (napkin_boards_view for visual appearance). It also notes the id source (napkin_boards_list), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_boards_listList Napkin sketchesARead-onlyInspect
Lists sketch boards in the active workspace — name, id, kind (sketch / deck), shape count, last activity. Boards are SPATIAL canvases (drawing, diagrams); the markdown DOCUMENTS are docs — napkin_docs_list. Archived boards are hidden unless includeArchived; q narrows by name. When the user mentions a sketch, drawing, whiteboard, or napkin, find it here, then LOOK at it with napkin_boards_view before discussing its contents. When referring the user to a sketch in your reply, put [[board:ID]] alone on its own line — it renders as a clickable card with a live thumbnail.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Name contains (case-insensitive). | |
| kind | No | Only sketches or only decks. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| includeArchived | No | Also list archived boards (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/no-destructive, so safety is covered; the description adds the real behavioral detail that archived boards are hidden by default and that `q` matches on name. It also discloses a display side-effect (the [[board:ID]] card rendering), which goes beyond structured fields, though it doesn't cover pagination or result limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then layers routing, filtering, follow-up, and formatting guidance with no filler. Each sentence carries a distinct actionable instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool whose annotations carry the safety profile and whose schema documents all four params, the description supplies the missing pieces: sibling disambiguation, filter defaults, the required follow-up view call, and reply formatting. Nothing an agent needs to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description nonetheless adds meaning by stating the default visibility rule for archived boards and the name-matching semantics of `q`. It does not elaborate on the workspace parameter beyond what the schema says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Lists sketch boards') and enumerates the returned fields (name, id, kind, shape count, last activity). It explicitly distinguishes itself from the sibling napkin_docs_list by contrasting spatial canvases (boards) with markdown documents (docs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete trigger phrases ('sketch, drawing, whiteboard, or napkin') and an explicit alternative ('the markdown DOCUMENTS are docs — napkin_docs_list'). It also prescribes the follow-up flow: find it here, then LOOK at it with napkin_boards_view before discussing contents.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_boards_updateRename, describe, archive, or share a Napkin boardADestructiveInspect
Edits a board's details: name, description, visibility, or archived (true takes it out of the gallery; false brings it back — the recovery move for a board you created by mistake or the user no longer wants). Board ids come from napkin_boards_list. Boards are working material: no approval card, every change attributed.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| boardId | Yes | Napkin board id (from napkin_boards_list). | |
| archived | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation safety profile is covered. The description adds genuinely new context: exact archived semantics (true removes from gallery, false restores it) and the operational note that boards skip approval and every change is attributed, which is useful audit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Edits a board's details') before parenthetical detail. The archive aside and attribution note each earn their place, though the sentence is somewhat dense with em-dashes and parentheticals.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers the archive behavior, attribution, and id sourcing that an agent needs to call this correctly. It does not specify partial-update semantics (whether omitted fields are preserved) or permission requirements, leaving a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, but the description compensates by explaining the archived flag's two-way behavior and confirming boardId sourcing from napkin_boards_list. The visibility enum and workspace override are already well documented in the schema, so combined coverage is strong; only 'name'/'description' semantics are left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Edits') and resource ('a board's details') and enumerates the mutable fields (name, description, visibility, archived). An agent can distinguish it from napkin_boards_create/get/view by name, though the description never explicitly names a sibling to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a real when-to-use for the archive path ('the recovery move for a board you created by mistake or the user no longer wants') and points to napkin_boards_list for ids. However, it offers no exclusions, no guidance on when to prefer another sibling, and no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_boards_viewView a Napkin sketch (rendered image)ARead-onlyInspect
Renders the sketch board to an image and returns it so you can SEE what's drawn — layout, arrows, handwriting-style strokes, sticky notes — not just data about it. Use this before answering any question about a sketch's contents, and cite the board when you do — [[board:ID]] alone on its own line embeds the sketch card in your reply.
| Name | Required | Description | Default |
|---|---|---|---|
| boardId | Yes | Napkin board id to render. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful behavioral context beyond the schema: it renders content to an image, returns that image, and specifies a citation convention for embedding the result. It does not discuss rate limits or permissions, but those are not central for a safe read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the key behavioral distinction (rendering an image vs. returning data) and then the usage/citation rule. There is no filler, and the format guidance is compact and actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only view tool with no output schema, the description tells the agent what it receives (a rendered image), when to use it, and how to cite the result. Annotations cover safety, and the schema covers parameters, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the description adds no parameter-level detail beyond what the schema already documents for boardId and workspace. With the schema carrying full parameter semantics, a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (renders) and resource (sketch board) and explains that it returns an image so the agent can SEE the drawing, not just data. It distinguishes itself from data-oriented siblings like napkin_boards_get by saying 'not just data about it.' An agent can tell this tool apart from get/list/update without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage guidance: 'Use this before answering any question about a sketch's contents' and instructs the agent to cite the board with a specific embed syntax. It does not name an alternative sibling or state when-not to use it, but the contrast with 'not just data' implicitly routes the agent away from napkin_boards_get.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_add_filesAdd pictures to a brand kitAInspect
Adds pictures to a brand kit section — logo versions, product marks, example screenshots, imagery. Each file comes from a public https url (fetched by the server) or, for a small file that isn't online, base64 data with a filename. PNG, JPEG, WebP, GIF and SVG, up to 20 MB each. Every file becomes a workspace file and an asset on the section, appended after the ones already there. Give each a name and a note saying what it's for; on logo sections set backdrop to the hex of the ground it's made for (#ffffff for a dark logo, #000000 for a reversed one), which is how pages pick the right version. Same rule as napkin_brand_draft: you can change any kit except the workspace's brand. Files that fail are reported and the rest still land.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Only when no section has that title yet: the kind of section to make. | |
| files | Yes | ||
| kitId | Yes | The kit to add to (napkin_brand_list). | |
| section | Yes | Title of the section to add them to, like Logo, Imagery or Examples. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint false, openWorldHint true, destructiveHint false) by disclosing server-side fetching of URLs, 20 MB per-file cap, append-after-existing ordering, partial-failure semantics ('files that fail are reported and the rest still land'), and the permission rule that any kit is editable except the workspace's brand. That is exactly the operational context an agent needs for a mutating, open-world call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause, then constraints, then the operational caveat. The prose is dense but nearly every sentence carries a rule; the parenthetical hex examples add precision rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no output schema, the description covers sourcing, size limits, ordering, partial failure, branding permissions, and the logo backdrop convention. It could be tighter on section auto-creation (the `kind` parameter's role), but no output schema is needed since failure reporting is described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 80%, but the description adds real meaning: `url` is fetched by the server, `data` requires `filename`, `backdrop` is the ground color a logo is designed for and is how pages select the right version, and `name`/`note` convey purpose. This exceeds the baseline for well-documented schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — adding pictures to a brand kit section — and enumerates the concrete content types (logo versions, product marks, screenshots, imagery). It clearly distinguishes itself from siblings by referencing napkin_brand_draft and napkin_brand_list, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use: adding files to an existing kit section, with the per-file rules and the backdrop convention called out for the logo case. It names the related sibling napkin_brand_draft for the permission rule but does not explicitly contrast alternatives such as napkin_brand_apply or state when NOT to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_applyApply the brand to a pieceADestructiveInspect
Restyles a piece in the brand — the workspace's brand unless you name a kit. deck (a board id): every slide gets the brand's background, typefaces, and text colors; sizes are kept. interface (an interface id): rewrites its brand.css, so anything styled with var(--brand-*) updates. accent makes one of the kit's colors the piece's accent — use it when the piece is about one product or campaign that has its own color in the brand (read the kit's Colors section to see which). Docs need nothing — they read the brand when exported. Check a deck afterwards with napkin_slide_view.
| Name | Required | Description | Default |
|---|---|---|---|
| deck | No | Board id of a deck. | |
| kitId | No | A brand kit id from napkin_brand_list. Omit for the workspace's brand. | |
| accent | No | Name of one of the kit's colors to use as this piece's accent. | |
| interface | No | Interface id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint annotation by describing exactly what mutates (slides get background/typefaces/text colors, interface brand.css is rewritten) and what is preserved ('sizes are kept'), plus the downstream effect on var(--brand-*) styling. This is substantial disclosure of mutation semantics for a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then spends each following clause on one parameter's distinct effect with no filler. Dense but every sentence carries actionable detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with no output schema, the description covers effects and verification well. It does not, however, clarify that at least one of deck/interface/kitId must be supplied (all params are optional), leaving an ambiguity an agent could trip on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all five parameters (baseline 3). The description adds real value beyond it by explaining accent's runtime meaning and kitId's default-to-workspace-brand behavior, though it says nothing extra about workspace or interface beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Restyles a piece in the brand') and enumerates the distinct effect per target type (deck, interface, accent, docs). This lets an agent distinguish it from siblings like napkin_deck_set_theme or napkin_brand_get without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditional guidance for the accent parameter ('use it when the piece is about one product or campaign that has its own color') and explicitly says docs need nothing. It also routes verification to napkin_slide_view, but does not name alternative tools for setting deck styling, so no explicit when-not guidance against siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_checkCheck writing against the brandARead-onlyInspect
Finds words the brand says to avoid in a piece of writing, with what to say instead and why. Pass text, or a deck (boardId) or doc (docId) to check its words. Run it on your own drafts before handing them back.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Text to check. | |
| docId | No | A doc to check. | |
| kitId | No | A brand kit id from napkin_brand_list. Omit for the workspace's brand. | |
| boardId | No | A deck or sketch to check. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a safe read (readOnly=true, destructive=false, openWorld=false), so the safety profile is covered. The description usefully adds that the result includes suggested replacements and the reason, which compensates for the absent output schema. It doesn't disclose limits (e.g., what happens with an invalid boardId or size caps), so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core behavior, then inputs, then a usage nudge. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, zero required params, and 100% schema coverage, the description covers what the tool returns and how to point it at content. The optional kitId/workspace path and error/edge behavior are left to the schema, which is acceptable but slightly incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema. The description adds only which inputs are mutually usable (text/boardId/docId) and omits kitId and workspace entirely, so it does not meaningfully exceed the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: it finds brand-discouraged words in writing and returns replacements plus rationale. This is clearly distinct from sibling brand tools like napkin_brand_apply (which applies a brand) and napkin_brand_get/list (which read the kit), so an agent can select it without opening sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use ('Run it on your own drafts before handing them back') and enumerates the three input modes (raw text, deck via boardId, doc via docId). It stops short of naming when to prefer a sibling tool or when not to use this one, so it doesn't reach the explicit-alternatives bar of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_draftDraft a brand kitADestructiveInspect
Creates a new brand kit, or edits one that is NOT the workspace's brand. Use it when someone asks you to put their brand together (from their site, a deck, or what they tell you) or to try a variation. You can't change the workspace's brand itself — make a draft and tell the user they can choose it in Napkin's Brand section. A kit is a document of sections, each with title, kind (overview, voice, colors, typography, logo, imagery, motion, layout, examples, custom), markdown body, and rules ({rule, why}); colors sections hold swatches ({name, value: 6-digit hex, note}), typography sections hold faces ({name: what it's for, family: a font name or sans / serif / mono / rounded / marker, note}), voice sections hold use and avoid ({term, instead, why}), and any section can hold assets ({fileId: an image already in the workspace's files, name, note, backdrop: hex}) — logo versions, example images — and links ({url, title, note}). An examples section holds work that gets the brand right: screenshots as assets and pages as links, each with a note on what makes it good. Exact numbers (timings, spacing) go in the section's prose. New kits start with empty standard sections plus the workspace's colors. content.sections upserts by title (or id): fields you pass replace, lists inside replace, sections you don't mention are kept; removeSections drops by title. content.uses says how Napkin uses the brand — color roles (ink, muted, background, surface, accent, accent2) and fonts (heading, body, mono) — by swatch or face NAME. Give every rule a why: the reason is what lets people decide cases the rule doesn't cover.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| kitId | No | Kit to edit. Omit to create a new kit. | |
| content | No | ||
| fromKitId | No | New kits only: start from this kit's values. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say destructiveHint=true, readOnlyHint=false, openWorldHint=false; the description goes far beyond by disclosing the merge semantics that make this destructive — 'fields you pass replace, lists inside replace, sections you don't mention are kept', plus removeSections drops by title, and new kits inherit the workspace's colors. This is exactly the kind of behavioral context that prevents accidental data loss.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the create/edit decision and the workspace-brand prohibition before the dense content-model exposition, which is the right ordering. The remaining prose is a long em-dash-heavy run of clauses, but nearly every sentence carries information an agent needs to build a valid kit payload, so the length is largely earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must carry the full burden — and it covers the mutation semantics, the content model, the workspace-brand boundary, and the 'give every rule a why' convention. It omits the return shape/identifier of a newly created kit, which is a minor gap for a creation tool whose caller likely needs the new kitId.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema coverage over a deeply nested six-parameter model, the description compensates well: it defines the section content model (title/kind/body/rules/swatches/faces/use/avoid/assets/links/layouts), explains the upsert-by-title behavior of content.sections, and clarifies that content.uses references swatch or face NAMES. It does not document name, description, workspace, or fromKitId directly, but the schema does cover those.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair ('Creates a new brand kit, or edits one that is NOT the workspace's brand') and immediately scopes it against the workspace brand, which is exactly the confusion an agent would have. An agent can distinguish this from napkin_brand_apply / napkin_brand_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering contexts ('when someone asks you to put their brand together... or to try a variation') and an explicit prohibition with a workaround ('You can't change the workspace's brand itself — make a draft and tell the user they can choose it'). It stops short of naming sibling tools (napkin_brand_apply, napkin_brand_examples) as alternatives, so it is clear context without explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_examplesLook at the brand's examplesARead-onlyInspect
Shows you the pictures in a kit's Examples sections — screenshots of work that gets the brand right — with each one's note on what makes it good, plus the example links. Look before you design a deck, page or interface in the brand, and hold your work to the same bar: the layouts, density, type sizes and use of color you see there.
| Name | Required | Description | Default |
|---|---|---|---|
| kitId | No | Kit to look at. Omit for the workspace's brand. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, and destructiveHint, so safety is covered. The description adds that it returns pictures, notes, and links, but it does not disclose authorization needs, rate limits, or other behavioral traits beyond the annotation-covered basics. A 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded and free of filler; the first explains what is shown and the second explains why to use it. It is appropriately sized for a tool with no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only viewer with rich annotations and full schema coverage, the description is complete: it explains the return content (pictures, notes, links) and the usage context. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both kitId and workspace fully documented in the schema. The description adds no parameter-level guidance, so baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (shows) and resource (brand kit Examples sections), and clarifies the payload (screenshots, notes, links). It does not explicitly name a sibling or distinguish itself from napkin_brand_view/get, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use: 'Look before you design a deck, page or interface in the brand.' No when-not or alternative tools are mentioned, so it's a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_fontsGet a brand's font filesAInspect
Gets the font files for a kit's typefaces from Google Fonts and keeps them in the kit, so decks, slide pictures and the canvas draw the brand's real typefaces instead of a stand-in. Only typefaces that name a family (like Poppins) and don't have files yet are fetched. Works on any kit, including the workspace's brand: it adds the files for the typefaces the kit already names and changes nothing else. A typeface Google doesn't have is reported back; the user can upload its files in the kit's typography section.
| Name | Required | Description | Default |
|---|---|---|---|
| kitId | No | Kit to fetch fonts for. Omit for the workspace's brand. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavior beyond the annotations: it is idempotent (skips typefaces with existing files or no family), additive-only ('changes nothing else'), reaches an external source (consistent with openWorldHint), and reports back typefaces Google lacks. This clarifies the non-destructive nature and the failure path in ways annotations alone (readOnlyHint=false, destructiveHint=false) do not.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core action and benefit, then the fetch condition, scope, and failure handling. No filler, though it is mildly wordy and could tighten the first sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and full schema coverage, the description supplies the missing pieces: what gets added, what is left untouched, idempotency, and the not-found path with an upload remedy. Safety is covered by annotations, so remaining gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so kitId and workspace semantics are already fully documented in the schema. The description restates the kit/workspace scope but adds no format or edge-case detail beyond what the schema provides. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fetches a kit's typeface font files from Google Fonts and stores them in the kit. The effect (decks, slides and canvas render the real brand typefaces) makes the purpose concrete. It does not explicitly name the sibling it differs from (e.g. napkin_brand_add_files), though the upload fallback hints at the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage via scope ('Works on any kit, including the workspace's brand') and a fetch condition ('only typefaces that name a family and don't have files yet are fetched'), which tells the agent when the call is a useful/likely no-op. However, it never states when to call this versus napkin_brand_add_files or other brand tools, so selection guidance is inferred rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_getRead the brandARead-onlyInspect
Reads a brand kit — the workspace's brand unless you name another. READ THIS BEFORE writing or designing anything for the workspace: a deck, a doc, an interface, copy. Returns brief (the whole kit as markdown: every section — overview, voice, colors, typography, imagery, motion, and whatever else the team wrote — with the reason behind each rule; follow the reasons when a case isn't covered), css (the --brand-* variables; interfaces already link it as brand.css, so style with var(--brand-accent) etc. rather than hex values), and tokens (the resolved values Napkin uses). Before any kit exists it returns values from the workspace's colors with kit: null.
| Name | Required | Description | Default |
|---|---|---|---|
| kitId | No | A brand kit id from napkin_brand_list. Omit for the workspace's brand. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish a safe read (readOnlyHint=true, destructiveHint=false), so the bar is lower; the description still adds real value by disclosing the three return fields (brief/css/tokens) and the pre-kit fallback where values come from workspace colors with kit: null. It omits any note on auth scope or rate limits, so it stops short of the top mark.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The imperative is front-loaded ahead of the return-value detail, and each subsequent clause (css variable usage, fallback behavior) carries operational information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return contract and it does: it names brief, css, and tokens, explains their contents and intended use, and covers the no-kit-yet case. Nothing an agent needs to call and consume this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description restates the default-workspace semantics that the schema already carries for kitId and adds no new syntax or format detail for either parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (reads) plus resource (brand kit) and an explicit default scope: the workspace's brand unless a kit is named. An agent can distinguish it from napkin_brand_list (which enumerates kits) without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a direct when-to-use imperative ('READ THIS BEFORE writing or designing anything for the workspace: a deck, a doc, an interface, copy') and points to napkin_brand_list as the source of kitId, so the routing decision against siblings is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_listList brand kitsARead-onlyInspect
Lists the workspace's brand kits: id, name, description, and which one is the workspace's brand (isDefault). Most workspaces have one. Use napkin_brand_get to read a kit.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds value beyond that by disclosing the return shape and the meaning of isDefault as the workspace's brand. No mention of pagination or auth, which is a minor gap for a list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste. The core purpose and returned fields are front-loaded, and the routing hint to napkin_brand_get comes last.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the return fields and explains isDefault, which an agent needs to interpret results. Coverage is complete for a simple list tool, with only pagination/ordering left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single workspace parameter's default/override/ignored-for-API-key semantics are fully documented in the schema. The description adds nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (the workspace's brand kits), then enumerates the returned fields (id, name, description, isDefault). It explicitly names napkin_brand_get as the reader tool, so an agent can tell the two apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Routes the agent clearly: use this to list, use napkin_brand_get to read a specific kit. Adds prevalence context ('Most workspaces have one'). No explicit when-not guidance, but the alternative is named with its selecting condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_brand_viewLook at a brand's picturesARead-onlyInspect
Look at the pictures in a brand kit — logos, product marks, footage stills, layout references, examples — before you use them. Without section or names it lists every picture by section with its note, so you know what exists. With section (a section title, like Imagery) or names (pictures' names as the brief lists them) it shows you those pictures, up to 8 at a time. Look before you build a design around a picture: whether a photo has room for text, which part of it carries the color, and whether it suits texture inside type are things you can only tell by seeing it.
| Name | Required | Description | Default |
|---|---|---|---|
| kitId | No | A kit from napkin_brand_list. Omit for the workspace's brand. | |
| names | No | Show these pictures, by name. | |
| section | No | Show the pictures in this section. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/non-destructive, so the bar is lower; the description still adds real behavior: with no filters it enumerates everything by section with notes, and with filters it caps results at 8 at a time. No disclosure of image format, resolution, or failure modes, hence 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the two behavioral modes, then the rationale. The closing sentence about text room and texture is somewhat motivational but does justify why seeing the image matters, so it earns most of its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey returns — and it does (list of pictures by section with notes, or up to 8 shown images). Combined with full schema coverage this is nearly complete, though it says nothing about how images are delivered (URL vs. inline).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so 3 is the baseline; the description goes beyond it by explaining the semantics of `section` (a section title like 'Imagery') and `names` (as the brief lists them) and by clarifying their mutual effect on output (list vs. show, max 8).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Look at the pictures in a brand kit') and enumerates what a kit contains (logos, product marks, footage stills, layout references, examples), which clearly separates it from data-returning siblings like napkin_brand_get or napkin_brand_examples.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the timing ('before you use them', 'Look before you build a design around a picture') and gives the rationale for doing so. It doesn't name a competing sibling tool to route against, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_deck_composeBuild a designed deckADestructiveInspect
Builds a designed deck: each slide is a layout with its slots filled, or HTML and CSS you write; a real browser lays it out, and it becomes an ordinary Napkin deck people can edit. Use this whenever a deck should look designed — it's how to get layout, big numbers, color blocks, and pictures right.
LAYOUTS FIRST
For ads, posts, carousels and title slides, start from a layout (napkin_layouts_list): give
layoutandslotsinstead ofhtml. Layouts adapt to every size, fit their text, keep to each size's safe zones, and use the brand's own logos. A slide made from a layout remembers it, so when someone resizes it, it's laid out again instead of scaled.A set of ads is one slide with
sizes: ["1:1","4:5","9:16","1.91:1"] — each size gets its own arrangement.Write HTML only for something no layout does.
HOW TO WRITE A SLIDE
htmlis the inside of one slide: a.slidebox 960×540 px (the deck's own size if it has one). Margins are reset; lay out with flexbox or grid, padding and gap. Put each slide's CSS in a block inside its html, or shared CSS incss.Read the brand first (napkin_brand_get). Style ONLY with its variables: var(--brand-ink), var(--brand-muted), var(--brand-background), var(--brand-surface), var(--brand-accent), var(--brand-accent-2), var(--brand-color-), var(--brand-font-heading), var(--brand-font-body), var(--brand-radius). Headings already use the heading face.
Pictures from the brand kit: , sized with CSS; object-fit: cover crops to fill, contain fits whole. Any workspace image: src="file:".
What carries over: boxes with a solid background, border, and rounded corners; text, including bold, italic, color, and size changes inside a line; pictures. What doesn't: gradients, shadows, background images, inline SVG — use solid color boxes instead.
Design like a designer: one idea per slide, a clear hierarchy, big numbers set large in the accent color, generous space, text 18px or larger (never under 14px), and something visual on most slides — a color panel, a photo, the logo. Keep every slide on the brand's own palette and voice.
Use the brand's own pictures. Put the logo on the title and closing slides (the reversed one on dark color), and use the kit's photos and icons where they carry the point — the brief lists every file under "Files", by name.
Say only what's true. Facts, figures, prices, eligibility, and how things work come from the brand kit and from what the user told you — never invent them. When a slide needs a fact you don't have, leave it out or ask.
HOW TO USE IT
New deck: pass
nameand allslides. Rebuild a deck: passboardIdand allslides(replaces everything on it).Ads, posts and other non-slide pieces: pass
size— a key like1:1(square post, 1080×1080),4:5(portrait post),9:16(story or reel),1.91:1(LinkedIn or Facebook landscape),og(link preview),300x250or728x90(display ads). The.slidebox is then that shape with its long side 960 px, so type sizes mean what they do on a slide, and the deck exports at the size's real pixels. A single ad is a one-slide deck; a carousel is a deck at1:1or4:5.A set of ads — the same ad at several sizes, or several headlines — is one deck: give each slide its own
size. Lay each size out for its own shape rather than squeezing one layout into all of them (a 9:16 story stacks what a 1.91:1 landscape puts side by side), and keep text out of the outer 7% of every edge and the top and bottom 14% of a 9:16. The deck's PNG export names each file by its pixels. Fix one slide:boardId,slide(1-based), and a single slide. Add slides:append: true.The result lists each slide's
warnings— text that overflows its box or the slide, overlapping text, text too small, low contrast, pictures that didn't load — and shows you a picture of every slide it made. Fix every warning and anything that reads badly in the pictures by recomposing just that slide. The contrast check is the standard one (4.5:1 for small text, 3:1 for large or bold text), so a contrast warning is real: darken the color, lighten the background, or set the text larger; never explain it away. You don't need napkin_slide_view afterwards: the pictures are the slides.Cite the deck with [[board:ID]] alone on its own line.
| Name | Required | Description | Default |
|---|---|---|---|
| css | No | CSS shared by every slide. | |
| name | No | Name for a new deck. | |
| size | No | The size of every slide, for a new deck or a whole rebuild: 16:9 = Widescreen slide 1920×1080; 16:10 = 16:10 slide 1920×1200; 4:3 = Standard slide 1024×768; 1:1 = Square 1080×1080; 4:5 = Portrait post 1080×1350; 9:16 = Story and reel 1080×1920; 1.91:1 = Landscape post 1200×628; x-post = X post 1600×900; og = Link preview 1200×630; pin = Pin 1000×1500; youtube-thumb = YouTube thumbnail 1280×720; linkedin-page = LinkedIn page cover 1128×191; linkedin-profile = LinkedIn profile banner 1584×396; x-header = X header 1500×500; youtube-banner = YouTube channel banner 2560×1440; facebook-cover = Facebook cover 851×315; 300x250 = Medium rectangle 300×250; 336x280 = Large rectangle 336×280; 728x90 = Leaderboard 728×90; 970x250 = Billboard 970×250; 300x600 = Half page 300×600; 160x600 = Wide skyscraper 160×600; 320x50 = Mobile banner 320×50; 320x100 = Large mobile banner 320×100; email-header = Email header 600×200; letter = Letter page 2550×3300; a4 = A4 page 2480×3508. Omit for a 16:9 deck or to keep a deck's size. | |
| slide | No | Replace just this slide (1-based) with the one slide given. | |
| accent | No | Name of one of the brand's colors to use as var(--brand-accent) for this deck — for a deck about one product or campaign that has its own color in the brand. | |
| append | No | Add these slides after the existing ones. | |
| slides | Yes | ||
| boardId | No | Deck to rebuild or change. Omit to make a new deck. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description earns that by disclosing the destructive semantics directly: passing boardId with all slides 'replaces everything on it'. It goes well beyond annotations with the carry-over contract (solid backgrounds/borders/radius/text/pictures survive; gradients, shadows, background images and inline SVG do not), the size/shape model, and the returned per-slide warnings plus rendered previews.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then organized under clear headers (LAYOUTS FIRST / HOW TO WRITE A SLIDE / HOW TO USE IT), and nearly every line is actionable. It is nonetheless very long for a tool description and repeats the size-set idea twice ('a set of ads is one slide with sizes' and again in HOW TO USE IT), which costs a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates: it explains that the result lists per-slide warnings (overflow, overlap, too-small text, low contrast, failed images) and returns a picture of every slide, plus the exact contrast thresholds. For a 9-parameter composition tool with destructive rebuild semantics, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is high (89%), so the schema already documents size keys, layout, slots and slide addressing. The description still adds real semantics the schema does not: layout-vs-html as an either/or decision, how `sizes` produces one arrangement per shape, how per-slide `size` composes a multi-size ad set, and the meaage that `css` is shared across slides. It leaves `accent`, `notes` and `picture` to the schema, hence not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Builds a designed deck') and then defines what a slide actually is (a layout with slots filled, or HTML/CSS laid out by a browser, becoming an editable Napkin deck). The opening line 'Use this whenever a deck should look designed' plus the layouts-vs-HTML framing cleanly separates it from sibling tooling like napkin_layouts_list and the slide-level tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing rules: start from a layout for ads/posts/carousels/title slides, write HTML only for something no layout does, read the brand first via napkin_brand_get, and it explicitly says 'You don't need napkin_slide_view afterwards'. It also defines when to pass name vs boardId vs slide vs append, which covers new/replace/append/single-slide-fix paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_deck_exportExport a deck as imagesAInspect
Exports a deck's slides as PNG files at each slide's real pixels — a square post at 1080×1080, a story at 1080×1920, a leaderboard at 728×90 — ready to upload to an ad platform or a social post. One slide comes back as a PNG; several as a .zip named by deck, slide and size. Hidden slides are left out. The file is saved to the workspace's files; give the user downloadUrl (it opens for anyone signed in to the workspace) and say it's also in their files. Slide numbers and the workspace mark aren't drawn. For a PDF or PowerPoint, the user exports from the deck's File menu.
| Name | Required | Description | Default |
|---|---|---|---|
| slides | No | Just these slides (1-based, as the deck shows them). Omit for every slide. | |
| boardId | Yes | The deck to export. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint=false, destructiveHint=false): it discloses that hidden slides are omitted, that slide numbers and the workspace mark are not drawn, that a single slide returns a PNG while multiple return a named .zip, and where the file is saved. This is exactly the extra context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and resolutions, then covers return shape, exclusions, and the post-call user instruction. It is dense but every sentence carries distinct information; the resolution list is slightly verbose but justified as concrete examples.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema export tool, the description supplies the return shape (PNG vs .zip naming), the delivery channel (workspace files plus downloadUrl), and the exact agent behavior expected (give the user the link and note it is in their files). Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents boardId, workspace, and slides including 1-based indexing and the omit-for-all default. The description reinforces the multi-vs-single-slide behavior but adds little syntactic detail beyond the schema, so it sits above the baseline 3 but not at the top.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (exports) and resource (a deck's slides as PNG images), and immediately grounds it in concrete output formats and resolutions. The final sentence explicitly redirects PDF/PowerPoint needs to the deck's File menu, distinguishing this tool from that alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it (exporting deck slides as PNGs ready for ad platforms or social posts) and when not to (PDF or PowerPoint — use the File menu instead). It also clarifies the agent-follow-up action of handing the user the downloadUrl.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_deck_outlineRead a deck's outline (markdown)ARead-onlyInspect
Returns a Napkin deck as a markdown outline — one section per slide in presentation order, with slide titles, text content (bold + bullets preserved), empty layout stubs still awaiting content, and a summary of drawn marks. This is the cheap way to know what a deck SAYS; use napkin_slide_view when you need to see how a specific slide LOOKS. Works on any board that has slides. Cite the deck with [[board:ID]] alone on its own line.
| Name | Required | Description | Default |
|---|---|---|---|
| boardId | Yes | Napkin board id (a deck). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the read-only, non-destructive, non-open-world profile, so the safety bar is covered. The description goes further by detailing the return shape (one section per slide, preserved bold/bullets, empty layout stubs, drawn-marks summary) and characterizing this as the 'cheap' read path, which is useful behavioral context in the absence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the return format, then the routing rule, then the citation note. Every sentence carries weight, though the trailing citation sentence is slightly tangential to invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema, the description fully carries the burden: it describes the return structure, scoping, and its relationship to the sibling viewer. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both boardId and workspace are fully documented in the schema, making 3 the baseline. The description adds no syntax or format detail beyond that, though the citation convention touches on deck identity rather than parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Returns a Napkin deck as a markdown outline') and enumerates the content it includes (slide titles, text, layout stubs, drawn marks). It explicitly distinguishes itself from the sibling napkin_slide_view by contrasting what the deck SAYS vs how a slide LOOKS.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit routing rule: use this to learn what a deck says, use napkin_slide_view when you need to see how a specific slide looks. Adds a scope note ('Works on any board that has slides') so the agent knows when it is applicable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_deck_set_themeSet a deck's themeADestructiveInspect
Paints every slide of a deck with a theme — slide background, heading and body typefaces, and text colors — in one call, the same as picking it in the deck's theme menu. Font sizes are kept. Prefer "brand" (the workspace's own brand kit) unless the user asks for something else. Built-in themes: plain (Helvetica on white. Gets out of the way.); editorial (Garamond headings on cream. Reads like a document.); stage (Helvetica, light on dark, for a projected room.); blueprint (Consolas headings on mist. For technical decks.); warm (Rounded on amber. Softer than it sounds.); notebook (Marker headings. Keeps the sketchbook feel.). For one-off colors or sizes on a single shape, use napkin_draw's updates instead. Check the result with napkin_slide_view.
| Name | Required | Description | Default |
|---|---|---|---|
| theme | Yes | `brand` for the workspace's brand kit, `brand:<kit id>` for a particular kit (napkin_brand_list), or a built-in theme key. | |
| accent | No | Brand themes only: the name of one of the kit's colors to use as this deck's accent instead of the kit's own — for a deck about one product or campaign that has its own color in the brand. | |
| boardId | Yes | Napkin board id (a deck). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the destructive, non-read-only profile, so the burden is lower. The description still adds real behavioral detail beyond them — 'Font sizes are kept' tells the agent exactly what is preserved when the theme is repainted, which is not derivable from the schema or annotations. It stops short of explicitly saying the previous theme styling is replaced/irreversible, so it is not a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and the alternate-tool routing are front-loaded in the first sentences; the theme catalog and verification hint follow. The six-item theme list is dense but earns its place because no enum exists in the schema. Minor padding in the flavor text of each theme name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a bulk-mutation tool with no output schema and full schema-description coverage, the definition supplies the affected scope, the preservation rule, the recommended default, the value catalog, an alternative tool, and a verification step. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by enumerating the built-in theme keys (plain, editorial, stage, blueprint, warm, notebook) with a semantic gloss for each — effectively supplying the enum that the schema lacks (0 enums declared). This meaningfully constrains the required `theme` string. The `accent` and `workspace` semantics remain schema-only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Paints every slide of a deck with a theme') and enumerates exactly what is affected — slide background, headings, body typefaces, text colors. It is clearly distinguishable from the sibling write paths (napkin_draw updates, napkin_slide_update) that it explicitly routes away from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-prefer guidance ('Prefer "brand" ... unless the user asks for something else'), names the alternative for the adjacent use case ('For one-off colors or sizes on a single shape, use napkin_draw's `updates` instead'), and prescribes a verification step ('Check the result with napkin_slide_view'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_deck_writeWrite a whole deck from a markdown outlineAInspect
Creates a deck from a markdown outline — the SAME dialect napkin_deck_outline reads back, so any deck you've read shows the format. Decks are SLIDES for presenting; a prose document ('write this up', 'draft a doc') is napkin_docs_create instead. Rules: ## Heading starts a slide (heading becomes the slide title and its biggest text block); optional trailing [layout: title | section | title-body | title-lead | two-column | three-column | comparison | statement | quote | closing | blank]; body lines fill the layout's remaining text areas in order, split into areas by a line containing only ---; - bullets are kept as bullets. Slides default to title-body (or title, when bodyless). Layout text areas you don't fill stay as visible prompts for the user. Example:
Q3 Review [layout: title]
What happened, what's next
Revenue [layout: title-body]
up 40% QoQ
churn flat
Bets [layout: two-column]
Double down on decks
Sunset legacy plans
After writing, cite the deck with [[board:ID]] alone on its own line — the user opens it from there.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Deck name. | |
| outline | Yes | The markdown outline (dialect above). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare write/non-destructive/non-open-world, so the description carries the real burden and does: it documents the full outline dialect, the layout enum, the `---` area-splitting rule, that unfilled layout areas stay as user-visible prompts, and the required `[[board:ID]]` citation after writing. It stops short of stating failure/overwrite behavior or limits, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose and sibling routing before the format spec, and every sentence does work — the rules and example are load-bearing for a format-heavy tool. It is long, but far less could not specify this dialect; only mild trimming is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param, no-output-schema writer, the definition is complete: authoring format, layout vocabulary, sibling disambiguation, and the post-write citation step that tells the user how to open the result. The remaining workspace/auth nuance is already carried by the schema's `workspace` description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description goes well beyond the schema's 'The markdown outline (dialect above)' by fully specifying the dialect, layout values, and bullet handling. It adds little for `name` and `workspace`, which the schema already covers, so it tops out at 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ('Creates a deck from a markdown outline') and immediately anchors it against two siblings: it is the writing counterpart to napkin_deck_outline and is explicitly NOT napkin_docs_create. An agent can route to it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-not with the alternative named: a prose document ('write this up', 'draft a doc') belongs to napkin_docs_create instead. It also fixes the round-trip contract with napkin_deck_outline so the agent knows this is the write side of that dialect. Nothing essential is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_diagrams_createCreate a Napkin diagramAInspect
Create a Napkin diagram — an interactive FLOWCHART (process, decision tree, system map), not a doc (napkin_docs_create) or drawing canvas (napkin_boards_create). Pass the mermaid source. Use for process maps the user asked you to draw. After creating, cite it as [[diagram:ID]] alone on its own line — it renders as a clickable card. NEVER cite it as [[board:ID]].
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| source | No | Mermaid flowchart source. Start with `flowchart TD` (top-down) or `flowchart LR` (left-right). Declare shaped nodes — `A["Rect"]`, `B("Rounded")`, `C{"Decision"}`, `D(["Stadium"])`, `E(("Circle"))`, `F{{"Hexagon"}}` — and link them with `A -->|label| B`, `A -.-> B` (dotted), or `A ==> B` (thick). Group steps with `subgraph Name["Title"] ... end`. Color a node with a trailing `style A fill:#fee2e2` line. Only the flowchart subset is supported. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnly=false, destructive=false, openWorld=false). The description adds substantial behavioral context beyond that: the output renders as an interactive clickable card, must be cited as [[diagram:ID]] on its own line, and must never be cited as [[board:ID]]. It omits any mention of auth/scoping limits, but the added rendering and citation semantics are genuinely useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the identity/scope distinction, followed by the input requirement and then output/citation behavior. No filler and no repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by specifying the citation format the agent must emit. Combined with annotations covering the safety profile and the rich 'source' schema, the definition is near-complete for a create tool; only auth/workspace scoping behavior is left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the 'source' parameter is already heavily documented in the schema (mermaid syntax, node shapes, subgraphs, styling). The description only restates 'Pass the mermaid source', adding no syntax or constraint detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a Napkin diagram — an interactive FLOWCHART') and immediately draws the boundary against the two nearest siblings by name (napkin_docs_create, napkin_boards_create). An agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternatives and what this tool is NOT ('not a doc... or drawing canvas'), then gives the positive trigger condition ('Use for process maps the user asked you to draw'). The mermaid input requirement further constrains selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_diagrams_getRead a Napkin diagramARead-onlyInspect
Fetch one Napkin diagram's mermaid source by id (from napkin_diagrams_list). Read before editing — napkin_diagrams_update replaces the whole source.
| Name | Required | Description | Default |
|---|---|---|---|
| diagramId | Yes | Diagram id from napkin_diagrams_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds return-content information ('mermaid source'), which matters because there is no output schema, plus a warning about update's whole-source replacement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the primary action front-loaded and the editing caveat second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with no output schema, the description discloses the return type and the read-before-write workflow, which is what an agent needs. The workspace/token nuance lives fully in the schema, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 2 parameters and 100% schema description coverage, the schema already documents diagramId and the workspace/token rules. The description only implies the id provenance, adding little beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch one Napkin diagram's mermaid source by id'), and pins the id's origin to the sibling napkin_diagrams_list. An agent can distinguish this from napkin_diagrams_list and napkin_diagrams_update without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing guidance ('Read before editing') and names the sibling that makes it necessary ('napkin_diagrams_update replaces the whole source'). The when-to-use and the reason are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_diagrams_listList Napkin diagramsARead-onlyInspect
Lists the workspace's Napkin diagrams (interactive flowcharts — processes, decision trees, system maps) newest first. Archived diagrams are hidden unless includeArchived; q narrows by title. Fetch one's mermaid source with napkin_diagrams_get. When referring the user to a diagram in your reply, put [[diagram:ID]] alone on its own line — it renders as a clickable card. NEVER cite a diagram as [[board:ID]]; boards are sketches.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Title contains (case-insensitive). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| includeArchived | No | Also list archived diagrams (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description still adds real behavior: default sort order (newest first), default filtering of archived items, and the output-rendering contract that `[[diagram:ID]]` alone on a line becomes a clickable card. Minor gap: no pagination or result-count behavior mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and scope, followed by filtering and routing. No redundancy; the citation-format sentence earns its place because it governs how the agent must emit output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-shape burden and does so adequately — it states what is listed, in what order, what is excluded by default, and how to retrieve per-item detail. Safety is covered by annotations, so nothing material is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (q, workspace, includeArchived) are already documented in structured data; baseline is 3. The description restates q's title-matching and includeArchived's archived-reveal behavior without adding format or syntax beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the workspace's Napkin diagrams') and defines the artifact inline as 'interactive flowcharts — processes, decision trees, system maps', which separates it from the sibling napkin_boards_list (sketches) without opening either schema. 'Newest first' adds ordering scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditionals: archived hidden unless `includeArchived`, `q` narrows by title, and routes to napkin_diagrams_get for a single diagram's mermaid source. It also states the exclusion — never cite a diagram as [[board:ID]] because boards are sketches — which is a genuine when-not rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_diagrams_updateEdit a Napkin diagramADestructiveInspect
Edit a Napkin diagram: retitle, describe, REPLACE the whole mermaid source, set visibility, or archived (true takes it out of the gallery, false brings it back — the recovery move for a diagram created by mistake). Read the diagram first (napkin_diagrams_get) so your new source keeps the parts the user wants. Diagram ids come from napkin_diagrams_list.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| source | No | Mermaid flowchart source. Start with `flowchart TD` (top-down) or `flowchart LR` (left-right). Declare shaped nodes — `A["Rect"]`, `B("Rounded")`, `C{"Decision"}`, `D(["Stadium"])`, `E(("Circle"))`, `F{{"Hexagon"}}` — and link them with `A -->|label| B`, `A -.-> B` (dotted), or `A ==> B` (thick). Group steps with `subgraph Name["Title"] ... end`. Color a node with a trailing `style A fill:#fee2e2` line. Only the flowchart subset is supported. | |
| archived | No | ||
| diagramId | Yes | Diagram id from napkin_diagrams_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | Short gallery summary. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, and the description earns credit by explaining the nature of that destruction: 'REPLACE the whole mermaid source' signals a full overwrite rather than a merge, and archived true/false is described as the undo path for a mistaken creation. The approvalId round-trip is only hinted at via the schema, not elaborated here, keeping this from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the enumeration of editable fields and then attaches the read-before-edit prerequisite and id source, with no filler sentences. Dense but every clause carries a distinct instruction; minor markdown/backtick noise slightly hurts readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, destructive mutation tool with no output schema and no annotations beyond the safety hints, it covers the replacement semantics, the recovery move, and where ids come from. The approval/confirmation flow and return behavior are left partly implicit, but nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most parameters carry their own descriptions. The description adds real meaning for `source` (whole-source replacement) and `archived` (gallery removal/recovery), but leaves title, description, workspace, and approvalId to the schema — baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edit a Napkin diagram') and enumerates the exact editable fields (retitle, describe, source, visibility, archived). It also names the sibling tools that supply the diagram id and current state (napkin_diagrams_get, napkin_diagrams_list), so it is distinguishable without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete workflow prerequisite ('Read the diagram first (napkin_diagrams_get) so your new source keeps the parts the user wants') and explains the recovery scenario for archived=true/false. It stops short of explicitly contrasting update against napkin_diagrams_create, so it is clear context rather than a full when/when-not map.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_docs_createCreate a Napkin docAInspect
Create a Napkin doc — a markdown DOCUMENT (prose, structure, code blocks), not a drawing canvas (that's napkin_boards_create) and not slides (napkin_deck_write). Optionally pass a title and initial markdown body. Use for drafts the user asked you to start. After creating, cite it as [[doc:ID]] alone on its own line — it renders as a clickable card. NEVER cite it as [[board:ID]].
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Initial markdown body. Write each paragraph as ONE long line — the editor wraps text to the reader's column, and hard-wrapped source shows up as narrow ragged lines when someone opens the paragraph to edit it. | |
| title | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description adds genuinely non-structured behavior: the post-creation citation contract ([[doc:ID]] on its own line, renders as a card, never [[board:ID]]). It does not mention auth/workspace constraints, but the schema carries those, so the remaining gap is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the artifact type, then the disambiguation, then the citation rule. Every sentence carries decision-relevant information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with no output schema, the description covers the type, the optional inputs, and the required follow-up citation format — enough to call it correctly. It stops short of stating the returned identifier shape explicitly (the [[doc:ID]] hint implies it), a small remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%. The body parameter is richly documented in the schema itself (single-long-line markdown guidance), and the description only says it is optional. The title parameter has no schema description at all, and the description only adds 'optionally pass a title.' Marginal added meaning over structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a Napkin doc — a markdown DOCUMENT') and immediately disambiguates from the two most confusable siblings by name: napkin_boards_create and napkin_deck_write. An agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage condition ('Use for drafts the user asked you to start') and explicit alternatives for the neighboring artifact types (drawing canvas -> napkin_boards_create, slides -> napkin_deck_write). This is when-to-use plus when-not, which is the top of the scale.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_docs_getRead a Napkin docARead-onlyInspect
Fetch one Napkin doc's full markdown body by id (from napkin_docs_list). Read before editing — napkin_docs_update replaces or appends against the current body.
| Name | Required | Description | Default |
|---|---|---|---|
| docId | Yes | Doc id from napkin_docs_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description still adds useful behavior beyond them: the response is the full markdown body, and the read-before-edit relationship to napkin_docs_update's replace/append semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, then the workflow caveat. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple id-based read with fully documented schema params and no output schema, the description supplies the needed return-format hint (full markdown body) and the editing workflow context. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both params are documented there, including the workspace/token-default nuance, so baseline 3 applies. The description only reinforces docId provenance ('from napkin_docs_list') and adds nothing about the workspace parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch), resource (one Napkin doc), scope (full markdown body) and the lookup key (id). This clearly separates it from napkin_docs_list (enumerates) and napkin_docs_update (mutates), so an agent can select it without inspecting siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit workflow condition — 'Read before editing' — and names napkin_docs_update as the reason, plus points to napkin_docs_list as the id source. It stops short of stating any exclusion or fallback case, so it is strong context rather than full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_docs_listList Napkin docsARead-onlyInspect
Lists the workspace's Napkin docs (quick markdown documents — drafts, notes, working text; the LINEAR surface, distinct from boards which are drawing canvases) newest first, with excerpts. Archived docs are hidden unless includeArchived; q narrows by title. Fetch one with napkin_docs_get. When referring the user to a doc in your reply, put [[doc:ID]] alone on its own line — it renders as a clickable card. NEVER cite a doc as [[board:ID]]; boards are sketches.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Title contains (case-insensitive). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| includeArchived | No | Also list archived docs (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read (readOnlyHint=true, destructiveHint=false), and the description still adds real behavior: default ordering (newest first), the return shape (excerpts), the default filtering of archived docs, and the output contract for citing docs ([[doc:ID]] on its own line, never [[board:ID]]). That is substantive disclosure beyond what the structured fields provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose and scope, then the filtering behavior, then the citation contract, with no filler sentences. The closing board guardrail ('NEVER cite a doc as [[board:ID]]; boards are sketches') partly echoes the earlier board distinction, but it functions as a distinct output-format constraint rather than pure repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of describing returns, and it does so (newest first, with excerpts) while also covering default filtering and the ID-citation format. For a three-parameter read-only list tool, nothing an agent needs to call or report results correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents q as case-insensitive title matching, workspace auth semantics, and includeArchived's default false. The description largely restates these ('q narrows by title', 'Archived docs are hidden unless includeArchived'), so it adds framing but no new syntax or meaning, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the workspace's Napkin docs') and explicitly delineates the resource from its closest sibling surface ('the LINEAR surface, distinct from boards which are drawing canvases'). It also distinguishes listing from retrieval by pointing to napkin_docs_get, so an agent can separate it from both napkin_boards_list and napkin_docs_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: use this to enumerate docs newest-first with excerpts, then 'Fetch one with napkin_docs_get' for a single doc, and names the conditions that change results (includeArchived, q). It does not, however, address when to prefer this over search-style siblings like search_docs or list_docs, so the routing guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_docs_updateEdit a Napkin docADestructiveInspect
Edit a Napkin doc. For a change to existing prose use edits — a list of exact find/replace pairs, which is the SAFEST option and the one to reach for by default: it leaves everything you did not target untouched, and it fails loudly rather than clobbering a doc somebody else is typing in. append adds a section at the end. body replaces the WHOLE document and should be a last resort. Also retitle, describe, set visibility, or archived (true takes it out of the gallery, false brings it back). Read the doc with napkin_docs_get first — edits match the current text exactly. Doc ids come from napkin_docs_list.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Full replacement markdown body. Prefer `edits` — this overwrites everything. Write each paragraph as ONE long line; the editor wraps it for the reader. | |
| docId | Yes | Doc id from napkin_docs_list. | |
| edits | No | Targeted find/replace pairs, applied in order. All must match or none are applied. | |
| title | No | ||
| append | No | Markdown appended to the end (own paragraph). | |
| archived | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | Short gallery summary. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description backs this with real consequences: edits 'fail loudly rather than clobbering a doc somebody else is typing in,' body 'replaces the WHOLE document,' archived toggles gallery visibility. It doesn't cover failure modes like concurrent-edit conflicts or the needs_confirmation/approvalId flow its annotations imply, but the destructive-risk context is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the default recommendation (edits) before the riskier alternatives, and each mode gets a compact clause. Slightly dense in one paragraph, but every sentence either routes a decision or warns of a consequence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the mutation modes, the ordering constraint for edits ('applied in order. All must match or none are applied'), and the prerequisite read. For a 10-param destructive tool with no output schema, this is adequate, though it omits the approvalId/needs_confirmation path and workspace-token nuances the schema describes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema carries most param detail. The description still adds value by naming which param does what ('`append` adds a section at the end', archived semantics), but it doesn't explain workspace/visibility/approvalId quirks beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (edit) and resource (Napkin doc), then enumerates the exact mutation modes (edits, append, body, retitle/describe/visibility/archive). An agent can distinguish this from napkin_docs_create and napkin_docs_get immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit prioritization among alternatives: `edits` is 'the SAFEST option and the one to reach for by default,' `body` is 'a last resort.' It also mandates reading the doc first via napkin_docs_get, which is a concrete precondition no other tool provides.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_drawDraw on a Napkin boardADestructiveInspect
Draws shapes, strokes, and labels on a board — diagram on a sketch, annotate a deck slide — and, via updates / deletes, changes or removes shapes that are already there (ids from napkin_boards_get; the fix for something you drew wrong). For NEW elements never compute absolute board coordinates: with slide, coordinates are slide-local — (0, 0) top-left to (960, 540) bottom-right of that slide. Without slide (sketches), (0, 0) is the top-left of the existing content — the same area napkin_boards_view renders, so place by what you saw there; on an empty board just start at (0, 0). Example — circle a slide's title and margin-note it:
{"boardId": "…", "slide": 2, "elements": [ {"type": "ellipse", "x": 60, "y": 40, "w": 400, "h": 90, "color": "#dc2626"}, {"type": "arrow", "x": 560, "y": 140, "x2": 470, "y2": 90, "color": "#dc2626"}, {"type": "text", "x": 575, "y": 130, "w": 260, "text": "tighten this claim", "color": "#dc2626"} ]}
After drawing, cite the board with [[board:ID]] alone on its own line.
| Name | Required | Description | Default |
|---|---|---|---|
| slide | No | 1-based slide number (decks only) — makes coordinates slide-local. | |
| boardId | Yes | Napkin board id. | |
| deletes | No | Ids of existing shapes to remove (from napkin_boards_get). | |
| updates | No | Existing shapes to change — move, resize, retext, restyle. Read napkin_boards_get first for ids and current geometry. | |
| elements | No | New elements to draw, painted in order. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and readOnlyHint=false; the description is consistent and adds real context on top — that deletes are the fix for a mis-drawn element, that ids must be sourced from napkin_boards_get, and how coordinate frames differ for slides vs sketches. It stops short of stating whether deletions are reversible or how many ops can be batched atomically.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the draw/create capability, then mutation, then the coordinate rules, then a compact example. Dense but each part carries information; the trailing citation instruction is short and actionable. Slightly heavy for a single description, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no output schema, it covers the two things an agent would most likely get wrong — where to get shape ids and which coordinate space applies. It omits batching limits (maxItems 100) and the workspace-slug requirement noted in the schema, so slightly incomplete rather than fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents parameters and the baseline is 3. The description goes beyond it with origin/coordinate-frame semantics (slide-local 0,0→960,540; board-relative for sketches), a worked example of an ellipse+arrow+text cluster, and the rule that ids come from napkin_boards_get.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs and resource: draws shapes/strokes/labels on a board, and via updates/deletes changes or removes existing shapes. Explicitly distinguishes itself from the read-side siblings napkin_boards_get and napkin_boards_view, so an agent can tell what this tool owns versus its neighbors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage contexts ('diagram on a sketch, annotate a deck slide') and routes the agent to napkin_boards_get for ids and napkin_boards_view for the coordinate frame the renderer uses. It does not, however, explicitly rule out nearby draw-capable siblings such as napkin_diagrams_create or napkin_deck_compose, leaving some alternative-selection inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_interfaces_createCreate a Napkin interfaceAInspect
Create an interface — a small screen of its own the user can open and send to people. Use when someone asks for a custom chat UI, a little tool, or a prototype of how something would feel in their product. It starts with a working index.html you then edit. It can reach NOTHING in the workspace until a flow is granted with napkin_interfaces_grant. Cite it in your reply as [[interface:ID]] alone on its own line — it renders as a card the user can click. Do NOT paste an image URL or try to embed the picture yourself; that renders as a broken image. An interface is plain HTML, CSS and JavaScript in one or more files, served from its own origin. There is NO build step and NO library: no React, no Tailwind, no CDN. A to anywhere but this origin is blocked, so write vanilla JS and put styles in a block. index.html is the entry point and must exist. THE BRAND: every interface carries the workspace brand as files — link it with and style with its variables: var(--brand-ink), var(--brand-muted), var(--brand-background), var(--brand-surface), var(--brand-accent), var(--brand-accent-2), var(--brand-color-), var(--brand-font-heading), var(--brand-font-body), var(--brand-radius). Never hard-code the brand's colours or font names. brand.css already loads the brand's own font files; web fonts (Google Fonts or any other) are BLOCKED, so linking one leaves the page in a system font. The brand's logos are files under brand/ (brand.css lists them at the top): . Use the logo rather than typing the name, and real icons (inline SVG) rather than emoji. Read the brand with napkin_brand_get first, and say only what the brand kit or the user tells you. PICTURES: a picture from anywhere on the web is BLOCKED and draws as a broken image, so never use an external image URL. Show a picture from the workspace's files with (in CSS, url(file:)) — PNG, JPEG, GIF or WebP, up to 5 MB. Write the reference literally; a file id assembled in JavaScript is not found. When the user gives you a picture in chat, save it with workspace_files_save_from_chat and use the file id it returns. A picture that's only on a website has to be attached in chat first, or uploaded in the interface's Files panel. LOOK before you hand it over: napkin_interfaces_view draws the page as it is now. Fix anything that doesn't look like the brand and look again. Load the client with . Then zw.ready() resolves with { viewer, flows }, and zw.flows.run(flowId, input) runs a flow the interface was granted. Input is {kind:"chat", messages:[{role:"user", content:"…"}]} or {kind:"form", values:{…}} — the same shapes the public API takes. zw.replyText(result) pulls the assistant text out of a chat result. A SHIM is not a flow and takes a different call: zw.shims.run(shimId, text), which decides on the viewer's own device with no network and no cost. It resolves with the same envelope a flow does, so read the decision with zw.decision(result) — NOT result.decision, which is undefined and makes an interface show one answer for every input. The decision is { answer, confidence, familiarity, action, probs, gates }; branch on action, the shim's own call about whether it was sure enough. Calling zw.flows.run with a shim id is refused. ctx.grants tells you which you have: each entry carries a kind of "flow" or "shim" alongside its id and name. Running a shim on every keystroke is fine — it costs nothing and there is no rate limit. Debounce ~150ms and COALESCE: remember the latest text and run it when the current call finishes, so the answer matches what is on screen. Clear any in-flight guard on failure as well as success, or one call that doesn't come back wedges the interface. d.action is "act" | "suggest" | "refuse" — those three strings, nothing else. Branch on it rather than on a confidence threshold you invent; "refuse" means the shim doesn't recognise the input well enough to answer, so say so rather than showing a low-confidence guess. If you run a requestAnimationFrame loop, remove CSS transitions from any property it writes — the two fight and the property looks frozen. A flow answers in markdown: render it with el.replaceChildren(zw.markdown(text)). It builds DOM nodes, so model output is never treated as markup. STREAM chat answers: zw.flows.run(id, input, { onEvent: fn }) delivers the text token by token, and a flow that takes several seconds reads as broken without it. Use zw.textDelta(event) for each chunk — it returns null for anything that isn't text, so pass it every event — and accumulate. The promise still settles at the end with the whole result; take the final text from there. onEvent also sees node_start / node_complete / node_error if you want to name the step. Don't name a top-level variable history, name, status, length, origin or top: those are already window properties, so var history = [] leaves you with the browser's History object and history.push fails. Prefix it, or keep it inside a function. Style it plainly and legibly: a system font stack, generous spacing, one column unless there's a reason. It runs on phones too.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| starter | No | Which worked example to start from. `chat` is a conversation against one flow; `form` is labelled fields sent as a form run; `decision` is text in and a shim's answer out. `blank` (the default) is a title and the client library — take it when none of the others is the shape you want, rather than deleting one you didn't need. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only declaring the generic safe-mutation profile, the description carries real behavioral weight: the interface starts from a working index.html, reaches nothing in the workspace until a flow is granted, has no build step or libraries, blocks cross-origin scripts/webfonts/external images, returns an ID to cite as [[interface:ID]], and requires a view pass before delivery. These are concrete operational traits an agent could not infer from readOnlyHint/destructiveHint/openWorldHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and when-to-use lead, so it is front-loaded, and each sentence is individually useful. But it is a ~700-word unstructured block that blends creation guidance with client-library API reference (shim envelopes, streaming, rAF conflicts, window-name collisions) that would be more appropriate elsewhere, imposing significant reading cost on tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with no output schema, the description covers everything an agent needs: what gets produced, the starter files, the grant/brand/view workflow, how the result is cited, and which sibling tools to call next. Nothing material about invoking or using the created artifact is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, with the starter enum thoroughly documented in-schema and workspace explained in-schema; only `name` is bare. The description adds nothing about the three parameters themselves (notably the meaningful blank/chat/form/decision choice), so it does not go beyond structured data. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource plus a plain-language gloss ('Create an interface — a small screen of its own the user can open and send to people'), and the surrounding text implicitly separates this from napkin_interfaces_write (edits after creation), napkin_interfaces_grant (permissions), and napkin_interfaces_view (rendering). An agent knows exactly what artifact it is producing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the triggering requests ('a custom chat UI, a little tool, or a prototype of how something would feel in their product') and routes the agent to prerequisites and follow-ups (napkin_brand_get first, napkin_interfaces_grant for access, napkin_interfaces_view before handing over). It never states when NOT to use it or contrasts with the sibling napkin_interfaces_write/create-vs-edit boundary explicitly, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_interfaces_getRead a Napkin interfaceARead-onlyInspect
Read one interface: its name, the flows it may run (with the published version each is pinned to), and the contents of its files. READ BEFORE EDITING — napkin_interfaces_write matches text exactly against what is there now. Pass a path to read one file when the interface has several. An interface is plain HTML, CSS and JavaScript in one or more files, served from its own origin. There is NO build step and NO library: no React, no Tailwind, no CDN. A to anywhere but this origin is blocked, so write vanilla JS and put styles in a block. index.html is the entry point and must exist. THE BRAND: every interface carries the workspace brand as files — link it with and style with its variables: var(--brand-ink), var(--brand-muted), var(--brand-background), var(--brand-surface), var(--brand-accent), var(--brand-accent-2), var(--brand-color-), var(--brand-font-heading), var(--brand-font-body), var(--brand-radius). Never hard-code the brand's colours or font names. brand.css already loads the brand's own font files; web fonts (Google Fonts or any other) are BLOCKED, so linking one leaves the page in a system font. The brand's logos are files under brand/ (brand.css lists them at the top): . Use the logo rather than typing the name, and real icons (inline SVG) rather than emoji. Read the brand with napkin_brand_get first, and say only what the brand kit or the user tells you. PICTURES: a picture from anywhere on the web is BLOCKED and draws as a broken image, so never use an external image URL. Show a picture from the workspace's files with (in CSS, url(file:)) — PNG, JPEG, GIF or WebP, up to 5 MB. Write the reference literally; a file id assembled in JavaScript is not found. When the user gives you a picture in chat, save it with workspace_files_save_from_chat and use the file id it returns. A picture that's only on a website has to be attached in chat first, or uploaded in the interface's Files panel. LOOK before you hand it over: napkin_interfaces_view draws the page as it is now. Fix anything that doesn't look like the brand and look again. Load the client with . Then zw.ready() resolves with { viewer, flows }, and zw.flows.run(flowId, input) runs a flow the interface was granted. Input is {kind:"chat", messages:[{role:"user", content:"…"}]} or {kind:"form", values:{…}} — the same shapes the public API takes. zw.replyText(result) pulls the assistant text out of a chat result. A SHIM is not a flow and takes a different call: zw.shims.run(shimId, text), which decides on the viewer's own device with no network and no cost. It resolves with the same envelope a flow does, so read the decision with zw.decision(result) — NOT result.decision, which is undefined and makes an interface show one answer for every input. The decision is { answer, confidence, familiarity, action, probs, gates }; branch on action, the shim's own call about whether it was sure enough. Calling zw.flows.run with a shim id is refused. ctx.grants tells you which you have: each entry carries a kind of "flow" or "shim" alongside its id and name. Running a shim on every keystroke is fine — it costs nothing and there is no rate limit. Debounce ~150ms and COALESCE: remember the latest text and run it when the current call finishes, so the answer matches what is on screen. Clear any in-flight guard on failure as well as success, or one call that doesn't come back wedges the interface. d.action is "act" | "suggest" | "refuse" — those three strings, nothing else. Branch on it rather than on a confidence threshold you invent; "refuse" means the shim doesn't recognise the input well enough to answer, so say so rather than showing a low-confidence guess. If you run a requestAnimationFrame loop, remove CSS transitions from any property it writes — the two fight and the property looks frozen. A flow answers in markdown: render it with el.replaceChildren(zw.markdown(text)). It builds DOM nodes, so model output is never treated as markup. STREAM chat answers: zw.flows.run(id, input, { onEvent: fn }) delivers the text token by token, and a flow that takes several seconds reads as broken without it. Use zw.textDelta(event) for each chunk — it returns null for anything that isn't text, so pass it every event — and accumulate. The promise still settles at the end with the whole result; take the final text from there. onEvent also sees node_start / node_complete / node_error if you want to name the step. Don't name a top-level variable history, name, status, length, origin or top: those are already window properties, so var history = [] leaves you with the browser's History object and history.push fails. Prefix it, or keep it inside a function. Style it plainly and legibly: a system font stack, generous spacing, one column unless there's a reason. It runs on phones too.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | One file to read, e.g. index.html. Omit for every file. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| interfaceId | Yes | Interface id from napkin_interfaces_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered without the description. The description does add real behavioral value by spelling out the read payload (files, flow grants and their pinned versions) and the single-file path mode, but the bulk of its behavioral text concerns authoring, rendering, shims and streaming — traits of sibling tools, not of this read call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, which is good. But the description then runs several hundred words on brand variables, blocked CDNs, image sources, shim decision envelopes, streaming callbacks and reserved global variable names — almost none of which can be acted on when reading an interface. Most sentences do not earn their place in this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the job of explaining what comes back, and it does so concretely (name, granted flows with pinned versions, file contents, single-file mode). For a three-parameter read tool that is sufficient; the only shortfall is that the excess authoring guidance dilutes rather than completes the picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (interfaceId, path, workspace) are already documented in the schema, including 'Omit for every file'. The description's 'Pass a path to read one file' merely restates what the schema's path description already says, adding no syntax, format or edge-case detail. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource and enumerates exactly what is returned: the name, the flows it may run with the published version each is pinned to, and the contents of its files. It also distinguishes itself from the sibling it overlaps with ('READ BEFORE EDITING — napkin_interfaces_write matches text exactly against what is there now'), so an agent can separate the read from the write without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is clearly framed as a prerequisite step ('READ BEFORE EDITING') and it names the alternative it pairs with, plus a per-call scope rule ('Pass a path to read one file when the interface has several'). It does not, however, state when NOT to reach for this tool — e.g. that napkin_interfaces_list is what you call to enumerate interfaces rather than one-by-one reads.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_interfaces_grantChange what a Napkin interface can reachADestructiveInspect
Set what an interface may run — flows and shims, each pinned to a published version. This is the interface's ONLY reach into the workspace, and it replaces the whole list, so include everything it should keep. Ids and published versions come from workbench_flows_list / workbench_shim_list and their revisions. Ask the user before widening this.
| Name | Required | Description | Default |
|---|---|---|---|
| grants | Yes | The complete list. Passing an empty list removes all access. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| interfaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations flag destructive/openWorld, and the description adds crucial detail beyond them: this replaces the entire list, an empty list removes all access, and widening requires user consent. These are exactly the behaviors an agent must know before calling a destructive setter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and scope, then the replacement warning, then the consent rule. No padding or restated field names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, no-output mutation tool this covers purpose, full-replace semantics, empty-list behavior, id sourcing, and consent. It omits what happens to in-flight runs or whether a prior state can be recovered, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and already documents grants/version/workspace/approvalId; the description adds sourcing guidance (where targetIds and published versions come from) that the schema does not state, which is meaningful beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — setting what an interface may run by granting flows and shims pinned to published versions. The scope phrase 'the interface's ONLY reach into the workspace' pins down exactly what this tool governs, distinct from napkin_interfaces_write/publish siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the key when-to-use context (replace the whole grant list, include everything to keep) plus where the required ids/versions come from (workbench_flows_list / workbench_shim_list and revisions), and a caution to ask the user before widening. It does not name a sibling alternative or a when-not-to-use case, keeping it below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_interfaces_listList Napkin interfacesBRead-onlyInspect
List the interfaces in a workspace — small self-contained screens (a custom chat, a control panel, a prototype) that run on their own origin and can call the flows they've been granted. Returns ids, names and how many flows each may run.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real value by explaining what the returned entities are and that results include ids, names and per-interface flow counts. It stops short of describing pagination, ordering, or how archived items affect results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded sentences with a definition and a return summary; every clause earns its place and nothing is padded. Minor cost is that the parenthetical examples lengthen the definition without aiding selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list with full annotation coverage, the description is nearly sufficient: it defines the resource and summarizes the return shape, which matters since there is no output schema. The one material hole is the undocumented includeArchived parameter, which the description should have covered given the schema gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: the workspace parameter is documented in the schema, but includeArchived carries no description anywhere. The prose mentions neither parameter, so it does not compensate for the gap — an agent gets no signal on what includeArchived does or whether workspace is optional.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (interfaces in a workspace) and even defines what an interface is — a self-contained screen running on its own origin — which is genuinely clarifying. It does not, however, distinguish this from sibling reads like napkin_interfaces_get or napkin_interfaces_view, so an agent still infers the list-vs-single distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no statement of preconditions, and no reference to the sibling retrieval tools (get/view) that an agent would need to choose between. Usage is only weakly implied by the verb 'List'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_interfaces_publishSave a version of a Napkin interfaceAInspect
Save where an interface is now as a numbered version. Share links point at a published version, so what someone was sent doesn't change while work continues. Do this when the user says it's ready to show people, not after every edit.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | One line on what changed. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| interfaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a non-destructive but mutating write (readOnlyHint=false, destructiveHint=false). The description adds genuine behavioral context beyond that: published versions are what share links point at, so a sent link is frozen while work continues. It does not cover permission/auth requirements or what the returned version looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action and then the rationale plus usage condition. No filler, no repetition of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity 3-parameter tool with no output schema and annotations covering the safety profile, the description supplies the conceptual model (versioned publish, frozen share links) and usage timing. The only real gap is that the required interfaceId is never referenced anywhere in prose, but the schema's required list covers it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: note and workspace are documented in the schema, but the required interfaceId has no description, and the description itself never mentions any parameter. With the schema already carrying most of the param meaning, the baseline of 3 applies and the description adds nothing to compensate for the undocumented required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Save where an interface is now as a numbered version" states a specific verb (save/publish) and a specific resource outcome (a numbered interface version), which is a meaningfully different operation from siblings like napkin_interfaces_write or napkin_interfaces_create. It does not name any sibling explicitly, so an agent must infer the boundary itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit trigger ("when the user says it's ready to show people") and an explicit exclusion ("not after every edit"), which is strong when-to-use guidance. It stops short of naming the alternative tool to use for ordinary edits, leaving that inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_interfaces_viewLook at a Napkin interfaceARead-onlyInspect
See what an interface LOOKS like right now: the whole page, drawn from its current files at desktop width under the same rules as the live page (no web fonts, no remote pictures). Use it after every write and before saying an interface is finished: reading the code tells you what you wrote, not what it renders. Look for text that's too small or runs together, stray system fonts, emoji standing in for icons, and anything that doesn't look like the brand.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| interfaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read-only, non-destructive operation. The description adds real behavioral context beyond that: rendering is desktop width only, uses the live page's rules, and excludes web fonts and remote pictures. It stops short of describing the return artifact (image vs. HTML), which would matter most for a view tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then usage timing, then a concrete inspection checklist. Three sentences with no filler, though the 'look for' list is slightly long relative to the rest.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of saying what comes back, and it never states the return format (screenshot, image URL, or markup). It covers rendering rules and the inspection intent well, but leaves that output-shape gap and the undocumented interfaceId for an agent to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, so the description is expected to compensate, and it says nothing about either parameter. interfaceId is left entirely undocumented and workspace constraints are not restated from the schema. The 50% gap is not closed by the prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (see/look at) and resource (a Napkin interface), and immediately distinguishes itself from reading code by contrasting render versus source. The sibling set includes interfaces_get/write/list/get, and 'see what an interface LOOKS like right now' clearly separates this render/view tool from those.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing guidance: 'Use it after every write and before saying an interface is finished.' That is a clear when-to-use rule. It does not name sibling alternatives (e.g., interfaces_get for metadata) or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_interfaces_writeEdit a Napkin interface's filesADestructiveInspect
Write one file of an interface. Use edits — exact find/replace pairs — which is the SAFEST option and the one to reach for by default: it leaves everything you did not target untouched, and it fails loudly rather than clobbering a file somebody is editing. body replaces the WHOLE file and is for a new file or a deliberate rewrite. deleteFile removes one (index.html can't be removed). Read with napkin_interfaces_get first — edits match the current text exactly. Also rename, describe, set visibility, or archive. An interface is plain HTML, CSS and JavaScript in one or more files, served from its own origin. There is NO build step and NO library: no React, no Tailwind, no CDN. A to anywhere but this origin is blocked, so write vanilla JS and put styles in a block. index.html is the entry point and must exist. THE BRAND: every interface carries the workspace brand as files — link it with and style with its variables: var(--brand-ink), var(--brand-muted), var(--brand-background), var(--brand-surface), var(--brand-accent), var(--brand-accent-2), var(--brand-color-), var(--brand-font-heading), var(--brand-font-body), var(--brand-radius). Never hard-code the brand's colours or font names. brand.css already loads the brand's own font files; web fonts (Google Fonts or any other) are BLOCKED, so linking one leaves the page in a system font. The brand's logos are files under brand/ (brand.css lists them at the top): . Use the logo rather than typing the name, and real icons (inline SVG) rather than emoji. Read the brand with napkin_brand_get first, and say only what the brand kit or the user tells you. PICTURES: a picture from anywhere on the web is BLOCKED and draws as a broken image, so never use an external image URL. Show a picture from the workspace's files with (in CSS, url(file:)) — PNG, JPEG, GIF or WebP, up to 5 MB. Write the reference literally; a file id assembled in JavaScript is not found. When the user gives you a picture in chat, save it with workspace_files_save_from_chat and use the file id it returns. A picture that's only on a website has to be attached in chat first, or uploaded in the interface's Files panel. LOOK before you hand it over: napkin_interfaces_view draws the page as it is now. Fix anything that doesn't look like the brand and look again. Load the client with . Then zw.ready() resolves with { viewer, flows }, and zw.flows.run(flowId, input) runs a flow the interface was granted. Input is {kind:"chat", messages:[{role:"user", content:"…"}]} or {kind:"form", values:{…}} — the same shapes the public API takes. zw.replyText(result) pulls the assistant text out of a chat result. A SHIM is not a flow and takes a different call: zw.shims.run(shimId, text), which decides on the viewer's own device with no network and no cost. It resolves with the same envelope a flow does, so read the decision with zw.decision(result) — NOT result.decision, which is undefined and makes an interface show one answer for every input. The decision is { answer, confidence, familiarity, action, probs, gates }; branch on action, the shim's own call about whether it was sure enough. Calling zw.flows.run with a shim id is refused. ctx.grants tells you which you have: each entry carries a kind of "flow" or "shim" alongside its id and name. Running a shim on every keystroke is fine — it costs nothing and there is no rate limit. Debounce ~150ms and COALESCE: remember the latest text and run it when the current call finishes, so the answer matches what is on screen. Clear any in-flight guard on failure as well as success, or one call that doesn't come back wedges the interface. d.action is "act" | "suggest" | "refuse" — those three strings, nothing else. Branch on it rather than on a confidence threshold you invent; "refuse" means the shim doesn't recognise the input well enough to answer, so say so rather than showing a low-confidence guess. If you run a requestAnimationFrame loop, remove CSS transitions from any property it writes — the two fight and the property looks frozen. A flow answers in markdown: render it with el.replaceChildren(zw.markdown(text)). It builds DOM nodes, so model output is never treated as markup. STREAM chat answers: zw.flows.run(id, input, { onEvent: fn }) delivers the text token by token, and a flow that takes several seconds reads as broken without it. Use zw.textDelta(event) for each chunk — it returns null for anything that isn't text, so pass it every event — and accumulate. The promise still settles at the end with the whole result; take the final text from there. onEvent also sees node_start / node_complete / node_error if you want to name the step. Don't name a top-level variable history, name, status, length, origin or top: those are already window properties, so var history = [] leaves you with the browser's History object and history.push fails. Prefix it, or keep it inside a function. Style it plainly and legibly: a system font stack, generous spacing, one column unless there's a reason. It runs on phones too.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Full replacement contents for `path`. Creates the file when it doesn't exist. Prefer `edits` on a file that already has something in it. | |
| name | No | ||
| path | No | Which file to write. Defaults to index.html. Use lowercase names like app.js or style.css. | |
| edits | No | Targeted find/replace pairs, applied in order. All must match or none are applied. | |
| archived | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| deleteFile | No | Remove `path` from the interface. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | ||
| interfaceId | Yes | Interface id from napkin_interfaces_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag destructive=true; the description adds substantial operational context beyond that: edits fail loudly rather than clobbering, are atomic, preserve untargeted content, and index.html is undeletable. It further discloses the environment constraints (no build step, no libraries, blocked external scripts/fonts/images, 5 MB image cap) that determine whether a write will actually work.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The write mechanics are correctly front-loaded, but the description runs far past what the tool does, embedding a full authoring manual (brand variables, shim/flow runtime semantics, streaming, even top-level variable naming). A large share of the text is reference material an agent only needs after the call succeeds, which dilutes scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with no output schema, the description leaves nothing critical unstated: it covers mode selection, atomicity, the file/asset environment, brand integration, and where to read current state before writing. An agent has everything needed to call it correctly on the first attempt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 70% and the description reinforces the key semantics: edits are exact find/replace against current text, body replaces the whole file and can create it, deleteFile removes path. It adds the dual-mode framing the schema only implies, though parameters like name, description, archived and workspace get no description-side elaboration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Front sentence 'Write one file of an interface' gives a specific verb and resource, then immediately distinguishes the three write modes (edits, body, deleteFile). It also names the sibling to read first (napkin_interfaces_get), so an agent can place it in the workflow without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit default guidance ('use `edits` ... the one to reach for by default'), a named exception for `body` ('a new file or a deliberate rewrite'), and a hard exclusion ('index.html can't be removed'). It also states the prerequisite (read with napkin_interfaces_get first) and warns edits match current text exactly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_layouts_listList slide layoutsARead-onlyInspect
Lists the slide layouts napkin_deck_compose can build from — the brand kit's own first, then the built-in ones — with when to use each and the slots it takes (needs must be filled; takes are optional). Every layout adapts to every size, fits its text, and uses the brand's colors, faces and logos; the logo fills itself in. Pick by what the slide has to say: one line (statement), a number (stat), a quote, a list (points, steps), an event, a carousel (cover, inner, closing), a picture (image-top, image-full, split, corner, product-shot), or a display ad (banner).
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/non-destructive, and the description adds real behavioral context: ordering (brand's own first, then built-ins), the needs-vs-takes slot contract, that layouts adapt to every size, auto-fit text, brand color/face/logo application, and self-filling logo. It stops short of describing output shape, but for a read-only list tool this is rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool returns and the needs/takes distinction, then a scannable enumeration of layout families. Dense but every clause adds information; the long single paragraph is slightly heavy but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must carry return-value burden, and it does explain what each layout entry conveys (when to use, needs, takes) and the adaptive behavior. The layout-family list is illustrative rather than exhaustive, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single workspace parameter, so the schema already carries the semantics. The description adds nothing about the workspace argument, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('lists the slide layouts napkin_deck_compose can build from') and clarifies scope (brand kit first, then built-in). It is clearly distinguishable from siblings like napkin_deck_compose or napkin_slide_add, which consume rather than enumerate layouts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit selection guidance ('Pick by what the slide has to say') with a concrete mapping of content type to layout family. It implies this is a lookup step before composing, though it never explicitly states when NOT to call it or names an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_add_chartPin a chart onto a sheetAInspect
Adds a LIVE chart to a sheet — its SQL re-runs against the tabs on every view, so it never goes stale. Write the query with napkin_sheets_query first to confirm the shape (first column = x axis, numeric columns = series, or set x/series explicitly). Prefer aggregated queries (GROUP BY) — a chart of raw rows is rarely the answer.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| sql | Yes | The SELECT to chart. | |
| type | Yes | ||
| title | Yes | Chart title. | |
| series | No | ||
| sheetId | Yes | Sheet id from napkin_sheets_list. | |
| stacked | No | Stack the series (bar/area only). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, destructiveHint=false) establish it as a non-destructive write; the description adds the critical behavioral fact that the chart is LIVE and its SQL re-runs on every view. It omits idempotency behavior (what happens when adding a duplicate chart) and error behavior on bad SQL, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the defining trait (LIVE chart), then the workflow prerequisite, then the best practice. No filler and every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter write tool with no output schema and 63% coverage, the description covers the crucial query-authoring semantics and the live-query model. It leaves return value and duplicate-handling unaddressed, but the annotations carry the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 63% schema coverage, the description adds real meaning for the undocumented x/series mapping ('first column = x axis, numeric columns = series, or set x/series explicitly'). It does not clarify type, stacked, or workspace semantics, which the schema mostly handles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Adds a LIVE chart to a sheet') and the title reinforces it. It is clearly distinguishable from napkin_sheets_view_chart, though it never names a sibling explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete workflow prerequisite ('Write the query with napkin_sheets_query first to confirm the shape') and a usage preference ('Prefer aggregated queries'). It lacks an explicit when-not-to-use or a named alternative for related tasks, so it is clear context without full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_createCreate a Napkin sheetAInspect
Creates a Napkin sheet (multi-tab spreadsheet), optionally titled. Then populate it with napkin_sheets_set_cells (headers in row 1). Tell the user where it landed — the Sheets tab in Napkin.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose a non-destructive write (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the safety profile is covered. The description adds modest behavioral context ('optionally titled', 'Tell the user where it landed', the Sheets tab location) but says nothing about permissions/auth requirements or what the operation returns, which is the bulk of the remaining burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the purpose, then the next-step workflow, then the UX instruction. Each sentence earns its place with minimal waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description does not say what the call returns (e.g., a sheet reference/ID needed by set_cells), which the agent needs to chain the follow-up call. It covers the workflow and UX hint well but leaves the return value implicit, so it is adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'workspace' is fully documented in the schema (including token/API-key rules), while 'title' has no schema description. The description's 'optionally titled' compensates partially by signaling title is an optional label, but adds no format or constraint detail beyond that. Baseline 3 is appropriate given the schema carries half the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a precise verb+resource ('Creates a Napkin sheet') and usefully clarifies that a sheet is a 'multi-tab spreadsheet', preventing confusion with a single-tab concept. It is distinguishable from siblings like napkin_sheets_list/query/update by the create semantics, though the description never explicitly contrasts them, so the differentiation leans on the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear follow-up workflow: 'Then populate it with napkin_sheets_set_cells (headers in row 1)', naming the alternative tool to use next and the convention for row 1. It provides clear context for the create-then-populate sequence, but stops short of explicit when-not-to-use or exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_listList Napkin sheetsARead-onlyInspect
Lists the workspace's Napkin sheets (multi-tab spreadsheets — the small-data surface) newest first with tab/chart counts. Archived sheets are hidden unless includeArchived; q narrows by title. Inspect one with napkin_sheets_schema, then query it with napkin_sheets_query.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Title contains (case-insensitive). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| includeArchived | No | Also list archived sheets (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/openWorldHint=false/destructiveHint=false, so safety is covered structurally. The description adds real context the annotations don't: newest-first ordering, default hiding of archived sheets, and the tab/chart counts returned. Return format detail (pagination, exact shape) is still absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero filler; the resource definition and default scoping come first, then the parameter effects, then the next-step routing. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with full schema coverage and no output schema, the description supplies the return ordering and summary fields an agent would otherwise guess at, plus the workflow handoff to schema/query. Only pagination or result-size expectations are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema, including the workspace-slug auth nuance. The description restates `includeArchived` and `q` behavior without adding syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (the workspace's Napkin sheets), and even defines the resource ('multi-tab spreadsheets — the small-data surface'), distinguishing it from sibling surfaces like boards, docs, and diagrams. The sort order and returned summary fields (tab/chart counts) are front-loaded.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit behavioral context ('Archived sheets are hidden unless `includeArchived`; `q` narrows by title') and names the follow-up path (napkin_sheets_schema to inspect, napkin_sheets_query to query). It lacks any explicit when-not-to-use guidance (e.g., vs search_workspace or napkin_docs_list), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_queryRun SQL against a sheetARead-onlyInspect
Runs a SQLite SELECT over a sheet's tabs-as-tables (get the table/column names from napkin_sheets_schema first). Full SQLite dialect: WHERE, GROUP BY, ORDER BY, JOINs across tabs, aggregates. Results cap at 200 rows. This is THE way to answer questions about a sheet's data — never eyeball cells when a query can answer precisely.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | A single SELECT. | |
| sheetId | Yes | Sheet id from napkin_sheets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/destructiveHint=false, but the description adds real behavioral context beyond them: the 200-row result cap and the full SQLite dialect surface (WHERE, GROUP BY, ORDER BY, JOINs, aggregates). It stops short of describing result shape or what happens when the cap truncates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, followed by prerequisites, capability surface, and a limit. The final emphatic sentence is slightly editorial but does routing work, so it mostly earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a query tool with no output schema, the definition covers input prerequisite, dialect capability, and the row cap. Column-level return shape and pagination/truncation behavior are unaddressed, but the referenced schema tool covers much of that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented, including the workspace override rules and sheetId provenance. The description only reinforces that sql must be a SELECT, which the schema itself already states, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Runs a SQLite SELECT) and a specific resource abstraction (a sheet's tabs-as-tables), which is unambiguous and distinct from siblings like napkin_sheets_schema (discovery) and napkin_sheets_set_cells (mutation). An agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the prerequisite tool ('get the table/column names from napkin_sheets_schema first') and an explicit when-to-use rule ('THE way to answer questions about a sheet's data — never eyeball cells'). Routing and exclusion are both handled.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_schemaA sheet's SQL schemaARead-onlyInspect
Returns a sheet's tabs as SQL table definitions — CREATE TABLE scripts with row counts (each tab is a table; its header row names the columns; formula cells contribute computed values). Read this BEFORE writing a napkin_sheets_query so your SQL matches the real tables.
| Name | Required | Description | Default |
|---|---|---|---|
| sheetId | Yes | Sheet id from napkin_sheets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavior about the returned content (row counts, header-row column naming, formula cells yielding computed values), which matters since there is no output schema. It stops short of noting auth/permission or size limits, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the return shape front-loaded and the usage instruction placed second. Every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly assumes the burden of explaining return content and does so precisely, while the 100%-covered input schema handles parameters. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (sheetId, workspace) are fully documented in the schema itself. The description adds nothing about parameter formats or constraints, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns a sheet's tabs as SQL table definitions') and immediately clarifies the shape of the result (CREATE TABLE scripts with row counts, header row naming columns, formula cells contributing computed values). An agent can distinguish it from napkin_sheets_query and napkin_sheets_list without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: 'Read this BEFORE writing a napkin_sheets_query so your SQL matches the real tables.' It names the sibling it pairs with and the ordering condition that selects it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_set_cellsWrite cells into a sheet tabADestructiveInspect
Sets cells in one tab of a sheet, by A1 address — values or formulas ('=SUM(B2:B9)'). Additive and surgical: only the addressed cells change. Check napkin_sheets_schema first so you know the tab names and where data ends; put headers in row 1 when creating a new region. To build a whole small table, write header cells + data cells in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| cells | Yes | A1 → raw value/formula, e.g. {"A1":"Region","B2":"=SUM(B3:B9)"}. | |
| sheetId | Yes | Sheet id from napkin_sheets_list. | |
| tabName | No | Tab name (defaults to the first tab). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, and the description adds meaningful scope: only the addressed cells are modified, which tells the agent the blast radius of a 'destructive' call. It does not describe the response or the needs_confirmation/approvalId flow, but that flow is documented in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and address format, followed by scoping behavior and workflow guidance. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param write tool with no output schema, the description covers the action, addressing, blast radius, and prerequisite lookup. The main gap is the return shape (e.g., which cells were written), though annotations already carry the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: cells, sheetId, tabName, workspace, and approvalId are all documented in the schema, including the A1→value map, default tab behavior, and approval flow. The description reiterates the A1/formula format and adds a header-row convention, but the schema already carries the parameter burden, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Sets cells in one tab of a sheet'), names the addressing scheme (A1), and the accepted value types (raw values or formulas with an example). The 'additive and surgical: only the addressed cells change' line distinguishes it from bulk write siblings like napkin_sheets_update and napkin_sheets_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use context: call napkin_sheets_schema first to learn tab names and where data ends, put headers in row 1 for a new region, and use one call for a whole small table. It does not explicitly contrast with the sibling napkin_sheets_update, so routing between the two write tools still requires inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_updateRename, describe, archive, or share a Napkin sheetADestructiveInspect
Edits a sheet's details: title, description, visibility, or archived (true takes it out of the gallery; false brings it back — the recovery move for a sheet created by mistake or no longer wanted). Cells are edited with napkin_sheets_set_cells, not here. Sheet ids come from napkin_sheets_list.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| sheetId | Yes | Sheet id from napkin_sheets_list. | |
| archived | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | Short gallery summary. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation profile is known. The description adds real behavioral context beyond that: archiving is reversible ('false brings it back') and removes the sheet from the gallery, which is exactly the kind of side-effect detail annotations cannot express. It does not clarify which specific edits are considered destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences: purpose first, then the archive behavior, then the sibling exclusion and id provenance. No filler, no restatement of the title, and the most decision-relevant content is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no output schema, the description covers purpose, required-id source, the tricky archived flag, and sibling routing. It does not discuss authorization requirements (covered by the workspace parameter description in the schema) or what changes are irreversible, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with the undocumented parameter being 'archived' — and the description compensates precisely there, explaining true/false semantics and the gallery removal effect. It also enumerates the other editable fields, adding marginal meaning over the already well-documented visibility enum and workspace slug.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (edits) plus the exact editable fields (title, description, visibility, archived) on a specific resource (a Napkin sheet), and explicitly separates itself from napkin_sheets_set_cells. An agent can distinguish this tool from the ~55 siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the negative case clearly ('Cells are edited with napkin_sheets_set_cells, not here') and gives the source of required input ('Sheet ids come from napkin_sheets_list'). It also frames archived=false as the recovery path, which tells the agent when to reach for this tool; it stops short of spelling out when updating title vs visibility is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_sheets_view_chartView a sheet chartARead-onlyInspect
Renders one of a sheet's pinned charts against the LIVE data and returns the image so you can SEE it — the exact chart the user sees. The text part carries imageUrl; to show the user the chart inline in your reply, put a markdown image on its own line: . Use after napkin_sheets_add_chart to confirm the chart reads well, or whenever the user asks about a chart.
| Name | Required | Description | Default |
|---|---|---|---|
| chartId | Yes | Chart id (napkin_sheets_list shows chart counts; the sheet's charts carry ids). | |
| sheetId | Yes | Sheet id from napkin_sheets_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safe-read behavior is covered. The description adds genuinely new behavioral context: rendering happens against LIVE data and the text payload carries the imageUrl. It stops short of covering edge cases (missing chartId, rendering failures), so a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus a short inline-rendering instruction; the core action and the primary use case are front-loaded. The markdown-display sentence is longer than strictly necessary but earns its place by specifying exactly how to present the result to the user.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain what comes back — and it does, describing both the image return and the imageUrl in the text part. Combined with the usage and rendering guidance, an agent has everything needed to call and present the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter (chartId, sheetId, workspace) is documented in the schema itself including the workspace-token nuance. The description adds no parameter-level syntax or format detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — renders a sheet's pinned chart against live data and returns the image — with the scope ('one of a sheet's pinned charts') pinned down. This clearly separates it from napkin_sheets_add_chart and the other sheet tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives when to use it: 'after napkin_sheets_add_chart to confirm the chart reads well, or whenever the user asks about a chart.' It names the related sibling tool and the condition that selects this one, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_slide_addAdd a slide to a deckAInspect
Appends one slide to an existing deck. Pick a layout (title | section | title-body | title-lead | two-column | three-column | comparison | statement | quote | closing | blank), give the title, and optionally content — markdown for the layout's remaining text areas, split by --- lines (two-column/comparison take one block per column). Check the deck first with napkin_deck_outline.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Slide title. | |
| layout | Yes | Layout key: title | section | title-body | title-lead | two-column | three-column | comparison | statement | quote | closing | blank. | |
| boardId | Yes | Napkin board id (a deck). | |
| content | No | Markdown body — blocks split on `---` lines. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, and the word "Appends" is consistent with an additive, non-destructive write. Beyond that, the description adds no auth/permission requirements, no note on what happens to existing slides, and no error/limit behavior, so it adds only modest behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action, then the layout/content rules, then the prerequisite. The inline enumeration of all eleven layout keys duplicates the schema's layout description verbatim, which is a minor waste of space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param additive write with no output schema, the description covers the required inputs (layout, boardId implied by "existing deck"), optional title/content, and a pre-flight check. It omits what the call returns (e.g., the new slide id) and the workspace-token nuance, which lives only in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description genuinely adds meaning: how `content` maps onto each layout, the `---` block splitting, and that two-column/comparison take one block per column. That is semantics the schema's terse "Markdown body — blocks split on `---` lines" does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Appends one slide to an existing deck" gives a specific verb (appends) and resource (slide within a deck), which is distinct from the sibling mutators napkin_slide_update and napkin_slide_fill. The behavior is clear, though it never names those siblings to explicitly disambiguate add-vs-fill-vs-update.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides one concrete workflow step — "Check the deck first with `napkin_deck_outline`" — but gives no guidance on when to use this versus napkin_slide_fill, napkin_slide_update, or napkin_deck_write/compose. Usage is implied rather than stated with alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_slide_fillFill a slide's empty text areasAInspect
Fills a slide's empty layout text areas by NAME — the names napkin_deck_outline surfaces as [empty text stub: "…"]. Pass texts mapping those exact names to content (markdown - bullets welcome). Read the outline first to see which stubs a slide still has open.
| Name | Required | Description | Default |
|---|---|---|---|
| slide | Yes | 1-based slide number in presentation order. | |
| texts | Yes | Placeholder name → content, e.g. {"Add a title": "Q3 Review"}. | |
| boardId | Yes | Napkin board id (a deck). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false), so the burden is lower. The description adds useful context that it targets only *empty* areas and draws names from the outline, but omits what happens with mismatched/duplicate names, partial fills, or whether the operation is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, verb front-loaded, with the name-linkage mechanism stated first and the outline-reading workflow last. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation with a nested object and no output schema, the description covers purpose, input key provenance, and workflow start point adequately. Minor gap: no guidance on error/partial-fill behavior, but nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description nonetheless adds real meaning beyond the schema by specifying that `texts` keys must exactly match the stub names surfaced by `napkin_deck_outline` and that markdown bullet syntax is accepted — the schema only says 'Placeholder name → content'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fills') and resource ('a slide's empty layout text areas') with a precise mechanism ('by NAME'). It distinguishes itself from siblings by explicitly tying its input to what `napkin_deck_outline` surfaces, so an agent knows exactly what this does versus napkin_slide_update or napkin_deck_write.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite sequence: 'Read the outline first to see which stubs a slide still has open,' which routes the agent to the correct sibling. It doesn't state a when-not condition (e.g. what to do if the slide has no open stubs), so it stops short of the top bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_slide_updateMove or change a deck slideADestructiveInspect
Changes one slide as a whole: move it to another position in the deck (moveTo, 1-based — the other slides shift around it), hide it from the show without deleting it (skipped), show or hide its slide number (pageNumber), or replace its speaker notes (notes, markdown). Slide numbers come from napkin_deck_outline. To change what's ON the slide, use napkin_slide_fill or napkin_draw.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | Speaker notes (markdown); replaces the current notes. | |
| slide | Yes | 1-based slide number in presentation order. | |
| moveTo | No | New 1-based position for this slide. | |
| boardId | Yes | Napkin board id (a deck). | |
| skipped | No | true hides the slide from present mode and exports. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| pageNumber | No | Show the slide number on this slide. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false; the description adds real behavioral context beyond them — skipped hides without deleting, moveTo causes other slides to shift, and notes replaces existing notes. It does not describe error behavior or whether moveTo/skipped can be combined in one call, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and scope, then a tight enumeration of the four mutation modes, then two routing sentences. No filler; every clause maps to a parameter or a sibling tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema, the description covers all operationally meaningful parameters and the destructive/visibility semantics. The only unaddressed item is the workspace parameter, which the schema already handles fully, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds semantics the schema does not: moveTo is 1-based and other slides shift around the moved slide, and skipped affects present mode and exports. The workspace parameter is left entirely to the schema, which is acceptable given its detailed description there.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Changes one slide as a whole') and then enumerates the exact scopes of change: position, visibility, slide number, speaker notes. It explicitly distinguishes itself from napkin_slide_fill and napkin_draw, so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when/when-not routing: use this for whole-slide properties, use napkin_slide_fill or napkin_draw to change slide content, and use napkin_deck_outline to obtain slide numbers. Both the alternative and the condition that selects it are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
napkin_slide_viewView one deck slide (rendered image)ARead-onlyInspect
Renders a single slide of a Napkin deck to an image — exactly what present mode shows, cropped to the slide. Navigate by 1-based slide number in presentation order (get the map from napkin_deck_outline first; the result echoes slideIndex/slideCount so you can step through a deck slide by slide). Use this to check visual layout, drawings, and images that the text outline can't carry.
| Name | Required | Description | Default |
|---|---|---|---|
| slide | Yes | 1-based slide number in presentation order. | |
| boardId | Yes | Napkin board id (a deck). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive=false, so safety is covered; the description adds the genuinely useful behavioral detail that output matches present mode and is cropped to the slide, plus that the response echoes slideIndex/slideCount to enable sequential stepping. It stops short of describing image format/resolution, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Renders a single slide... to an image') and packs navigation and use-case guidance into one tight paragraph with no filler. Slightly dense, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by noting the rendered image matches present mode and that the result echoes slideIndex/slideCount. Board/workspace scoping is left entirely to the schema, which is reasonable given 100% coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds navigation semantics beyond the schema: 1-based indexing in presentation order and that the result echoes slideIndex/slideCount for stepping. It does not add much on boardId/workspace, which the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Renders a single slide of a Napkin deck to an image') and scopes it precisely ('exactly what present mode shows, cropped to the slide'). An agent can distinguish it from napkin_deck_outline (text map) and napkin_deck_export without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: get the slide map from napkin_deck_outline first, then step through by slide number. States the use case ('check visual layout, drawings, and images that the text outline can't carry'), which is a clear when-to-use vs the outline sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_fields_createCreate a Prism fieldAInspect
Creates a field from an origin idea — the seed of an exploration canvas. Counts against the workspace's field quota. May return needs_confirmation — tell the user what you're proposing and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Display name (defaults to the idea). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| originIdea | Yes | The idea at the center — a product, concept, sentence. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare it is a non-read-only, non-destructive, closed-world operation. The description adds two genuinely useful behavioral facts beyond that: it consumes the workspace's field quota, and it may return a needs_confirmation envelope requiring user approval. It stops short of describing the created object or idempotency/error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, front-loaded with the core action, with no filler. The em-dash gloss on 'origin idea' and the confirmation warning each carry information; the sentences are tight and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 100% schema coverage and no output schema, the description's main remaining duty is to warn about the non-obvious needs_confirmation response, which it does. Quota consumption and the approval flow are covered; only the return shape and failure modes are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (including the originIdea seed, workspace slug rules, and visibility enum) are already documented in the schema. The description restates the originIdea concept but adds no syntax, format, or constraint detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates a field') and clarifies what a field is ('the seed of an exploration canvas'), which an agent needs since 'field' is ambiguous. This distinguishes it cleanly from sibling read/update/delete field tools and from the adjacent prism_interviews/series creation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool's role (field creation) but never states when to reach for this versus siblings like prism_fields_update or prism_interviews_create, nor covers prerequisites such as required quota or prior approval. Usage is inferable from the verb but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_fields_deleteDelete a Prism fieldADestructiveInspect
Removes a field and its whole canvas (soft delete; frees the workspace's field quota). Use when the user is done with an exploration or one was created by mistake — studies attached to its nodes are not deleted. Field ids come from prism_fields_list. May return needs_confirmation — name the field and wait.
| Name | Required | Description | Default |
|---|---|---|---|
| fieldId | Yes | Field id (from prism_fields_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds substantial context beyond them: it is a soft delete that frees quota, node-attached studies survive, and the call may return needs_confirmation requiring the agent to name the field and wait. That is exactly the behavioral detail an agent needs before invoking a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and scope, then layers caveats and the confirmation flow. Every sentence carries information, though the dash-embedded clauses make it slightly dense to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still discloses the meaningful return state (needs_confirmation) and the confirmation protocol. Combined with full param coverage and safety annotations, nothing needed to call this destructive tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so fieldId, workspace, and approvalId are already documented in the schema. The description reinforces the fieldId source and the needs_confirmation/approvalId loop but adds no new syntax or format semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Removes a field and its whole canvas') and immediately scopes it as a soft delete that frees the workspace's field quota. This is clearly distinguishable from prism_fields_update/create in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ('done with an exploration or one was created by mistake') plus a key side-effect caveat that studies are not deleted. It does not name an alternative sibling for the 'don't delete, deactivate' case, so it stops short of a full when-not/alternatives treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_fields_getRead a Prism field's nodesARead-onlyInspect
One field's full tree: every node (id, parent, axis, description, pinned, color, whether it has an image / desk research) plus the AI-clustered groups. The origin node has parentId=null; children sit along named axes. Use node ids with the expand / research / study tools.
| Name | Required | Description | Default |
|---|---|---|---|
| fieldId | Yes | Field id (from prism_fields_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds real behavioral value beyond that: it discloses the returned data shape (id, parent, axis, description, pinned, color, image/desk research), the AI-clustered groups, and the data model detail that the origin node has parentId=null. That is useful structural context with no output schema present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the resource and contents, followed by a practical pointer to downstream tools. Every clause carries information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values and does so thoroughly. Combined with readOnly annotations covering safety and the schema covering parameters, an agent has what it needs to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both fieldId and workspace are already documented in the schema, including the workspace-token nuance. The description adds only incidental meaning (node ids) and does not expand on parameter semantics. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (one field's full node tree) and enumerates its contents, distinguishing it from prism_fields_list and the node-level tools. It falls just short of explicitly naming which sibling it complements at the start, but the picture is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It hints at the workflow by saying 'Use node ids with the expand / research / study tools,' implying this is a discovery step, but never states when to prefer this over prism_fields_list or prism_nodes_expand. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_fields_listList Prism fieldsARead-onlyInspect
Lists Prism fields (idea-exploration canvases) in the active workspace: name, origin idea, node count. Fetch one with prism_fields_get for its nodes.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds value by disclosing the lightweight shape of the return (summary fields, not the full canvas), which tells the agent this is a cheap browse operation rather than a detail fetch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the scope and returned fields, closing with the alternative sibling. Every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned fields, and annotations cover the safety profile. The one parameter is fully documented in the schema, so nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, so the schema fully documents the workspace slug and its override/ignore rules. The description's mention of 'the active workspace' loosely aligns with that scope but adds no syntax or precedence detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists Prism fields') and even defines the domain concept ('idea-exploration canvases'), then names the exact fields returned. It distinguishes itself from prism_fields_get by noting that the latter provides nodes, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to `prism_fields_get` when nodes are needed, giving a clear alternative for the deeper-read case. It doesn't spell out a 'when not to use' for listing itself, but the browse-vs-fetch split is adequately conveyed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_fields_updateRename, describe, or share a Prism fieldADestructiveInspect
Edits a field's name (null reverts to the origin idea), description, or visibility. Nodes are untouched. Field ids come from prism_fields_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| fieldId | Yes | Field id (from prism_fields_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write/destructive profile, so the description earns credit for going beyond them: 'Nodes are untouched' bounds the blast radius of the edit, 'May return needs_confirmation' discloses an approval round-trip the agent must handle, and null-reverts-to-origin documents a surprising mutation. It doesn't say whether visibility changes are reversible or what the success response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the mutation scope, followed by the constraint (nodes untouched), the id source, and the confirmation caveat. No filler and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation tool with no output schema, the description covers scope, id provenance, and the confirmation workflow, and annotations carry the safety profile. It leaves the workspace-slug fallback and per-param formatting to the schema, which is reasonable. Nothing critical to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with the workspace, approvalId, fieldId, and visibility params already well documented in-schema. The description adds meaning the schema lacks: that passing null for name reverts to the origin idea, and that approvalId ties to the needs_confirmation envelope. It says nothing about name/description length bounds, which the schema handles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (edits) and resource (field) and enumerates exactly which attributes are mutable: name, description, visibility. That distinguishes it from prism_fields_create/delete/get without needing the schema. It stops short of explicitly routing against those siblings, so a 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied. It supplies two useful preconditions — field ids come from prism_fields_list and the call may return needs_confirmation — but never says when to reach for this tool versus prism_fields_create or prism_fields_delete. No exclusions or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_insights_promoteKeep a study insight in LedgerAInspect
Promotes an analysis theme (or, with no themeId, the analysis summary) into the workspace's Ledger as a belief, carrying the supporting verbatims as evidence with links back to the study. THE step that turns a finding into workspace memory — use when the user says an insight matters. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| themeId | No | A theme id from prism_studies_get; omit for the summary. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (write op, non-destructive, closed-world), so the bar is lower. The description still adds real behavior: it carries supporting verbatims as evidence with links back to the study, and it flags that the call "may return needs_confirmation" – a non-obvious outcome an agent must handle.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then the framing/value, then the return caveat. No wasted words; each sentence adds distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation tool with no output schema, the description covers purpose, side effects (verbatim evidence + study links), and the confirmation outcome. The only real omission is the relationship between approvalId and the needs_confirmation flow, which would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
At 50% schema coverage, the schema already documents themeId (including the omit-for-summary rule) and workspace, and the description reinforces the themeId fallback. However, approvalId is undocumented in both schema and description, leaving a meaningful gap – the likely tie-in to needs_confirmation is never explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (promotes) and resource (an analysis theme or, with no themeId, the analysis summary) and names the destination (workspace's Ledger as a belief). This is distinct from any sibling tool – none of the other prism_* tools perform promotion into Ledger memory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"use when the user says an insight matters" gives a concrete usage trigger, and the parenthetical clarifies the themeId-omitted fallback. It lacks explicit exclusions or named alternatives (e.g. research_findings_list), but the trigger condition is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_interviews_createMint a study interview inviteAInspect
Creates a one-on-one AI-led interview on a study and returns the invite link to forward — the guest needs no account. The transcript stays on the study and joins the next analysis run (it never lands in Compass). focusPrompt briefs the interviewer on what to dig into; omitted, it explores the study's territory. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| focusPrompt | No | ||
| intervieweeName | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=true); the description goes well beyond that by disclosing where the transcript lands ('stays on the study', 'never lands in Compass'), that it joins the next analysis run, that the guest needs no account, and that it may return `needs_confirmation`. It stops short of auth/permission requirements, rate limits, or idempotency, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences that front-load the action and the returned invite link, then layer the side effects and the optional parameter's behavior. No filler and no repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, no-output-schema mutation with 20% schema coverage, the description covers the return value (invite link) and the confirmation edge case, but the undocumented approvalId and intervieweeName leave an agent unable to judge what those inputs do or whether they are needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (just `workspace`), so the description must carry the load. It explains focusPrompt well ('briefs the interviewer on what to dig into; omitted, it explores the study's territory') but leaves approvalId, intervieweeName, and studyId entirely undefined, so most parameters remain opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates a one-on-one AI-led interview on a study') plus the concrete return artifact (invite link). This clearly separates it from the sibling prism_interviews_list, which lists rather than mints interviews.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context (forward the invite to a guest who needs no account) and explains what focusPrompt does when supplied or omitted, but it never states when to choose this tool over near-neighbors like prism_series_create or prism_series_launch_wave, nor any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_interviews_listA study's interviewsARead-onlyInspect
The study's one-on-one qual conversations: guest, status, focus, and — once completed — the structured memo (summary, key claims, tensions, verbatim quotes). Raw transcript included only while no memo exists. inviteUrl is the link to forward for interviews still waiting on their guest.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, non-open-world). The description adds genuinely useful behavior beyond that: field availability is conditional – the memo appears once completed and the raw transcript only while no memo exists, and inviteUrl is only meaningful for interviews awaiting a guest. These conditional-return rules are not in the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph, front-loaded on what the resource is, with the conditional field rules trailing. Every clause carries information about returned content, though the sentence is long enough to require close reading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the burden of describing return values and does so in detail, including the memo-vs-transcript conditionality. Missing only pagination/ordering and workspace scoping behavior, which are minor for a read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% – workspace is documented in the schema, studyId is not. The description implies studyId scoping ('The study's...') but adds no format or constraint detail for either parameter. Baseline 3 is appropriate given the partial schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource as a study's one-on-one interviews and enumerates the returned fields (guest, status, focus, memo, inviteUrl). The list verb is implied by the tool name rather than stated, and there's no explicit contrast with prism_interviews_create, but a reader grasps what it does immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: you call this to read a study's interviews. There is no statement of when to prefer it over prism_studies_get, prism_studies_results, or prism_interviews_create, and no prerequisites or exclusions are given. Adequate but with a clear guidance gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_nodes_expandExpand a Prism node along an axisAInspect
THE core Prism gesture: generate variations of a node along a named semantic axis ('more visceral', 'for the skeptic', 'stripped to essentials'…) in one of four directions. Runs generation against the workspace's inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| axis | Yes | The lens for the variations — short, evocative. | |
| count | No | How many (default 3). | |
| fieldId | Yes | Field id. | |
| parentId | Yes | Node to branch from. | |
| direction | No | Canvas direction (default right; pick a free side). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare non-destructive, closed-world, non-read-only; the description adds real behavioral context beyond them — that generation consumes the workspace's inference credit and that the call may return `needs_confirmation` (matching the approvalId param). This is valuable disclosure, though the return format is not fully elaborated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose and the axis mechanism, then cost and confirmation notes. Efficient with no obvious filler, though the opening all-caps framing is slightly rhetorical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param generative tool with no output schema, the description covers purpose, cost implication, and the confirmation flow, which are the key things an agent needs. It stops short of detailing return content, but that is largely minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents params (including count default, direction default, and workspace scoping rules). The description adds little param-level meaning beyond restating 'named semantic axis' and 'four directions', so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (generate variations), resource (a Prism node), and mechanism (named semantic axis, four directions) with concrete axis examples. An agent can clearly distinguish this generation gesture from siblings like prism_nodes_update and prism_nodes_research.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Positions the tool as 'THE core Prism gesture' and gives axis examples, implying when to reach for it, but never states when NOT to use it or contrasts with alternatives such as prism_nodes_research. Usage is implied rather than delineated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_nodes_researchRun desk research on a Prism nodeAInspect
Web-sourced secondary research on a node's idea — market context with citations, stored on the node. depth 'thorough' runs three angled passes (market, evidence, shifts) at ~3× the credit. focus steers what it goes after; without one it answers a generic brief off the node's own description, so pass it whenever the user has said what they actually want to know. Runs against the workspace's inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| focus | No | What to concentrate on, e.g. 'pricing and who already pays for this' or 'regulatory constraints in the EU'. | |
| nodeId | Yes | ||
| fieldId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds behavior the annotations don't: 'thorough' runs three passes at ~3x credit, the call runs against the workspace's inference credit (a real cost signal), and it may return 'needs_confirmation'. That is meaningful extra context, though latency/citation-failure behavior is unaddressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, then cost/behavior details follow. The prose is dense but each sentence carries information (depth, focus, credits, confirmation). Slightly long but no clear filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, credit-consuming, web-touching tool with no output schema, the description covers the essential stakes: what is produced, where it is stored, cost, and the needs_confirmation return. It stops short of explaining the required id parameters or the approval flow, but is largely complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 33% schema coverage the description must compensate, and it does explain the two most important optional params: what 'depth: thorough' entails and what 'focus' steers. However, required params fieldId and nodeId and the approvalId parameter receive no explanation, so the required inputs remain opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Web-sourced secondary research on a node's idea — market context with citations, stored on the node.' This clearly conveys what the tool produces and where it lands. It does not, however, name or differentiate itself from plausible siblings like prism_nodes_expand or prism_research_digest, so an agent still has to infer which research path to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives solid conditional guidance for parameter selection: pass 'focus' whenever the user has said what they want to know, otherwise a generic brief is answered, and use 'thorough' for broader coverage. But it offers no when-to-use/when-not guidance relative to sibling tools such as prism_nodes_expand, leaving tool selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_nodes_updatePin or color a Prism nodeCDestructiveInspect
Sets pinned state and/or the color tag on a node. Pins are the cross-field shortlist. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| color | No | Hex color, or null to clear. | |
| nodeId | Yes | ||
| pinned | No | ||
| fieldId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered; the description adds one genuinely useful behavioral fact, that the call 'May return `needs_confirmation`', signaling an approval flow. It still does not explain what the destructive hint means here (e.g. whether pinning overwrites existing pins/colors) or how approvalId participates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the core action leads and the caveat follows. It is efficient, though the middle sentence is a semantic aside rather than call-critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter destructive mutation with no output schema and 33% schema coverage, the description is too thin: the approval/confirmation mechanism, workspace behavior, and what can be overwritten are all left to the schema and annotations to imply.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%; pinned, nodeId, fieldId and approvalId carry no schema descriptions. The description only loosely gestures at 'pinned state and/or the color tag' and says nothing about approvalId's role in the needs_confirmation flow or fieldId's meaning, so it fails to compensate for the sparse schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Sets') plus the exact resources touched (pinned state, color tag), so the agent knows precisely what mutation occurs. It does not, however, differentiate itself from siblings such as entity_tags_set or prism_nodes_expand, leaving the agent to infer scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is given. 'Pins are the cross-field shortlist' hints at the semantic role of pinning but never states alternatives (e.g. entity_tags_set for tagging) or prerequisites such as workspace resolution.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_research_digestWhat the workspace's research has foundARead-onlyInspect
A cross-study digest: every study with responses — title, origin (standalone / canvas node / tracker wave), response count, latest analysis summary, top themes, and which insights were kept in Ledger. The starting point for 'what have we learned about X?' — follow up with prism_studies_get (insights) or prism_studies_results on the studies that matter.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description earns credit by disclosing the scope filter ('every study with responses') and the payload shape, including the non-obvious 'origin' taxonomy (standalone / canvas node / tracker wave) and the Ledger-kept insight field — details the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the payload enumeration and closing with the routing advice. The field list is long but justified because there is no output schema; nothing is wasted, though the enumeration could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return values and does so thoroughly. Remaining gaps are minor: no mention of ordering, pagination, or size/limit behavior for workspaces with many studies, and no note on behavior when a workspace has zero studies with responses.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter and schema description coverage is 100%, so the schema already explains that workspace is required for personal tokens with no default and ignored for workspace API keys. The description adds nothing about this parameter, so the baseline 3 for a fully-documented single-param tool is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific artifact ('a cross-study digest') and enumerates exactly what it contains — every study with responses, title, origin, response count, latest analysis summary, top themes, and Ledger-kept insights. It routes distinctly away from siblings by pointing to prism_studies_get and prism_studies_results for follow-up, so an agent can tell it apart from prism_studies_list without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the triggering scenario ('The starting point for "what have we learned about X?"') and names two concrete follow-up alternatives with the object they act on. It stops short of saying when NOT to use it — notably it never distinguishes itself from the sibling research_findings_list or prism_studies_list_workspace, which a naive agent might otherwise reach for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_series_createStart a recurring research seriesAInspect
Creates a research program that fields the SAME instrument on a schedule — brand tracking, a weekly pulse — each wave an ordinary study. The questions freeze once the first wave fields (that's the trendline), so get them right with the user first. Optionally bind a Ledger metric + score rule so every wave posts a reading (a scorer needs a metricId). Waves launch on schedule, or on demand with prism_series_launch_wave. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| scorer | No | How a wave becomes one number: { kind, questionIndex, option?, minN? }. | |
| interval | Yes | daily or weekly. | |
| metricId | No | Ledger metric the waves post readings to (from ledger_metrics_list). | |
| dayOfWeek | No | Weekly only: 0 = Sunday … 6 = Saturday. | |
| questions | Yes | The instrument: an ordered list of questions ({ type, prompt, options?, … } — read an existing study with prism_studies_get for the shape). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| scheduleTime | No | HH:MM local to scheduleTimezone (default 09:00). | |
| scheduleTimezone | No | IANA zone (default UTC). | |
| adaptiveFollowUps | No | AI follow-up questions on curious answers (default true). | |
| autoAnalyzeTarget | No | Analyze each wave automatically at this many completes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only establish that this is a non-destructive write in a closed world; the description adds the consequential behavior beyond them: questions become immutable once the first wave fields (an irreversible commitment), scorer requires a paired metricId, and the call may return needs_confirmation (implying an approvalId retry loop). These are exactly the traits an agent must know before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with what the tool does before the constraints. The em-dash asides carry real information rather than filler, though the prose is slightly chatty and could be tightened without losing content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does disclose the one non-obvious return condition (needs_confirmation). Combined with the freeze semantics and schedule behavior, an agent has enough to call it correctly; only the scheduled-wave failure/backfill behavior is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is already 92%, so the schema carries most parameter meaning. The description still adds cross-parameter semantics the schema cannot express — that a `scorer` needs a `metricId` to post readings — and frames `questions` as the frozen trendline instrument. It does not explain interval/dayOfWeek interaction beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — creating a recurring research series that fields the SAME instrument on a schedule — and immediately scopes it with concrete examples (brand tracking, a weekly pulse) plus the key distinction that each wave is an ordinary study. An agent can tell this apart from the study-creation siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear selection context: waves launch on schedule automatically, or on demand via prism_series_launch_wave, which routes the agent to the right sibling. It also advises getting questions right with the user before the first wave because they freeze. No explicit when-not-to-use guidance, but the alternative path is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_series_launch_waveField the next wave of a seriesADestructiveInspect
CLOSES the current wave (no more responses; it scores and posts to Ledger if bound) and fields the next one from the series' frozen instrument, readying its participant link. Fielding spends credit for follow-ups, closing chats, and analysis, and closing a live wave can't be undone — say both plainly before proposing. Manual launches don't move the schedule clock. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| seriesId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive/openWorld, but the description adds substantial context beyond them: it discloses that closing ends responses and posts to Ledger, that fielding spends credit for follow-ups/chats/analysis, that a live close is irreversible, and that it may return needs_confirmation. This is exactly the kind of behavioral disclosure that carries the burden annotations cannot.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the most important fact ('CLOSES the current wave') with no preamble, and every clause carries operational weight. It is dense with em-dash subordination but nothing is filler; slightly heavy for a single paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately surfaces the notable return (needs_confirmation) and covers destruction, credit cost, and irreversibility. The main gap is the absence of any guidance on the required seriesId and the approval/workspace parameters, which an agent must infer from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'workspace' is documented; seriesId and approvalId are bare), and the description adds no parameter-level meaning for any of the three — no distinction between seriesId semantics, workspace selection, or how approvalId relates to the needs_confirmation flow. With low coverage, the description was needed to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific dual verb+resource (closes the current wave, fields the next one from the frozen instrument) and gives the scope ('a series'). This clearly distinguishes it from siblings like prism_series_create and prism_series_update, which manage series metadata rather than advancing waves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an operational directive ('say both plainly before proposing') and notes manual launches don't move the schedule clock, which frames when to use it versus scheduler-driven launches. It does not explicitly name alternatives or state hard when-not conditions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_series_listList research seriesBRead-onlyInspect
Recurring research programs (ADR 0031): name, cadence, enabled state, next wave time, wave count, and the metric + score rule when the series posts readings. Waves themselves are ordinary studies.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds meaningful context beyond that: it explains what the returned records represent (recurring programs with cadence and score rules) and clarifies the domain boundary that waves are ordinary studies. It does not cover pagination, result volume, or whether series are workspace-scoped at all.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is compressed into a single sentence with an enumerated field list, front-loading the resource definition. The ADR 0031 citation is arguably noise for an invoking agent, and the sentence runs long, but there is no padding or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields, which is the main thing lacking elsewhere. However, for a list tool it omits pagination, ordering, and result-size behavior, and never states how many series are typically returned or what an empty result means.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is only one optional parameter (workspace), whose semantics — default-workspace behavior and API-key handling — are fully documented in the schema. The description adds no further meaning about the parameter, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource — recurring research programs (series) — and enumerates the fields it surfaces (name, cadence, enabled state, next wave time, wave count, metric/score rule). Combined with the title 'List research series' and the sibling names prism_series_create/update/launch_wave, an agent can infer this is the read/list counterpart. It never states the verb 'list' in the description text itself and does not explicitly contrast with prism_series_get-style siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit statement of when to call this versus alternatives, and no prerequisites or exclusions. The closing note 'Waves themselves are ordinary studies' is an implied routing hint (use prism_studies_* for waves rather than this tool), but it is a disambiguation aside rather than usable when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_series_updatePause, resume, or reschedule a research seriesADestructiveInspect
Edits a series: enabled: false pauses the schedule (true resumes), interval / scheduleTime / scheduleTimezone / dayOfWeek move the cadence, name and follow-up / auto-analyze settings change any time. The instrument (questions, scorer) is editable only while no wave has fielded — after that the server refuses with invalid, and the answer is a new series. Series ids come from prism_series_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| scorer | No | ||
| enabled | No | false pauses, true resumes. | |
| interval | No | ||
| metricId | No | ||
| seriesId | Yes | Series id (from prism_series_list). | |
| dayOfWeek | No | ||
| questions | No | Replacement instrument — only before the first wave fields. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| scheduleTime | No | HH:MM. | |
| scheduleTimezone | No | IANA zone. | |
| adaptiveFollowUps | No | ||
| autoAnalyzeTarget | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the bar is lower, yet the description adds real behavioral context: post-fielding instrument edits are refused with `invalid`, and the call may return `needs_confirmation` (implying the approvalId flow). It does not spell out what a successful mutation returns, but the refusal guard is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and then packs pause/resume, cadence, anytime settings, the instrument constraint, id provenance, and the confirmation hint into tightly worded clauses with no filler. The slash-and-backtick density is slightly hard to scan but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter mutation tool with no output schema, the description covers the mutability model, the hard refusal condition, the id source, and the confirmation envelope. The only meaningful omission is the unexplained metricId parameter, which is undocumented everywhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
At 50% schema coverage the description compensates by grouping parameters semantically: interval/scheduleTime/scheduleTimezone/dayOfWeek move the cadence, enabled toggles pause/resume, and name plus follow-up/auto-analyze settings change anytime. metricId remains unexplained in both schema and description, leaving one gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource ("Edits a series") and enumerates exactly which fields are mutable, distinguishing scheduling fields from broadcast instrument fields. It does not name a sibling tool directly, though it hints at the create path via "the answer is a new series."
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete when/when-not rule: the instrument is editable only while no wave has fielded, and afterwards the correct move is a new series. It also points to prism_series_list as the id source. It stops short of explicitly contrasting with prism_series_create or launch_wave by name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_analyzeRun or refresh a study's analysisADestructiveInspect
Kicks the theme analysis (incremental — reads only responses and interviews since the last run; the previous themes carry forward). Returns immediately with the run's phase; the run continues server-side, so wait a moment and re-read prism_studies_get — findings are fresh once its insights stop reporting an active run. upToDate: true means there was nothing new (a healthy no-op, not an error). Runs against inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| reanalyzeAll | No | Full clean-slate re-read of every response (costlier). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing the incremental read scope, that prior themes are preserved, that the call returns immediately while work continues server-side, that it consumes inference credit, and that it may return needs_confirmation. These are exactly the behavioral traits an agent needs and none are derivable from readOnlyHint/destructiveHint alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then layers the async/cost/edge-case details in a single dense paragraph where each sentence carries information. It is a touch dense but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the right thing by describing the return shape (phase, upToDate, needs_confirmation) and the polling workflow. It is nearly complete for a mutation tool, but omits how to act on a needs_confirmation result and how approvalId relates to it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (studyId and approvalId undocumented), so the description must carry more weight. It implicitly contrasts incremental mode with a full re-read, which maps onto reanalyzeAll, but it never names that parameter and says nothing about approvalId even though it warns about needs_confirmation. Adequate but leaves real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('kicks the theme analysis') and immediately scopes it as incremental with the prior themes carrying forward, which an agent can distinguish from a read tool like prism_studies_get. It also names prism_studies_get as the follow-up reader, so the purpose is unambiguous without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: run when you want fresh theme analysis, expect an async run, wait and re-read prism_studies_get, and treat upToDate:true as a healthy no-op rather than an error. It stops short of an explicit when-not-to-use or a comparison against the full clean-slate (reanalyzeAll) path, so it is strong context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_cutCrosstab a question by a cutARead-onlyInspect
Server-computed banner cut with significance: a target question (scale/number → group means + NPS on 0–10; choice/multi → per-option shares) split by acquisition source or by any single-choice question, each group tested against its complement at 95% (Welch t / two-proportion z). vsRest says higher/lower/not_significant — or not_tested when either side is under the 30-response floor; never present not_tested as 'no difference'. Use this for every 'does X differ by Y?' question instead of eyeballing raw rows.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| cutBySource | No | Cut by acquisition source (?src= tags). | |
| targetIndex | Yes | 0-based index of the question to measure. | |
| cutByQuestionIndex | No | 0-based index of a single-choice question to cut by. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnly/non-destructive annotations: it discloses the statistical tests (Welch t, two-proportion z), the 95% confidence level, the 30-response floor for testing, the exact vsRest output values (higher/lower/not_significant/not_tested), and an explicit caution never to present not_tested as 'no difference'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core operation and packs in only relevant details; the parentheticals and value enumerations earn their place, though the single dense paragraph is heavier than strictly necessary and could be split for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by enumerating vsRest outcomes and their meaning. Combined with the statistical method and sample-size floor, an agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80% and already documents each parameter, but the description adds real semantic value by explaining that the target question type determines the output (means + NPS on 0–10 for scale/number, per-option shares for choice/multi) and constraining the cut to acquisition source or a single-choice question.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific operation (server-computed crosstab of a target question against a cut) and specifies both the target measurement types and the split dimensions, making it clearly distinguishable from siblings like prism_studies_analyze or prism_studies_segments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this 'for every "does X differ by Y?" question instead of eyeballing raw rows,' giving clear positive guidance and a stated anti-pattern. It does not name a specific alternative sibling tool to use instead, so it falls just short of the top tier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_deleteDelete a studyADestructiveInspect
Permanently deletes a study with its responses, interviews, analysis, and share links. Right for a draft that won't be fielded or a duplicate; for a live study the user just wants to stop, prefer prism_studies_update closed: true — that keeps the data. Study ids come from prism_studies_list_workspace. May return needs_confirmation — name the study and its response count, then wait.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | Study id (from prism_studies_list_workspace). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description goes further: it enumerates exactly what is destroyed, notes the operation is permanent, and discloses the needs_confirmation round-trip and required follow-up behavior. That is real behavioral context beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the destructive cascade, then the alternative, then the id source and confirmation protocol. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema, the description supplies the cascade scope, reversibility, routing alternative, id provenance, and the needs_confirmation continuation path. An agent has everything required to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so studyId, workspace, and approvalId are already documented. The description still adds the provenance of the id (prism_studies_list_workspace) and the role of approvalId in the confirmation flow, which meaningfully complements the schema without re-teaching it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (permanently deletes) and resource (study), plus the exact cascade scope: responses, interviews, analysis, share links. An agent immediately knows this is a destructive, full-cascade delete rather than a soft delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names when to use it (a draft that won't be fielded, or a duplicate) and when not to, with the alternative sibling and parameter spelled out: 'prefer prism_studies_update closed: true — that keeps the data.' Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_draftDraft a feedback study from a nodeAInspect
Generates a short study (5-8 questions probing the node's problem space — respondents never see the idea itself) and creates it as a DRAFT. The user reviews, previews, and opens it for responses from the study page (or via prism_studies_update publish). Runs generation against inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| nodeId | Yes | ||
| fieldId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations establish it is a non-read-only, non-destructive, closed-world operation, and the description adds genuinely non-structured context: the output is a DRAFT not a live study, generation consumes inference credit (a cost side effect), and it may return `needs_confirmation`. These are useful behavioral disclosures beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and output type are front-loaded, and each sentence carries distinct information (what it generates, the review flow, the cost, the confirmation case). Slightly dense but no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a simple 4-parameter mutation, the description covers the creation flow, cost, and confirmation behavior adequately. The main gap is that required params (nodeId, fieldId) and approvalId are left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% — only `workspace` is documented. The description adds no meaning for `nodeId`, `fieldId`, or `approvalId`; notably, the mention of `needs_confirmation` is not linked to the `approvalId` parameter, which would have been the obvious place to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Generates a short study... and creates it as a DRAFT') plus scope detail ('5-8 questions probing the node's problem space'). It is clear what the tool produces, but it never explicitly distinguishes itself from the sibling tools prism_studies_draft_from_interviews and prism_studies_draft_standalone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It describes the downstream workflow ('user reviews, previews, and opens it... or via prism_studies_update publish'), which routes the agent to a follow-up step. However, it gives no guidance on when to choose this tool over the other two draft variants, leaving that selection to inference from the name/title.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_draft_from_interviewsDraft a survey from a study's interviewsADestructiveInspect
Qual-first design: drafts a survey GROUNDED in the study's completed interviews — the recurring claims become measurable questions, the tensions become the choices, in the interviewees' own words. Replaces the study's current questions with the draft (nothing fields until publish). Needs at least one completed interview (see prism_interviews_list). An optional steer biases the instrument ('focus on pricing'). Runs generation against inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| steer | No | ||
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: it clarifies that the tool replaces the study's current questions but nothing fields until publish (softening the destructiveHint), discloses that generation consumes inference credit, and warns of a possible `needs_confirmation` response. These are exactly the behavioral traits annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph that front-loads purpose, then mechanism, then prerequisites, then cost/return caveats. Every clause earns its place, though the em-dash appositives make it heavier than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the replacement semantics, the credit cost, the precondition, and the confirmation state. The main remaining gap is the undocumented approvalId parameter, which an agent may need in the needs_confirmation flow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 25% schema coverage across 4 parameters, the description must carry more of the load. It explains `steer` well with a worked example ('focus on pricing'), but says nothing about `approvalId` (only obliquely hinted by `needs_confirmation`) and nothing about `studyId` beyond implicitness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('drafts a survey') and immediately narrows it with 'GROUNDED in the study's completed interviews', which distinguishes it from the generic prism_studies_draft and prism_studies_draft_standalone siblings. The mechanism (recurring claims become questions, tensions become choices) tells an agent exactly what output to expect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete precondition ('needs at least one completed interview') and routes the agent to prism_interviews_list to check it. It does not explicitly contrast against prism_studies_draft or prism_studies_draft_standalone, so the when-to-prefer-this-over-siblings decision is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_draft_standaloneDraft a standalone studyAInspect
Generates and creates a DRAFT study that isn't tied to any canvas node — brand tracking, workspace-level research, anything you can describe in a sentence. It appears on Prism's Studies page; the user previews and opens it for responses there (or via prism_studies_update publish). Runs generation against inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| about | Yes | What the study should probe — a topic, question, or problem space in plain language. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it a non-destructive write, but the description adds meaningful traits beyond them: the artifact stays a DRAFT, it surfaces on the Studies page, generation consumes inference credit (a real cost signal), and it may return `needs_confirmation`. The confirmation flow is only hinted at, not explained, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause, with supporting detail following efficiently. Slightly dense with em-dash asides and a parenthetical cross-reference, but every sentence carries information useful to an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and three parameters, the description shoulders more burden; it does well by explaining the created draft's lifecycle, cost, and the possible `needs_confirmation` response. The one gap is the unexplained `approvalId` parameter and no hint of the success response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: `about` and `workspace` are documented in the schema, but `approvalId` is not, and the description never mentions any parameter by name. The 'May return needs_confirmation' line vaguely gestures at the approval flow that `approvalId` presumably serves, but the connection is never made explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Generates and creates a DRAFT study') and its defining constraint ('isn't tied to any canvas node'), which cleanly separates it from the sibling prism_studies_draft and prism_studies_draft_from_interviews. Concrete examples (brand tracking, workspace-level research) reinforce the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions for use — 'anything you can describe in a sentence' and explicitly not tied to a canvas node — which routes the agent away from the canvas-bound draft sibling. It stops short of naming the alternative to use when a canvas node IS involved, so it is strong context without explicit when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_getRead a study + its insightsARead-onlyInspect
One study's questions, response count, open/closed state, and the STORED analysis (summary + themes with strength; insights.activeRun set = an analysis is running right now — wait for it before citing). participantUrl is the live share link to forward when the study is open and one has been minted. Cite theme strengths when reporting results.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/non-destructive, and the description adds real beyond-schema context: the analysis is STORED (cached), activeRun signals a live analysis in progress requiring the caller to wait, and participantUrl only exists when the study is open and minted. These are non-obvious state conditions an agent must respect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource contents, and every clause is informative, but it crams several distinct ideas into one run-on sentence with nested parentheticals, making the activeRun and participantUrl conditions harder to scan than they need to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-shape burden and does it well, naming the fields, the analysis state signal, and the share-link conditional. The one gap is that it never clarifies the tool's relationship to the sibling read/report tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the description says nothing about either parameter — it describes return fields instead. studyId is self-evident and workspace is documented in the schema, so the burden is lightly handled, but no added meaning is provided for the parameters themselves.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates exactly what 'get one study' returns (questions, response count, open/closed state, stored analysis), and the 'STORED analysis' phrasing implicitly separates it from prism_studies_analyze, which runs analysis. It never names a sibling for explicit contrast, so differentiation is inferential rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives genuine conditional guidance ('insights.activeRun set = wait for it before citing', 'cite theme strengths when reporting results'), but only about post-retrieval behavior. It never tells the agent when to call this versus prism_studies_list, prism_studies_results, or prism_studies_report.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_listList a node's feedback studiesBRead-onlyInspect
Feedback studies attached to one Prism node: title, question count, response count, open/closed.
| Name | Required | Description | Default |
|---|---|---|---|
| nodeId | Yes | ||
| fieldId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered externally. The description adds that this is a node-scoped read returning summary fields, but says nothing about ordering, pagination, or what an empty result means. Adequate given the annotation coverage, not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence fragment with no filler; the scope comes first and the returned fields second. It is appropriately sized for a simple list tool, though the fragment style leaves no room for the usage signal the tool needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description usefully compensates by naming the returned fields. However, for a 3-parameter tool with two required, undocumented identifiers and no mention of pagination or result limits, the definition is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: only 'workspace' carries a description (default-workspace override rules), while nodeId and fieldId are bare strings. The description hints at nodeId via 'one Prism node' but never explains what fieldId is or why both are required, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (feedback studies) and scopes it to 'one Prism node', which separates it in spirit from the workspace-wide sibling. It also enumerates the fields returned (title, question count, response count, open/closed), so the agent knows what it gets. It stops short of naming prism_studies_list_workspace as the alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Attached to one Prism node' implies the usage context (fetch studies for a specific node rather than workspace-wide), but there is no explicit when-to-use, when-not-to-use, or named alternative despite prism_studies_list_workspace and prism_studies_get existing as plausible siblings. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_list_workspaceList every study in the workspaceARead-onlyInspect
All studies regardless of origin — standalone, canvas-node, and series waves — newest first: title, question count, response count, draft / live / closed state, and the field / node / series it hangs off. The starting point for 'which studies do we have?'; prism_studies_list is the per-node view, prism_research_digest the findings view. Read one with prism_studies_get.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=true, destructive=false, openWorld=false, so the safety profile is covered. The description adds real behavioral context beyond that: sort order (newest first) and the full set of returned fields. It stops short of disclosing pagination or result-size limits, which for a workspace-wide list would be the remaining useful detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with scope, then return fields, then routing to siblings — a logical ordering with no filler. The em-dash enumeration is dense but each item (origin types, sort order, returned fields) carries information; the only slight cost is that the field list makes the first sentence long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned fields, and with only one fully documented parameter the schema burden is low. What remains unaddressed is list mechanics — pagination, caps, or how large a workspace list can get — which is a minor but real gap for a listing endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single workspace parameter is described thoroughly in the schema itself (token-type behavior, override semantics, API-key exemption). The description adds nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) plus resource (studies) and explicitly enumerates the scope it covers — standalone, canvas-node, and series waves — which is precisely the boundary that separates it from the per-node sibling. It also names the returned attributes (title, question count, response count, state, parent field/node/series), so an agent knows exactly what it gets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the tool around a concrete question ('which studies do we have?') and contrasts it with two named alternatives: prism_studies_list for the per-node view and prism_research_digest for the findings view, plus prism_studies_get for retrieval of a single study. Routing is explicit with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_reportWrite a study's report narrativeBDestructiveInspect
An AI research analyst writes the report's narrative layer — executive summary, key findings with their numbers, recommendations — from the study's computed record (funnel, per-question aggregates, themes, quotes). Stored on the study; the printable report at the study page renders it above the charts. Runs against inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds genuinely useful context beyond the annotations: it consumes inference credit (a cost signal), may return `needs_confirmation`, is persisted on the study, and surfaces above the charts in the printable report. It stops short of warning what happens to a previously existing narrative, which matters given destructiveHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences that front-load the action and its inputs, then add operational facts. No filler, though the em-dash list of artifacts is on the edge of being a run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description has to carry more weight. It does explain where the result is stored and shown and flags the `needs_confirmation` return, but with a destructive write and a 33%-covered schema, the absence of any overwrite/approval detail leaves real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: studyId and approvalId are undocumented in the schema, and the description never mentions any parameter by name. The implicit sourcing 'from the study's computed record' hints at studyId, but workspace override and approvalId (plausibly tied to `needs_confirmation`) are left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (writes the report's narrative layer) and resource (the study's report), with an enumeration of the outputs produced: executive summary, key findings with numbers, recommendations. It is clearly distinguishable from the many prism_studies_draft* siblings, though it never names an alternative to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance, and no reference to the closely related siblings (prism_studies_draft, prism_studies_analyze, prism_studies_results) that an agent could confuse this with. The agent must infer the intended moment (post-analysis, on an existing study) entirely on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_resultsA study's computed resultsARead-onlyInspect
Everything the Results tab shows, as numbers you can trust without counting raw rows: fielding funnel (opens → starts → completes, median completion time, sources, per-question drop-off), per-question aggregates (scale distributions + means + NPS on 0–10, choice/multi counts, rank first-place + mean ranks, MaxDiff set scores, word frequencies, text samples), and quality flags (speeders, duplicate participant ids). ALWAYS use this for quantitative questions about a study — never tally answers yourself.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and closed-world, so the safety profile is covered. The description adds genuine behavioral value beyond that: it asserts the numbers are pre-computed and trustworthy ('without counting raw rows'), which is the key trait an agent needs in order to avoid redundant manual aggregation, and it discloses the quality-flag content (speeders, duplicate ids) that callers should expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The directive sentence is well front-loaded and the closing 'ALWAYS...never...' instruction is the right note to end on. But the middle is a single sprawling sentence with three levels of nested parentheses enumerating every aggregate type, which is harder to scan than a short bulleted structure would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of describing return values, and it does so comprehensively — funnel stages, per-question aggregate types, and quality flags. Combined with annotations covering safety, an agent has everything needed to decide to call this tool and interpret the payload.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: the workspace parameter is documented in the schema, but studyId carries no description anywhere. The tool description says nothing at all about either parameter — no format for studyId, no guidance on when the workspace slug is required. With low coverage, the description was expected to compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete resource (a study's computed results) and enumerates exactly what those results contain — funnel metrics, per-question aggregates, quality flags — so an agent knows precisely what this tool returns. It does not, however, differentiate itself from near-neighbors such as prism_studies_analyze, prism_studies_cut, or prism_studies_report, which an agent could easily confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit directive — 'ALWAYS use this for quantitative questions about a study — never tally answers yourself' — which tells the agent both when to use it and what not to substitute for it. It stops short of naming the sibling tools (report, analyze, cut) that would cover other question types, leaving that inference to the caller.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_segmentsA study's discovered segmentsARead-onlyInspect
The stored k-means segmentation over the study's numeric answers: named segments with size, share, and the distinguishing features (segment mean vs overall). Null when none computed yet — prism_studies_segments_compute discovers them (needs 30+ responses and 2+ numeric questions).
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description still adds real behavioral value: it discloses the null return state and the shape of what is returned (segment mean vs overall), which is not in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the coverage constraint front-loaded and the null behavior plus alternative named at the end. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description compensates by describing the return structure and the null case, and it explains how segments come to exist. Complete enough for a read tool, with only the parameter gap remaining.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, and the description mentions neither studyId nor workspace. The undocumented studyId parameter gets no compensating explanation, so an agent gains nothing about parameters from the description beyond the schema's own workspace note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (retrieves the stored k-means segmentation over a study's numeric answers) and details what the result contains (named segments, size, share, distinguishing features). It is clearly distinguishable from the compute sibling, which it explicitly names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the null-when-not-computed condition and routes the agent to prism_studies_segments_compute with its prerequisites (30+ responses, 2+ numeric questions). What is missing is explicit guidance on when an agent should read segments at all versus other study outputs like results or report.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_segments_computeDiscover a study's segmentsADestructiveInspect
Runs k-means over the study's numeric answers (silhouette-picked k) and names the discovered segments from their computed profiles. Needs 30+ responses and 2+ numeric questions; recompute overwrites. Runs naming against inference credit. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| studyId | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: the 30+/2+ data thresholds, that recompute overwrites existing results, that naming runs against inference credit (a cost), and that a needs_confirmation result is possible. This is exactly the behavioral disclosure the destructiveHint implies, made concrete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the operation and followed by preconditions and side effects. Every clause carries information; none is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a compute tool with no output schema, the description covers the operation, prerequisites, cost, destructive overwrite, and a possible confirmation outcome well. The remaining gap is parameter meaning, particularly approvalId, against only 33% schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% and the description adds no parameter meaning at all. The required studyId and the approvalId are undocumented in both places, leaving the agent to guess how approval relates to the mentioned needs_confirmation flow. A likely approvalId connection is never made explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and mechanism: 'Runs k-means over the study's numeric answers (silhouette-picked k) and names the discovered segments.' This clearly distinguishes the compute operation from generic study tools. It does not, however, explicitly differentiate itself from the closely named sibling prism_studies_segments, leaving the agent to infer the relationship.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete preconditions for use: 'Needs 30+ responses and 2+ numeric questions.' This is genuine when-to-use guidance that an agent can act on. It stops short of naming alternatives (e.g., prism_studies_segments or prism_studies_analyze) or stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prism_studies_updateEdit, publish, open, or close a studyADestructiveInspect
publish: true takes a draft live (one-way — the participant link starts working). closed toggles whether a live study accepts responses. Also edits the study: title, the questions instrument (whole replacement — read prism_studies_get first; LOCKED once responses exist, the server refuses with invalid and the answer is a new study), adaptiveFollowUps (AI probes on curious answers), the closing chat (endChatEnabled + endChatFocus), completionRedirectUrl for a BYO panel (null clears), and autoAnalyzeTarget (null = manual only). Study ids come from prism_studies_list_workspace. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| closed | No | ||
| publish | No | ||
| studyId | Yes | ||
| questions | No | Replacement instrument (drafts only): ordered list of { type, prompt, options?, … } — read prism_studies_get for the shape. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | ||
| endChatFocus | No | What the closing chat should be about; null clears. | |
| endChatEnabled | No | ||
| adaptiveFollowUps | No | ||
| autoAnalyzeTarget | No | ||
| completionRedirectUrl | No | http(s) URL respondents land on after submitting ({pid} = participant id); null clears. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructive/openWorld annotations by disclosing that publish is one-way (the participant link goes live), that questions replacement is whole and refused with `invalid` once responses exist, and that the call may return `needs_confirmation`. Null-clearing semantics for completionRedirectUrl and autoAnalyzeTarget are also spelled out, which the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the highest-stakes behavior (publish one-way, then closed) and every clause carries information. It is a single dense run-on rather than cleanly separated sentences, which slightly hurts scannability, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-param destructive mutation with no output schema, the description covers the consequential behaviors, error signaling (needs_confirmation, invalid), and id sourcing. It leaves approvalId unexplained and does not sketch the return payload, so it is strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 33% schema description coverage across 12 params, the description compensates heavily, explaining the meaning of title, questions (whole replacement), adaptiveFollowUps, endChatEnabled/endChatFocus, completionRedirectUrl, autoAnalyzeTarget, publish, closed, and where studyId comes from. It adds real semantics (null clears, one-way, lock behavior) rather than restating types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (publish, close, edit) and the exact resources touched (title, questions instrument, adaptiveFollowUps, closing chat, completionRedirectUrl, autoAnalyzeTarget). It also names siblings prism_studies_get (for reading the shape first) and prism_studies_list_workspace (for ids), so an agent can tell it apart from the read and list tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete conditional guidance: read prism_studies_get before replacing questions, and when the instrument is LOCKED the answer is a new study rather than this tool. That points to an alternative, but it never states the broad when-to-use vs prism_studies_delete or prism_studies_series_update, so it stops short of full coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_findings_listEverything we know about a research subjectARead-onlyInspect
Cross-instrument findings for one subject (today: a Prism node) — attached feedback studies with response counts + theme counts, and whether desk research exists. (Interviews are study-level now — read them per study with prism_interviews_list.) THE PLACE TO LOOK before proposing new research: cite what already exists.
| Name | Required | Description | Default |
|---|---|---|---|
| subjectId | Yes | Subject id (e.g. a Prism node id). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| subjectType | Yes | Subject type. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds real value beyond that by disclosing the return content (study list with response/theme counts, desk-research presence) and a scope change note that interviews have moved to study level. It doesn't cover auth/limits, but for a read tool this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first clause, followed by return contents and then routing guidance. The parenthetical asides are terse and each sentence carries distinct information (what it returns, the interview exception, when to reach for it). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey returns — and it does, listing the study-level counts and desk-research flag. Combined with annotations covering the read-only profile, an agent has enough to call it correctly, though pagination or full result shape is not described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning to the subjectType enum by clarifying it is 'today: a Prism node,' explaining why the enum has a single value and hinting at future expansion. It adds little for subjectId and workspace, but the schema already documents those.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('Cross-instrument findings for one subject') and enumerates exactly what is returned: attached feedback studies with response counts + theme counts, and desk-research existence. It also distinguishes itself from the sibling prism_interviews_list by stating interviews are now study-level. An agent can tell what this does and how it differs from neighbors without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance is provided ('THE PLACE TO LOOK before proposing new research: cite what already exists') and a concrete alternative is named for the interview case ('read them per study with prism_interviews_list'). It lacks a general when-not clause, but the routing intent is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rooms_add_flowBring a flow into a roomAInspect
Adds a Workbench flow to a room as a member, so people can @mention it and it can work on things with you there. The flow must be shared with the workspace and published. Use it after building a specialist flow for a room's work, or when someone asks for a flow to join. May return needs_confirmation; say which flow and why, and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Workbench flow id. | |
| roomId | Yes | Room id (from rooms_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this is a non-destructive, non-open-world write. The description adds value beyond them: the required preconditions (shared + published) and the fact that it 'may return needs_confirmation', with the expected agent behavior of stating the flow and reason and waiting for approval. It stops short of describing the success result or idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core action front-loaded, followed by prerequisites and the confirmation clause. No filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema mutation tool with full annotation coverage, the description supplies the prerequisites, usage triggers, and the confirmation workflow an agent needs to call it correctly. It leaves the post-success state and any room-membership side effects unstated, but the essential gaps are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the needs_confirmation/approval path, which contextualizes the optional approvalId parameter and how the two-required-param call behaves on confirmation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Adds a Workbench flow to a room as a member') plus the concrete effect ('people can @mention it and it can work on things with you there'). This clearly differentiates it from nearby mutation siblings such as rooms_post or rooms_start_agent_work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers ('after building a specialist flow for a room's work, or when someone asks for a flow to join') and prerequisites (flow must be shared with the workspace and published). It does not name a competing alternative tool, but no obvious sibling overlaps this operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rooms_listList roomsARead-onlyInspect
The workspace's rooms: where groups of people work together with zv1 and flows. Returns each room's id, name, people and flows, linked folders, open sidebar count, and when it last moved. Use it to find the room a user means before rooms_read or rooms_post.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| includeArchived | No | Include archived rooms (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false), so the bar is lower, and the description adds genuinely useful context by enumerating the returned fields (id, name, people, flows, linked folders, sidebar count, last-moved). It does not mention pagination, limits, or ordering, but with no annotations gap to fill this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the resource definition, then returns, then routing guidance, in three tight sentences. The 'where groups of people work together with zv1 and flows' clause is mild flavor but sets context cheaply.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately compensates by listing the return fields and stating the tool's role in the room workflow. Nothing critical for correct invocation is missing, though pagination/ordering behavior is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both workspace and includeArchived well documented in the schema itself, so baseline is 3. The description adds nothing about either parameter beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('The workspace's rooms') and clearly positions the tool among siblings by naming what it returns and how it relates to rooms_read and rooms_post. An agent can distinguish it from those siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to use it ('to find the room a user means before rooms_read or rooms_post'), naming the alternatives and the selecting condition. No explicit when-not or exclusion is given, but the routing intent is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rooms_postPost into a roomAInspect
Posts a message into a room's stream as zv1, where everyone in the room will see it. Use it when a user asks you to share something with a room, or when a routine's instruction says to deliver its result to a room other than the one the routine already posts in. Write for the room, not for the person who asked: say what it is and why it's here in the first line. May return needs_confirmation; summarize the post and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The message, in markdown. | |
| roomId | Yes | Room id (from rooms_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate a write operation (readOnlyHint=false) with non-destructive behavior (destructiveHint=false, openWorldHint=false). The description adds critical context: 'May return needs_confirmation; summarize the post and wait for approval.' This is valuable transparency beyond annotations, though it doesn't cover rate limits or exact return structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph that is front-loaded with the core action and usage guidelines. Every sentence adds value: the posting action, the two use cases, the writing style advice, and the approval note. It is slightly dense but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is a write operation with four parameters and no output schema, the description covers the essential: what it does, when to use it, and the potential need for approval. It could benefit from mentioning the 'zv1' agent identity or any additional constraints, but it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all parameters. The description mentions writing for the room but doesn't add syntax or format details for parameters like text or roomId beyond what the schema provides. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Posts a message into a room's stream.' It clearly distinguishes itself from siblings like rooms_read, comments_create, and rooms_start_agent_work by defining the action as a broadcast into a room. It is clear but doesn't explicitly contrast with every sibling, though the scene is well set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it: 'Use it when a user asks you to share something with a room, or when a routine's instruction says to deliver its result to a room other than the one the routine already posts in.' It provides clear conditions without needing to list alternatives, as the action is unique among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rooms_readRead a roomARead-onlyInspect
What a room has been saying: its catch-up summary (and whether messages have landed since), its sidebars with their outcomes, and the most recent messages as plain lines with who said them. Use it when a user asks what happened in a room, what a group decided, or to catch them up. Quote people by name; never present a summary as a decision unless a sidebar settled it.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many recent messages to return (default 40). | |
| roomId | Yes | Room id (from rooms_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: it reports whether messages have landed since the summary, and it sets an editorial rule ('never present a summary as a decision unless a sidebar settled it') that shapes how results should be interpreted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Contents are front-loaded before usage and editorial guidance, and every sentence carries weight. The opening sentence is dense with comma-separated clauses, which slightly slows parsing but remains purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on the burden of describing the return shape and does so at a conceptual level (summary, sidebars, recent messages as plain lines with speakers). It omits pagination/volume behavior, but the limit parameter in the schema partially covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with roomId, limit (default 40, max 200), and workspace/auth semantics all documented in the schema. The description's mention of 'most recent messages' loosely maps to limit but adds no syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (a room) and enumerates exactly what is returned: a catch-up summary, sidebars with outcomes, and recent messages with speakers. It reads clearly as a read operation, but never explicitly contrasts itself with siblings like rooms_list or rooms_post, so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete trigger conditions: use when a user asks what happened in a room, what a group decided, or to catch them up. There is no explicit when-not or named alternative, but the positive usage context is well specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rooms_start_agent_workPut agents on something in a roomAInspect
Starts zv1 and flows in a room working on something together in a sidebar: they take turns without a person between each, and the result comes back to the room's stream. The flows must already be in the room (rooms_add_flow). Use it when someone asks for agents to work something out, or to hand a specialist flow a job. Anyone in the room can read along or stop it. May return needs_confirmation; say what they'll work on and who, and wait for approval.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | What they should work on, as a clear task. | |
| mode | No | 'moderated' (zv1 picks who speaks) or 'in_order'. Default moderated. | |
| output | No | What to come back with, e.g. 'a one-paragraph recommendation'. | |
| roomId | Yes | Room id (from rooms_list). | |
| flowIds | No | Flows in the room to take part. | |
| maxTurns | No | Most turns before they write up (default 8). | |
| sidebarId | No | Work in this open sidebar instead of opening a new one. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| includeZv1 | No | Whether zv1 takes part too (default true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover mutability/open-world; the description adds substantial behavior: multi-agent turn-taking without a human in the loop, results landing in the room stream, that any room member can read or stop it, and a needs_confirmation envelope requiring an explicit approval wait before proceeding. These are concrete operational facts beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then usage, then the approval caveat. Each sentence carries content, though the first sentence is long and dense with clauses (zv1 + flows + sidebar + turn-taking + stream return).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 10-parameter mutation tool with no output schema, the description covers the mechanism, prerequisite, side effect (result to room stream), and the confirmation/approval path an agent must handle. Only the format of a successful return (beyond the stream) is left implicit, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all ten parameters (including the mode enum, maxTurns, sidebarId, approvalId, workspace) are already documented. The description restates the flowIds precondition and the goal/approval intent but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Starts') and the exact resources involved ('zv1 and flows in a room working on something together in a sidebar'), plus how they operate (take turns, result returns to the room stream). It also names the prerequisite sibling rooms_add_flow, so it is clearly distinguishable from workbench_flows_run and rooms_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use it when someone asks for agents to work something out, or to hand a specialist flow a job') and cites the rooms_add_flow prerequisite. It does not contrast against the closest alternative of running a single flow directly (workbench_flows_run), which is the main remaining gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
routines_createCreate a routineAInspect
Create a zv1 routine: a natural-language instruction run on a schedule ("scan for new competitors and propose a Compass page for each"). The instruction is executed by zv1 with full workspace context, and any writes it attempts during a run are proposed for human approval. Use when the user asks for recurring work ("every day…", "each Monday…"). Write the instruction in second person, as a directive to a future zv1. To run a SPECIFIC built Workbench flow on a schedule, use workbench_flows_schedule instead — a routine runs a free-form instruction, not a flow, so scheduling a flow as a routine makes zv1 re-derive and run it conversationally rather than executing the flow deterministically.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Short display name. | |
| roomId | No | With delivery "room": the room to post in (rooms_list gives ids). | |
| actsFor | No | Visibility: "workspace" (everyone sees it) or "personal" (just the owner). Default workspace. | |
| delivery | No | Where each run's briefing lands (default notification). "room" posts it in a team room for everyone there; pass roomId. The user must be in the room if it's private, and a "personal" routine can post only in a private room. Use it when the user names a room ("…and post it in #claims") or asks from inside one for the room. | |
| interval | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| instruction | Yes | What to do on each run, as a directive. | |
| scheduleTime | No | Wall-clock "HH:MM" anchor (default 09:00). | |
| scheduleTimezone | No | IANA timezone for the anchor (default UTC). | |
| scheduleDayOfWeek | No | Weekly only: 0=Sunday … 6=Saturday (default Monday). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a non-read-only, non-destructive, closed-world write, so the safety profile is partly covered. The description adds genuinely non-schema behavior: runs execute with full workspace context and any writes attempted during a run are proposed for human approval, which is the key reason this write is non-destructive. It stops short of describing what creation returns or scheduling limits, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition, then the trigger condition, then the authoring rule, then the sibling disambiguation — a sensible priority order. The workbench_flows_schedule clause is long, but it is doing real routing work rather than padding, so the length is defensible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 11 parameters, no output schema, and annotation-only safety disclosure, the description covers the concept, the write-approval behavior, the authoring convention, and the main sibling alternative. It leaves delivery/room/visibility and schedule anchor semantics entirely to the schema (where they are documented), and never says what a successful creation yields, but nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 91%, so the baseline is 3 and most parameter meaning is already documented. The description adds real formatting semantics the schema only gestures at: the instruction must be written in second person as a directive to a future zv1, which materially changes how an agent authors the required 'instruction' field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (create a routine) and defines the resource by its mechanics: a natural-language instruction run on a schedule. The inline example ('scan for new competitors and propose a Compass page for each') makes the concept unambiguous and distinguishes it from sibling creation tools like workbench_tasks_create or ledger_feeds_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use when the user asks for recurring work — "every day…", "each Monday…"') and names the alternative for a near-miss case: workbench_flows_schedule for a specific built flow, with the reason a routine is wrong there. This is exactly the when/when-not/alternative structure the dimension asks for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
routines_listList the workspace's routinesARead-onlyInspect
List zv1 routines — scheduled natural-language instructions run by zv1. Returns each routine's id, name, instruction, cadence, delivery, enabled state, and last-run outcome. Check this before creating a routine so you don't duplicate one that already exists.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered structurally. The description adds real value beyond that: it enumerates the returned fields (id, name, instruction, cadence, delivery, enabled state, last-run outcome), which compensates for the absent output schema. It omits pagination and auth behavior, keeping it at a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose followed by return fields and a usage note. Efficient with no filler, though the field enumeration is dense and could be slightly compressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only single-parameter list tool with annotations covering safety, the description supplies everything needed: identity, return shape, and a pre-create dedupe reason. The field list substitutes well for the missing output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter (workspace) with 100% schema description coverage, so the schema already documents its semantics fully. The description adds nothing about parameter usage, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) plus resource (zv1 routines) and immediately defines what a routine is: 'scheduled natural-language instructions run by zv1.' This lets an agent distinguish it from routines_create and routines_update without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use condition tied to a sibling: 'Check this before creating a routine so you don't duplicate one that already exists.' That routes the agent clearly to the create flow. It doesn't state when-not to call it or mention routines_update, so it falls just short of the full rubric.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
routines_updateUpdate a routineADestructiveInspect
Update a routine's instruction, cadence, delivery, or enabled switch (pause/resume). Only pass the fields being changed. Cadence changes recompute the next fire time.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Routine id (from routines_list). | |
| name | No | ||
| roomId | No | With delivery "room": the room to post in. | |
| enabled | No | ||
| delivery | No | ||
| interval | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| instruction | No | ||
| scheduleTime | No | Wall-clock "HH:MM" anchor (default 09:00). | |
| scheduleTimezone | No | ||
| scheduleDayOfWeek | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so safety is partially covered; the description adds genuine behavioral context beyond that: cadence edits recompute the next fire time, and enabled functions as a pause/resume switch. It still doesn't say whether omitted fields are preserved, whether changes need confirmation, or what the approval envelope requires.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, all front-loaded with the highest-value facts first (what can be updated, then partial-update rule, then side effect). No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter mutation with no output schema, the description covers the core concepts, the partial-update contract, and the main side effect, which is the bulk of what an agent needs. Gaps are the delivery=room/roomId dependency and the approval-confirmation flow, though both are partly documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 42% across 12 parameters. The description adds conceptual meaning for the undocumented enabled, delivery, and instruction fields, but it lumps four separate schedule parameters (interval, scheduleTime, scheduleDayOfWeek, scheduleTimezone) under the single word "cadence" without mapping them, and stays silent on name, roomId, workspace, and approvalId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource and enumerates the mutable surface (instruction, cadence, delivery, enabled/pause-resume), so an agent immediately knows what can be changed. However, it never points at the sibling it complements (routines_list to discover ids, routines_create as the alternative), so differentiation is left to inference from the verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Only pass the fields being changed" is an explicit and valuable partial-update invocation rule, and "(pause/resume)" clarifies the intended use of the enabled flag. There is no when-not guidance and no prerequisites around the approval/confirmation path that the approvalId parameter implies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsSearch ZeroWidth docsARead-onlyInspect
Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max number of results. Defaults to 10. | |
| query | Yes | Search query — keywords or natural-language phrase. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful non-annotation context: 'No authentication required — the docs corpus is public' and the shape of returned results (title, slug, public URL, snippet). It does not mention pagination or ordering, but it goes beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four short sentences, front-loaded with the core action, followed by return information, usage trigger, and auth note. Every sentence earns its place, and there is no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only search tool, the description covers purpose, return format, auth requirements, and usage triggers. It does not explain how this differs from search_workspace or list_docs/get_doc, and it lacks result-ordering or empty-result behavior, but it is otherwise sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters are documented in the schema itself. The description says the query returns a 'query-relevant snippet' but adds no syntax, format, or constraint details beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Search ZeroWidth product documentation.' It also names the return fields and the product scope with concrete examples (Compass, Workbench, Caliper, etc.), which lets an agent distinguish it from siblings like list_docs, get_doc, and search_workspace without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: 'Use this when the user asks about a ZeroWidth product ..., an API behavior, or a policy.' However, it does not name alternative tools (search_workspace, list_docs, get_doc) or state when not to use this tool, so it stops short of explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_workspaceSearch the whole workspaceARead-onlyInspect
Finds entities across every tool by name in one call — Workbench flows, Compass pages, Caliper datasets, evals, rubrics, reviews, specs and sources (apps sending agent traces), Ledger entries, Napkin sketches and decks. Use it FIRST when the user names something without saying where it lives ('the onboarding flow', 'that invoice page'); reach for a tool's own list only when you already know the tool. Each hit carries its id, kind, and workspace-relative path, so the id feeds the matching *_get tool and the path makes a link. Results only include what the user can see, and only kinds this token may read.
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Case-insensitive substring matched against names/titles. | |
| kinds | No | Restrict to these kinds (flow, page, dataset, eval, entry, board). Omit to search everything. | |
| limit | No | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context beyond that: results are filtered to what the user can see and to kinds the token may read, and it discloses the hit shape (id, kind, workspace-relative path).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences with no filler: purpose and coverage first, routing second, return/scope semantics last. Every clause carries information the agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden and does so by describing the id/kind/path tuple and how the id feeds *_get tools. Combined with the permission scoping note, an agent has everything needed to call and use this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most parameters are already documented structurally. The description reinforces name/title matching but adds no new syntax or format detail for kinds, limit, or workspace, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb (Finds) and resource (entities across every tool by name), then enumerates the concrete kinds covered (flows, pages, datasets, evals, rubrics, etc.). An agent can immediately distinguish this cross-tool search from the many per-tool list siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('Use it FIRST when the user names something without saying where it lives'), gives concrete examples ('the onboarding flow'), and names the alternative plus its selection condition ('reach for a tool's own list only when you already know the tool'). This is textbook when/when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_executions_getInspect a flow's executions (debug traces)ARead-onlyInspect
The flow's stack traces: recent executions with per-node timelines — what each node received, produced, how long it took, and the exact error when one failed. USE THIS when a run misbehaves instead of guessing: read the failing node's inputs/error, then propose a fix (workbench_flows_edit_text) grounded in what actually happened. Pass executionId to inspect one run, or just flowId for the most recent runs. Node inputs/outputs are truncated for transport — the full record is in the flow's dev drawer.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many recent executions (default 3, max 10). Ignored when executionId is set. | |
| flowId | Yes | Workbench flow id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic). | |
| executionId | No | Specific execution to inspect. Omit for the latest runs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive=false, and the description adds genuinely useful non-obvious behavior: node inputs/outputs are truncated for transport and the full record lives in the flow's dev drawer. It also implies no auth caveats beyond what the schema documents, so this is solid but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with what is returned before the usage guidance. Slightly editorial ('instead of guessing') but every sentence carries routing or return-shape information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description fully describes the return shape (per-node timeline, inputs, outputs, durations, error), the truncation caveat, and the flowId vs executionId selection. Nothing an agent needs to call or interpret it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds the key interaction rule: pass executionId to inspect one run, or just flowId for the most recent runs (echoing that limit is ignored when executionId is set). That interaction meaning goes beyond the per-parameter schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope: 'recent executions with per-node timelines' plus what each node received, produced, took, and its error. This is clearly distinguishable from siblings like workbench_flows_get or workbench_tasks_run_view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit triggering condition ('when a run misbehaves instead of guessing') and routes the agent forward to workbench_flows_edit_text for the fix. It does not, however, name alternative read tools (e.g. caliper_traces_get or workbench_tasks_run_view) or state when not to use this.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flow_authoring_guideWorkbench flow-authoring guideARead-onlyInspect
READ THIS FIRST before writing or editing any raw flow orchestration JSON. Covers the document shape, the settings-vs-input-ports rule (temperature, max_tokens, tools, response_format are PORTS, not settings — constants reach ports via value nodes), exact link format, plugin links, and three complete worked examples.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: it is a preparatory reading resource whose content includes three complete worked examples, implying a substantial reference payload. It says nothing about response size or format, keeping it at 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the imperative 'READ THIS FIRST' so the most actionable instruction leads. The second sentence enumerates the guide's contents tersely with no filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter reference tool with no output schema, the description adequately conveys what an agent will receive when it calls it. It could note the output modality (returned documentation text) or roughly how large the content is, but what an agent needs to decide to call it is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty with 100% coverage, so the baseline of 4 applies. There are no parameters for the description to illuminate, and it correctly does not invent any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states precisely what the tool is and returns: a reference covering document shape, the settings-vs-input-ports rule, link format, plugin links, and worked examples. It is unmistakably a flow-authoring documentation resource, distinct in kind from the flow CRUD siblings. It stops short of naming a specific sibling it competes with, hence 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'READ THIS FIRST before writing or editing any raw flow orchestration JSON' gives an explicit trigger condition (before authoring/editing raw JSON). It doesn't name the alternative tools (flows_edit_text, flows_scaffold, flows_update) or state when the guide is NOT needed, but the intended usage window is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_deleteDelete a Workbench flowADestructiveInspect
Deletes a flow. Its schedules stop, guest share links die, and evals targeting it can no longer run (their binding shows flowOk: false) — check caliper_flow_performance for evals and workbench_flows_schedules_list before proposing, and prefer workbench_flows_update with archived: true when the user just wants it out of the way. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Flow id, from workbench_flows_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations flag destructiveHint=true, but the description adds real substance beyond them: exactly what is destroyed (schedules stop, guest share links die, evals can no longer run, binding shows flowOk: false) and the confirmation round-trip ('May return needs_confirmation'). This is precisely the consequence-level detail an agent needs before a destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and its consequences, then appends the pre-check guidance and the re-confirmation note. Dense but every clause carries distinct, load-bearing information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the blast radius, the pre-flight checks, the softer alternative, and the confirmation protocol. An agent has everything needed to invoke it safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so flowId, workspace, and approvalId are all documented in the schema. The description reinforces the approval flow by mentioning the needs_confirmation response, but adds no syntax or format detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Deletes a flow') that unambiguously states the operation. It is clearly distinguishable from sibling tools like workbench_flows_update or workbench_flows_schedules_delete, which it even references by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('check caliper_flow_performance for evals and workbench_flows_schedules_list before proposing') and names a concrete alternative with the selecting condition ('prefer workbench_flows_update with archived: true when the user just wants it out of the way'). No routing decision is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_edit_textSurgically edit text inside a flowADestructiveInspect
Applies ONE precise text replacement to a flow's DRAFT orchestration — Edit-tool semantics: find must be the EXACT current text, copied character-for-character from workbench_flows_get, and must occur exactly once anywhere in the flow (system prompts, node settings, metadata, a node's type). find is text inside ONE value, not JSON: to change a model, find the model slug alone (anthropic-claude-haiku-4-5) and replace it with another llm slug from workbench_node_catalog_get. It can't add or remove nodes or links. Zero or multiple matches return an error instead of guessing. This is the improvement primitive: check receipts first with caliper_flow_performance, cite the run id in note, apply the edit after approval, then re-run the eval with caliper_evals_run and report the score delta — never claim improvement without the before/after. Edits land on the draft only; the published version changes when someone publishes (workbench_flows_publish, after the user has seen the result). May return needs_confirmation — show the user the exact find/replace diff and wait.
| Name | Required | Description | Default |
|---|---|---|---|
| find | Yes | EXACT current text (≥3 chars), copied verbatim — never reconstructed from memory. | |
| note | Yes | Why this edit. Cite eval run ids when it follows a measurement (e.g. 'refund answers scored 1/5 on accuracy in run cer_abc — adds the manager-approval rule'). Lands in the audit log. | |
| flowId | Yes | Flow to edit. | |
| replace | Yes | Replacement text. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations flag it as a destructive, non-read-only mutation, and the description goes well beyond that: edits land on the draft only, zero or multiple matches error rather than guess, it may return a `needs_confirmation` envelope requiring a diff to be shown, and it cannot add/remove nodes or links.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core find/replace constraint before the workflow narrative, and nearly every sentence carries operational weight. It is dense and long, but the length is justified by the number of constraints it must convey.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 6-param mutation with no output schema, the description covers the failure modes (zero/multiple matches, needs_confirmation), the scoping (draft vs published), and the surrounding measurement workflow. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value not in the schema: `find` is text inside ONE value (not JSON), must be copied character-for-character, must occur exactly once, with a concrete model-slug example for the replace flow.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Applies ONE precise text replacement to a flow's DRAFT orchestration') and immediately scopes it with Edit-tool semantics. It distinguishes itself from structural siblings by explicitly stating it can't add or remove nodes or links.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Prescribes the full workflow: measure with caliper_flow_performance first, cite the run id in `note`, apply after approval, then re-run caliper_evals_run and report the delta. Also states the draft/publish boundary and names workbench_flows_get and workbench_flows_publish as the surrounding tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_forkFork a public flow into the workspaceAInspect
Copies a PUBLIC flow from another workspace (a template) into this one as a new draft the user owns. Identify it by its flowUuid (the id on Workbench's public template pages and in workbench_flows_get output); optionally pin which published version to copy. Flows already in this workspace can't be forked — open them instead. Counts toward the plan's flow cap. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| flowUuid | Yes | The source flow's public uuid. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| forkedFromVersion | No | Published version label to copy. Default: the latest published. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description adds real context beyond them: the result is a user-owned draft, it counts toward the plan's flow cap, and it may return `needs_confirmation`. It stops short of describing the response shape or how to resolve the confirmation flow in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and constraints. Every sentence carries distinct information (identity, version pinning, non-forkable case, cap cost, confirmation signal) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation with no output schema, the description covers the critical agent-facing facts: public-source requirement, version selection, the already-in-workspace exclusion, quota impact, and the confirmation return. Nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds locating value for flowUuid ('the id on Workbench's public template pages and in `workbench_flows_get` output') and clarifies that forkedFromVersion optionally pins a published version. It goes beyond restating the schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Copies a PUBLIC flow from another workspace into this one as a new draft') and immediately clarifies the fork semantics (template -> owned draft). It distinguishes itself from siblings by noting flows already in the workspace can't be forked and should be opened instead, and by pointing to workbench_flows_get for the id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use (a PUBLIC/template flow from another workspace), explicit when-not ('Flows already in this workspace can't be forked — open them instead'), and it names the alternative path. Nothing about selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_getRead one Workbench flowARead-onlyInspect
One flow. Default view: summary lists its nodes and links so you can pick one; view: node with a nodeId returns that node's full settings and prompt — read THAT before workbench_flows_edit_text, and copy find text from it character-for-character, never from memory. view: full returns the whole orchestration body.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | `summary` (default): the flow's fields plus one line per node (id, type, label, and how long its prompt is) and the links — enough to pick a node. `node`: the full settings of one node (pass nodeId) — what you need before workbench_flows_edit_text. `full`: the whole orchestration body; large. | |
| flowId | Yes | Flow id. | |
| nodeId | No | With view `node`: which node. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds useful payload-size context ('full: large') and clarifies that node returns full settings and prompt, but does not discuss auth/rate limits or truncation, keeping it just under a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core fact ('One flow') then walks the three views in the order an agent would use them, with each clause earning its place and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain return values — and it does per view (node list and links, one node's settings/prompt, whole body). Complete for a read tool whose safety profile is already in annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the `view` enum is fully documented in the schema, so the schema already does the heavy lifting. The description restates view behavior and adds workflow context but no new format or syntax detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read/get) and resource (a single Workbench flow), and the three view modes make the scope precise. It is clearly distinguishable from workbench_flows_list (many) without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit, actionable guidance: use summary to pick a node, then view:node before calling workbench_flows_edit_text, and copy `find` text character-for-character rather than from memory. It names the downstream sibling and the precondition that selects this call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_listList Workbench flowsARead-onlyInspect
Lists every flow in the active workspace the caller can see. Returns summaries (id, name, visibility, updatedAt) — fetch one with workbench_flows_get for the full orchestration body.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, so the safety profile is covered. The description adds real value beyond that by disclosing the returned summary fields (id, name, visibility, updatedAt) and the visibility scoping ('the caller can see'), which an agent needs since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste; the core listing behavior is front-loaded and the get-for-detail note follows naturally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety, a single fully documented parameter, and no output schema, the description properly compensates by naming the return fields and pointing to the detail tool. Only the full shape of each summary is left implicit, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the workspace param already documents slug semantics, default/override behavior, and API-key treatment. The description adds nothing about the parameter, so baseline 3 applies since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists every flow in the active workspace') with clear scoping ('the caller can see'), and explicitly distinguishes itself from workbench_flows_get by noting it returns only summaries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context for use (enumerate flows, choose one to inspect) and routes to the alternative (`workbench_flows_get`) when the full orchestration body is needed. No explicit when-not guidance, but the routing hint is strong and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_publishPublish a Workbench flow versionADestructiveInspect
Snapshots the flow's current DRAFT as a named published version — the version schedules (workbench_flows_schedule), guest share links (workbench_flows_share_create), public-API runs, and 'published'-stage evals execute. The draft keeps evolving after this; runs on the published side don't change until the next publish. Publishing does NOT run the flow, but any eval set to runOnPublish starts a run (that spends credit — mention it when one exists). Publish only after the user has tested the draft (workbench_flows_run) or asked for it. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | What changed — a commit message. | |
| flowId | Yes | Flow to publish. | |
| version | Yes | Version label, e.g. '1.0.0' or '2026-09-04'. Re-using a label retags it to this snapshot. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare write/destructive/openWorld; the description adds the real behavior: the draft keeps evolving while published runs stay frozen until the next publish, publishing does not execute the flow, and runOnPublish evals will start a run that spends credit (with an instruction to surface that to the user). It also discloses the 'needs_confirmation' return path, which is beyond anything the annotations carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The snapshot definition is front-loaded in the first clause and every subsequent sentence carries distinct load-bearing information: consumers of the published version, draft/published divergence, the no-run clarification, credit-spending evals, and the publish precondition. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating tool with no output schema, the description covers the lifecycle semantics, side effects (eval runs and credit spend), the confirmation flow, and prerequisites. Nothing material an agent needs before calling it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents note, version retagging, workspace defaulting, and the approvalId provenance, so the description adds little parameter-level meaning. Its only addition is the reference to 'needs_confirmation', which is already tied back to approvalId in the schema itself. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a precise verb and resource — 'Snapshots the flow's current DRAFT as a named published version' — and immediately distinguishes this from nearby siblings by naming what consumes the published side (workbench_flows_schedule, workbench_flows_share_create, public-API runs, published-stage evals). An agent can separate this from workbench_flows_run and workbench_flows_update without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states an explicit precondition and alternative: 'Publish only after the user has tested the draft (workbench_flows_run) or asked for it.' It also clarifies the boundary versus running ('Publishing does NOT run the flow') and versus the draft, which is exactly the when/when-not routing an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_revisions_listList a flow's versions (publish history)ARead-onlyInspect
The flow's revision history, newest first: published versions carry a version label and note; entries with version null are draft autosaves. Use it to answer 'is this published?' (any entry with a version), to find a revision id for caliper_evals_create's flowRevisionId, or to see when the draft last changed.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Entries to return, default 20. | |
| flowId | Yes | Flow id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| publishedOnly | No | true = only labeled published versions (default false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and destructiveHint=false, so safety is covered. The description adds real behavioral context beyond that: reverse-chronological ordering and the semantic distinction between labeled published versions and null-versioned draft autosaves. It stops short of noting pagination or how limit interacts with ordering, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, output shape front-loaded before the use cases. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the return shape (ordering, version label, null for drafts) and anchors the tool to its downstream consumer. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already carries param semantics and baseline is 3. The description's published/null distinction illuminates what publishedOnly=true filters to, but adds nothing about limit or workspace beyond what the schema documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and what it returns: the flow's revision history, newest first, with published vs draft entries differentiated. An agent can distinguish this from workbench_flows_list and workbench_flows_get without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three explicit when-to-use scenarios: answering 'is this published?', sourcing a revision id for caliper_evals_create's flowRevisionId, and checking when the draft last changed. It even names the consuming sibling tool by parameter, so the routing decision is fully resolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_runRun a Workbench flowAInspect
Execute a flow and return its outputs synchronously. Runs the draft by default; pass source:"published" (optionally a version) to run the live published version. Spends workspace LLM budget, bounded by the token's cost cap. input is the flow's input envelope, e.g. {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}.
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | Flow input envelope (RunFlowInput). Omit for a flow that needs none. | |
| flowId | Yes | Flow id to run. | |
| source | No | Which version to run. Default: draft. | |
| version | No | Published version label (only with source:"published"). | |
| timeoutMs | No | Max runtime in ms (also the hard cost ceiling). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this. Ignored for workspace API keys. | |
| approvalId | No | Re-call with the approvalId from a needs_confirmation response after the user approves. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations: it discloses that execution 'Spends workspace LLM budget, bounded by the token's cost cap' and that results come back 'synchronously'. The cost/budget disclosure is exactly the kind of non-obvious operational consequence annotations (readOnlyHint=false, destructiveHint=false) don't capture, and it is consistent with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action and result, then source semantics, cost, and input examples. Every sentence carries information; the only slight sprawl is the inline input examples, but they are useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers what matters: synchronous return of outputs, budget/cap implications, draft vs published selection, and input envelope shape. The confirmation/approval re-call flow (approvalId) is only documented in the schema, leaving a minor gap for a 7-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3. The description adds real meaning on top: it explains the `source` semantics (default draft vs published-with-optional-version) and illustrates the `input` envelope with concrete examples ({"kind":"chat",...}, {"kind":"form",...}) that the schema does not provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Execute a flow') and adds the key behavioral result ('return its outputs synchronously'). Combined with the draft-vs-published distinction, an agent can distinguish this from sibling tools like workbench_flows_get or workbench_flows_publish without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Conveys the default behavior ('Runs the draft by default; pass source:"published" ... to run the live published version') and the cost context, which implicitly guides selection. However, it never states when to prefer this tool over alternatives (e.g. schedules_run_now, tasks_run_view) or any exclusions, so guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_scaffoldScaffold a first-draft Workbench flowAInspect
Creates a runnable first-draft flow from a spec you author: pick the simplest pattern that fits (classifier for read-and-bucket, structurer for transform/extract/draft-for-review, agent for genuinely conversational), write a production-quality system prompt grounded in what the user told you, and mark anything stubbed with [STUB: ...] markers plus stubNotes. The draft opens in Workbench's simple editor at /w//flows/ — give the user that path. May return needs_confirmation; show the user what you're proposing and wait for their approval, then re-call with the approvalId. When the draft comes from a Compass change, pass its opportunityId so the flow and the change point at each other (Compass shows 'Open in Workbench'; the flow shows where it came from). Then: workbench_flows_run to try it, workbench_flows_publish when it's ready for schedules and share links.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | Pattern editor the draft opens in. Pick interviewer for conversations that LEARN something from a person (research calls, intake, retros) — pair it with workbench_flows_share_create so non-account humans can talk to it. | |
| model | No | The model node slug the draft runs on — an `llm` entry from workbench_node_catalog_get that the workspace allows (e.g. an OpenAI, Anthropic, or Google model). Defaults to Claude Haiku 4.5. Set it here rather than editing the flow afterwards. | |
| flowName | Yes | Imperative + specific, ≤80 chars. | |
| stubNotes | No | What's stubbed / missing / worth wiring next. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| categories | No | classifier only: 2-12 buckets. | |
| description | No | 1-2 sentences: what it does, where it fits. | |
| systemPrompt | No | The draft's behavior — production-quality first attempt with [STUB: ...] / [FILL IN: ...] markers where reality is missing. Required for structurer and classifier. Leave it out of an agent to get the bare model with no instructions (a 'promptless' agent). | |
| opportunityId | No | Compass change (from compass_opportunities_list) this draft implements. Sets the flow ↔ change link; refused when the change already has a draft. | |
| inputDescription | No | What the flow's single input carries. | |
| responseSchemaJson | No | structurer only: the output JSON Schema, JSON-encoded as a string (object root, required fields, additionalProperties false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the generic safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description goes well beyond that: the draft opens in Workbench's simple editor at a concrete path, may return needs_confirmation requiring user approval and re-call with approvalId, and opportunityId is refused when the change already has a draft. Rich behavioral disclosure the annotations cannot supply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and then structured as author-guide → return-path → follow-on tools. Dense and long for a single description, but nearly every clause is actionable and nothing reads as filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter mutation tool with no output schema, the description covers the critical operational context: what gets produced, where it opens, the approval loop, the linking behavior, and the recommended next tools. An agent has enough to call it correctly on the first attempt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning on top: it explains how to choose `mode` (pattern → use case), calls for [STUB: ...] markers paired with stubNotes, and describes the flow↔change linking role of opportunityId. It adds value beyond the schema, though it does not touch every parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates a runnable first-draft flow from a spec you author') and frames the scope as scaffolding a draft rather than editing, running, or publishing one. An agent can distinguish it from siblings like workbench_flows_update, workbench_flows_fork, and workbench_flows_run without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit pattern-selection guidance (classifier for read-and-bucket, structurer for transform/extract, agent for conversational) and names the follow-on tools with their conditions: workbench_flows_run to try it, workbench_flows_publish when ready for schedules/share links. It also covers the needs_confirmation edge case and the Compass opportunityId path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_scheduleSchedule a Workbench flowAInspect
Set a specific built flow to run automatically on a cadence — the deterministic counterpart to a zv1 routine. Each run executes the flow's latest PUBLISHED revision with the fixed input envelope and routes the output per delivery. USE THIS when the user wants a flow they've built to run on a schedule ("run my digest flow every morning"). Do NOT create a routine for this — a routine runs a free-form instruction, not a built flow. The schedule runs the PUBLISHED flow, so the flow must be published (check workbench_flows_revisions_list; publish with workbench_flows_publish if it isn't). input is fixed for every run, so the flow itself should fetch anything time-varying at run time. Manage existing schedules with workbench_flows_schedules_list / workbench_flows_schedules_update.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | Fixed input envelope every run executes with — {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}, the same shape workbench_flows_run takes. Constant across runs. | |
| label | No | ||
| flowId | Yes | Flow to schedule (id from workbench_flows_list). | |
| delivery | No | ||
| interval | Yes | Cadence: hourly, daily, or weekly. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| scheduleTime | No | ||
| scheduleTimezone | No | ||
| scheduleDayOfWeek | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, so the mutation/safety profile is already known. The description adds real context beyond that: it runs the latest PUBLISHED revision (not draft), the input envelope is fixed for every run, output is routed per delivery, and the flow must be published first. It stops short of explaining approval/needs_confirmation behavior despite an approvalId parameter, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then usage, then prerequisites, then management siblings. Dense but each sentence carries routing or behavioral information. Slightly long, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation with no output schema and an approval flow, the description covers prerequisites, execution semantics, and sibling management well. The notable omission is scheduling cadence detail (time-of-day, timezone, day-of-week) which an agent needs to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, so the description must compensate and it partly does: it clarifies that `input` is a constant envelope matching workbench_flows_run's shape and that `delivery` routes output. However it says nothing about the timing parameters (scheduleTime, scheduleTimezone, scheduleDayOfWeek) or label, leaving half the parameters to bare schema. Baseline 3 given the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Set a built flow to run automatically on a cadence') and immediately contrasts it with the nearest sibling ('the deterministic counterpart to a zv1 routine'). An agent can distinguish it from routines_create and workbench_flows_run without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit USE THIS trigger ('run my digest flow every morning') plus an explicit exclusion ('Do NOT create a routine for this'). It also routes the agent to the exact prerequisite and lifecycle siblings: workbench_flows_revisions_list, workbench_flows_publish, workbench_flows_schedules_list/update.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_schedules_deleteDelete a flow scheduleADestructiveInspect
Removes a schedule for good. Prefer workbench_flows_schedules_update with enabled: false when the user might want it back. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Flow the schedule belongs to. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| scheduleId | Yes | Schedule id, from workbench_flows_schedules_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds meaningful context beyond them: permanence ('for good'), the safer soft-delete alternative, and the fact that the call may return `needs_confirmation` (which pairs with the approvalId parameter). It stops short of describing idempotency or side effects on dependents, so it's strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and permanence, then the alternative, then the confirmation caveat. No filler; every sentence carries a distinct decision-relevant fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return burden and handles it by flagging `needs_confirmation` and tying it to approvalId. For a destructive mutation with a confirmation round-trip this is nearly sufficient; it could still note what happens to a running schedule or in-flight executions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so flowId, scheduleId, workspace, and approvalId are all documented in the schema itself. The description hints at the confirmation flow (approvalId) but adds no syntax or value details beyond what the schema already provides, matching the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Removes a schedule') with the permanence qualifier ('for good'), clearly distinguishing it from the sibling update tool. An agent can tell what this does and how it differs from workbench_flows_schedules_update without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (workbench_flows_schedules_update with enabled: false) and the exact condition that selects it ('when the user might want it back'). This is a textbook when-not-to-use-the-destructive-tool guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_schedules_listList a flow's schedulesARead-onlyInspect
Every schedule on one flow: cadence, whether it's paused (enabled: false), the next fire time, delivery, and the last run's status. Read this before workbench_flows_schedule (don't create a duplicate) and to get the scheduleId for workbench_flows_schedules_update / workbench_flows_schedules_delete / workbench_flows_schedules_run_now.
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Flow id, from workbench_flows_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, non-destructive, closed-world behavior, so the bar is lower. The description goes further by disclosing the returned fields, including how paused state surfaces as 'enabled: false' — signaling the actual response shape in the absence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the payload contents are front-loaded ahead of the dependency guidance. Every clause carries a distinct routing or content signal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields, and it chains correctly to the create/update/delete/run_now siblings. Nothing an agent needs to call this correctly and thread the scheduleId forward is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both flowId and workspace are already documented in the schema (flowId's provenance, workspace's auth caveats). The description adds no parameter-level detail beyond implying a single flow, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Every schedule on one flow') and enumerates the payload (cadence, paused state, next fire time, delivery, last run status), which distinguishes it cleanly from sibling schedule mutators. An agent knows exactly what it gets back without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use: read before workbench_flows_schedule to avoid creating a duplicate, and to obtain the scheduleId required by workbench_flows_schedules_update / _delete / _run_now. Alternatives and the condition selecting them are named outright.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_schedules_run_nowFire a flow schedule once, nowAInspect
Runs one iteration of a schedule immediately through the real scheduled pipeline (same input, same delivery — the email or channel it normally posts to) without moving its cadence. Works on paused schedules. Use it to test a schedule the user just set up. Spends workspace credit and delivers for real, so it sits behind the approval gate — may return needs_confirmation. The run is fire-and-forget: check workbench_flows_schedules_list for lastStatus, or workbench_executions_get for the trace.
| Name | Required | Description | Default |
|---|---|---|---|
| flowId | Yes | Flow the schedule belongs to. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| scheduleId | Yes | Schedule id, from workbench_flows_schedules_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, destructiveHint=false, openWorldHint=true. The description adds the real behavioral weight: it spends workspace credit, delivers for real to the email/channel, sits behind an approval gate that may return needs_confirmation, and is fire-and-forget with pointers to check lastStatus or the trace. This is exactly the mutation/credit context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then layers the operational constraints (credit, real delivery, approval) and the follow-up path in tight sentences. No filler; every clause carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, credit-spending tool with no output schema, the description covers the important gaps: async fire-and-forget semantics and where to read status (schedules_list) or trace (executions_get). An agent has everything needed to call it correctly and handle the confirmation flow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including workspace and approvalId. The description only indirectly relates to approvalId via its needs_confirmation note and adds no format or usage detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource: 'Runs one iteration of a schedule immediately through the real scheduled pipeline.' It explicitly contrasts with the cadence-altering sibling behavior ('without moving its cadence'), so an agent can distinguish it from workbench_flows_schedules_update or workbench_flows_run without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete when-to-use ('Use it to test a schedule the user just set up') and a scope note ('Works on paused schedules'), plus the approval-gate prerequisite. It does not explicitly name the alternative delete/update siblings or state when NOT to use it, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_schedules_updatePause / resume / re-time a flow scheduleADestructiveInspect
Changes one schedule: enabled: false pauses it (configuration kept), enabled: true resumes and recomputes the next fire time from now; interval / scheduleTime / scheduleTimezone / scheduleDayOfWeek change the cadence; label, delivery, and input can change too. Pass only what changes. Get the scheduleId from workbench_flows_schedules_list. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | New fixed input envelope — {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}. | |
| label | No | ||
| flowId | Yes | Flow the schedule belongs to. | |
| enabled | No | false pauses, true resumes. | |
| delivery | No | ||
| interval | No | New cadence: hourly, daily, or weekly. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| scheduleId | Yes | Schedule id, from workbench_flows_schedules_list. | |
| scheduleTime | No | ||
| scheduleTimezone | No | ||
| scheduleDayOfWeek | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations flag destructiveHint=true and readOnlyHint=false, so the mutation profile is already known. The description goes beyond that by explaining that pausing keeps configuration, that resuming recomputes the next fire time from now, and that the call can return needs_confirmation (a two-step approval flow). It leaves some gaps (permission requirements, effects on in-flight runs) but adds real behavioral value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the tool's core behavior and dense with useful field mapping; every clause carries information. The semicolon-chained field list is efficient though a touch packed, and the scheduleId/confirmation notes are properly appended.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter destructive mutation with no output schema and 58% schema coverage, the description covers the confirmation flow, the partial-update contract, and the key field semantics. Minor uncovered params (workspace) and the absence of return-value detail are acceptable given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 58% schema description coverage across 12 params, the description compensates by naming enabled, interval, scheduleTime, scheduleTimezone, scheduleDayOfWeek, label, delivery, and input and describing their effects. It doesn't mention workspace or approvalId directly, though the needs_confirmation note implies the latter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (changes) plus resource (one schedule) and then enumerates exactly what each field does, including the pause/resume semantics. An agent can distinguish this from workbench_flows_schedules_list, _delete, and _run_now without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational guidance: 'Pass only what changes' (partial update), where to obtain the scheduleId, and that the call may return needs_confirmation. It does not explicitly contrast with alternatives (delete a schedule, run now, create a new schedule via workbench_flows_schedule), so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_flows_updateRename / describe / archive / re-scope a Workbench flowADestructiveInspect
Changes a flow's name, description, visibility, tags, or archived state — the metadata around the flow, NOT its orchestration (prompts and nodes change through workbench_flows_edit_text). Pass only what changes. Archiving hides the flow from the default gallery but leaves it runnable, scheduled, and shared; pass archived: false to restore. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Replacement tag set (lowercase labels). The whole set, not a diff. | |
| flowId | Yes | Flow id, from workbench_flows_list. | |
| archived | No | true archives, false restores. | |
| flowName | No | New name. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No | New description; null clears it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the mutation/destructive profile, so the description earns credit for going further: it explains that archiving hides the flow from the default gallery while leaving it runnable, scheduled, and shared, that archived:false restores, and that the call may return `needs_confirmation`. These are real behavioral facts an agent needs that the schema does not spell out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operation and its scope, then the exclusion, then partial-update rule, then the archive semantics. Every sentence carries information and none is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a metadata-mutation tool with no output schema this is complete: it covers scope, exclusion, update semantics, the destructive-ish archive behavior, and even the confirmation round-trip. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning the schema lacks — the whole-set partial-update contract ('Pass only what changes') and the restore semantics of archived:false. This is useful framing beyond the field-level descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (changes) and resource (flow metadata: name, description, visibility, tags, archived state) and explicitly scopes out what it does NOT do — orchestration changes route to workbench_flows_edit_text. An agent can distinguish this from the sibling edit tool without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative (workbench_flows_edit_text) with the condition that selects it (prompts/nodes = orchestration) and gives the partial-update rule 'Pass only what changes.' No explicit when-NOT beyond the orchestration split, but routing guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_add_documentsAdd documents to a knowledge baseAInspect
Adds documents to a knowledge base and starts ingestion (chunk + embed; graph KBs also extract entities; tabular KBs load CSV as tables). Text goes inline (Markdown, plain text, CSV — encoding utf8); binary files (PDF, .docx) go base64-encoded with their mimeType. Up to 50 sources per call, ~10 MB each. Ingestion runs in the BACKGROUND — the result carries a runId; check it with workbench_kb_ingestion_status when the user asks, don't poll. workbench_kb_search works once it finishes. Content from blocks is ideal source material. Check workbench_kb_documents_list first so you don't add a document twice. Embedding spends workspace inference credit, so this sits behind the approval gate: may return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| kbId | Yes | Knowledge base id (from workbench_kb_list or workbench_kb_create). | |
| sources | Yes | Documents to ingest (1-50). | |
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the basic safety profile (readOnly=false, destructive=false, openWorld=false). The description goes well beyond: ingestion runs in the BACKGROUND and returns a runId, embedding spends workspace inference credit, and the call sits behind an approval gate that may return a needs_confirmation envelope. This is exactly the extra behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and ingestion behavior, then constraints and routing. It is dense and somewhat long, but nearly every clause carries operational value (limits, encoding, background semantics, credit spend); the density is justified rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still closes the loop by explaining that the result carries a runId, how to check it, and the possible needs_confirmation envelope. Combined with the limits and encoding rules, an agent has everything required to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: text goes inline as Markdown/plain text/CSV in utf8, binary files go base64-encoded with their mimeType, and it restates the 50-source / ~10 MB limits. It also points to <attached-file> blocks as ideal source material.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Adds documents to a knowledge base and starts ingestion') and immediately differentiates by ingestion mode (text vs binary vs graph vs tabular). It also names the sibling tools it interacts with, so an agent can place it against workbench_kb_search and workbench_kb_documents_list without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when/when-not guidance and alternatives: check workbench_kb_documents_list first to avoid duplicates, use workbench_kb_ingestion_status to check progress and explicitly 'don't poll', and workbench_kb_search only works once ingestion finishes. Names the approval gate as a condition that may return needs_confirmation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_createCreate a knowledge baseAInspect
Creates an empty knowledge base. Recipe 'docs' (default) for reference material an agent searches at runtime; 'tabular' for spreadsheet-style data queried with SQL; 'graph' for entity/relationship extraction. After creating, add content with workbench_kb_add_documents, then attach the KB to an agent in the Workbench editor (or tell the user to). May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Knowledge base name (1-200 chars). | |
| recipe | No | Default 'docs'. | |
| workspace | No | Workspace slug override. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. | |
| visibility | No | Who can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE. | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare this is a non-destructive write (readOnlyHint=false, destructiveHint=false). The description adds meaningful context beyond the annotations: the KB starts empty, and it notes the 'May return `needs_confirmation`' flow, which signals a possible approval step. It does not detail permissions or the confirmation payload, but the added disclosure is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by recipe semantics, then the follow-up workflow. Three sentences, each carrying distinct, useful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param write tool with 83% schema coverage and no output schema, the description covers the essentials: what it creates, how to select a recipe, the next steps, and the confirmation possibility. Minor gaps remain (workspace defaulting, error behavior), but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so most parameters are self-documented and the schema covers visibility in depth. The description adds genuine meaning for the recipe values ('docs'/'tabular'/'graph'), which the schema only labels as 'Default docs'. This elevates it above the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates an empty knowledge base') and immediately differentiates itself from the sibling workbench_kb_add_documents by describing the create-then-add workflow. The recipe enumeration further clarifies the tool's scope and variants.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear lifecycle: create, then add content with workbench_kb_add_documents, then attach to an agent in the editor. The recipe descriptions also guide selection ('docs' for runtime-searched reference material, 'tabular' for SQL, 'graph' for entity extraction). It lacks explicit when-not-to-use conditions, but the context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_deleteDelete a knowledge baseADestructiveInspect
Deletes a knowledge base and everything in it. Flows that search it will find nothing — check which agents use it (the user knows; the KB page in Workbench lists them) and say so before proposing. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| kbId | Yes | Knowledge base id, from workbench_kb_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, but the description adds substantive behavior: the delete is cascading ("everything in it"), has downstream effects ("Flows that search it will find nothing"), and may short-circuit into an approval flow returning needs_confirmation. That is exactly the extra context a destructive tool's annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and its scope before the cautionary guidance. The parenthetical "(the user knows; the KB page in Workbench lists them)" is slightly clunky but earns its place by telling the agent where the consumer list lives.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema destructive tool, the description covers scope, downstream impact, and the confirmation handshake, which is most of what an agent needs. It omits whether deletion is permanent/irreversible and what permissions are required, leaving a small but real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so kbId, workspace, and approvalId are all documented in the schema itself. The description only indirectly gestures at approvalId via "May return needs_confirmation"; it adds no format or constraint detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Deletes a knowledge base") and immediately scopes the blast radius ("everything in it"), which distinguishes it from the sibling workbench_kb_document_delete that removes a single document. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete precondition — check which agents use the KB and warn the user before proposing deletion — plus the confirmation path ("May return needs_confirmation"). It does not explicitly contrast against sibling alternatives such as workbench_kb_document_delete, so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_document_deleteRemove a document from a knowledge baseADestructiveInspect
Removes one ingested document and all its chunks; searches stop returning it at once. Get the documentId from workbench_kb_documents_list. To replace a document, remove it then workbench_kb_add_documents the new version. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| kbId | Yes | Knowledge base id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| documentId | Yes | Document id, from workbench_kb_documents_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag this as destructive and non-read-only; the description adds real behavior: chunks are removed wholesale, search results stop returning the document immediately, and the call may return `needs_confirmation`. That confirmation signal pairs with the approvalId parameter and is the kind of context annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each load-bearing: effect, id sourcing, replacement path, and confirmation caveat. The destructive effect is front-loaded rather than buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the key things an agent must know: irreversibility of chunk removal, immediate search impact, and the needs_confirmation/approvalId handshake. It omits any note on failure modes or whether removal is reversible, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so kbId, workspace, approvalId, and documentId are already documented in the schema, including the workspace-token nuance. The description only reinforces that documentId comes from workbench_kb_documents_list, adding marginal value over the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Removes one ingested document') and adds scope detail ('and all its chunks'), which cleanly separates it from workbench_kb_delete (whole KB) and workbench_flows_delete. An agent can pick it out of the densely populated workbench_kb_* family without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete preconditions ('Get the documentId from workbench_kb_documents_list') and a workflow alternative for replacement ('remove it then workbench_kb_add_documents the new version'). It lacks an explicit when-not statement (e.g., when to use KB-level delete instead), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_documents_listList the documents in a knowledge baseARead-onlyInspect
Every document ingested into the KB's current content: id, name, type, size, chunk count. Empty for a KB with nothing ingested yet (or one whose ingestion is still running — check workbench_kb_ingestion_status). Use it before workbench_kb_add_documents to avoid re-adding a document, and to get the documentId for workbench_kb_document_delete.
| Name | Required | Description | Default |
|---|---|---|---|
| kbId | Yes | Knowledge base id, from workbench_kb_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, closed-world), and the description adds genuinely new behavior: the result is empty both when nothing has been ingested and when ingestion is still running, with a pointer to workbench_kb_ingestion_status. It stops short of pagination or ordering behavior, but the empty-state disclosure is the operationally important one.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler, and the critical scope and return fields come first. The parenthetical about the empty state is dense but placed immediately after the claim it qualifies, so nothing important is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by naming the returned fields. Combined with the empty-state semantics, the ingestion cross-reference, and the two downstream use cases, an agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so kbId and workspace are already documented with their constraints (workspace slug rules, personal-token requirement). The description adds nothing about either parameter. Baseline 3 applies when the schema carries the full parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list documents in a KB) and pins the exact scope: 'Every document ingested into the KB's current content.' It even enumerates the returned fields (id, name, type, size, chunk count), which no sibling tool claims, so an agent can separate it from workbench_kb_search or workbench_kb_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives two concrete when-to-use conditions: call before workbench_kb_add_documents to avoid re-adding a document, and call to obtain the documentId required by workbench_kb_document_delete. It also routes the empty-result case to workbench_kb_ingestion_status, so no ambiguous path is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_getRead one knowledge baseARead-onlyInspect
One knowledge base: name, recipe (docs / tabular / graph), status, document and chunk counts, embedding model, and ingestion settings. Read it to confirm a KB has content before pointing a flow at it. Get the kbId from workbench_kb_list.
| Name | Required | Description | Default |
|---|---|---|---|
| kbId | Yes | Knowledge base id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds value beyond that by disclosing what the read returns (name, recipe, status, document/chunk counts, embedding model, ingestion settings). It does not address auth/scope behavior, but the safety profile is already handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the returned fields and following with the usage condition. No filler. The field enumeration is a slightly long list but each item is informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden, and it does so by listing the concrete fields an agent would receive. For a simple single-record read with full schema coverage and safety annotations, this is nearly complete; only auth/scope behavior is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both kbId and workspace (including the workspace-auth nuance). The description adds only provenance guidance ('Get the kbId from workbench_kb_list'), so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (one knowledge base) and enumerates the fields it returns, and the title supplies the read verb. It implicitly distinguishes itself from the plural sibling workbench_kb_list by pointing to it for the kbId, though the verb itself is left to the title. Clear and unambiguous, just not a fully self-contained verb+resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit usage trigger: 'Read it to confirm a KB has content before pointing a flow at it,' which tells the agent when this tool is the right call. It also routes to workbench_kb_list to obtain the id. There is no explicit when-not-to-use or statement of alternatives for the read itself, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_ingestion_statusCheck an ingestion runARead-onlyInspect
The state of one ingestion run started by workbench_kb_add_documents: PENDING / RUNNING / DONE / FAILED, chunks written, and the error when it failed. Read it when the user asks whether their documents are in yet, or before a search that needs them — don't poll in a loop; ingestion of a few documents takes under a minute. Pass the runId the add-documents call returned.
| Name | Required | Description | Default |
|---|---|---|---|
| kbId | Yes | Knowledge base id. | |
| runId | Yes | Ingestion run id, from workbench_kb_add_documents. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this a read-only, non-destructive, closed-world operation, but the description adds context annotations cannot carry: the exact state machine (PENDING/RUNNING/DONE/FAILED), what is reported on success (chunks written) and failure (the error), and the latency expectation that discourages polling. It stops short of describing pagination or retention of run records, keeping it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: what is returned, when to call it (plus the polling warning), and where the key parameter comes from. The return shape is front-loaded ahead of the routing advice.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned status values, the chunk count, and the error field, and it covers a 3-parameter tool whose schema is fully documented. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so kbId, runId, and workspace are already documented in the schema; the description only adds provenance for runId ('the add-documents call returned'). Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (the state of one ingestion run) and ties it explicitly to its producer, workbench_kb_add_documents, which separates it from every sibling in the kb family. An agent can identify this as a status-polling tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers ('when the user asks whether their documents are in yet, or before a search that needs them') and an explicit anti-pattern with the reason ('don't poll in a loop; ingestion of a few documents takes under a minute'). It also names the sibling that supplies the required runId.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_listList knowledge basesARead-onlyInspect
Lists every knowledge base in the active workspace the caller can see. Returns summaries (id, name, description, status, doc/chunk counts, embedding model). Use an id with workbench_kb_search to retrieve content.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds real value beyond that: it discloses that results are visibility-filtered to the caller's workspace and enumerates the summary payload (id, name, description, status, doc/chunk counts, embedding model). It doesn't address pagination or result-size limits, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler, with the listing scope front-loaded and the follow-on routing placed last. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned summary fields, and the single parameter is fully documented in the schema. What remains unstated is pagination/ordering behavior and what an empty result means, which are minor for a read-only list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single workspace parameter, including the token-vs-API-key nuance, so the schema carries the burden. The description only says 'active workspace' and adds no format or default-override detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists every knowledge base') plus the visibility scope ('in the active workspace the caller can see'), which separates it from workbench_kb_get and workbench_kb_search. It even names the return fields, so an agent knows this is the discovery/list tool rather than a content fetch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for it (enumerating KBs visible to the caller) and explicitly routes the follow-on action: 'Use an id with workbench_kb_search to retrieve content.' It does not contrast with workbench_kb_get or explain when listing is preferable to searching, so it stops short of full alternative coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_kb_searchSearch a knowledge baseARead-onlyInspect
Retrieves the most relevant chunks from one knowledge base — the core RAG primitive. mode:"semantic" (default) embeds the query and ranks by meaning (spends a small amount of workspace credit); mode:"keyword" is a free case-insensitive substring match. Each hit carries the chunk text, its document, and a score (cosine similarity 0-1 for semantic, match count for keyword). Get a kbId from workbench_kb_list.
| Name | Required | Description | Default |
|---|---|---|---|
| kbId | Yes | Knowledge base id (from workbench_kb_list). | |
| mode | No | semantic (embed + rank by meaning, default) or keyword (free substring match). | |
| limit | No | Max hits to return, 1-50. Default 10. | |
| query | Yes | What to search for. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and non-destructive status, and the description adds genuinely new context: semantic mode spends workspace credit, keyword is free and case-insensitive, and hits carry text, document, and score. It stops short of describing limits on behavior such as indexing freshness or empty-result handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler; the core primitive statement, mode behavior, and kbId source are all front-loaded and each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the return shape (chunk text, document, score, and score scale per mode). Combined with the workspace-credential note in the schema, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents kbId, mode, limit, query, and workspace. The description restates mode defaults and clarifies score semantics, but adds little parameter detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retrieves) and resource (most relevant chunks from one knowledge base) and frames it as the core RAG primitive. The scoping to 'one knowledge base' distinguishes it from the broader search_docs/search_workspace siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains the mode tradeoff (semantic embeds and ranks by meaning, keyword is free substring match) and points to workbench_kb_list for obtaining a kbId. It does not, however, say when to prefer this tool over search_docs or search_workspace, leaving that boundary to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_node_catalog_getSearch the Workbench node catalogARead-onlyInspect
The ground truth for what nodes exist and what their ports actually are — verify against this instead of recalling. Search by keyword/category for summaries; pass slug for one node's full detail (inputs, outputs, settings). Inputs are PORTS fed by links; settings live on the node — see workbench_flow_authoring_guide. The catalog is the WORKSPACE'S: model nodes the workspace's inference policy forbids come back with allowed: false — never author with those; pick an allowed model. Pass flowId to include the flow's pinned imports as nodes.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Keyword over slug / name / tagline / description. | |
| slug | No | Exact node slug → full detail for that one node (ports + settings). | |
| limit | No | Results per page, 1-50. Default 20. | |
| flowId | No | A flow whose pinned imports should appear as `imported-<uuid>` nodes. | |
| offset | No | Pagination offset. | |
| category | No | Filter by category. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; the catalog's policy annotation is per workspace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish a safe read-only, closed-world profile, but the description adds real behavioral context: the catalog is workspace-scoped, policy-forbidden model nodes are returned with `allowed: false`, and personal tokens without a default workspace must pass `workspace`. These are non-obvious traits an agent could not infer from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and dense with no filler: the ground-truth claim comes first, then search/detail modes, then the workspace policy caveat. It is packed tightly and borders on over-compressed, but every sentence conveys distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, 7-parameter tool with no output schema, the description covers purpose, key parameter roles, workspace policy behavior, and the cross-reference an agent needs. It does not address the paginated result shape or how to interpret `limit`/`offset` behavior, which is a minor residual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds conceptual meaning the schema lacks – that `q` returns summaries while `slug` returns full detail, `flowId` surfaces pinned imports as `imported-<uuid>` nodes, and `inputs` are PORTS fed by links while settings live on the node. Pagination semantics (limit/offset) stay in the schema only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource with scope: it is the ground truth for what nodes exist and what their ports are, searchable by keyword/category for summaries or by `slug` for full per-node detail. It also routes away from the sibling `workbench_flow_authoring_guide` for settings-vs-ports guidance, so an agent can distinguish it from related workbench tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance ('verify against this instead of recalling'), an exclusion ('never author with those [allowed: false] nodes'), and a named cross-reference for the related concern in `workbench_flow_authoring_guide`. This is actionable routing, not implied usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_shim_add_answerAdd an answer to a shimBInspect
Adds one answer (a label and, optionally, when it applies). The shim needs three examples of it before it builds again with the new answer. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| answer | Yes | ||
| shimId | Yes | Shim id, from workbench_shim_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| description | No | When this answer applies. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only covering safety flags (readOnly=false, destructive=false, openWorld=false), the description adds real behavioral context: the confirmation gate ('May return `needs_confirmation`') and the prerequisite of three examples before the shim rebuilds. It still doesn't say what state the answer is added in or whether it can be undone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, with no filler. Slightly dense wording around 'three examples of it' but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param mutation tool with no output schema, the description hints at the confirmation return path but does not explain what happens after needs_confirmation or how approvalId closes the loop. It omits any note on workspace scoping (left to the schema) and does not describe the effect of adding an answer before the rebuild threshold is met.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents shimId, workspace, approvalId, and description. The description loosely maps to the `answer` and `description` params ('a label and, optionally, when it applies') but adds no syntax or format detail beyond the schema's maxLength/minLength constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ('Adds one answer' to a shim), so the operation is identifiable. However, it does not distinguish itself from the adjacent sibling workbench_shim_add_examples, and the phrase 'a label and, optionally, when it applies' is ambiguous about what an 'answer' actually is (the schema calls it a plain string).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus workbench_shim_add_examples or workbench_shim_create. The only contextual hint is the behavioral note that three examples are needed before a rebuild, which is not a usage-selection rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_shim_add_examplesAdd examples to an answerAInspect
Adds example inputs to one answer. Write inputs a real person would actually type in the shim's setting — varied in length, register and specifics, each one clearly this answer and not another — never paraphrases of the answer's name. Read the shim first so new examples fit alongside the existing ones and land where recall is low. Exact duplicates are skipped. The shim rebuilds on its own afterwards; read it again for the new report. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| answer | Yes | The answer these examples belong to, exactly as labelled. | |
| shimId | Yes | Shim id, from workbench_shim_list. | |
| examples | Yes | ||
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (readOnly=false, destructive=false, openWorld=false); the description adds substantial behavior beyond that — exact duplicates are silently skipped, the shim rebuilds asynchronously on its own, the shim must be re-read for the new report, and a `needs_confirmation` path exists. That is exactly the mutation/confirmation detail an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then guidance, then the confirmation caveat. Sentences are dense but each carries distinct information; the middle exemplar sentence is long but earns its length by defining the value contract.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly covers the post-call state (rebuild, re-read, possible needs_confirmation). Combined with 80% schema coverage and annotations, an agent can call and handle the result correctly; only the absence of an explicit alternative-routing statement keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the params are largely documented. The description still adds real meaning: what a good value for `examples` looks like (varied length, register, specifics, distinct from other answers, never paraphrases of the answer's name), which the schema's minLength/maxItems constraints cannot convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource ('Adds example inputs to one answer'), which cleanly separates it from sibling workbench_shim_add_answer. It does not explicitly name the sibling it differs from, but the resource scoping ('one answer') makes the target unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operating context: read the shim first so examples fit alongside existing ones and 'land where recall is low', which tells the agent both sequencing and intent. It stops short of naming when-not-to-use conditions or alternatives to add_answer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_shim_createCreate a shimAInspect
Creates a shim from a name, the question it answers, and its answers. Add examples afterwards with workbench_shim_add_examples — three or more per answer and it builds on its own. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | A short name, e.g. "Support triage". | |
| answers | Yes | ||
| question | Yes | The question, e.g. "Which team should read this message?" | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. | |
| inputsComeFrom | No | One line on where the inputs come from, in the user's words. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-destructive, closed-world write, and the description adds real behavior beyond that: the tool may return a `needs_confirmation` state, and three or more examples per answer cause the shim to "build on its own." The auto-build behavior and confirmation interlock are not present in the annotations or schema, so this is genuine added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the creation act and the inputs, followed by the follow-up step and the confirmation caveat. Nothing is redundant, though the phrase "it builds on its own" is slightly vague for a behavioral claim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on return-value disclosure and does flag the `needs_confirmation` outcome, which the schema's approvalId field complements. It stops short of describing the happy-path response or what state the shim is in after creation (draft vs active), which an agent planning follow-up calls might want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so parameters are largely self-documenting; the description names the three required parameters (name, question, answers) in natural language but adds no format, syntax, or constraint detail beyond the schema. The optional workspace, approvalId, and inputsComeFrom parameters are never mentioned. Baseline 3 is appropriate when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Creates a shim") and immediately names the three inputs that constitute it, so the agent knows what it will be building. It is distinguishable from siblings like workbench_shim_add_examples or workbench_shim_get. It never defines what a 'shim' actually is, relying on the question/answers structure to imply a routing/classification artifact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit sequencing guidance ("Add examples afterwards with workbench_shim_add_examples"), which usefully directs the agent's next call. It does not, however, say when to reach for shim_create versus other workbench primitives such as flows or tasks, nor state prerequisites beyond the approval flow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_shim_getRead one shimARead-onlyInspect
One shim in full: the question, where its inputs come from, every answer with its description and examples, what the shim still needs before it can build, and the newest build's report — held-out accuracy, recall per answer, which answers get mistaken for which, and the compiler's notes. Read it before writing examples: match the setting and the existing examples' register, and put new examples where recall is low or the compiler says an answer is thin.
| Name | Required | Description | Default |
|---|---|---|---|
| shimId | Yes | Shim id, from workbench_shim_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (readOnly, non-destructive, closed-world), and the description adds substantial context beyond that: it discloses the full return shape including held-out accuracy, per-answer recall, confusion pairs, the compiler's notes, and what the shim still needs before it can build. With no output schema, this disclosure is doing real work and reveals behavior an agent would otherwise have to discover empirically.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the payload contents, then a second sentence of prescriptive guidance. It is dense and the first sentence is a long enumeration, but every clause maps to a real part of the returned report, so little is wasted; a light trim would make it a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-resource read with no output schema, the description fully covers what comes back, why to call it, and how to act on it. Parameters are covered by the schema and safety by the annotations, so nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — shimId is documented as coming from workbench_shim_list and workspace covers default/override semantics for personal vs API-key tokens. The description adds nothing about either parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific scope statement, 'One shim in full,' and then enumerates exactly what a shim is composed of (the question, input sources, answers with descriptions/examples, build requirements, newest build report). This lets an agent distinguish it from siblings like workbench_shim_list (plural/roster) and the shim mutation tools (create/add_answer/add_examples) without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage trigger, 'Read it before writing examples,' and even prescribes what to do with the result ('match the setting and the existing examples' register,' target low-recall answers). That is strong when-to-use and workflow guidance, though it does not name the sibling action tool explicitly or state a when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_shim_listList shimsARead-onlyInspect
Every shim the user can see: id, name, the question it answers, how many answers and examples it has, and the published version's held-out accuracy (null until something is published). A shim is a small decision model that runs inside an app with no model call. Start here to get a shimId.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, closed-world). The description adds real value beyond them: visibility scope ('the user can see'), the null-accuracy-until-published semantic, and the fact that a shim runs with no model call. It does not mention pagination or result limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the returned-field enumeration, then the definition of a shim, then the entry-point cue. The field list is dense but earns its place by compensating for the absent output schema; only minor trimming possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, enumerating returned fields is the right move and is done well; annotations cover safety. The only gap is that nothing indicates result size, pagination, or ordering for a list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema fully explains the workspace slug including the personal-token default/override rules. The description says nothing about the workspace parameter, so it adds no meaning beyond the schema — the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Every shim the user can see') and enumerates the exact fields returned, including a mini-definition of what a shim is. 'Start here to get a shimId' distinguishes it from workbench_shim_get, which needs an id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions itself as the entry point ('Start here to get a shimId'), which tells the agent when to use this over workbench_shim_get. No explicit when-not or exclusion conditions are given, but the entry-point framing covers the main routing decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_actAct on a Task gate (record a verdict)AInspect
Records THIS user's verdict on a run that is waiting on them — the labels come from the gate (workbench_tasks_run_view shows them via the plan; a wrong label errors listing the real options). Echo the verdict + note back to the user in one line BEFORE calling; that echo is the confirmation. Refuses when the gate is assigned to someone else.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | The user's reasoning, attached to the verdict. | |
| runId | Yes | Task run id that is waiting_human. | |
| verdict | Yes | One of the gate's verdict labels, e.g. "confirmed". | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare mutation/no-destruct/open-world. The description adds real behavioral context: it refuses when the gate belongs to another user, a wrong label errors listing the real options, and it introduces the two-step approvalId/needs_confirmation flow plus the required echo-as-confirmation pattern.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Records THIS user's verdict...') in one dense paragraph. The echo instruction is packed with clauses but every sentence earns its place; slightly convoluted but no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with an approval/confirmation flow, the description covers ownership refusal, error behavior on bad labels, and the echo confirmation step. With no output schema and full schema coverage, this is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all five parameters (including note, runId, verdict, workspace, approvalId) are already documented. The description reinforces that verdict labels originate from the gate, but adds little syntactic or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: records THIS user's verdict on a task run waiting on them. It distinguishes itself from siblings by referencing workbench_tasks_run_view as where the gate labels are shown, so an agent can tell it apart from workbench_tasks_start/freeze/get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context — use it on a run that is waiting on the user — and names the sibling that surfaces the gate labels. It notes the refusal condition (gate assigned to someone else), though it doesn't explicitly contrast with other task lifecycle tools like start or freeze.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_compileCompile a sentence into a Task planAInspect
Runs the compiler over a sentence and returns the PROPOSED plan (or repair-grade problems — fix by refining the sentence with the user, not by guessing). Nothing persists. Present the proposal in plain terms: which clauses run flows, which you (zv1) will handle, who decides, what the compiler noted. Then iterate or freeze. Compile may take ~15s.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | No | The task this sentence belongs to, when compiling an edit — keeps it out of the child-task candidates. | |
| sentence | Yes | The sentence to compile. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Re-call with the approvalId from a needs_confirmation response after the user approves. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds context annotations do not carry: nothing is persisted, the call can return repair-grade errors rather than a plan, and it may take ~15s. There is mild tension between “Nothing persists” and readOnlyHint=false, but the schema's approvalId / needs_confirmation flow plausibly accounts for that, so this reads as added detail rather than a contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded: the first sentence gives the action and the return shape, followed by guidance, then the latency note. Dense but every clause is doing work; the “which you (zv1) will handle” phrasing is a touch cryptic but purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully characterizes the return (proposed plan, compiler notes, repair problems) and the iteration/approval loop, plus latency. It stops short of explaining the plan's structure or the confirmation response shape, but is sufficient to call and act on it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the four parameters (taskId, sentence, workspace, approvalId) is already documented in the schema. The description only reinforces the sentence-refinement loop and adds nothing about formats or limits, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — “runs the compiler over a sentence” — and names the artifact it returns: the PROPOSED plan (or repair-grade problems). The “Nothing persists” clause and the “iterate or freeze” ending implicitly separate it from workbench_tasks_freeze/workbench_tasks_create, which commit state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context (“fix by refining the sentence with the user, not by guessing”, “Then iterate or freeze”) and tells the agent what to do after the call. It does not name the sibling that commits the plan (freeze/start), so a little inference is still required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_createCreate a draft TaskAInspect
Creates a DRAFT Task from a name + sentence. Use this once the user's sentence feels settled after you've refined it together — then compile it. The sentence should be THEIR words: trigger first ('When a claim arrives…'), then steps, branches, and who decides what.
| Name | Required | Description | Default |
|---|---|---|---|
| sentence | Yes | The process sentence, in the user's words. | |
| taskName | Yes | Short task name. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare write-but-non-destructive (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the safety bar is lower. The description adds genuine context beyond that: the created object is a DRAFT rather than an active task, and it encodes a readiness gate ('sentence feels settled') before creation. It doesn't describe what is returned or the confirmation/approval round-trip, but the schema covers the latter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope, followed by the usage condition and then format guidance, in three tight sentences. The em-dash aside and the 'who decides what' clause are slightly conversational but each carries usable instruction rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still conveys enough to call it correctly: what is created, its draft state, the readiness condition, and the downstream compile step. Minor gaps remain — what the call returns (e.g., a task id needed for the subsequent compile) and the two-phase approvalId flow are left to the schema rather than narrated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents all four parameters, including 'in the user's words.' The description partly repeats that phrasing, but it adds real semantic guidance the schema lacks: the sentence structure should be trigger-first ('When a claim arrives…'), then steps, branches, and decision ownership. That is actionable authoring guidance for the most important parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb + resource + lifecycle scope ('Creates a DRAFT Task from a name + sentence'), which immediately separates it from the many sibling task operations (compile, start, freeze, update). The word 'DRAFT' plus 'then compile it' tells an agent exactly where this sits in the task lifecycle without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear readiness condition — use it 'once the user's sentence feels settled after you've refined it together' — and points to the natural next step ('then compile it'). It does not state what to do instead when the sentence is NOT settled (e.g., keep refining, or use update on an existing task), so there are no explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_deleteDelete a Workbench TaskADestructiveInspect
Deletes a task. Runs still in flight are stopped (a gate waiting on someone is closed), and its schedule trigger stops firing; past runs stay in history. Check workbench_tasks_get for live runs and say so before proposing. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | Task id, from workbench_tasks_list. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite destructiveHint=true already flagging the risk, the description adds substantial context: in-flight runs are stopped, a waiting gate is closed, the schedule trigger stops firing, and past runs remain in history. It also discloses the needs_confirmation return, which is exactly the kind of behavior an agent must anticipate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action, then concisely covers consequences, the recommended pre-check, and the confirmation flow. Every sentence carries information, though the 'say so before proposing' phrasing is slightly informal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully discloses the possible needs_confirmation return and the fate of in-flight vs. past runs. Complete enough for a destructive tool, though the confirmation flow could be spelled out a touch more.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so taskId, workspace, and approvalId are already documented in the schema. The description's mention of needs_confirmation gives indirect context for approvalId, but adds no syntax or format detail beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Deletes a task'), so an agent immediately knows the operation. It does not differentiate from siblings like workbench_tasks_freeze or workbench_tasks_update, but the purpose itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit precondition workflow: 'Check workbench_tasks_get for live runs and say so before proposing.' It names the alternative tool and the condition that matters before deletion, though it stops short of stating when-not to delete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_freezeFreeze a compiled plan (activate the Task)ADestructiveInspect
Freezes a reviewed plan as the Task's next version and activates it. Pass the EXACT plan object a compile returned — never hand-edit it (change the sentence and recompile instead). Freeze ONLY after the user has seen the proposal and said go; echo what you're freezing in one line first. In-flight runs keep their pinned versions. After freezing, offer to start a first run.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | Yes | The plan object from workbench_tasks_compile, verbatim. | |
| taskId | Yes | The task id (create a draft first if none). | |
| sentence | Yes | The sentence this plan was compiled from. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description goes well beyond that by scoping the damage: it is a versioned activation, and in-flight runs keep their pinned versions. It also surfaces a confirmation discipline ('echo what you're freezing in one line first') that the destructive annotation alone would not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each carrying a distinct obligation: what freeze does, how to pass the plan, when it is permitted, and what follows. The irreversible action and its precondition are front-loaded with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the essential agent concerns: the source of the plan argument, the confirmation gate, the effect on running work, and the natural next step. Nothing an agent needs in order to call it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning for the plan/sentence pair: the plan must be the verbatim object a compile returned, and if the sentence changes the plan must be recompiled rather than edited. That is a coupling constraint the schema does not express. It says nothing extra about workspace or approvalId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (freezes) plus the exact outcome: turns a reviewed plan into the Task's next version and activates it. It also distinguishes itself from workbench_tasks_compile by requiring output from that tool and by describing the activation step, which no sibling performs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions (only after the user has seen the proposal and said go), an explicit anti-pattern with the correct alternative (never hand-edit the plan; change the sentence and recompile), and a post-action recommendation (offer to start a first run). Nothing about when to reach for this tool is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_getRead one Workbench Task (definition + recent runs)ARead-onlyInspect
One Task in full: its sentence, status, the latest saved plan (the clauses — which run flows, which zv1 handles, who decides at each gate) with the flow and people names the plan refers to, and its recent runs. Read this before editing a task's sentence (workbench_tasks_compile with taskId) and when the user asks what a task does. Get the taskId from workbench_tasks_list.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | Task id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds genuine value by disclosing what the read returns and how deep it goes (plan clauses, gate deciders, recent runs) — not just 'gets a task.' It stops short of describing run recency limits or whether the plan may be absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource and its contents in a single dense sentence, then two short routing sentences. Slightly heavy on parenthetical enumeration but every clause serves differentiation or elicitation; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey return shape — and it does, naming the sentence, status, latest plan with clause-level detail, referenced flow/people names, and recent runs. Combined with annotations and a fully-described 2-param schema, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both taskId and the workspace/auth-override semantics are already documented. The description adds only the provenance hint for taskId (from workbench_tasks_list) and says nothing about the workspace parameter; baseline 3 fits when the schema carries the weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('One Task in full') and enumerates what comes back: sentence, status, latest saved plan/clauses, flow and people names, recent runs. It contrasts clearly with siblings workbench_tasks_list (source of ids) and workbench_tasks_compile (edit path), so an agent can distinguish it without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use: 'Read this before editing a task's sentence (workbench_tasks_compile with taskId)' and 'when the user asks what a task does.' It also names the prerequisite source for the required id ('Get the taskId from workbench_tasks_list'), leaving no routing inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_listList Workbench TasksARead-onlyInspect
Lists the workspace's Tasks — sentence-orchestrated processes — with their sentence, status, and recent runs (id, status, who a waiting run is on). Use this to find a task the user names ('claims intake') and to see what's in flight; workbench_tasks_get reads one in full.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, closed-world), so the bar is lower. The description adds genuine behavioral context by listing the returned fields (sentence, status, recent runs) and notably the 'who a waiting run is on' detail. It stops short of disclosing pagination, limits, or volume behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and return contents before the routing guidance. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one documented parameter, the description is nearly complete: it covers contents, use case, and sibling routing. Without an output schema it reasonably sketches returned fields, though it doesn't address result volume or pagination.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single workspace parameter is already fully documented in the schema, including default-token and API-key behavior. The description adds nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (the workspace's Tasks), and defines what a Task is (sentence-orchestrated processes) rather than assuming the reader knows. It distinguishes itself from the sibling workbench_tasks_get, which reads one task in full.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions: finding a task the user names (with a concrete example, 'claims intake') and seeing what's in flight. It names the alternative (workbench_tasks_get) and the condition that selects it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_run_viewView a Task run (the lit sentence)ARead-onlyInspect
The run's current state rendered as its sentence: what ran (with timings), which path each gate took, what's waiting on whom and for how long, what's coming up. Use whenever the user asks where something stands, and after acting on a gate.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | Task run id. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/destructiveHint=false, so the safety profile is covered. The description adds real value by disclosing the shape of the rendered result (what ran with timings, which path each gate took, what is waiting and on whom), which goes beyond the annotation set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with what the tool returns and followed by the usage trigger; no filler. The metaphor "rendered as its sentence / the lit sentence" is slightly opaque jargon but does not bloat the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the burden of explaining return content, and it does so via the enumerated categories. Combined with 100% param coverage and read-only annotations, an agent has enough to call it correctly, though the unaddressed sibling overlap is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the runId and workspace parameters are fully documented in the schema; the description adds no syntax, format, or semantic detail. Baseline 3 applies when structured data does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource (view a Task run) and enumerates the content categories it surfaces: timings, gate paths, blockers, upcoming steps. However, it never distinguishes itself from the similarly named sibling workbench_tasks_get, leaving the agent to infer which read tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use whenever the user asks where something stands, and after acting on a gate" gives a clear triggering condition tied to a concrete workflow moment. It stops short of naming an alternative or exclusion (e.g., when to use workbench_tasks_get instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_startStart a Task runAInspect
Kicks off a run of a Task. When the task expects an input document (its trigger has an input), pass the user's material as input — pasted text, extracted attachment text, or JSON. Cite the run id back and tell the user what happens next (the run may immediately be waiting on someone).
| Name | Required | Description | Default |
|---|---|---|---|
| input | No | Trigger payload — the document text / JSON the task starts from. | |
| taskId | Yes | Task id (from workbench_tasks_list). | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=true, so the mutation/side-effect profile is covered. The description adds the useful behavioral note that the run 'may immediately be waiting on someone', but omits the approval/needs_confirmation flow and any auth or rate-limit context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded sentences with the purpose first and no filler. The trailing instruction to cite the run id and describe next steps is slightly prescriptive but still earns its place as post-call guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the primary input semantics and hints at the async run behavior. The approvalId and workspace parameters are handled by the 100%-covered schema, so nothing critical is missing, though the confirmation flow could be clearer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real value on the `input` parameter: it explains WHEN to supply it (task trigger has an input) and what forms it takes (pasted text, extracted attachment text, JSON) beyond the schema's terse 'Trigger payload' text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('kicks off a run') and resource ('a Task'), so an agent knows this initiates a run. It does not explicitly distinguish itself from nearby siblings like workbench_tasks_act or workbench_tasks_run_view, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives conditional guidance for the input (pass material when the trigger has an input), which is useful. However, it never states when to use this over alternatives such as workbench_tasks_act or workbench_tasks_run_view, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbench_tasks_updateRename a Workbench TaskADestructiveInspect
Changes a task's name — the gallery label only. The sentence and plan are versioned and change through workbench_tasks_compile + workbench_tasks_freeze, never here. May return needs_confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | Task id, from workbench_tasks_list. | |
| taskName | Yes | New name. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=true). The description adds genuinely non-structured behavior: the two-step needs_confirmation/approvalId flow and which fields are versioned. It does not explain reversibility or the confirmation trigger condition, but it goes well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: scope, exclusions with named alternatives, and an edge-case return. The most important constraint is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity rename with a rich schema and no output schema, the definition is nearly sufficient — it covers the special needs_confirmation return and the routing to other tools. Minor gap: no explicit statement that the rename itself is applied without confirmation or how long the needs_confirmation state persists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with each parameter (including approvalId's 'omit on the first call' and workspace's token rules) documented inline, so the baseline is 3. The description only alludes to approvalId via 'needs_confirmation' and adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Changes a task's name') and immediately bounds its scope ('the gallery label only'). It also names the two siblings that handle the other fields (workbench_tasks_compile, workbench_tasks_freeze), so the agent can distinguish it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when NOT to use this tool — sentence and plan changes go through compile/freeze, 'never here'. It stops short of spelling out the positive-use conditions or the workspace/approval prerequisites, but the routing guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_credit_getWorkspace credit postureARead-onlyInspect
The workspace's inference-credit position: plan, this month's grant and how much of it is spent, and the purchased balance runs draw on past the grant. Check this before proposing spend-heavy work (eval runs, batch expansions) when credit has come up — and cite the numbers. Read-only; top-ups and plan changes happen in accounts, by a human.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds non-obvious operational context: that credit changes are not possible here and must go through accounts with human involvement. It also instructs the agent to cite the returned numbers, which is behavioral guidance not captured anywhere in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, well front-loaded: return shape first, then the pre-spend trigger, then the read-only boundary. Minimal waste, though the 'when credit has come up' clause is slightly loose and 'cite the numbers' is an instruction rather than a description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still enumerates the returned fields and the scope of the lookup, and it covers the mutation boundary. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'workspace' parameter already documents slug/default/override/API-key behavior fully. The description adds no parameter detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names the specific resource (the workspace's inference-credit position) and enumerates exactly what is returned: plan, this month's grant and spend against it, and the purchased balance that overage draws on. No sibling tool covers credit, so the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the trigger condition ('before proposing spend-heavy work (eval runs, batch expansions) when credit has come up') and routes the agent elsewhere for the mutation path ('top-ups and plan changes happen in accounts, by a human'). Both when-to-use and when-not-to-use are explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_files_readRead a file from the workspaceARead-onlyInspect
Read an uploaded file (PDF, Word, PowerPoint, spreadsheet, text) as markdown. Find file ids with workspace_folders_browse (a folder's items of kind workspace_file) or search_workspace. For anything long, call with outline: true first: it returns the headings with their character offsets and the total length. Then read only the sections you need with offset and maxChars. Spreadsheets come back as tables. A file that can't be read as text says so; don't guess at its contents.
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | Yes | ||
| offset | No | Character offset to start from. Default 0. | |
| outline | No | Return the headings and their offsets instead of text. | |
| maxChars | No | How much to return. Default 12000. | |
| workspace | No | Workspace slug. Ignored when the workspace is already set. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds valuable behavior: outline returns headings with offsets and total length, spreadsheets render as tables, and unreadable files report failure rather than being guessed at. It stops short of describing pagination limits or exact return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five dense sentences, front-loaded with purpose, then id lookup, then the outline-to-section strategy, then format edge cases. No filler; every sentence changes how the agent calls the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry return-value meaning, and it does explain markdown output, heading/offset outlines, and table rendering. Minor gaps remain around default sizes and maxChars cap, which are inferable from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80% (fileId is undocumented anywhere). The description nonetheless gives operational meaning to outline, offset, and maxChars by explaining they form a progressive-read workflow, adding intent beyond the schema's terse field notes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) plus resource (uploaded file) and the output format (as markdown), even enumerating supported file types. An agent can immediately distinguish this from save/list/browse siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: file ids come from workspace_folders_browse or search_workspace, and for long files it prescribes outline:true first, then targeted reads with offset/maxChars. This is when-and-how guidance rather than mere context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_files_save_from_chatSave a picture from this chat to your filesAInspect
Save a picture the user attached in THIS conversation to the workspace's files, in the Home folder "From chat", and get back its file id. Use it when the user wants a picture they sent you used somewhere — an interface shows it as (CSS: url(file:)). Pick the picture by name: the name in its [attached image: <name>] marker. Omit name only when the latest message with pictures has exactly one. Saving the same picture again returns the file it was already saved as. Only pictures from this conversation can be saved; a picture on a website can't be fetched, so ask the user to attach it here.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | The picture's name, as in `[attached image: <name>]`. | |
| workspace | No | Workspace slug. Ignored when the workspace is already set for this chat. | |
| approvalId | No | Approval id from a prior needs_confirmation response. Omit on the first call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), and the description adds real context: idempotency ("Saving the same picture again returns the file it was already saved as"), the storage location, and the returned file id. It stops short of stating auth/permission needs or the confirmation flow, which the schema's approvalId only implies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense, front-loaded paragraph that leads with the action and destination before qualifying conditions. Every sentence carries information, though the naming/omission rules and the idempotency note stack up enough that trimming could help.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter write tool with no output schema, the description covers the action, destination folder, name-resolution rule, idempotent behavior, and return value (file id). Nothing an agent needs to invoke this correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds selection logic the schema lacks: pick `name` from the `[attached image: <name>]` marker and omit it only when the latest picture-bearing message has exactly one image. It also implicitly frames `workspace` as optional-when-already-set, matching the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+destination: save a picture attached in THIS conversation into the workspace's "From chat" Home folder and return its file id. The scope (this conversation, attached pictures) is precise and easily distinguished from the surrounding napkin_* document/sheet tooling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger ("Use it when the user wants a picture they sent you used somewhere") and a hard exclusion ("Only pictures from this conversation can be saved; a picture on a website can't be fetched, so ask the user to attach it here"). It also tells the agent how to resolve the `name` argument and when it may be omitted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_folders_browseBrowse the workspace's foldersARead-onlyInspect
Without a folder: every folder in the workspace as a tree, each with its id, full path, and how many things are filed directly in it. With a folder (by id or by name): what's inside it — subfolders, then the items themselves with their kind, id, title, tags, and path. Folders are the engagement axis: a client or project gathers work from every tool, so a folder is how you answer 'what do we have for Acme' when the answer spans flows, pages, and datasets. A folder's NAME is not written on the things inside it, so reach for this rather than search_workspace whenever the user names a folder or asks what's in one. Only shows items the user can see and kinds this token may read.
| Name | Required | Description | Default |
|---|---|---|---|
| folder | No | A folder id, or its name (case-insensitive, exact). Omit to list the whole tree. | |
| workspace | No | Workspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/safe/non-open-world, and the description adds substantial context beyond them: it discloses permission scoping ('Only shows items the user can see and kinds this token may read'). It does not mention pagination or result-size limits, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the two modes before the rationale, and every sentence carries information (mode behavior, engagement-axis rationale, alternative routing, permission scoping). It is on the longer side, but the length is justified by genuinely distinct facts.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by enumerating returned fields per mode (id, path, count / kind, id, title, tags, path). Combined with permission scoping and explicit sibling routing, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (folder, workspace) are already fully documented in the schema. The description echoes the id-or-name behavior and workspace semantics but adds no syntax, format, or edge-case detail beyond what the schema provides — baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific verb (browse) and resource (workspace folders), then precisely delineates the two behaviors: no folder → full tree with id, path, and count; with folder → subfolders plus items with kind, id, title, tags, path. An agent can tell it apart from siblings like search_workspace purely from this text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: 'reach for this rather than search_workspace whenever the user names a folder or asks what's in one,' and explains the reasoning (a folder's name is not written on its contents). Also clarifies the two invocation modes and when each applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_setupSet up folders and rooms for their workADestructiveInspect
Creates folders for the work someone described, rooms linked to those folders (with your first message in each, the concrete next step), and starting pieces filed inside. Records what the workspace is for, which orders what home offers first. Use it once you know what they're working on, typically near the end of setting a workspace up. Say in one sentence what you're setting up and call it; the approval card lists the folders, rooms, and pieces, so don't list them in words too. Every workspace already has a Start here folder and a General room, so don't recreate those. May return needs_confirmation; then wait for approval.
Starting shapes by kind of work (adapt, drop, or combine; name everything after what they told you):
process_redesign: a folder per process being changed ("Claims intake redesign"). Pieces: a Compass page with pageType WORKFLOW describing how it works today, a Napkin diagram of current and proposed steps, a doc for the change plan. Room: " redesign". After setup, offer a Ledger expectation for the metric they want to move.
research: a folder per study or question. Pieces: a research-plan doc (question, who, method), a synthesis doc. Room: " research". After setup, offer a Prism field for the question.
ai_design: a folder per AI feature. Pieces: a spec doc (what it does, for whom, what good looks like), a Caliper dataset for test cases. Room: the feature's name, with the flow in it once one exists (file an existing flow via existing). After setup, offer to scaffold the flow and a rubric.
knowledge: a folder per domain. Pieces: Compass pages for the main topics. Room: "Ask about ".
exploring: usually nothing beyond Start here and General; offer one small first thing instead.
| Name | Required | Description | Default |
|---|---|---|---|
| folders | Yes | ||
| workKinds | No | What the workspace is for, from what they said. | |
| workspace | No | Workspace slug. Ignored for workspace API keys. | |
| approvalId | No | Approval id from a prior needs_confirmation envelope. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=false, so the description is not the sole source of the risk profile. It still adds real behavioral context beyond the annotations: the needs_confirmation return envelope and the requirement to wait for approval, plus the note that the approval card enumerates the created objects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded, but the definition is long, with a five-bullet starter-shape catalogue and meta-instructions ("Say in one sentence what you're setting up and call it..."). Much of it earns its place, but it is heavier than a description needs to be and leans toward a playbook rather than a definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-resource setup tool with no output schema, the description covers what to build, when to call it, naming guidance, exclusions, the per-kind templates, and the approval/needs_confirmation flow. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so most parameters are documented in the schema, and the description still adds meaning: naming conventions ("name everything after what they told you"), how workKinds should be derived, and what the nested pieces/rooms should contain per work kind. It goes beyond restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and several concrete resources ("Creates folders... rooms linked to those folders... and starting pieces filed inside") and explains the scope (records what the workspace is for). This clearly distinguishes it from sibling single-resource creators like compass_pages_create or rooms_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit timing ("Use it once you know what they're working on, typically near the end of setting a workspace up") and an explicit exclusion ("Every workspace already has a Start here folder and a General room, so don't recreate those"). The per-kind starter shapes further route the agent on what to build when.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
248 tool updates
- First observed
caliper_datasets_add_items - First observed
caliper_datasets_create - First observed
caliper_datasets_delete - First observed
caliper_datasets_generate - First observed
caliper_datasets_generate_items - First observed
caliper_datasets_get - First observed
caliper_datasets_list - First observed
caliper_datasets_remove_item - First observed
caliper_datasets_update - First observed
caliper_datasets_update_item - First observed
caliper_evals_create - First observed
caliper_evals_delete - First observed
caliper_evals_get - First observed
caliper_evals_list - First observed
caliper_evals_run - First observed
caliper_evals_runs_cancel - First observed
caliper_evals_runs_get - First observed
caliper_evals_runs_item_execution - First observed
caliper_evals_runs_list - First observed
caliper_evals_update - First observed
caliper_flow_performance - First observed
caliper_rubrics_create - First observed
caliper_rubrics_delete - First observed
caliper_rubrics_get - First observed
caliper_rubrics_list - First observed
caliper_rubrics_update - First observed
caliper_source_feeds_list - First observed
caliper_source_feeds_stop - First observed
caliper_sources_list - First observed
caliper_sources_over_time - First observed
caliper_sources_to_dataset - First observed
caliper_starters_install - First observed
caliper_starters_list - First observed
caliper_traces_get - First observed
caliper_traces_list - First observed
comments_create - First observed
comments_list - First observed
comments_resolve - First observed
compass_changes_upsert - First observed
compass_gaps_create - First observed
compass_gaps_list - First observed
compass_gaps_resolve - First observed
compass_gaps_update - First observed
compass_inbox_list - First observed
compass_interview_invite - First observed
compass_interview_targets - First observed
compass_interviews_create - First observed
compass_interviews_get - First observed
compass_interviews_list - First observed
compass_interviews_revoke - First observed
compass_links_create - First observed
compass_links_delete - First observed
compass_links_update - First observed
compass_map_view - First observed
compass_opportunities_accept - First observed
compass_opportunities_create - First observed
compass_opportunities_delete - First observed
compass_opportunities_list - First observed
compass_opportunities_mark_implemented - First observed
compass_opportunities_propose - First observed
compass_opportunities_register_expectation - First observed
compass_opportunities_set_status - First observed
compass_opportunities_update - First observed
compass_page_links_list - First observed
compass_pages_create - First observed
compass_pages_delete - First observed
compass_pages_get - First observed
compass_pages_list - First observed
compass_pages_restore - First observed
compass_pages_search - First observed
compass_pages_update - First observed
compass_statuses_list - First observed
entity_tags_browse - First observed
entity_tags_get - First observed
entity_tags_set - First observed
get_doc - First observed
get_post - First observed
ledger_entries_add_evidence - First observed
ledger_entries_create - First observed
ledger_entries_draft - First observed
ledger_entries_get - First observed
ledger_entries_list - First observed
ledger_entries_retract - First observed
ledger_entries_settle - First observed
ledger_entries_update - First observed
ledger_feed_sources_list - First observed
ledger_feeds_backfill - First observed
ledger_feeds_create - First observed
ledger_feeds_delete - First observed
ledger_feeds_list - First observed
ledger_feeds_preview - First observed
ledger_feeds_run - First observed
ledger_feeds_update - First observed
ledger_metric_presets_adopt - First observed
ledger_metric_presets_list - First observed
ledger_metric_starters_adopt - First observed
ledger_metric_starters_list - First observed
ledger_metrics_archive - First observed
ledger_metrics_create - First observed
ledger_metrics_get - First observed
ledger_metrics_list - First observed
ledger_metrics_record_reading - First observed
ledger_metrics_update - First observed
ledger_plan_add - First observed
ledger_plan_delete - First observed
ledger_plan_list - First observed
ledger_plan_update - First observed
list_docs - First observed
list_posts - First observed
napkin_boards_create - First observed
napkin_boards_get - First observed
napkin_boards_list - First observed
napkin_boards_update - First observed
napkin_boards_view - First observed
napkin_brand_add_files - First observed
napkin_brand_apply - First observed
napkin_brand_check - First observed
napkin_brand_draft - First observed
napkin_brand_examples - First observed
napkin_brand_fonts - First observed
napkin_brand_get - First observed
napkin_brand_list - First observed
napkin_brand_view - First observed
napkin_deck_compose - First observed
napkin_deck_export - First observed
napkin_deck_outline - First observed
napkin_deck_set_theme - First observed
napkin_deck_write - First observed
napkin_diagrams_create - First observed
napkin_diagrams_get - First observed
napkin_diagrams_list - First observed
napkin_diagrams_update - First observed
napkin_docs_create - First observed
napkin_docs_get - First observed
napkin_docs_list - First observed
napkin_docs_update - First observed
napkin_draw - First observed
napkin_interfaces_create - First observed
napkin_interfaces_get - First observed
napkin_interfaces_grant - First observed
napkin_interfaces_list - First observed
napkin_interfaces_publish - First observed
napkin_interfaces_view - First observed
napkin_interfaces_write - First observed
napkin_layouts_list - First observed
napkin_sheets_add_chart - First observed
napkin_sheets_create - First observed
napkin_sheets_list - First observed
napkin_sheets_query - First observed
napkin_sheets_schema - First observed
napkin_sheets_set_cells - First observed
napkin_sheets_update - First observed
napkin_sheets_view_chart - First observed
napkin_slide_add - First observed
napkin_slide_fill - First observed
napkin_slide_update - First observed
napkin_slide_view - First observed
prism_fields_create - First observed
prism_fields_delete - First observed
prism_fields_get - First observed
prism_fields_list - First observed
prism_fields_update - First observed
prism_insights_promote - First observed
prism_interviews_create - First observed
prism_interviews_list - First observed
prism_nodes_expand - First observed
prism_nodes_research - First observed
prism_nodes_update - First observed
prism_research_digest - First observed
prism_series_create - First observed
prism_series_launch_wave - First observed
prism_series_list - First observed
prism_series_update - First observed
prism_studies_analyze - First observed
prism_studies_cut - First observed
prism_studies_delete - First observed
prism_studies_draft - First observed
prism_studies_draft_from_interviews - First observed
prism_studies_draft_standalone - First observed
prism_studies_get - First observed
prism_studies_list - First observed
prism_studies_list_workspace - First observed
prism_studies_report - First observed
prism_studies_results - First observed
prism_studies_segments - First observed
prism_studies_segments_compute - First observed
prism_studies_update - First observed
research_findings_list - First observed
rooms_add_flow - First observed
rooms_list - First observed
rooms_post - First observed
rooms_read - First observed
rooms_start_agent_work - First observed
routines_create - First observed
routines_list - First observed
routines_update - First observed
search_docs - First observed
search_workspace - First observed
workbench_executions_get - First observed
workbench_flow_authoring_guide - First observed
workbench_flows_delete - First observed
workbench_flows_edit_text - First observed
workbench_flows_fork - First observed
workbench_flows_get - First observed
workbench_flows_list - First observed
workbench_flows_publish - First observed
workbench_flows_revisions_list - First observed
workbench_flows_run - First observed
workbench_flows_scaffold - First observed
workbench_flows_schedule - First observed
workbench_flows_schedules_delete - First observed
workbench_flows_schedules_list - First observed
workbench_flows_schedules_run_now - First observed
workbench_flows_schedules_update - First observed
workbench_flows_share_create - First observed
workbench_flows_shares_list - First observed
workbench_flows_shares_revoke - First observed
workbench_flows_update - First observed
workbench_kb_add_documents - First observed
workbench_kb_create - First observed
workbench_kb_delete - First observed
workbench_kb_document_delete - First observed
workbench_kb_documents_list - First observed
workbench_kb_get - First observed
workbench_kb_ingestion_status - First observed
workbench_kb_list - First observed
workbench_kb_search - First observed
workbench_node_catalog_get - First observed
workbench_shim_add_answer - First observed
workbench_shim_add_examples - First observed
workbench_shim_create - First observed
workbench_shim_get - First observed
workbench_shim_list - First observed
workbench_tasks_act - First observed
workbench_tasks_compile - First observed
workbench_tasks_create - First observed
workbench_tasks_delete - First observed
workbench_tasks_freeze - First observed
workbench_tasks_get - First observed
workbench_tasks_list - First observed
workbench_tasks_run_view - First observed
workbench_tasks_start - First observed
workbench_tasks_update - First observed
workspace_credit_get - First observed
workspace_files_read - First observed
workspace_files_save_from_chat - First observed
workspace_folders_browse - First observed
workspace_setup
Related MCP Connectors
Build, run and publish Workbench flows, tasks, shims and knowledge bases.
551Query geospatial datasets, manage workflow jobs, deploy published releases, and search Tilebox docs.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
- RunableOAuthcom.runable
Run agent tasks, track progress, retrieve files, and search meeting notes with Runable.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients and coding agents to search a local-first catalogue of existing software, packages, MCP servers, patterns, and UI components, inspect source evidence, prepare adoption briefs, and optionally edit or export React trials in a shared workbench.MIT

yotta-dev-mcpofficial
AlicenseNot gradedqualityBmaintenanceProvides deterministic local development tools over stdio MCP, including repository mapping, code review, secret and dependency scanning, release checks, whitelisted checks, scaffolding, and workflow state tracking.718 npmMIT- AlicenseNot gradedqualityCmaintenanceEnables MCP clients to inspect operational status and channels, lint human-reviewed plans, search and query ledger data, and run queue-based writes with dry-run defaults and safeguards against unverified or unapproved campaigns.MIT
- FlicenseNot gradedqualityCmaintenanceEnables querying enterprise records and retention policies from any MCP client over stdio, with read-only tools for searching records, fetching retention verdicts, identifying archival candidates, summarizing departments, forecasting retentions, and viewing audit history.-
Glama MCP Gateway
Add one secure layer between your agents and this server.