Skip to main content
Glama

ZeroWidth Workbench

Server Details

Build, run and publish Workbench flows, tasks, shims and knowledge bases.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

A3.8/5.0

Scored across 55 tools

Disambiguation4/5

Tools are largely distinct: comments, tags, docs, flows, schedules, shares, KBs, shims, and tasks each form a clear domain with verb-specific actions. A few pairs need the descriptions to separate them (workbench_flows_run vs workbench_flows_schedules_run_now, workbench_flows_schedule vs workbench_flows_schedules_update, KB search vs search_workspace vs search_docs), but the rich descriptions resolve these well.

Naming Consistency4/5

Most tools follow a predictable prefix_resource_action pattern (workbench_flows_get, workbench_kb_create, workbench_tasks_act, entity_tags_set). Deviations are minor: the docs/search tools (list_docs, get_doc, search_docs, search_workspace) drop the domain prefix and use verb_noun order, and singular/plural is mixed (workbench_kb vs workbench_flows, workbench_flow_authoring_guide).

Tool Count2/5

55 tools is very heavy; while the platform is genuinely broad (flows, schedules, shares, KBs, shims, tasks, comments, tags, docs), most sub-domains could likely be consolidated (e.g., schedule CRUD spread across five tools). The sheer surface area raises selection and context costs.

Completeness4/5

Coverage is strong: full lifecycle for flows (get/list/run/publish/fork/edit/update/delete/revisions), schedules, shares, KBs with ingestion, shims, tasks, comments, tags, and cross-cutting search. Minor gaps include no KB update/rename tool and no plain blank-flow create (only scaffold/fork), but agents can work around these.

Available Tools

55 tools
comments_createComment on an entityAInspect

Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
rootIdNoReply into this thread; omit to start a new one.
entityIdYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNo
entityKindYesWhat the thread hangs on.
mentionedUserIdsNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the write/safety profile is covered. The description adds value beyond that by disclosing a side effect annotations cannot express: mentioning users rings their notification bell, with an accompanying social caution. It omits permission/auth requirements for posting.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and its two modes, then layers usage and mention etiquette compactly. Two sentences, minimal waste, though the trailing 'never mention someone who didn't ask' is advisory padding rather than invocation-critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation with 43% schema coverage and no output schema, the description covers the social/mention dimension well but leaves the entity-targeting parameters and the undocumented approvalId unexplained, which an agent would need to call this reliably in all cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, so the description must compensate, and it partially does: rootId (reply target), mentionedUserIds (workspace member ids, notification behavior), and body are implied. However entityKind's 17-value enum, entityId, and especially approvalId are unexplained in both schema and description, leaving real gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ('Posts a comment on a workspace entity') and distinguishes the two modes of operation: new thread vs. reply when rootId is given. It never names its closest siblings (comments_list, comments_resolve), so it stops just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context ('leave findings where the discussion already lives') with two illustrative scenarios (eval result on a debated flow, summary on a long thread). It does not state when not to use it or point to a sibling alternative, so it lacks the explicit routing of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

comments_listRead an entity's comment threadsA
Read-only
Inspect

Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityIdYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
entityKindYesWhat the thread hangs on.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavior: threads are returned open-before-resolved and include authors and timestamps, which shapes how an agent interprets output. It stops short of pagination or volume limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with the key facts front-loaded and no redundancy. The second sentence is motivational framing that carries mild value but is slightly softer than a hard routing rule.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema and a mostly documented schema, the description covers ordering, content, and the entity scoping needed to call it. Pagination/result-size behavior is the only notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the schema itself handles the workspace slug nuances and the enum list. The description only says 'one workspace entity', adding little beyond the structured fields, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (lists) and resource (comment threads on one workspace entity), plus the ordering rule (open first, then resolved) and payload (authors, timestamps). This clearly separates it from comments_create and comments_resolve, which mutate rather than read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives a soft situational cue ('read this before weighing in on contested work'), which implies when the tool is useful. However it names no alternatives and gives no explicit when-not or prerequisite guidance, so usage remains inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

comments_resolveResolve or reopen a threadA
Destructive
Inspect

Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootIdYes
resolvedYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=false, so the mutation/safety profile is covered. The description adds real behavioral context beyond that: the precondition (a reply explaining what settled the thread) and the reopen semantics. It doesn't clarify reversibility or how the state change affects existing replies, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler; the action and the primary precondition are front-loaded before the reopening clause. Every clause carries decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a small toggle tool with no output schema, the description covers the action, the target, and the conditions well. It leaves approvalId and workspace behavior unexplained, which matters for a destructive write, but overall an agent has enough to act correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25%, and the description compensates for rootId only (root comment id). 'resolved' is implied by resolve/reopen framing, but 'workspace' is documented only in the schema and 'approvalId' is explained nowhere — a notable gap for a mutation tool with an approval parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (sets resolved state) plus resource (comment thread), and covers both directions — resolve and reopen. The parenthetical 'rootId = the thread's root comment id' disambiguates the target, making it clearly distinct from comments_create/comments_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gates the action: resolve ONLY when the human asked or the question is demonstrably settled, and reply first explaining what settled it; reopen is for new evidence. This is genuine when/when-not guidance with a required prior step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

entity_tags_browseBrowse the workspace's tagsA
Read-only
Inspect

Without a tag: every tag in use across the workspace with how many entities carry it, most-used first — the vocabulary the team already organizes by. With a tag: everything filed under it across every tool, each with its kind, id, title, and path. Use it to reuse existing labels instead of inventing near-duplicates, and to answer 'show me everything about X' when X is a label.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoA tag to expand into its items. Omit to list tags.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/destructive/openWorld, so safety is covered, and the description adds real behavioral detail: result ordering (most-used first), entity counts per tag, and the per-item fields returned in tag mode (kind, id, title, path). No pagination or size-limit behavior is mentioned, keeping this from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and cleanly parallel: 'Without a tag:' then 'With a tag:' then the usage clause. Every sentence carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the return-value burden for both modes, and annotations cover the safety profile while the schema covers both parameters. Nothing an agent needs to select or call this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by characterizing what each mode of the tag parameter actually returns, turning a bare 'omit to list tags' into the two distinct result shapes. The workspace parameter semantics remain entirely schema-borne.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific dual operation (list every tag in the workspace, or expand one tag into all entities filed under it) with the exact resource and scope. It is immediately distinguishable from the write-oriented siblings entity_tags_get and entity_tags_set because the browse mode and cross-tool aggregation are spelled out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use guidance: reuse existing labels rather than inventing near-duplicates, and answer 'show me everything about X' when X is a label. It stops short of naming sibling alternatives or stating when-not-to-use (e.g. versus search_workspace), so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

entity_tags_getRead the tags on entitiesA
Read-only
Inspect

Returns the tags on a batch of entities of one kind — the labels galleries organize by. Ids come from the kind's list/get tool or from search_workspace. Use it before entity_tags_set so you replace the full set knowingly, and to answer 'what is this filed under'. Entities the user can't see are omitted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
entityKindYesWhich kind the ids belong to.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered; the description adds the non-obvious behavioral fact that 'Entities the user can't see are omitted,' i.e. results are permission-filtered rather than erroring. It does not mention limits (e.g. the 100-id cap or ordering), but the value-add beyond annotations is genuine.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all front-loaded: what it returns first, then usage, then the permission caveat. No filler and each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does state what comes back (tags) and the omission behavior. It leaves minor gaps — tag value format, ordering, and whether unknown ids are silently dropped or error — but nothing that blocks correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, so the schema documents most params, but the description adds real provenance for the hardest parameter: ids 'come from the kind's list/get tool or from search_workspace.' That tells an agent where to obtain valid ids, which the bare array schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Returns the tags on a batch of entities of one kind.' The parenthetical 'the labels galleries organize by' distinguishes tags from other entity metadata, and the named siblings (entity_tags_set, search_workspace) make the boundary clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage contexts: 'Use it before entity_tags_set so you replace the full set knowingly, and to answer what is this filed under.' That is a real when-to-use plus a stated alternative. It does not mention entity_tags_browse, the other obvious read-side sibling, so it falls short of fully disambiguating the read alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

entity_tags_setSet an entity's tagsA
Destructive
Inspect

Replaces the FULL tag set on one entity (an empty list clears it). Read the current tags with entity_tags_get first and pass the merged list — this is not additive. Tags are lowercase letters, numbers, spaces, and hyphens; prefer labels already in use (entity_tags_browse) so the workspace's vocabulary stays small. The id comes from the kind's list/get tool or search_workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsYes
entityIdYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
entityKindYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, and the description reinforces this with the replace-not-merge semantics and the empty-list-clears behavior. It does not mention the needs_confirmation/approvalId retry flow that the schema implies, which is the one notable behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, all front-loaded with the destructive replace semantics first, followed by workflow and vocabulary guidance. Every sentence earns its place; the parentheticals make it slightly busy but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description covers destructive semantics, prerequisites, id sourcing, and tag format well enough to invoke the tool correctly. The confirmation/approval retry path is left to the schema field description rather than the tool description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40%, and the description compensates by documenting the tag character set (lowercase letters, numbers, spaces, hyphens), the merge requirement for the tags array, and the origin of entityId. The workspace and approvalId parameters are only explained in the schema, so coverage is good but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replaces the FULL tag set on one entity') and immediately clarifies the destructive scope with '(an empty list clears it)'. This cleanly separates it from the sibling readers entity_tags_get and entity_tags_browse.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit workflow: read current tags with entity_tags_get first, pass the merged list, and warns 'this is not additive'. Also routes to entity_tags_browse for vocabulary reuse and names where the entityId comes from (kind's list/get tool or search_workspace), so the agent knows both prerequisites and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_docFetch ZeroWidth doc by slugA
Read-only
Inspect

Fetch the full Markdown body of a specific docs page by its slug. Use this after search_docs when the user needs the complete content of a page. No authentication required.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesPage slug. Accepts 'compass/api', '/compass/api', or 'docs/compass/api'.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, covering the safety profile. The description adds valuable context by stating 'No authentication required' and specifying the return body as full Markdown, which agents need to know beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core purpose ('Fetch the full Markdown body...'), followed by usage routing and a salient behavioral note. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with full schema coverage and annotations covering safety, the description is complete: it states the return format, the required input type, usage context, and authentication status. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single `slug` parameter is fully documented in the schema with accepted formats. The description adds no syntax or format details beyond what the schema provides, so it meets the baseline rather than exceeding it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch'), resource ('full Markdown body of a specific docs page'), and retrieval key ('by its slug'). It distinguishes itself from the search-oriented sibling by positioning as the step after `search_docs` for complete page content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use this after `search_docs` when the user needs the complete content of a page, which names the alternative and the condition that selects it. It does not state when not to use it (e.g., for listing or metadata), but the context is clear enough for a simple read tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_docsList ZeroWidth docs pagesA
Read-only
Inspect

Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling get_doc. No authentication required.

ParametersJSON Schema
NameRequiredDescriptionDefault
productNoOptional product slug filter (e.g. 'compass', 'legal', 'overview').

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds context beyond them with 'No authentication required,' a genuinely useful operational fact for callers, though it says nothing about pagination or result size for a full enumeration.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with purpose, routing guidance, and the auth fact front-loaded in order of importance. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-optional-param listing tool whose annotations cover safety, the description is nearly sufficient. Without an output schema it could note the return shape (e.g. that results are slugs/pages), but the 'slugs' reference largely covers that, so only a small gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the single 'product' parameter. The description's example values ('compass', 'legal', 'overview') duplicate the schema description verbatim, adding no meaning beyond it. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Enumerate') and resource ('all available docs pages') with scope, plus the optional product filter. It names the sibling get_doc and frames itself as the discovery step before retrieval, letting an agent distinguish it from single-doc fetch without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: 'to discover what slugs exist before calling get_doc,' giving a clear dependency flow and naming the alternative. It does not address when NOT to use it or whether search_docs is a better discovery path, leaving a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_docsSearch ZeroWidth docsA
Read-only
Inspect

Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of results. Defaults to 10.
queryYesSearch query — keywords or natural-language phrase.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful non-annotation context: 'No authentication required — the docs corpus is public' and the shape of returned results (title, slug, public URL, snippet). It does not mention pagination or ordering, but it goes beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four short sentences, front-loaded with the core action, followed by return information, usage trigger, and auth note. Every sentence earns its place, and there is no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only search tool, the description covers purpose, return format, auth requirements, and usage triggers. It does not explain how this differs from search_workspace or list_docs/get_doc, and it lacks result-ordering or empty-result behavior, but it is otherwise sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters are documented in the schema itself. The description says the query returns a 'query-relevant snippet' but adds no syntax, format, or constraint details beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search ZeroWidth product documentation.' It also names the return fields and the product scope with concrete examples (Compass, Workbench, Caliper, etc.), which lets an agent distinguish it from siblings like list_docs, get_doc, and search_workspace without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: 'Use this when the user asks about a ZeroWidth product ..., an API behavior, or a policy.' However, it does not name alternative tools (search_workspace, list_docs, get_doc) or state when not to use this tool, so it stops short of explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_workspaceSearch the whole workspaceA
Read-only
Inspect

Finds entities across every tool by name in one call — Workbench flows, Compass pages, Caliper datasets, evals, rubrics, reviews, specs and sources (apps sending agent traces), Ledger entries, Napkin sketches and decks. Use it FIRST when the user names something without saying where it lives ('the onboarding flow', 'that invoice page'); reach for a tool's own list only when you already know the tool. Each hit carries its id, kind, and workspace-relative path, so the id feeds the matching *_get tool and the path makes a link. Results only include what the user can see, and only kinds this token may read.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYesCase-insensitive substring matched against names/titles.
kindsNoRestrict to these kinds (flow, page, dataset, eval, entry, board). Omit to search everything.
limitNo
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context beyond that: results are filtered to what the user can see and to kinds the token may read, and it discloses the hit shape (id, kind, workspace-relative path).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense, front-loaded sentences with no filler: purpose and coverage first, routing second, return/scope semantics last. Every clause carries information the agent can act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden and does so by describing the id/kind/path tuple and how the id feeds *_get tools. Combined with the permission scoping note, an agent has everything needed to call and use this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so most parameters are already documented structurally. The description reinforces name/title matching but adds no new syntax or format detail for kinds, limit, or workspace, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb (Finds) and resource (entities across every tool by name), then enumerates the concrete kinds covered (flows, pages, datasets, evals, rubrics, etc.). An agent can immediately distinguish this cross-tool search from the many per-tool list siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('Use it FIRST when the user names something without saying where it lives'), gives concrete examples ('the onboarding flow'), and names the alternative plus its selection condition ('reach for a tool's own list only when you already know the tool'). This is textbook when/when-not/alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_executions_getInspect a flow's executions (debug traces)A
Read-only
Inspect

The flow's stack traces: recent executions with per-node timelines — what each node received, produced, how long it took, and the exact error when one failed. USE THIS when a run misbehaves instead of guessing: read the failing node's inputs/error, then propose a fix (workbench_flows_edit_text) grounded in what actually happened. Pass executionId to inspect one run, or just flowId for the most recent runs. Node inputs/outputs are truncated for transport — the full record is in the flow's dev drawer.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many recent executions (default 3, max 10). Ignored when executionId is set.
flowIdYesWorkbench flow id.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic).
executionIdNoSpecific execution to inspect. Omit for the latest runs.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/destructive=false, and the description adds genuinely useful non-obvious behavior: node inputs/outputs are truncated for transport and the full record lives in the flow's dev drawer. It also implies no auth caveats beyond what the schema documents, so this is solid but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with what is returned before the usage guidance. Slightly editorial ('instead of guessing') but every sentence carries routing or return-shape information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description fully describes the return shape (per-node timeline, inputs, outputs, durations, error), the truncation caveat, and the flowId vs executionId selection. Nothing an agent needs to call or interpret it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds the key interaction rule: pass executionId to inspect one run, or just flowId for the most recent runs (echoing that limit is ignored when executionId is set). That interaction meaning goes beyond the per-parameter schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and scope: 'recent executions with per-node timelines' plus what each node received, produced, took, and its error. This is clearly distinguishable from siblings like workbench_flows_get or workbench_tasks_run_view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit triggering condition ('when a run misbehaves instead of guessing') and routes the agent forward to workbench_flows_edit_text for the fix. It does not, however, name alternative read tools (e.g. caliper_traces_get or workbench_tasks_run_view) or state when not to use this.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flow_authoring_guideWorkbench flow-authoring guideA
Read-only
Inspect

READ THIS FIRST before writing or editing any raw flow orchestration JSON. Covers the document shape, the settings-vs-input-ports rule (temperature, max_tokens, tools, response_format are PORTS, not settings — constants reach ports via value nodes), exact link format, plugin links, and three complete worked examples.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: it is a preparatory reading resource whose content includes three complete worked examples, implying a substantial reference payload. It says nothing about response size or format, keeping it at 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the imperative 'READ THIS FIRST' so the most actionable instruction leads. The second sentence enumerates the guide's contents tersely with no filler; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter reference tool with no output schema, the description adequately conveys what an agent will receive when it calls it. It could note the output modality (returned documentation text) or roughly how large the content is, but what an agent needs to decide to call it is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty with 100% coverage, so the baseline of 4 applies. There are no parameters for the description to illuminate, and it correctly does not invent any.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool is and returns: a reference covering document shape, the settings-vs-input-ports rule, link format, plugin links, and worked examples. It is unmistakably a flow-authoring documentation resource, distinct in kind from the flow CRUD siblings. It stops short of naming a specific sibling it competes with, hence 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'READ THIS FIRST before writing or editing any raw flow orchestration JSON' gives an explicit trigger condition (before authoring/editing raw JSON). It doesn't name the alternative tools (flows_edit_text, flows_scaffold, flows_update) or state when the guide is NOT needed, but the intended usage window is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_deleteDelete a Workbench flowA
Destructive
Inspect

Deletes a flow. Its schedules stop, guest share links die, and evals targeting it can no longer run (their binding shows flowOk: false) — check caliper_flow_performance for evals and workbench_flows_schedules_list before proposing, and prefer workbench_flows_update with archived: true when the user just wants it out of the way. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowIdYesFlow id, from workbench_flows_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations flag destructiveHint=true, but the description adds real substance beyond them: exactly what is destroyed (schedules stop, guest share links die, evals can no longer run, binding shows flowOk: false) and the confirmation round-trip ('May return needs_confirmation'). This is precisely the consequence-level detail an agent needs before a destructive call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and its consequences, then appends the pre-check guidance and the re-confirmation note. Dense but every clause carries distinct, load-bearing information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers the blast radius, the pre-flight checks, the softer alternative, and the confirmation protocol. An agent has everything needed to invoke it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so flowId, workspace, and approvalId are all documented in the schema. The description reinforces the approval flow by mentioning the needs_confirmation response, but adds no syntax or format detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Deletes a flow') that unambiguously states the operation. It is clearly distinguishable from sibling tools like workbench_flows_update or workbench_flows_schedules_delete, which it even references by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance ('check caliper_flow_performance for evals and workbench_flows_schedules_list before proposing') and names a concrete alternative with the selecting condition ('prefer workbench_flows_update with archived: true when the user just wants it out of the way'). No routing decision is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_edit_textSurgically edit text inside a flowA
Destructive
Inspect

Applies ONE precise text replacement to a flow's DRAFT orchestration — Edit-tool semantics: find must be the EXACT current text, copied character-for-character from workbench_flows_get, and must occur exactly once anywhere in the flow (system prompts, node settings, metadata, a node's type). find is text inside ONE value, not JSON: to change a model, find the model slug alone (anthropic-claude-haiku-4-5) and replace it with another llm slug from workbench_node_catalog_get. It can't add or remove nodes or links. Zero or multiple matches return an error instead of guessing. This is the improvement primitive: check receipts first with caliper_flow_performance, cite the run id in note, apply the edit after approval, then re-run the eval with caliper_evals_run and report the score delta — never claim improvement without the before/after. Edits land on the draft only; the published version changes when someone publishes (workbench_flows_publish, after the user has seen the result). May return needs_confirmation — show the user the exact find/replace diff and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
findYesEXACT current text (≥3 chars), copied verbatim — never reconstructed from memory.
noteYesWhy this edit. Cite eval run ids when it follows a measurement (e.g. 'refund answers scored 1/5 on accuracy in run cer_abc — adds the manager-approval rule'). Lands in the audit log.
flowIdYesFlow to edit.
replaceYesReplacement text.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this.
approvalIdNoApproval id from a prior needs_confirmation envelope.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag it as a destructive, non-read-only mutation, and the description goes well beyond that: edits land on the draft only, zero or multiple matches error rather than guess, it may return a `needs_confirmation` envelope requiring a diff to be shown, and it cannot add/remove nodes or links.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core find/replace constraint before the workflow narrative, and nearly every sentence carries operational weight. It is dense and long, but the length is justified by the number of constraints it must convey.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 6-param mutation with no output schema, the description covers the failure modes (zero/multiple matches, needs_confirmation), the scoping (draft vs published), and the surrounding measurement workflow. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value not in the schema: `find` is text inside ONE value (not JSON), must be copied character-for-character, must occur exactly once, with a concrete model-slug example for the replace flow.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Applies ONE precise text replacement to a flow's DRAFT orchestration') and immediately scopes it with Edit-tool semantics. It distinguishes itself from structural siblings by explicitly stating it can't add or remove nodes or links.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Prescribes the full workflow: measure with caliper_flow_performance first, cite the run id in `note`, apply after approval, then re-run caliper_evals_run and report the delta. Also states the draft/publish boundary and names workbench_flows_get and workbench_flows_publish as the surrounding tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_forkFork a public flow into the workspaceAInspect

Copies a PUBLIC flow from another workspace (a template) into this one as a new draft the user owns. Identify it by its flowUuid (the id on Workbench's public template pages and in workbench_flows_get output); optionally pin which published version to copy. Flows already in this workspace can't be forked — open them instead. Counts toward the plan's flow cap. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowUuidYesThe source flow's public uuid.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
forkedFromVersionNoPublished version label to copy. Default: the latest published.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description adds real context beyond them: the result is a user-owned draft, it counts toward the plan's flow cap, and it may return `needs_confirmation`. It stops short of describing the response shape or how to resolve the confirmation flow in detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action and constraints. Every sentence carries distinct information (identity, version pinning, non-forkable case, cap cost, confirmation signal) with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-param mutation with no output schema, the description covers the critical agent-facing facts: public-source requirement, version selection, the already-in-workspace exclusion, quota impact, and the confirmation return. Nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds locating value for flowUuid ('the id on Workbench's public template pages and in `workbench_flows_get` output') and clarifies that forkedFromVersion optionally pins a published version. It goes beyond restating the schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Copies a PUBLIC flow from another workspace into this one as a new draft') and immediately clarifies the fork semantics (template -> owned draft). It distinguishes itself from siblings by noting flows already in the workspace can't be forked and should be opened instead, and by pointing to workbench_flows_get for the id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (a PUBLIC/template flow from another workspace), explicit when-not ('Flows already in this workspace can't be forked — open them instead'), and it names the alternative path. Nothing about selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_getRead one Workbench flowA
Read-only
Inspect

One flow. Default view: summary lists its nodes and links so you can pick one; view: node with a nodeId returns that node's full settings and prompt — read THAT before workbench_flows_edit_text, and copy find text from it character-for-character, never from memory. view: full returns the whole orchestration body.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNo`summary` (default): the flow's fields plus one line per node (id, type, label, and how long its prompt is) and the links — enough to pick a node. `node`: the full settings of one node (pass nodeId) — what you need before workbench_flows_edit_text. `full`: the whole orchestration body; large.
flowIdYesFlow id.
nodeIdNoWith view `node`: which node.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds useful payload-size context ('full: large') and clarifies that node returns full settings and prompt, but does not discuss auth/rate limits or truncation, keeping it just under a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core fact ('One flow') then walks the three views in the order an agent would use them, with each clause earning its place and no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must explain return values — and it does per view (node list and links, one node's settings/prompt, whole body). Complete for a read tool whose safety profile is already in annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the `view` enum is fully documented in the schema, so the schema already does the heavy lifting. The description restates view behavior and adds workflow context but no new format or syntax detail, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read/get) and resource (a single Workbench flow), and the three view modes make the scope precise. It is clearly distinguishable from workbench_flows_list (many) without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit, actionable guidance: use summary to pick a node, then view:node before calling workbench_flows_edit_text, and copy `find` text character-for-character rather than from memory. It names the downstream sibling and the precondition that selects this call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_listList Workbench flowsA
Read-only
Inspect

Lists every flow in the active workspace the caller can see. Returns summaries (id, name, visibility, updatedAt) — fetch one with workbench_flows_get for the full orchestration body.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, so the safety profile is covered. The description adds real value beyond that by disclosing the returned summary fields (id, name, visibility, updatedAt) and the visibility scoping ('the caller can see'), which an agent needs since no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste; the core listing behavior is front-loaded and the get-for-detail note follows naturally.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety, a single fully documented parameter, and no output schema, the description properly compensates by naming the return fields and pointing to the detail tool. Only the full shape of each summary is left implicit, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the workspace param already documents slug semantics, default/override behavior, and API-key treatment. The description adds nothing about the parameter, so baseline 3 applies since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists every flow in the active workspace') with clear scoping ('the caller can see'), and explicitly distinguishes itself from workbench_flows_get by noting it returns only summaries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for use (enumerate flows, choose one to inspect) and routes to the alternative (`workbench_flows_get`) when the full orchestration body is needed. No explicit when-not guidance, but the routing hint is strong and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_publishPublish a Workbench flow versionA
Destructive
Inspect

Snapshots the flow's current DRAFT as a named published version — the version schedules (workbench_flows_schedule), guest share links (workbench_flows_share_create), public-API runs, and 'published'-stage evals execute. The draft keeps evolving after this; runs on the published side don't change until the next publish. Publishing does NOT run the flow, but any eval set to runOnPublish starts a run (that spends credit — mention it when one exists). Publish only after the user has tested the draft (workbench_flows_run) or asked for it. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoWhat changed — a commit message.
flowIdYesFlow to publish.
versionYesVersion label, e.g. '1.0.0' or '2026-09-04'. Re-using a label retags it to this snapshot.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare write/destructive/openWorld; the description adds the real behavior: the draft keeps evolving while published runs stay frozen until the next publish, publishing does not execute the flow, and runOnPublish evals will start a run that spends credit (with an instruction to surface that to the user). It also discloses the 'needs_confirmation' return path, which is beyond anything the annotations carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The snapshot definition is front-loaded in the first clause and every subsequent sentence carries distinct load-bearing information: consumers of the published version, draft/published divergence, the no-run clarification, credit-spending evals, and the publish precondition. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating tool with no output schema, the description covers the lifecycle semantics, side effects (eval runs and credit spend), the confirmation flow, and prerequisites. Nothing material an agent needs before calling it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents note, version retagging, workspace defaulting, and the approvalId provenance, so the description adds little parameter-level meaning. Its only addition is the reference to 'needs_confirmation', which is already tied back to approvalId in the schema itself. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a precise verb and resource — 'Snapshots the flow's current DRAFT as a named published version' — and immediately distinguishes this from nearby siblings by naming what consumes the published side (workbench_flows_schedule, workbench_flows_share_create, public-API runs, published-stage evals). An agent can separate this from workbench_flows_run and workbench_flows_update without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states an explicit precondition and alternative: 'Publish only after the user has tested the draft (workbench_flows_run) or asked for it.' It also clarifies the boundary versus running ('Publishing does NOT run the flow') and versus the draft, which is exactly the when/when-not routing an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_revisions_listList a flow's versions (publish history)A
Read-only
Inspect

The flow's revision history, newest first: published versions carry a version label and note; entries with version null are draft autosaves. Use it to answer 'is this published?' (any entry with a version), to find a revision id for caliper_evals_create's flowRevisionId, or to see when the draft last changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoEntries to return, default 20.
flowIdYesFlow id.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
publishedOnlyNotrue = only labeled published versions (default false).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and destructiveHint=false, so safety is covered. The description adds real behavioral context beyond that: reverse-chronological ordering and the semantic distinction between labeled published versions and null-versioned draft autosaves. It stops short of noting pagination or how limit interacts with ordering, which keeps it out of the top band.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, output shape front-loaded before the use cases. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the return shape (ordering, version label, null for drafts) and anchors the tool to its downstream consumer. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already carries param semantics and baseline is 3. The description's published/null distinction illuminates what publishedOnly=true filters to, but adds nothing about limit or workspace beyond what the schema documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and what it returns: the flow's revision history, newest first, with published vs draft entries differentiated. An agent can distinguish this from workbench_flows_list and workbench_flows_get without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three explicit when-to-use scenarios: answering 'is this published?', sourcing a revision id for caliper_evals_create's flowRevisionId, and checking when the draft last changed. It even names the consuming sibling tool by parameter, so the routing decision is fully resolved.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_runRun a Workbench flowAInspect

Execute a flow and return its outputs synchronously. Runs the draft by default; pass source:"published" (optionally a version) to run the live published version. Spends workspace LLM budget, bounded by the token's cost cap. input is the flow's input envelope, e.g. {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputNoFlow input envelope (RunFlowInput). Omit for a flow that needs none.
flowIdYesFlow id to run.
sourceNoWhich version to run. Default: draft.
versionNoPublished version label (only with source:"published").
timeoutMsNoMax runtime in ms (also the hard cost ceiling).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this. Ignored for workspace API keys.
approvalIdNoRe-call with the approvalId from a needs_confirmation response after the user approves.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: it discloses that execution 'Spends workspace LLM budget, bounded by the token's cost cap' and that results come back 'synchronously'. The cost/budget disclosure is exactly the kind of non-obvious operational consequence annotations (readOnlyHint=false, destructiveHint=false) don't capture, and it is consistent with them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action and result, then source semantics, cost, and input examples. Every sentence carries information; the only slight sprawl is the inline input examples, but they are useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description covers what matters: synchronous return of outputs, budget/cap implications, draft vs published selection, and input envelope shape. The confirmation/approval re-call flow (approvalId) is only documented in the schema, leaving a minor gap for a 7-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3. The description adds real meaning on top: it explains the `source` semantics (default draft vs published-with-optional-version) and illustrates the `input` envelope with concrete examples ({"kind":"chat",...}, {"kind":"form",...}) that the schema does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Execute a flow') and adds the key behavioral result ('return its outputs synchronously'). Combined with the draft-vs-published distinction, an agent can distinguish this from sibling tools like workbench_flows_get or workbench_flows_publish without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Conveys the default behavior ('Runs the draft by default; pass source:"published" ... to run the live published version') and the cost context, which implicitly guides selection. However, it never states when to prefer this tool over alternatives (e.g. schedules_run_now, tasks_run_view) or any exclusions, so guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_scaffoldScaffold a first-draft Workbench flowAInspect

Creates a runnable first-draft flow from a spec you author: pick the simplest pattern that fits (classifier for read-and-bucket, structurer for transform/extract/draft-for-review, agent for genuinely conversational), write a production-quality system prompt grounded in what the user told you, and mark anything stubbed with [STUB: ...] markers plus stubNotes. The draft opens in Workbench's simple editor at /w//flows/ — give the user that path. May return needs_confirmation; show the user what you're proposing and wait for their approval, then re-call with the approvalId. When the draft comes from a Compass change, pass its opportunityId so the flow and the change point at each other (Compass shows 'Open in Workbench'; the flow shows where it came from). Then: workbench_flows_run to try it, workbench_flows_publish when it's ready for schedules and share links.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYesPattern editor the draft opens in. Pick interviewer for conversations that LEARN something from a person (research calls, intake, retros) — pair it with workbench_flows_share_create so non-account humans can talk to it.
modelNoThe model node slug the draft runs on — an `llm` entry from workbench_node_catalog_get that the workspace allows (e.g. an OpenAI, Anthropic, or Google model). Defaults to Claude Haiku 4.5. Set it here rather than editing the flow afterwards.
flowNameYesImperative + specific, ≤80 chars.
stubNotesNoWhat's stubbed / missing / worth wiring next.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
categoriesNoclassifier only: 2-12 buckets.
descriptionNo1-2 sentences: what it does, where it fits.
systemPromptNoThe draft's behavior — production-quality first attempt with [STUB: ...] / [FILL IN: ...] markers where reality is missing. Required for structurer and classifier. Leave it out of an agent to get the bare model with no instructions (a 'promptless' agent).
opportunityIdNoCompass change (from compass_opportunities_list) this draft implements. Sets the flow ↔ change link; refused when the change already has a draft.
inputDescriptionNoWhat the flow's single input carries.
responseSchemaJsonNostructurer only: the output JSON Schema, JSON-encoded as a string (object root, required fields, additionalProperties false).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the generic safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description goes well beyond that: the draft opens in Workbench's simple editor at a concrete path, may return needs_confirmation requiring user approval and re-call with approvalId, and opportunityId is refused when the change already has a draft. Rich behavioral disclosure the annotations cannot supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and then structured as author-guide → return-path → follow-on tools. Dense and long for a single description, but nearly every clause is actionable and nothing reads as filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter mutation tool with no output schema, the description covers the critical operational context: what gets produced, where it opens, the approval loop, the linking behavior, and the recommended next tools. An agent has enough to call it correctly on the first attempt.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning on top: it explains how to choose `mode` (pattern → use case), calls for [STUB: ...] markers paired with stubNotes, and describes the flow↔change linking role of opportunityId. It adds value beyond the schema, though it does not touch every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Creates a runnable first-draft flow from a spec you author') and frames the scope as scaffolding a draft rather than editing, running, or publishing one. An agent can distinguish it from siblings like workbench_flows_update, workbench_flows_fork, and workbench_flows_run without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit pattern-selection guidance (classifier for read-and-bucket, structurer for transform/extract, agent for conversational) and names the follow-on tools with their conditions: workbench_flows_run to try it, workbench_flows_publish when ready for schedules/share links. It also covers the needs_confirmation edge case and the Compass opportunityId path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_scheduleSchedule a Workbench flowAInspect

Set a specific built flow to run automatically on a cadence — the deterministic counterpart to a zv1 routine. Each run executes the flow's latest PUBLISHED revision with the fixed input envelope and routes the output per delivery. USE THIS when the user wants a flow they've built to run on a schedule ("run my digest flow every morning"). Do NOT create a routine for this — a routine runs a free-form instruction, not a built flow. The schedule runs the PUBLISHED flow, so the flow must be published (check workbench_flows_revisions_list; publish with workbench_flows_publish if it isn't). input is fixed for every run, so the flow itself should fetch anything time-varying at run time. Manage existing schedules with workbench_flows_schedules_list / workbench_flows_schedules_update.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesFixed input envelope every run executes with — {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}, the same shape workbench_flows_run takes. Constant across runs.
labelNo
flowIdYesFlow to schedule (id from workbench_flows_list).
deliveryNo
intervalYesCadence: hourly, daily, or weekly.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
scheduleTimeNo
scheduleTimezoneNo
scheduleDayOfWeekNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, so the mutation/safety profile is already known. The description adds real context beyond that: it runs the latest PUBLISHED revision (not draft), the input envelope is fixed for every run, output is routed per delivery, and the flow must be published first. It stops short of explaining approval/needs_confirmation behavior despite an approvalId parameter, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then usage, then prerequisites, then management siblings. Dense but each sentence carries routing or behavioral information. Slightly long, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter mutation with no output schema and an approval flow, the description covers prerequisites, execution semantics, and sibling management well. The notable omission is scheduling cadence detail (time-of-day, timezone, day-of-week) which an agent needs to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%, so the description must compensate and it partly does: it clarifies that `input` is a constant envelope matching workbench_flows_run's shape and that `delivery` routes output. However it says nothing about the timing parameters (scheduleTime, scheduleTimezone, scheduleDayOfWeek) or label, leaving half the parameters to bare schema. Baseline 3 given the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Set a built flow to run automatically on a cadence') and immediately contrasts it with the nearest sibling ('the deterministic counterpart to a zv1 routine'). An agent can distinguish it from routines_create and workbench_flows_run without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit USE THIS trigger ('run my digest flow every morning') plus an explicit exclusion ('Do NOT create a routine for this'). It also routes the agent to the exact prerequisite and lifecycle siblings: workbench_flows_revisions_list, workbench_flows_publish, workbench_flows_schedules_list/update.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_schedules_deleteDelete a flow scheduleA
Destructive
Inspect

Removes a schedule for good. Prefer workbench_flows_schedules_update with enabled: false when the user might want it back. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowIdYesFlow the schedule belongs to.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
scheduleIdYesSchedule id, from workbench_flows_schedules_list.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds meaningful context beyond them: permanence ('for good'), the safer soft-delete alternative, and the fact that the call may return `needs_confirmation` (which pairs with the approvalId parameter). It stops short of describing idempotency or side effects on dependents, so it's strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and permanence, then the alternative, then the confirmation caveat. No filler; every sentence carries a distinct decision-relevant fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return burden and handles it by flagging `needs_confirmation` and tying it to approvalId. For a destructive mutation with a confirmation round-trip this is nearly sufficient; it could still note what happens to a running schedule or in-flight executions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so flowId, scheduleId, workspace, and approvalId are all documented in the schema itself. The description hints at the confirmation flow (approvalId) but adds no syntax or value details beyond what the schema already provides, matching the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Removes a schedule') with the permanence qualifier ('for good'), clearly distinguishing it from the sibling update tool. An agent can tell what this does and how it differs from workbench_flows_schedules_update without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative (workbench_flows_schedules_update with enabled: false) and the exact condition that selects it ('when the user might want it back'). This is a textbook when-not-to-use-the-destructive-tool guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_schedules_listList a flow's schedulesA
Read-only
Inspect

Every schedule on one flow: cadence, whether it's paused (enabled: false), the next fire time, delivery, and the last run's status. Read this before workbench_flows_schedule (don't create a duplicate) and to get the scheduleId for workbench_flows_schedules_update / workbench_flows_schedules_delete / workbench_flows_schedules_run_now.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowIdYesFlow id, from workbench_flows_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, non-destructive, closed-world behavior, so the bar is lower. The description goes further by disclosing the returned fields, including how paused state surfaces as 'enabled: false' — signaling the actual response shape in the absence of an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, and the payload contents are front-loaded ahead of the dependency guidance. Every clause carries a distinct routing or content signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the returned fields, and it chains correctly to the create/update/delete/run_now siblings. Nothing an agent needs to call this correctly and thread the scheduleId forward is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both flowId and workspace are already documented in the schema (flowId's provenance, workspace's auth caveats). The description adds no parameter-level detail beyond implying a single flow, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Every schedule on one flow') and enumerates the payload (cadence, paused state, next fire time, delivery, last run status), which distinguishes it cleanly from sibling schedule mutators. An agent knows exactly what it gets back without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use: read before workbench_flows_schedule to avoid creating a duplicate, and to obtain the scheduleId required by workbench_flows_schedules_update / _delete / _run_now. Alternatives and the condition selecting them are named outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_schedules_run_nowFire a flow schedule once, nowAInspect

Runs one iteration of a schedule immediately through the real scheduled pipeline (same input, same delivery — the email or channel it normally posts to) without moving its cadence. Works on paused schedules. Use it to test a schedule the user just set up. Spends workspace credit and delivers for real, so it sits behind the approval gate — may return needs_confirmation. The run is fire-and-forget: check workbench_flows_schedules_list for lastStatus, or workbench_executions_get for the trace.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowIdYesFlow the schedule belongs to.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
scheduleIdYesSchedule id, from workbench_flows_schedules_list.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, destructiveHint=false, openWorldHint=true. The description adds the real behavioral weight: it spends workspace credit, delivers for real to the email/channel, sits behind an approval gate that may return needs_confirmation, and is fire-and-forget with pointers to check lastStatus or the trace. This is exactly the mutation/credit context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then layers the operational constraints (credit, real delivery, approval) and the follow-up path in tight sentences. No filler; every clause carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, credit-spending tool with no output schema, the description covers the important gaps: async fire-and-forget semantics and where to read status (schedules_list) or trace (executions_get). An agent has everything needed to call it correctly and handle the confirmation flow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters, including workspace and approvalId. The description only indirectly relates to approvalId via its needs_confirmation note and adds no format or usage detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource: 'Runs one iteration of a schedule immediately through the real scheduled pipeline.' It explicitly contrasts with the cadence-altering sibling behavior ('without moving its cadence'), so an agent can distinguish it from workbench_flows_schedules_update or workbench_flows_run without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-to-use ('Use it to test a schedule the user just set up') and a scope note ('Works on paused schedules'), plus the approval-gate prerequisite. It does not explicitly name the alternative delete/update siblings or state when NOT to use it, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_schedules_updatePause / resume / re-time a flow scheduleA
Destructive
Inspect

Changes one schedule: enabled: false pauses it (configuration kept), enabled: true resumes and recomputes the next fire time from now; interval / scheduleTime / scheduleTimezone / scheduleDayOfWeek change the cadence; label, delivery, and input can change too. Pass only what changes. Get the scheduleId from workbench_flows_schedules_list. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputNoNew fixed input envelope — {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}.
labelNo
flowIdYesFlow the schedule belongs to.
enabledNofalse pauses, true resumes.
deliveryNo
intervalNoNew cadence: hourly, daily, or weekly.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
scheduleIdYesSchedule id, from workbench_flows_schedules_list.
scheduleTimeNo
scheduleTimezoneNo
scheduleDayOfWeekNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations flag destructiveHint=true and readOnlyHint=false, so the mutation profile is already known. The description goes beyond that by explaining that pausing keeps configuration, that resuming recomputes the next fire time from now, and that the call can return needs_confirmation (a two-step approval flow). It leaves some gaps (permission requirements, effects on in-flight runs) but adds real behavioral value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the tool's core behavior and dense with useful field mapping; every clause carries information. The semicolon-chained field list is efficient though a touch packed, and the scheduleId/confirmation notes are properly appended.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter destructive mutation with no output schema and 58% schema coverage, the description covers the confirmation flow, the partial-update contract, and the key field semantics. Minor uncovered params (workspace) and the absence of return-value detail are acceptable given no output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 58% schema description coverage across 12 params, the description compensates by naming enabled, interval, scheduleTime, scheduleTimezone, scheduleDayOfWeek, label, delivery, and input and describing their effects. It doesn't mention workspace or approvalId directly, though the needs_confirmation note implies the latter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (changes) plus resource (one schedule) and then enumerates exactly what each field does, including the pause/resume semantics. An agent can distinguish this from workbench_flows_schedules_list, _delete, and _run_now without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational guidance: 'Pass only what changes' (partial update), where to obtain the scheduleId, and that the call may return needs_confirmation. It does not explicitly contrast with alternatives (delete a schedule, run now, create a new schedule via workbench_flows_schedule), so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_share_createMint a guest share link for a flowAInspect

Creates a no-account guest link where anyone can chat with the flow's latest PUBLISHED revision at workbench's /s/ page — the fastest way to put a working flow in a stakeholder's hands ('here, try it'). The flow must have a published version (workbench_flows_revisions_list shows it; publish with workbench_flows_publish if not). Conversations are capped per guest; the link can expire, and you can list links with workbench_flows_shares_list and kill one with workbench_flows_shares_revoke. May return needs_confirmation — say who the link is for and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNoWho/what this link is for, e.g. 'support pilot'.
flowIdYesFlow to share.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this.
approvalIdNoApproval id from a prior needs_confirmation envelope.
expiresInDaysNoDays until the link dies. Omit = revoke-only.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavioral context beyond the annotations: per-guest conversation caps, link expiry, the possible `needs_confirmation` return envelope, and the required response ('say who the link is for and wait'). The annotations confirm a non-read-only, open-world, non-destructive operation, and the description enriches this with concrete operational constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first clause, then layers precondition, behavioral traits, and the confirmation protocol. Every sentence carries distinct, load-bearing information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still covers the meaningful return signals (link, expiry, needs_confirmation) and the full lifecycle (list/revoke). An agent has everything needed to invoke it correctly and handle the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning to `label` ('say who the link is for') and `approvalId` (implicitly tied to the needs_confirmation flow) that the schema alone only lightly conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Creates a no-account guest link' to the flow's latest PUBLISHED revision) and even names the surface (/s/<token> page). It is immediately distinguishable from siblings like workbench_flows_publish and workbench_flows_shares_list, which it explicitly references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit usage context ('the fastest way to put a working flow in a stakeholder's hands'), a precondition (flow must have a published version, checked via workbench_flows_revisions_list, published via workbench_flows_publish), and names the alternatives for managing/removing links (shares_list, shares_revoke).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_shares_listList a flow's guest share linksA
Read-only
Inspect

Every guest link on one flow, newest first: label, whether it's still live, the shareUrl (live links only — dead ones have no URL), expiry, and how many guests and messages it has seen. Use it to answer 'who has access', to recover a link the user lost, and to get the shareId for workbench_flows_shares_revoke.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowIdYesFlow id, from workbench_flows_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavior beyond that: results are ordered newest first, and dead links deliberately omit a shareUrl. It does not discuss pagination or scoping limits, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the scope and return shape before the usage sentence, and every clause carries information. The first sentence is dense with a parenthetical, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so by naming the fields (label, live status, shareUrl, expiry, guest/message counts). Combined with the annotation-covered read-only profile, nothing needed to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so flowId and workspace are already fully documented in the schema, including the per-call override rules. The description only implies flowId via 'one flow' and adds no syntax or format detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (guest share links on one flow) and enumerates the returned fields, so the agent knows exactly what it gets back. It also distinguishes itself from the sibling shares_revoke by noting it supplies the shareId that revoke needs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three explicit use cases ('who has access', recovering a lost link, obtaining a shareId for workbench_flows_shares_revoke) and names the alternative sibling. Nothing is left to inference about when to reach for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_shares_revokeRevoke a flow's guest share linkA
Destructive
Inspect

Kills one guest link immediately — anyone holding it gets a closed page from then on. Conversations already had stay in the flow's history. Get the shareId from workbench_flows_shares_list. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowIdYesFlow the link belongs to.
shareIdYesShare id, from workbench_flows_shares_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds real behavioral value beyond that: immediate effect ("anyone holding it gets a closed page from then on"), what is preserved ("Conversations already had stay in the flow's history"), and the two-step confirmation flow ("May return needs_confirmation") tied to the approvalId parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each carrying distinct information, with the destructive effect front-loaded. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, no-output-schema action, the definition covers effect, data preservation, id source, and confirmation behavior. It is essentially complete for correct invocation, though it doesn't touch permission/auth requirements beyond what the workspace param implies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented. The description restates the shareId source (already in the schema) and only implicitly touches approvalId via the needs_confirmation note, adding little beyond the structured fields. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Kills one guest link immediately" gives a specific verb (revoke) and a precisely scoped resource (a single guest share link), and the description clearly separates this from workbench_flows_shares_list and workbench_flows_share_create. An agent can distinguish it from siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It directs the caller to obtain shareId from workbench_flows_shares_list and notes the confirmation path, giving clear context for invocation. It doesn't state when NOT to use this or name alternatives, so it falls short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_flows_updateRename / describe / archive / re-scope a Workbench flowA
Destructive
Inspect

Changes a flow's name, description, visibility, tags, or archived state — the metadata around the flow, NOT its orchestration (prompts and nodes change through workbench_flows_edit_text). Pass only what changes. Archiving hides the flow from the default gallery but leaves it runnable, scheduled, and shared; pass archived: false to restore. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoReplacement tag set (lowercase labels). The whole set, not a diff.
flowIdYesFlow id, from workbench_flows_list.
archivedNotrue archives, false restores.
flowNameNoNew name.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
visibilityNoWho can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE.
descriptionNoNew description; null clears it.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the mutation/destructive profile, so the description earns credit for going further: it explains that archiving hides the flow from the default gallery while leaving it runnable, scheduled, and shared, that archived:false restores, and that the call may return `needs_confirmation`. These are real behavioral facts an agent needs that the schema does not spell out.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the operation and its scope, then the exclusion, then partial-update rule, then the archive semantics. Every sentence carries information and none is redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a metadata-mutation tool with no output schema this is complete: it covers scope, exclusion, update semantics, the destructive-ish archive behavior, and even the confirmation round-trip. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning the schema lacks — the whole-set partial-update contract ('Pass only what changes') and the restore semantics of archived:false. This is useful framing beyond the field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (changes) and resource (flow metadata: name, description, visibility, tags, archived state) and explicitly scopes out what it does NOT do — orchestration changes route to workbench_flows_edit_text. An agent can distinguish this from the sibling edit tool without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (workbench_flows_edit_text) with the condition that selects it (prompts/nodes = orchestration) and gives the partial-update rule 'Pass only what changes.' No explicit when-NOT beyond the orchestration split, but routing guidance is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_add_documentsAdd documents to a knowledge baseAInspect

Adds documents to a knowledge base and starts ingestion (chunk + embed; graph KBs also extract entities; tabular KBs load CSV as tables). Text goes inline (Markdown, plain text, CSV — encoding utf8); binary files (PDF, .docx) go base64-encoded with their mimeType. Up to 50 sources per call, ~10 MB each. Ingestion runs in the BACKGROUND — the result carries a runId; check it with workbench_kb_ingestion_status when the user asks, don't poll. workbench_kb_search works once it finishes. Content from blocks is ideal source material. Check workbench_kb_documents_list first so you don't add a document twice. Embedding spends workspace inference credit, so this sits behind the approval gate: may return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
kbIdYesKnowledge base id (from workbench_kb_list or workbench_kb_create).
sourcesYesDocuments to ingest (1-50).
workspaceNoWorkspace slug override.
approvalIdNoApproval id from a prior needs_confirmation envelope.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the basic safety profile (readOnly=false, destructive=false, openWorld=false). The description goes well beyond: ingestion runs in the BACKGROUND and returns a runId, embedding spends workspace inference credit, and the call sits behind an approval gate that may return a needs_confirmation envelope. This is exactly the extra behavioral context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and ingestion behavior, then constraints and routing. It is dense and somewhat long, but nearly every clause carries operational value (limits, encoding, background semantics, credit spend); the density is justified rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still closes the loop by explaining that the result carries a runId, how to check it, and the possible needs_confirmation envelope. Combined with the limits and encoding rules, an agent has everything required to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: text goes inline as Markdown/plain text/CSV in utf8, binary files go base64-encoded with their mimeType, and it restates the 50-source / ~10 MB limits. It also points to <attached-file> blocks as ideal source material.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Adds documents to a knowledge base and starts ingestion') and immediately differentiates by ingestion mode (text vs binary vs graph vs tabular). It also names the sibling tools it interacts with, so an agent can place it against workbench_kb_search and workbench_kb_documents_list without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when/when-not guidance and alternatives: check workbench_kb_documents_list first to avoid duplicates, use workbench_kb_ingestion_status to check progress and explicitly 'don't poll', and workbench_kb_search only works once ingestion finishes. Names the approval gate as a condition that may return needs_confirmation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_createCreate a knowledge baseAInspect

Creates an empty knowledge base. Recipe 'docs' (default) for reference material an agent searches at runtime; 'tabular' for spreadsheet-style data queried with SQL; 'graph' for entity/relationship extraction. After creating, add content with workbench_kb_add_documents, then attach the KB to an agent in the Workbench editor (or tell the user to). May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesKnowledge base name (1-200 chars).
recipeNoDefault 'docs'.
workspaceNoWorkspace slug override.
approvalIdNoApproval id from a prior needs_confirmation envelope.
visibilityNoWho can see it: PRIVATE (only the user), WORKSPACE (every member, the default), or SHARED (specific people, granted afterwards). Say 'make it private' → PRIVATE.
descriptionNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare this is a non-destructive write (readOnlyHint=false, destructiveHint=false). The description adds meaningful context beyond the annotations: the KB starts empty, and it notes the 'May return `needs_confirmation`' flow, which signals a possible approval step. It does not detail permissions or the confirmation payload, but the added disclosure is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by recipe semantics, then the follow-up workflow. Three sentences, each carrying distinct, useful information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param write tool with 83% schema coverage and no output schema, the description covers the essentials: what it creates, how to select a recipe, the next steps, and the confirmation possibility. Minor gaps remain (workspace defaulting, error behavior), but nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so most parameters are self-documented and the schema covers visibility in depth. The description adds genuine meaning for the recipe values ('docs'/'tabular'/'graph'), which the schema only labels as 'Default docs'. This elevates it above the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Creates an empty knowledge base') and immediately differentiates itself from the sibling workbench_kb_add_documents by describing the create-then-add workflow. The recipe enumeration further clarifies the tool's scope and variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear lifecycle: create, then add content with workbench_kb_add_documents, then attach to an agent in the editor. The recipe descriptions also guide selection ('docs' for runtime-searched reference material, 'tabular' for SQL, 'graph' for entity extraction). It lacks explicit when-not-to-use conditions, but the context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_deleteDelete a knowledge baseA
Destructive
Inspect

Deletes a knowledge base and everything in it. Flows that search it will find nothing — check which agents use it (the user knows; the KB page in Workbench lists them) and say so before proposing. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
kbIdYesKnowledge base id, from workbench_kb_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, but the description adds substantive behavior: the delete is cascading ("everything in it"), has downstream effects ("Flows that search it will find nothing"), and may short-circuit into an approval flow returning needs_confirmation. That is exactly the extra context a destructive tool's annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and its scope before the cautionary guidance. The parenthetical "(the user knows; the KB page in Workbench lists them)" is slightly clunky but earns its place by telling the agent where the consumer list lives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema destructive tool, the description covers scope, downstream impact, and the confirmation handshake, which is most of what an agent needs. It omits whether deletion is permanent/irreversible and what permissions are required, leaving a small but real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so kbId, workspace, and approvalId are all documented in the schema itself. The description only indirectly gestures at approvalId via "May return needs_confirmation"; it adds no format or constraint detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Deletes a knowledge base") and immediately scopes the blast radius ("everything in it"), which distinguishes it from the sibling workbench_kb_document_delete that removes a single document. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete precondition — check which agents use the KB and warn the user before proposing deletion — plus the confirmation path ("May return needs_confirmation"). It does not explicitly contrast against sibling alternatives such as workbench_kb_document_delete, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_document_deleteRemove a document from a knowledge baseA
Destructive
Inspect

Removes one ingested document and all its chunks; searches stop returning it at once. Get the documentId from workbench_kb_documents_list. To replace a document, remove it then workbench_kb_add_documents the new version. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
kbIdYesKnowledge base id.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
documentIdYesDocument id, from workbench_kb_documents_list.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag this as destructive and non-read-only; the description adds real behavior: chunks are removed wholesale, search results stop returning the document immediately, and the call may return `needs_confirmation`. That confirmation signal pairs with the approvalId parameter and is the kind of context annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each load-bearing: effect, id sourcing, replacement path, and confirmation caveat. The destructive effect is front-loaded rather than buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers the key things an agent must know: irreversibility of chunk removal, immediate search impact, and the needs_confirmation/approvalId handshake. It omits any note on failure modes or whether removal is reversible, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so kbId, workspace, approvalId, and documentId are already documented in the schema, including the workspace-token nuance. The description only reinforces that documentId comes from workbench_kb_documents_list, adding marginal value over the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Removes one ingested document') and adds scope detail ('and all its chunks'), which cleanly separates it from workbench_kb_delete (whole KB) and workbench_flows_delete. An agent can pick it out of the densely populated workbench_kb_* family without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete preconditions ('Get the documentId from workbench_kb_documents_list') and a workflow alternative for replacement ('remove it then workbench_kb_add_documents the new version'). It lacks an explicit when-not statement (e.g., when to use KB-level delete instead), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_documents_listList the documents in a knowledge baseA
Read-only
Inspect

Every document ingested into the KB's current content: id, name, type, size, chunk count. Empty for a KB with nothing ingested yet (or one whose ingestion is still running — check workbench_kb_ingestion_status). Use it before workbench_kb_add_documents to avoid re-adding a document, and to get the documentId for workbench_kb_document_delete.

ParametersJSON Schema
NameRequiredDescriptionDefault
kbIdYesKnowledge base id, from workbench_kb_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, closed-world), and the description adds genuinely new behavior: the result is empty both when nothing has been ingested and when ingestion is still running, with a pointer to workbench_kb_ingestion_status. It stops short of pagination or ordering behavior, but the empty-state disclosure is the operationally important one.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero filler, and the critical scope and return fields come first. The parenthetical about the empty state is dense but placed immediately after the claim it qualifies, so nothing important is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by naming the returned fields. Combined with the empty-state semantics, the ingestion cross-reference, and the two downstream use cases, an agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so kbId and workspace are already documented with their constraints (workspace slug rules, personal-token requirement). The description adds nothing about either parameter. Baseline 3 applies when the schema carries the full parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (list documents in a KB) and pins the exact scope: 'Every document ingested into the KB's current content.' It even enumerates the returned fields (id, name, type, size, chunk count), which no sibling tool claims, so an agent can separate it from workbench_kb_search or workbench_kb_get without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives two concrete when-to-use conditions: call before workbench_kb_add_documents to avoid re-adding a document, and call to obtain the documentId required by workbench_kb_document_delete. It also routes the empty-result case to workbench_kb_ingestion_status, so no ambiguous path is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_getRead one knowledge baseA
Read-only
Inspect

One knowledge base: name, recipe (docs / tabular / graph), status, document and chunk counts, embedding model, and ingestion settings. Read it to confirm a KB has content before pointing a flow at it. Get the kbId from workbench_kb_list.

ParametersJSON Schema
NameRequiredDescriptionDefault
kbIdYesKnowledge base id.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds value beyond that by disclosing what the read returns (name, recipe, status, document/chunk counts, embedding model, ingestion settings). It does not address auth/scope behavior, but the safety profile is already handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the returned fields and following with the usage condition. No filler. The field enumeration is a slightly long list but each item is informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden, and it does so by listing the concrete fields an agent would receive. For a simple single-record read with full schema coverage and safety annotations, this is nearly complete; only auth/scope behavior is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both kbId and workspace (including the workspace-auth nuance). The description adds only provenance guidance ('Get the kbId from workbench_kb_list'), so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource (one knowledge base) and enumerates the fields it returns, and the title supplies the read verb. It implicitly distinguishes itself from the plural sibling workbench_kb_list by pointing to it for the kbId, though the verb itself is left to the title. Clear and unambiguous, just not a fully self-contained verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit usage trigger: 'Read it to confirm a KB has content before pointing a flow at it,' which tells the agent when this tool is the right call. It also routes to workbench_kb_list to obtain the id. There is no explicit when-not-to-use or statement of alternatives for the read itself, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_ingestion_statusCheck an ingestion runA
Read-only
Inspect

The state of one ingestion run started by workbench_kb_add_documents: PENDING / RUNNING / DONE / FAILED, chunks written, and the error when it failed. Read it when the user asks whether their documents are in yet, or before a search that needs them — don't poll in a loop; ingestion of a few documents takes under a minute. Pass the runId the add-documents call returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
kbIdYesKnowledge base id.
runIdYesIngestion run id, from workbench_kb_add_documents.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a read-only, non-destructive, closed-world operation, but the description adds context annotations cannot carry: the exact state machine (PENDING/RUNNING/DONE/FAILED), what is reported on success (chunks written) and failure (the error), and the latency expectation that discourages polling. It stops short of describing pagination or retention of run records, keeping it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: what is returned, when to call it (plus the polling warning), and where the key parameter comes from. The return shape is front-loaded ahead of the routing advice.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the returned status values, the chunk count, and the error field, and it covers a 3-parameter tool whose schema is fully documented. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so kbId, runId, and workspace are already documented in the schema; the description only adds provenance for runId ('the add-documents call returned'). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (the state of one ingestion run) and ties it explicitly to its producer, workbench_kb_add_documents, which separates it from every sibling in the kb family. An agent can identify this as a status-polling tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit triggers ('when the user asks whether their documents are in yet, or before a search that needs them') and an explicit anti-pattern with the reason ('don't poll in a loop; ingestion of a few documents takes under a minute'). It also names the sibling that supplies the required runId.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_kb_listList knowledge basesA
Read-only
Inspect

Lists every knowledge base in the active workspace the caller can see. Returns summaries (id, name, description, status, doc/chunk counts, embedding model). Use an id with workbench_kb_search to retrieve content.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys (workspace is intrinsic).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds real value beyond that: it discloses that results are visibility-filtered to the caller's workspace and enumerates the summary payload (id, name, description, status, doc/chunk counts, embedding model). It doesn't address pagination or result-size limits, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, zero filler, with the listing scope front-loaded and the follow-on routing placed last. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the returned summary fields, and the single parameter is fully documented in the schema. What remains unstated is pagination/ordering behavior and what an empty result means, which are minor for a read-only list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single workspace parameter, including the token-vs-API-key nuance, so the schema carries the burden. The description only says 'active workspace' and adds no format or default-override detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists every knowledge base') plus the visibility scope ('in the active workspace the caller can see'), which separates it from workbench_kb_get and workbench_kb_search. It even names the return fields, so an agent knows this is the discovery/list tool rather than a content fetch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to reach for it (enumerating KBs visible to the caller) and explicitly routes the follow-on action: 'Use an id with workbench_kb_search to retrieve content.' It does not contrast with workbench_kb_get or explain when listing is preferable to searching, so it stops short of full alternative coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_node_catalog_getSearch the Workbench node catalogA
Read-only
Inspect

The ground truth for what nodes exist and what their ports actually are — verify against this instead of recalling. Search by keyword/category for summaries; pass slug for one node's full detail (inputs, outputs, settings). Inputs are PORTS fed by links; settings live on the node — see workbench_flow_authoring_guide. The catalog is the WORKSPACE'S: model nodes the workspace's inference policy forbids come back with allowed: false — never author with those; pick an allowed model. Pass flowId to include the flow's pinned imports as nodes.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoKeyword over slug / name / tagline / description.
slugNoExact node slug → full detail for that one node (ports + settings).
limitNoResults per page, 1-50. Default 20.
flowIdNoA flow whose pinned imports should appear as `imported-<uuid>` nodes.
offsetNoPagination offset.
categoryNoFilter by category.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; the catalog's policy annotation is per workspace.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish a safe read-only, closed-world profile, but the description adds real behavioral context: the catalog is workspace-scoped, policy-forbidden model nodes are returned with `allowed: false`, and personal tokens without a default workspace must pass `workspace`. These are non-obvious traits an agent could not infer from the schema or annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and dense with no filler: the ground-truth claim comes first, then search/detail modes, then the workspace policy caveat. It is packed tightly and borders on over-compressed, but every sentence conveys distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, 7-parameter tool with no output schema, the description covers purpose, key parameter roles, workspace policy behavior, and the cross-reference an agent needs. It does not address the paginated result shape or how to interpret `limit`/`offset` behavior, which is a minor residual gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds conceptual meaning the schema lacks – that `q` returns summaries while `slug` returns full detail, `flowId` surfaces pinned imports as `imported-<uuid>` nodes, and `inputs` are PORTS fed by links while settings live on the node. Pagination semantics (limit/offset) stay in the schema only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource with scope: it is the ground truth for what nodes exist and what their ports are, searchable by keyword/category for summaries or by `slug` for full per-node detail. It also routes away from the sibling `workbench_flow_authoring_guide` for settings-vs-ports guidance, so an agent can distinguish it from related workbench tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use guidance ('verify against this instead of recalling'), an exclusion ('never author with those [allowed: false] nodes'), and a named cross-reference for the related concern in `workbench_flow_authoring_guide`. This is actionable routing, not implied usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_shim_add_answerAdd an answer to a shimBInspect

Adds one answer (a label and, optionally, when it applies). The shim needs three examples of it before it builds again with the new answer. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
answerYes
shimIdYesShim id, from workbench_shim_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
descriptionNoWhen this answer applies.

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only covering safety flags (readOnly=false, destructive=false, openWorld=false), the description adds real behavioral context: the confirmation gate ('May return `needs_confirmation`') and the prerequisite of three examples before the shim rebuilds. It still doesn't say what state the answer is added in or whether it can be undone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action, with no filler. Slightly dense wording around 'three examples of it' but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-param mutation tool with no output schema, the description hints at the confirmation return path but does not explain what happens after needs_confirmation or how approvalId closes the loop. It omits any note on workspace scoping (left to the schema) and does not describe the effect of adding an answer before the rebuild threshold is met.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents shimId, workspace, approvalId, and description. The description loosely maps to the `answer` and `description` params ('a label and, optionally, when it applies') but adds no syntax or format detail beyond the schema's maxLength/minLength constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('Adds one answer' to a shim), so the operation is identifiable. However, it does not distinguish itself from the adjacent sibling workbench_shim_add_examples, and the phrase 'a label and, optionally, when it applies' is ambiguous about what an 'answer' actually is (the schema calls it a plain string).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus workbench_shim_add_examples or workbench_shim_create. The only contextual hint is the behavioral note that three examples are needed before a rebuild, which is not a usage-selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_shim_add_examplesAdd examples to an answerAInspect

Adds example inputs to one answer. Write inputs a real person would actually type in the shim's setting — varied in length, register and specifics, each one clearly this answer and not another — never paraphrases of the answer's name. Read the shim first so new examples fit alongside the existing ones and land where recall is low. Exact duplicates are skipped. The shim rebuilds on its own afterwards; read it again for the new report. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
answerYesThe answer these examples belong to, exactly as labelled.
shimIdYesShim id, from workbench_shim_list.
examplesYes
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (readOnly=false, destructive=false, openWorld=false); the description adds substantial behavior beyond that — exact duplicates are silently skipped, the shim rebuilds asynchronously on its own, the shim must be re-read for the new report, and a `needs_confirmation` path exists. That is exactly the mutation/confirmation detail an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then guidance, then the confirmation caveat. Sentences are dense but each carries distinct information; the middle exemplar sentence is long but earns its length by defining the value contract.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly covers the post-call state (rebuild, re-read, possible needs_confirmation). Combined with 80% schema coverage and annotations, an agent can call and handle the result correctly; only the absence of an explicit alternative-routing statement keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the params are largely documented. The description still adds real meaning: what a good value for `examples` looks like (varied length, register, specifics, distinct from other answers, never paraphrases of the answer's name), which the schema's minLength/maxItems constraints cannot convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource ('Adds example inputs to one answer'), which cleanly separates it from sibling workbench_shim_add_answer. It does not explicitly name the sibling it differs from, but the resource scoping ('one answer') makes the target unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete operating context: read the shim first so examples fit alongside existing ones and 'land where recall is low', which tells the agent both sequencing and intent. It stops short of naming when-not-to-use conditions or alternatives to add_answer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_shim_createCreate a shimAInspect

Creates a shim from a name, the question it answers, and its answers. Add examples afterwards with workbench_shim_add_examples — three or more per answer and it builds on its own. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesA short name, e.g. "Support triage".
answersYes
questionYesThe question, e.g. "Which team should read this message?"
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.
inputsComeFromNoOne line on where the inputs come from, in the user's words.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-read-only, non-destructive, closed-world write, and the description adds real behavior beyond that: the tool may return a `needs_confirmation` state, and three or more examples per answer cause the shim to "build on its own." The auto-build behavior and confirmation interlock are not present in the annotations or schema, so this is genuine added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the creation act and the inputs, followed by the follow-up step and the confirmation caveat. Nothing is redundant, though the phrase "it builds on its own" is slightly vague for a behavioral claim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description takes on return-value disclosure and does flag the `needs_confirmation` outcome, which the schema's approvalId field complements. It stops short of describing the happy-path response or what state the shim is in after creation (draft vs active), which an agent planning follow-up calls might want.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, so parameters are largely self-documenting; the description names the three required parameters (name, question, answers) in natural language but adds no format, syntax, or constraint detail beyond the schema. The optional workspace, approvalId, and inputsComeFrom parameters are never mentioned. Baseline 3 is appropriate when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Creates a shim") and immediately names the three inputs that constitute it, so the agent knows what it will be building. It is distinguishable from siblings like workbench_shim_add_examples or workbench_shim_get. It never defines what a 'shim' actually is, relying on the question/answers structure to imply a routing/classification artifact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit sequencing guidance ("Add examples afterwards with workbench_shim_add_examples"), which usefully directs the agent's next call. It does not, however, say when to reach for shim_create versus other workbench primitives such as flows or tasks, nor state prerequisites beyond the approval flow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_shim_getRead one shimA
Read-only
Inspect

One shim in full: the question, where its inputs come from, every answer with its description and examples, what the shim still needs before it can build, and the newest build's report — held-out accuracy, recall per answer, which answers get mistaken for which, and the compiler's notes. Read it before writing examples: match the setting and the existing examples' register, and put new examples where recall is low or the compiler says an answer is thin.

ParametersJSON Schema
NameRequiredDescriptionDefault
shimIdYesShim id, from workbench_shim_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (readOnly, non-destructive, closed-world), and the description adds substantial context beyond that: it discloses the full return shape including held-out accuracy, per-answer recall, confusion pairs, the compiler's notes, and what the shim still needs before it can build. With no output schema, this disclosure is doing real work and reveals behavior an agent would otherwise have to discover empirically.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the payload contents, then a second sentence of prescriptive guidance. It is dense and the first sentence is a long enumeration, but every clause maps to a real part of the returned report, so little is wasted; a light trim would make it a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-resource read with no output schema, the description fully covers what comes back, why to call it, and how to act on it. Parameters are covered by the schema and safety by the annotations, so nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — shimId is documented as coming from workbench_shim_list and workspace covers default/override semantics for personal vs API-key tokens. The description adds nothing about either parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific scope statement, 'One shim in full,' and then enumerates exactly what a shim is composed of (the question, input sources, answers with descriptions/examples, build requirements, newest build report). This lets an agent distinguish it from siblings like workbench_shim_list (plural/roster) and the shim mutation tools (create/add_answer/add_examples) without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage trigger, 'Read it before writing examples,' and even prescribes what to do with the result ('match the setting and the existing examples' register,' target low-recall answers). That is strong when-to-use and workflow guidance, though it does not name the sibling action tool explicitly or state a when-not-to-use condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_shim_listList shimsA
Read-only
Inspect

Every shim the user can see: id, name, the question it answers, how many answers and examples it has, and the published version's held-out accuracy (null until something is published). A shim is a small decision model that runs inside an app with no model call. Start here to get a shimId.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, closed-world). The description adds real value beyond them: visibility scope ('the user can see'), the null-accuracy-until-published semantic, and the fact that a shim runs with no model call. It does not mention pagination or result limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the returned-field enumeration, then the definition of a shim, then the entry-point cue. The field list is dense but earns its place by compensating for the absent output schema; only minor trimming possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, enumerating returned fields is the right move and is done well; annotations cover safety. The only gap is that nothing indicates result size, pagination, or ordering for a list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema fully explains the workspace slug including the personal-token default/override rules. The description says nothing about the workspace parameter, so it adds no meaning beyond the schema — the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Every shim the user can see') and enumerates the exact fields returned, including a mini-definition of what a shim is. 'Start here to get a shimId' distinguishes it from workbench_shim_get, which needs an id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly positions itself as the entry point ('Start here to get a shimId'), which tells the agent when to use this over workbench_shim_get. No explicit when-not or exclusion conditions are given, but the entry-point framing covers the main routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_actAct on a Task gate (record a verdict)AInspect

Records THIS user's verdict on a run that is waiting on them — the labels come from the gate (workbench_tasks_run_view shows them via the plan; a wrong label errors listing the real options). Echo the verdict + note back to the user in one line BEFORE calling; that echo is the confirmation. Refuses when the gate is assigned to someone else.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoThe user's reasoning, attached to the verdict.
runIdYesTask run id that is waiting_human.
verdictYesOne of the gate's verdict labels, e.g. "confirmed".
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare mutation/no-destruct/open-world. The description adds real behavioral context: it refuses when the gate belongs to another user, a wrong label errors listing the real options, and it introduces the two-step approvalId/needs_confirmation flow plus the required echo-as-confirmation pattern.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action ('Records THIS user's verdict...') in one dense paragraph. The echo instruction is packed with clauses but every sentence earns its place; slightly convoluted but no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an approval/confirmation flow, the description covers ownership refusal, error behavior on bad labels, and the echo confirmation step. With no output schema and full schema coverage, this is complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all five parameters (including note, runId, verdict, workspace, approvalId) are already documented. The description reinforces that verdict labels originate from the gate, but adds little syntactic or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: records THIS user's verdict on a task run waiting on them. It distinguishes itself from siblings by referencing workbench_tasks_run_view as where the gate labels are shown, so an agent can tell it apart from workbench_tasks_start/freeze/get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context — use it on a run that is waiting on the user — and names the sibling that surfaces the gate labels. It notes the refusal condition (gate assigned to someone else), though it doesn't explicitly contrast with other task lifecycle tools like start or freeze.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_compileCompile a sentence into a Task planAInspect

Runs the compiler over a sentence and returns the PROPOSED plan (or repair-grade problems — fix by refining the sentence with the user, not by guessing). Nothing persists. Present the proposal in plain terms: which clauses run flows, which you (zv1) will handle, who decides, what the compiler noted. Then iterate or freeze. Compile may take ~15s.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdNoThe task this sentence belongs to, when compiling an edit — keeps it out of the child-task candidates.
sentenceYesThe sentence to compile.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoRe-call with the approvalId from a needs_confirmation response after the user approves.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds context annotations do not carry: nothing is persisted, the call can return repair-grade errors rather than a plan, and it may take ~15s. There is mild tension between “Nothing persists” and readOnlyHint=false, but the schema's approvalId / needs_confirmation flow plausibly accounts for that, so this reads as added detail rather than a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded: the first sentence gives the action and the return shape, followed by guidance, then the latency note. Dense but every clause is doing work; the “which you (zv1) will handle” phrasing is a touch cryptic but purposeful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully characterizes the return (proposed plan, compiler notes, repair problems) and the iteration/approval loop, plus latency. It stops short of explaining the plan's structure or the confirmation response shape, but is sufficient to call and act on it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each of the four parameters (taskId, sentence, workspace, approvalId) is already documented in the schema. The description only reinforces the sentence-refinement loop and adds nothing about formats or limits, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — “runs the compiler over a sentence” — and names the artifact it returns: the PROPOSED plan (or repair-grade problems). The “Nothing persists” clause and the “iterate or freeze” ending implicitly separate it from workbench_tasks_freeze/workbench_tasks_create, which commit state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context (“fix by refining the sentence with the user, not by guessing”, “Then iterate or freeze”) and tells the agent what to do after the call. It does not name the sibling that commits the plan (freeze/start), so a little inference is still required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_createCreate a draft TaskAInspect

Creates a DRAFT Task from a name + sentence. Use this once the user's sentence feels settled after you've refined it together — then compile it. The sentence should be THEIR words: trigger first ('When a claim arrives…'), then steps, branches, and who decides what.

ParametersJSON Schema
NameRequiredDescriptionDefault
sentenceYesThe process sentence, in the user's words.
taskNameYesShort task name.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare write-but-non-destructive (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the safety bar is lower. The description adds genuine context beyond that: the created object is a DRAFT rather than an active task, and it encodes a readiness gate ('sentence feels settled') before creation. It doesn't describe what is returned or the confirmation/approval round-trip, but the schema covers the latter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, followed by the usage condition and then format guidance, in three tight sentences. The em-dash aside and the 'who decides what' clause are slightly conversational but each carries usable instruction rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still conveys enough to call it correctly: what is created, its draft state, the readiness condition, and the downstream compile step. Minor gaps remain — what the call returns (e.g., a task id needed for the subsequent compile) and the two-phase approvalId flow are left to the schema rather than narrated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema already documents all four parameters, including 'in the user's words.' The description partly repeats that phrasing, but it adds real semantic guidance the schema lacks: the sentence structure should be trigger-first ('When a claim arrives…'), then steps, branches, and decision ownership. That is actionable authoring guidance for the most important parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb + resource + lifecycle scope ('Creates a DRAFT Task from a name + sentence'), which immediately separates it from the many sibling task operations (compile, start, freeze, update). The word 'DRAFT' plus 'then compile it' tells an agent exactly where this sits in the task lifecycle without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear readiness condition — use it 'once the user's sentence feels settled after you've refined it together' — and points to the natural next step ('then compile it'). It does not state what to do instead when the sentence is NOT settled (e.g., keep refining, or use update on an existing task), so there are no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_deleteDelete a Workbench TaskA
Destructive
Inspect

Deletes a task. Runs still in flight are stopped (a gate waiting on someone is closed), and its schedule trigger stops firing; past runs stay in history. Check workbench_tasks_get for live runs and say so before proposing. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesTask id, from workbench_tasks_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite destructiveHint=true already flagging the risk, the description adds substantial context: in-flight runs are stopped, a waiting gate is closed, the schedule trigger stops firing, and past runs remain in history. It also discloses the needs_confirmation return, which is exactly the kind of behavior an agent must anticipate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the action, then concisely covers consequences, the recommended pre-check, and the confirmation flow. Every sentence carries information, though the 'say so before proposing' phrasing is slightly informal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully discloses the possible needs_confirmation return and the fate of in-flight vs. past runs. Complete enough for a destructive tool, though the confirmation flow could be spelled out a touch more.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so taskId, workspace, and approvalId are already documented in the schema. The description's mention of needs_confirmation gives indirect context for approvalId, but adds no syntax or format detail beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Deletes a task'), so an agent immediately knows the operation. It does not differentiate from siblings like workbench_tasks_freeze or workbench_tasks_update, but the purpose itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit precondition workflow: 'Check workbench_tasks_get for live runs and say so before proposing.' It names the alternative tool and the condition that matters before deletion, though it stops short of stating when-not to delete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_freezeFreeze a compiled plan (activate the Task)A
Destructive
Inspect

Freezes a reviewed plan as the Task's next version and activates it. Pass the EXACT plan object a compile returned — never hand-edit it (change the sentence and recompile instead). Freeze ONLY after the user has seen the proposal and said go; echo what you're freezing in one line first. In-flight runs keep their pinned versions. After freezing, offer to start a first run.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYesThe plan object from workbench_tasks_compile, verbatim.
taskIdYesThe task id (create a draft first if none).
sentenceYesThe sentence this plan was compiled from.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, and the description goes well beyond that by scoping the damage: it is a versioned activation, and in-flight runs keep their pinned versions. It also surfaces a confirmation discipline ('echo what you're freezing in one line first') that the destructive annotation alone would not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each carrying a distinct obligation: what freeze does, how to pass the plan, when it is permitted, and what follows. The irreversible action and its precondition are front-loaded with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the essential agent concerns: the source of the plan argument, the confirmation gate, the effect on running work, and the natural next step. Nothing an agent needs in order to call it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning for the plan/sentence pair: the plan must be the verbatim object a compile returned, and if the sentence changes the plan must be recompiled rather than edited. That is a coupling constraint the schema does not express. It says nothing extra about workspace or approvalId.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (freezes) plus the exact outcome: turns a reviewed plan into the Task's next version and activates it. It also distinguishes itself from workbench_tasks_compile by requiring output from that tool and by describing the activation step, which no sibling performs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit preconditions (only after the user has seen the proposal and said go), an explicit anti-pattern with the correct alternative (never hand-edit the plan; change the sentence and recompile), and a post-action recommendation (offer to start a first run). Nothing about when to reach for this tool is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_getRead one Workbench Task (definition + recent runs)A
Read-only
Inspect

One Task in full: its sentence, status, the latest saved plan (the clauses — which run flows, which zv1 handles, who decides at each gate) with the flow and people names the plan refers to, and its recent runs. Read this before editing a task's sentence (workbench_tasks_compile with taskId) and when the user asks what a task does. Get the taskId from workbench_tasks_list.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesTask id.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds genuine value by disclosing what the read returns and how deep it goes (plan clauses, gate deciders, recent runs) — not just 'gets a task.' It stops short of describing run recency limits or whether the plan may be absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and its contents in a single dense sentence, then two short routing sentences. Slightly heavy on parenthetical enumeration but every clause serves differentiation or elicitation; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey return shape — and it does, naming the sentence, status, latest plan with clause-level detail, referenced flow/people names, and recent runs. Combined with annotations and a fully-described 2-param schema, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both taskId and the workspace/auth-override semantics are already documented. The description adds only the provenance hint for taskId (from workbench_tasks_list) and says nothing about the workspace parameter; baseline 3 fits when the schema carries the weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('One Task in full') and enumerates what comes back: sentence, status, latest saved plan/clauses, flow and people names, recent runs. It contrasts clearly with siblings workbench_tasks_list (source of ids) and workbench_tasks_compile (edit path), so an agent can distinguish it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use: 'Read this before editing a task's sentence (workbench_tasks_compile with taskId)' and 'when the user asks what a task does.' It also names the prerequisite source for the required id ('Get the taskId from workbench_tasks_list'), leaving no routing inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_listList Workbench TasksA
Read-only
Inspect

Lists the workspace's Tasks — sentence-orchestrated processes — with their sentence, status, and recent runs (id, status, who a waiting run is on). Use this to find a task the user names ('claims intake') and to see what's in flight; workbench_tasks_get reads one in full.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, closed-world), so the bar is lower. The description adds genuine behavioral context by listing the returned fields (sentence, status, recent runs) and notably the 'who a waiting run is on' detail. It stops short of disclosing pagination, limits, or volume behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and return contents before the routing guidance. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with one documented parameter, the description is nearly complete: it covers contents, use case, and sibling routing. Without an output schema it reasonably sketches returned fields, though it doesn't address result volume or pagination.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single workspace parameter is already fully documented in the schema, including default-token and API-key behavior. The description adds nothing about the parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lists) and resource (the workspace's Tasks), and defines what a Task is (sentence-orchestrated processes) rather than assuming the reader knows. It distinguishes itself from the sibling workbench_tasks_get, which reads one task in full.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions: finding a task the user names (with a concrete example, 'claims intake') and seeing what's in flight. It names the alternative (workbench_tasks_get) and the condition that selects it, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_run_viewView a Task run (the lit sentence)A
Read-only
Inspect

The run's current state rendered as its sentence: what ran (with timings), which path each gate took, what's waiting on whom and for how long, what's coming up. Use whenever the user asks where something stands, and after acting on a gate.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesTask run id.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/destructiveHint=false, so the safety profile is covered. The description adds real value by disclosing the shape of the rendered result (what ran with timings, which path each gate took, what is waiting and on whom), which goes beyond the annotation set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with what the tool returns and followed by the usage trigger; no filler. The metaphor "rendered as its sentence / the lit sentence" is slightly opaque jargon but does not bloat the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the burden of explaining return content, and it does so via the enumerated categories. Combined with 100% param coverage and read-only annotations, an agent has enough to call it correctly, though the unaddressed sibling overlap is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the runId and workspace parameters are fully documented in the schema; the description adds no syntax, format, or semantic detail. Baseline 3 applies when structured data does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource (view a Task run) and enumerates the content categories it surfaces: timings, gate paths, blockers, upcoming steps. However, it never distinguishes itself from the similarly named sibling workbench_tasks_get, leaving the agent to infer which read tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use whenever the user asks where something stands, and after acting on a gate" gives a clear triggering condition tied to a concrete workflow moment. It stops short of naming an alternative or exclusion (e.g., when to use workbench_tasks_get instead).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_startStart a Task runAInspect

Kicks off a run of a Task. When the task expects an input document (its trigger has an input), pass the user's material as input — pasted text, extracted attachment text, or JSON. Cite the run id back and tell the user what happens next (the run may immediately be waiting on someone).

ParametersJSON Schema
NameRequiredDescriptionDefault
inputNoTrigger payload — the document text / JSON the task starts from.
taskIdYesTask id (from workbench_tasks_list).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=true, so the mutation/side-effect profile is covered. The description adds the useful behavioral note that the run 'may immediately be waiting on someone', but omits the approval/needs_confirmation flow and any auth or rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with the purpose first and no filler. The trailing instruction to cite the run id and describe next steps is slightly prescriptive but still earns its place as post-call guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the primary input semantics and hints at the async run behavior. The approvalId and workspace parameters are handled by the 100%-covered schema, so nothing critical is missing, though the confirmation flow could be clearer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real value on the `input` parameter: it explains WHEN to supply it (task trigger has an input) and what forms it takes (pasted text, extracted attachment text, JSON) beyond the schema's terse 'Trigger payload' text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('kicks off a run') and resource ('a Task'), so an agent knows this initiates a run. It does not explicitly distinguish itself from nearby siblings like workbench_tasks_act or workbench_tasks_run_view, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives conditional guidance for the input (pass material when the trigger has an input), which is useful. However, it never states when to use this over alternatives such as workbench_tasks_act or workbench_tasks_run_view, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbench_tasks_updateRename a Workbench TaskA
Destructive
Inspect

Changes a task's name — the gallery label only. The sentence and plan are versioned and change through workbench_tasks_compile + workbench_tasks_freeze, never here. May return needs_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesTask id, from workbench_tasks_list.
taskNameYesNew name.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=true). The description adds genuinely non-structured behavior: the two-step needs_confirmation/approvalId flow and which fields are versioned. It does not explain reversibility or the confirmation trigger condition, but it goes well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: scope, exclusions with named alternatives, and an edge-case return. The most important constraint is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity rename with a rich schema and no output schema, the definition is nearly sufficient — it covers the special needs_confirmation return and the routing to other tools. Minor gap: no explicit statement that the rename itself is applied without confirmation or how long the needs_confirmation state persists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter (including approvalId's 'omit on the first call' and workspace's token rules) documented inline, so the baseline is 3. The description only alludes to approvalId via 'needs_confirmation' and adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Changes a task's name') and immediately bounds its scope ('the gallery label only'). It also names the two siblings that handle the other fields (workbench_tasks_compile, workbench_tasks_freeze), so the agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when NOT to use this tool — sentence and plan changes go through compile/freeze, 'never here'. It stops short of spelling out the positive-use conditions or the workspace/approval prerequisites, but the routing guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 55 tool updates
    • First observedcomments_create
    • First observedcomments_list
    • First observedcomments_resolve
    • First observedentity_tags_browse
    • First observedentity_tags_get
    • First observedentity_tags_set
    • First observedget_doc
    • First observedlist_docs
    • First observedsearch_docs
    • First observedsearch_workspace
    • First observedworkbench_executions_get
    • First observedworkbench_flow_authoring_guide
    • First observedworkbench_flows_delete
    • First observedworkbench_flows_edit_text
    • First observedworkbench_flows_fork
    • First observedworkbench_flows_get
    • First observedworkbench_flows_list
    • First observedworkbench_flows_publish
    • First observedworkbench_flows_revisions_list
    • First observedworkbench_flows_run
    • First observedworkbench_flows_scaffold
    • First observedworkbench_flows_schedule
    • First observedworkbench_flows_schedules_delete
    • First observedworkbench_flows_schedules_list
    • First observedworkbench_flows_schedules_run_now
    • First observedworkbench_flows_schedules_update
    • First observedworkbench_flows_share_create
    • First observedworkbench_flows_shares_list
    • First observedworkbench_flows_shares_revoke
    • First observedworkbench_flows_update
    • First observedworkbench_kb_add_documents
    • First observedworkbench_kb_create
    • First observedworkbench_kb_delete
    • First observedworkbench_kb_document_delete
    • First observedworkbench_kb_documents_list
    • First observedworkbench_kb_get
    • First observedworkbench_kb_ingestion_status
    • First observedworkbench_kb_list
    • First observedworkbench_kb_search
    • First observedworkbench_node_catalog_get
    • First observedworkbench_shim_add_answer
    • First observedworkbench_shim_add_examples
    • First observedworkbench_shim_create
    • First observedworkbench_shim_get
    • First observedworkbench_shim_list
    • First observedworkbench_tasks_act
    • First observedworkbench_tasks_compile
    • First observedworkbench_tasks_create
    • First observedworkbench_tasks_delete
    • First observedworkbench_tasks_freeze
    • First observedworkbench_tasks_get
    • First observedworkbench_tasks_list
    • First observedworkbench_tasks_run_view
    • First observedworkbench_tasks_start
    • First observedworkbench_tasks_update

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables managing orchestration workflows through tools for creating, preflighting, linting, and running sequences, checking instance status and outputs, sending signals, retrying instances, and inspecting DLQ and usage.
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Enables AI agents to inspect saved flows and their input schemas, start and monitor runs, and manage agents, sessions, skills, Brain knowledge, artifacts and account administration through 92 tools. Operations are confirmation-gated and use explicitly selected private credential profiles for account and team access.
    92
    463 npm
    AGPL 3.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources