Skip to main content
Glama

Server Details

Coding agents are fast, but they forget the plan when a session ends. Pizza Developer keeps it for them.

You write down Ideas. Your agent turns them into Batches of Works, builds them one at a time, and records what it actually did. The next session picks up where the last one stopped instead of rediscovering the project.

You stay in charge of how the work happens. Instructions tell the agent how you want things built. Workflows set the route each kind of request takes. Rules refuse a write that breaks them. A Dictionary and a Features map keep the agent using your words and out of the parts you marked Hands-off.

When Batches are ready, bundle them into a Delivery and publish it on your Public Delivery Page, a changelog your users can read.

Ownership verified
Status
Healthy
OAuth
Works in Glama
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

B3.4/5.0

Scored across 77 tools

Disambiguation4/5

Tools target distinct resources and lifecycle stages, and detailed descriptions clarify most overlaps. Minor ambiguity remains between update_delivery and assign_batch_to_delivery for managing delivery batches, and between create_batch_plan and start_workflow for planning batches.

Naming Consistency5/5

All names use snake_case with a consistent verb_noun pattern (create_*, get_*, list_*, update_*, archive_*, delete_*). No camelCase or mixed conventions; deviations like reorder are still clear verbs.

Tool Count1/5

77 tools is far beyond the typical 3–15 range and exceeds even the 50+ extreme-mismatch threshold. While the domain is broad, the surface is unwieldy for an agent to navigate.

Completeness4/5

CRUD/lifecycle coverage is extensive across projects, ideas, batches, works, deliveries, workflows, rules, dictionary, and features. Minor gaps include no tool-level get/list for instructions (read via resources) and no image upload for public pages.

Available Tools

77 tools
archive_deliveryArchive deliveryA
Idempotent
Inspect

Archive a delivery. It keeps its record and batches but leaves the Deliveries page, the Public Delivery Page, and the showcase, and remains readable through list_deliveries with archived=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and the description is consistent with all three while adding real side-effect detail the annotations cannot convey: which surfaces the delivery disappears from and which data survives. The agent knows exactly what changes and what remains accessible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the action first, then the consequences. No filler, no repetition of the title, and the most decision-relevant fact (what is preserved) is placed where an agent will read it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, non-destructive mutation with no output schema, the description covers the action, the visible effects, and the retrieval path. It stops short of noting reversibility via unarchive_delivery or any permission requirements, which would have made it fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter ('id') with 100% schema description coverage, so the schema already carries the parameter documentation. The description adds no format, sourcing, or lookup guidance for the id, making this baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Archive a delivery') and immediately scopes it: the record and batches are kept, but the item leaves the Deliveries page, Public Delivery Page, and showcase. That scope statement is what separates it from the adjacent delete_delivery and unarchive_delivery siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when this is the right operation (soft removal that preserves records and batches) and tells the agent how to still read the item via list_deliveries with archived=true. It never explicitly names delete_delivery as the destructive alternative or unarchive_delivery as the undo path, so routing guidance is strong but not fully spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_dictionary_termArchive dictionary termA
Idempotent
Inspect

Archive a dictionary term. It leaves the active dictionary but keeps its content and remains readable through list_dictionary_terms with archived=true. Only archived terms can be deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Dictionary term.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false. The description adds real value beyond that: the term is removed from the active dictionary but its content is preserved and it remains discoverable through list_dictionary_terms with archived=true, and it becomes deletable. It does not spell out whether the operation is reversible (the sibling unarchive_dictionary_term implies it, but the description doesn't say so).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then the state consequence, then the downstream constraint. Nothing is redundant and every sentence adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, non-destructive, idempotent state change with no output schema, the description covers what an agent needs: effect on visibility, content retention, discoverability, and the delete prerequisite. Only the reverse operation (unarchive) is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, fully documented in the schema (100% coverage). The description adds no syntax or format detail for 'id'. Baseline 3 is appropriate when the schema already carries the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Archive a dictionary term') and immediately differentiates it from the sibling delete_dictionary_term by explaining the term survives archival. An agent can distinguish archive / unarchive / delete / update at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: the term leaves the active dictionary but stays readable via list_dictionary_terms with archived=true, and notes that only archived terms can be deleted (an implicit prerequisite for delete_dictionary_term). It stops short of an explicit when-not-to-use statement (e.g. 'use delete_dictionary_term to remove permanently').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_ideaArchive ideaA
Idempotent
Inspect

Archive a completed idea. It leaves the active planning order but keeps its content and remains available through list_ideas with state archived.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Idea.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false, but the description adds real value beyond them: it explains the effect (leaves the active planning order), the preservation guarantee (keeps its content), and the resulting state ('state archived' visible via list_ideas). It does not mention reversibility via restore_idea or any permission requirements, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste, front-loaded with the action and then the post-condition. Every clause earns its place by describing a distinct behavioral consequence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, no-output-schema mutation with annotations covering safety, the description supplies the essential state-transition and retrieval context. It stops short of naming the reverse operation (restore_idea), a minor gap for an agent deciding between archive and delete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'id' parameter, so the schema already documents it fully. The description adds no format, sourcing, or ID-resolution guidance beyond that, which is the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Archive) and resource (idea) and immediately scopes it to completed ideas. The clause about leaving the active planning order while keeping content distinguishes it from the destructive delete_idea sibling, so an agent can tell them apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies the condition for use ('a completed idea') but names no alternative explicitly, so the agent must infer that restore_idea reverses it and delete_idea removes it permanently. The when-to-use is present; the when-not-to and alternatives are not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_instructionArchive instructionA
Idempotent
Inspect

Archive an instruction document. It leaves the active list and the listed instruction resources but keeps its content and revision history. Only archived instructions can be deleted. During an active workflow run, instruction changes are accepted only in its self-improvement stage.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Instruction.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations establish the mutation/safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false), and the description adds genuinely new behavior: what leaves the active list, that content and revision history are retained, that it is a precondition for deletion, and a workflow-stage restriction on changes. That is substantive context beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and effect, followed by the retention guarantee and the workflow-stage caveat. No filler and no repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still tells the agent what the tool does, what state changes result, what is preserved, and when the operation is permitted. Nothing needed to invoke it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage ('Id of the Instruction.'), so the schema fully documents it. The description adds nothing about the identifier's format or source, which yields the baseline of 3 for a well-covered single-param schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Archive an instruction document') and immediately clarifies the exact effect: it leaves the active list and instruction resources while keeping content and revision history. This also implicitly distinguishes it from delete_instruction, which requires prior archiving.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a real usage constraint — during an active workflow run, instruction changes are accepted only in the self-improvement stage — plus a routing hint that only archived instructions can be deleted. It does not explicitly name update_instruction or delete_instruction as alternatives, so guidance is strong but not fully exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_projectArchive projectAInspect

Archive a project. Hides it from the default dashboard view but keeps it reachable (list_projects and search still find it).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Project.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, i.e. a non-destructive mutation. The description adds real behavioral content beyond that: the project disappears from the default dashboard yet remains discoverable through list_projects and search. It does not state reversibility or idempotency, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler, and the core action plus its visibility consequence are front-loaded. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, no-output-schema mutation with annotations already covering the safety profile, the description supplies the key effect (hidden from default views, still searchable). The main omission is that it never points to unarchive_project as the undo path, which would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single 'id' parameter, so the schema already documents it fully. The description adds nothing about the parameter's format or provenance, which is the expected baseline of 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Archive a project') plus a clarifying statement of what archiving does: hides from the default dashboard but keeps the project reachable via list_projects and search. This implicitly separates it from delete_project, though no sibling is named outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the reader can infer archive is preferred over delete when the project should be retained. There is no explicit when-to-use/when-not guidance, and the counterpart unarchive_project is never mentioned as the reversal path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_workflowArchive workflowAInspect

Archive a workflow the vibe coder no longer uses. It stops being offered for new runs; an open run of it continues and its history stays.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesId of the Project.
workflowIdYesId of the Workflow.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, and the description usefully elaborates on that non-destructive profile: new runs stop, in-flight runs continue, history is preserved. It doesn't cover reversibility (unarchive_workflow exists) or auth/permission needs, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action. The second sentence carries only non-redundant behavioral detail about runs and history.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-param archival tool with full schema coverage and no output schema, this covers the essentials plus the key non-destructive behavior. The main missing piece is routing guidance versus delete/cancel/unarchive siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both projectId and workflowId documented in the schema, so the baseline is 3. The description adds no syntax, format, or scoping detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (archive) and resource (workflow) and frames the intent (no longer used). It does not explicitly distinguish itself from close siblings like delete_workflow, cancel_workflow, or unarchive_workflow, so the agent must infer the boundary from naming alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies when to use it via 'no longer uses', but never states when-not-to-use or names an alternative such as delete_workflow. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assign_batch_to_deliveryAssign batch to deliveryA
Destructive
Inspect

Put a ready batch in a delivery, or take it out. Pass deliveryId to assign, omit to unassign. Only a ready batch no delivery carries can be assigned. Take it out of its current delivery first. At a workflow run's Delivery stage, assigning every ready Batch of the run to one existing Delivery completes the stage instead of create_delivery. In a run of a workflow that starts at Delivery, adding any ready Batch no delivery carries to an existing Delivery completes the stage.

ParametersJSON Schema
NameRequiredDescriptionDefault
batchIdYesReady Batch to assign or unassign.
metricsNoOptional client measurements, stored on the workflow step this call completes.
deliveryIdNoDelivery to assign to; omit to unassign.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the mutation risk is flagged. The description adds genuine behavioral context beyond that: assignment preconditions and the fact that assigning all ready batches completes a workflow Delivery stage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and toggle, then preconditions, then workflow-stage nuance. The two workflow-stage sentences are long but carry real semantic weight rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description covers operation mode, preconditions, and workflow-stage side effects. The metrics parameter's storage behavior is left to the schema, which documents it, so the definition is essentially complete for a mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents batchId, deliveryId, and the metrics object. The description's 'Pass deliveryId to assign, omit to unassign' largely restates the schema's own deliveryId description, adding little new parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Put a ready batch in a delivery, or take it out') and covers both directions of the operation (assign/unassign). It clearly distinguishes itself from the sibling create_delivery by describing the stage-completion alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative (create_delivery) and the condition selecting it, states the toggle behavior for deliveryId, and lays out preconditions ('Only a ready batch no delivery carries can be assigned. Take it out of its current delivery first'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

begin_workBegin workAInspect

Begin repository implementation in one boundary call. Pass workId to claim that Work, or projectId to resume the active Work or claim the next ordered one. Each open workflow run builds its own Batch; pass runId to work in a specific run when several are open. If nothing is developing there, the next planned batch is started automatically when it contains implementable work. Returns the Work brief and the project's own instruction documents that apply to it.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdNoWorkflow run to work in, when several are open.
workIdNoWork to claim.
projectIdNoProject whose active or next Work to begin.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only reveal write/non-destructive; the description adds meaningful behavior beyond that, disclosing that it claims a Work, auto-starts the next planned batch when implementable work exists, and returns Work brief plus applicable instruction documents. No mention of auth, failure modes, or side effects on concurrent runs, but the automatic batch-start disclosure is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, purpose front-loaded and followed by parameter routing and side-effect behavior. Every sentence carries information; no padding, though it front-loads the boundary-call framing before the more actionable parameter guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully covers return values (Work brief and instruction documents) plus the auto-start behavior and the zero-required-parameter model. Adequate for a mutation tool, with only the sibling relationship to start_workflow/resume_workflow left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds relational meaning the schema lacks: workId actively claims, projectId carries the resume-or-claim-next semantics, and runId selects among concurrent runs. This goes beyond the terse per-parameter schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Begin repository implementation') and clarifies the consolidated 'one boundary call' scope. An agent can distinguish it from the paired finish_work, though the description never explicitly names or contrasts with that sibling or start_workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit conditional guidance for selecting between the argument modes: pass workId to claim, projectId to resume/claim-next, runId when several runs are open. That is unusually clear routing logic, but it stops short of telling the agent when to prefer begin_work over start_workflow or resume_workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_workflowCancel workflow runA
Destructive
Inspect

Cancel an open workflow run without deleting its recorded route or measurements. A truthful reason is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesId of the Workflow run.
reasonYesTruthful reason for cancelling the run.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=true and readOnlyHint=false, which do not reveal that the run's recorded route and measurements survive the operation. That preservation guarantee is real added context not available in structured fields. It leaves out reversibility and post-cancel state, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the key non-obvious fact (data is preserved) front-loaded ahead of the requirement reminder. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema, the description covers the essential outcome difference from deletion. It omits what state the run ends in, whether the action is reversible, and any error/failure conditions, which would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both runId and reason documented in the schema itself, so the baseline is 3. The description's 'a truthful reason is required' restates the schema's own wording rather than adding format or constraint detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (cancel) and resource (workflow run), and adds a scoping clause ('without deleting its recorded route or measurements') that behaviorally separates it from delete_workflow. It does not name that sibling explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'open workflow run' implies a precondition (only open runs are cancellable), which is useful implied guidance. However, there is no explicit when-to-use guidance, no mention of resume_workflow as a related/opposite action, and no statement about what to do for already-completed runs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_rulesCheck rulesA
Read-only
Inspect

Check a draft against the rulesets on the current step of an open workflow run, without saving it and without refusing. Rules run only inside workflow runs: runId names the run, and record.kind picks idea, batch, work, or delivery, which must be the kind of write that step's connection checks. Pass a record's id to start from the saved record; any fields you pass override the saved values. Batch rules check the plan, so a saved Batch can be rechecked only until one of its Works is done. Returns passed and failed rules, each with its probability (a rule passes at 0.75 or above), any rules that could not be checked with the reason, and how many rules were sent to TypeSafe (sentToTypeSafe) or answered from saved verdicts (fromSavedVerdicts). Unchanged records reuse earlier verdicts.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesId of the Workflow run.
recordYesDraft to check; kind is idea, batch, work, or delivery.
projectIdYesId of the Project.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint/openWorldHint; the description adds substantial behavioral context: it is a dry-run that never refuses, the 0.75 pass threshold, that unchecked rules report a reason, the sentToTypeSafe/fromSavedVerdicts counters, and that unchanged records reuse prior verdicts. This is exactly the value-add expected above annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, and each sentence carries information (preservation of drafts, kind constraint, id-override behavior, return shape). It is dense and slightly long, but almost nothing is filler given the absent output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema tool it thoroughly explains the return values (passed/failed with probability, unchecked with reason, sentToTypeSafe/fromSavedVerdicts) and the key preconditions, so an agent has enough to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds genuine meaning: passing record.id starts from the saved record and any supplied fields override saved values, plus the kind constraint tied to the step's connection. That override/merge semantics isn't in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope: checking a draft against rulesets on a step of a workflow run, explicitly "without saving it and without refusing." This clearly distinguishes it from the mutation siblings (update_idea, update_work) and from rule-management siblings (create_rule, update_rule).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real preconditions: rules only run inside workflow runs, runId names the run, record.kind must match the kind of write the step's connection checks, and a saved Batch can only be rechecked until one of its Works is done. It doesn't explicitly name an alternative tool for validating-then-saving, but the when-to-use context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_batch_planPlan batchAInspect

Create a planned batch and its planned works in one call. This is the normal way to plan a batch once its scope is agreed. The batch name is 1 to 3 lowercase words joined by single hyphens, like dark-mode-toggle. The batch takes the next sequential number and queue position; works start planned in the order given, and the result carries each work's batch-scoped reference (like 4-1), ready for commit messages. create_work adds a Work to an existing Batch. When several workflow runs wait for a Batch, pass runId to plan it for that run. In a workflow run, this completes the planning step (and an open Evaluate-the-idea step), and edgeId can name the Batch stage route so no separate resume_workflow call is needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesBatch name: 1 to 3 lowercase words joined by single hyphens.
runIdNoWorkflow run this plan belongs to.
worksYesPlanned Works in order.
edgeIdNoId of the Workflow connection to take.
metricsNoOptional client measurements, stored on the workflow step this call completes.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (readOnlyHint=false, destructiveHint=false): it discloses that the batch takes the next sequential number and queue position, that works start planned in the given order, what the result contains (batch-scoped references like `4-1`), and that in a workflow run it completes the planning step and any open Evaluate-the-idea step. These are real side effects an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and keeps to about five sentences, each carrying distinct information (naming convention, numbering/ordering, sibling contrast, runId, workflow completion). Some sentences are long and pack multiple clauses, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers naming constraints, ordering, numbering side effects, the shape of the returned references, and workflow integration. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: the ordering of `works` matters ('works start planned in the order given'), the returned reference format ties to `works`, and runId/edgeId are explained in workflow terms. It does not add much for `metrics` or `requestId`, which the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (create a planned batch and its planned works) and explicitly scopes it as a single combined call. It also distinguishes itself from the sibling create_work ('create_work adds a Work to an existing Batch'), so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Says this is 'the normal way to plan a batch once its scope is agreed', names the alternative create_work, and gives the condition for runId ('When several workflow runs wait for a Batch, pass runId'). The workflow-run context (completes the planning step, edgeId avoids a separate resume_workflow call) is explicit when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_deliveryCreate deliveryAInspect

Create a shipped public delivery from selected undelivered batches: batches no other delivery carries, including ready Batches of several finished workflow runs. Supply its version label (free-form; semver like 2.4.0 is common but not required) and a factual summary. To record a release that shipped before the project used Pizza Developer, omit batchIds and pass shippedAt, the real past ship date (ISO date or date-time); without shippedAt it ships now. A delivery without batches never completes a workflow run's Delivery stage. In a run of a workflow that starts at Delivery, any ready Batches no delivery carries can go in, and the call completes the run's Delivery stage. When the project has a workflow that starts at Delivery, create_delivery outside a workflow run is refused and the refusal names that workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
metricsNoOptional client measurements, stored on the workflow step this call completes.
summaryYesFactual summary of what shipped.
versionYesVersion label, such as 2.4.0.
batchIdsNoUndelivered Batches to include.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.
shippedAtNoShip date, as an ISO date or date-time.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=false), the description discloses real behavioral nuance: the refusal case names the offending workflow, a batchless delivery never completes a Delivery stage, and omitting shippedAt ships immediately. It stops short of describing the returned record or what happens to the consumed batches, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most important scoping and version-label facts are front-loaded, and the text is dense with no filler. However, the workflow-stage sentences are convoluted and the two clauses about 'a run of a workflow that starts at Delivery' repeat the same idea, costing some readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the key edge cases an agent must handle: historical backfill, immediate ship, the batchless case, and workflow-run refusal. It does not describe the response shape or how to recover the created delivery's id, which is the one notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all seven parameters, but the description still adds meaning: version is free-form and semver is optional, shippedAt is the 'real past ship date', and batchIds must be batches no other delivery carries. These are semantic constraints the schema strings do not fully convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb and resource ('Create a shipped public delivery') and constrains the source material ('from selected undelivered batches'), which cleanly separates it from assign_batch_to_delivery, update_delivery, and archive_delivery. An agent can identify what this tool produces without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit conditional routing: omit batchIds and pass shippedAt to backfill a pre-Pizza-Developer release, omit shippedAt to ship now, and it states that create_delivery outside a workflow run is refused when the project has a workflow starting at Delivery. When-to-use, when-not-to-use, and the workflow-run path are all spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_dictionary_termCreate dictionary termBInspect

Create a project-scoped dictionary term: a word, its aliases, and its definition.

ParametersJSON Schema
NameRequiredDescriptionDefault
wordYesThe Dictionary word.
aliasesNoOther words that mean the same.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.
definitionYesThe approved meaning of the word.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the write/non-destructive profile is known. The description adds only the 'project-scoped' context; it does not mention idempotency/duplicate handling (which is buried in the schema's requestId) or required permissions, so it adds limited value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with zero wasted words and the key noun (dictionary term) up front. Efficient, though arguably too terse to be maximally useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 100% schema coverage, no output schema, and annotation-covered safety, an agent has enough to invoke this correctly. Idempotency is explained in the schema's requestId, so the missing usage/prerequisite context is the only real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents word, aliases, definition, projectId, and requestId. The description merely echoes the word/aliases/definition fields and adds no syntax, format, or constraint detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Create') and resource ('dictionary term') and adds the scope ('project-scoped') plus the fields it accepts (word, aliases, definition). This clearly separates it from update/get/archive/delete_dictionary_term siblings, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus update_dictionary_term (e.g., for existing terms) or what prerequisites/permissions are needed. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_ideaCreate ideaAInspect

Create a project idea with a required short name of at most 20 characters. Summary and content are optional and are never generated. Impact is semantic product scope, not priority. A new idea joins at the back of the order; move it forward with reorder. A name matching a declined idea is refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesShort name, at most 20 characters.
impactNoProduct scope of the Idea: minor or major.
contentNoFull detail.
summaryNoShort summary.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false, so the description's extra facts are genuinely additive: summary/content are never auto-generated, a new idea is appended at the back of the order, and a name colliding with a declined idea is refused. That last point is a real precondition an agent would otherwise discover only by failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the creation action and its hard constraint, then constraints and the ordering rule. No filler and every sentence carries a distinct fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter write tool with no output schema, it covers the creation constraints, ordering behavior, and the declined-name refusal well. The one remaining gap is that nothing states what a successful call returns (the created idea object), which matters since no output schema exists to fall back on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by clarifying that 'impact' is semantic product scope and explicitly not priority, and by restating that summary/content are optional and never generated. That disambiguation is not present in the field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Create a project idea') and immediately bounds it with the required short-name constraint, so an agent knows exactly what is produced. It does not, however, explicitly differentiate itself from sibling creation or lifecycle tools such as update_idea, decline_idea, or restore_idea.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete routing hint – new ideas land at the back of the order and 'move it forward with reorder' – which points at the follow-up tool. There is no guidance on when to create an idea versus updating or restoring one, so usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_instructionCreate instructionAInspect

Create an instruction document on a project, freely titled (there is no fixed set), with its full text when known. Titles are unique within the project. During an active workflow run, instruction changes are accepted only in its self-improvement stage. Returns its id for update_instruction.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesInstruction title, unique in the Project.
contentNoFull Instruction text.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false, so the description carries most of the burden and does well: it discloses the project-level title uniqueness constraint, the workflow-stage gating rule, and that an id is returned. It omits failure behavior on duplicate titles, but the coverage is above average.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: creation, then title semantics, then workflow gating, then return value. Slightly dense/run-on in the third and fourth clauses, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with no output schema, the description covers the key facts an agent needs: uniqueness of titles, when writes are allowed during a run, and what is returned. Missing only error/duplicate-handling detail, which is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including requestId idempotency. The description only mildly extends this by saying titles are free-form and content is supplied 'when known' (i.e. optional), which is the baseline level of added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create an instruction document on a project') and clarifies the free-form title space, which distinguishes it from the many create_* siblings. It also points to update_instruction as the follow-up path, giving partial sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Adds a real usage constraint: instruction changes during an active workflow run are accepted only in its self-improvement stage. It also implies the create-then-update flow by noting the returned id feeds update_instruction. It does not, however, state exclusions against archive/delete siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_projectCreate projectAInspect

Create a project from its details, with any Ideas the vibe coder supplied. Only the name is required; Ideas, Batches, Works, Deliveries, a Workflow, a repository, and a live URL can all be added later. Supplied Ideas keep their order, and up to the first two become planning input. Everything is validated before anything is written.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesProject name.
tagsNoProject tags.
ideasNoIdeas to create with the Project, in order.
liveUrlNoLive URL of the Project.
repoUrlNoRepository URL of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.
descriptionNoWhat the Project is.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-destructive write. The description adds genuinely new behavioral context: validation happens before any write (atomicity), and supplied Ideas preserve their order with the first two feeding planning input. These are meaningful traits not present in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and immediately followed by the required-vs-optional constraint and validation guarantee. No filler, though the Ideas-ordering detail is somewhat dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter create tool with no output schema and only basic annotations, the description covers validation, optionality, nested-idea ordering and planning semantics. It does not describe what the call returns (e.g., the created project identifier), which is a minor gap given no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds semantics the schema cannot: that Ideas retain order and that only the first two become planning input. It also reinforces that everything but name is optional, going beyond the raw field docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Create a project from its details") and clarifies that supplied Ideas are attached to the project, which distinguishes it from the sibling create_idea. It does not explicitly name alternatives, but the scope is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by noting only the name is required and that Ideas, Batches, Works, Deliveries, workflows, repo and live URL "can all be added later," which helps an agent choose a minimal create. However, it never states when to create a project here versus creating its child entities via their own tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_ruleCreate ruleAInspect

Add a rule to the end of a ruleset: one narrow statement and level must (default) or should. Rules run only inside workflow runs, where the judge is given the connection's Dictionary and Features context and the run's Idea automatically.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe Rule as one narrow statement.
levelNomust refuses a failing write; should only reports it.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.
rulebookIdYesId of the Ruleset.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the safety profile is covered. The description adds genuine behavioral context beyond that: rules execute only inside workflow runs, and the judge receives the connection's Dictionary and Features plus the run's Idea automatically. It omits auth/permission needs, but with annotations present the bar is lower.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and placement, then the execution context. Dense but no wasted sentences; the second sentence is the only arguably tangential part and it earns its place as behavioral context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-destructive create tool with annotations covering safety, a fully documented schema, and no output schema, the description covers purpose, level semantics, and execution context adequately. Missing only return/response behavior and error cases, which are minor given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema documents all four parameters. The description nevertheless adds value by disclosing that 'must' is the default level, which the schema's enum description does not state, and by reinforcing the one-narrow-statement constraint on text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add a rule to the end of a ruleset'), and the placement detail ('to the end') plus the level concept make it distinguishable from create_rulebook/update_rule. It does not explicitly name a sibling or contrast with one, so it falls short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains where rules apply ('Rules run only inside workflow runs') and the judge context, but gives no explicit when-to-use-this-tool-vs-alternatives guidance or prerequisites for calling create_rule itself. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_rulebookCreate rulesetAInspect

Create a ruleset with a name and optionally its rules. Write each rule as one narrow statement the record either meets or not. Rules run only inside workflow runs: a ruleset is checked when it is attached to a workflow connection (rulebookIds in update_workflow) and the write that completes that step happens.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesRuleset name.
rulesNoRules to create with the Ruleset, in order.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnlyHint=false, destructiveHint=false); the description goes further by disclosing that rules are evaluated only inside workflow runs and only against the write that completes an attached step. It still omits permission requirements and name-uniqueness behavior, but adds real behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, with no filler. The rule-writing sentence partially restates schema guidance, but it is short and framed as authoring advice rather than repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with no output schema and full schema coverage, the description supplies the missing operational context: where rulesets run and what triggers a check. Return values and auth requirements are not covered, but nothing critical to invoking it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters including the must/should level semantics and the requestId idempotency retry. The description reinforces the intended shape of a rule ('one narrow statement the record either meets or not') but adds no syntax or format detail the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a ruleset with a name and optionally its rules'), and clarifies the scope of the nested rules. It does not explicitly distinguish itself from the sibling create_rule, so an agent must infer the boundary between creating a ruleset and creating an individual rule.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the operational context ('Rules run only inside workflow runs') and how a ruleset becomes active (attached via rulebookIds in update_workflow), which implies when it matters. It never states when to prefer create_rulebook over create_rule or list_rulebooks, so the routing guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_workCreate workAInspect

Create a Work in a Batch that is not ready. It always starts planned; begin_work starts it. In a workflow run, the Batch plan with the new Work is checked against the rulesets on the connection the run entered Batch on, until one of its Works is done; a Must rule that does not pass refuses the write.

ParametersJSON Schema
NameRequiredDescriptionDefault
sizeYesWork size.
titleYesWork title.
labelsNoLabels on the Work.
batchIdYesId of the Batch.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.
descriptionNoWhat changes and where.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare it is a non-destructive write; the description goes further by disclosing the initial state (always planned), the ruleset validation that occurs during a workflow run, and the refusal semantics when a Must rule fails. It does not cover permissions or return behavior, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core action and precondition before the ruleset nuance. No filler, though the ruleset clause is long and would benefit from slight tightening.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with no output schema, the description covers precondition, initial state, and validation behavior. It leaves the return payload unspecified, but the requestId schema field partly compensates on the idempotency side, so the gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the requestId idempotency contract and the size enum, so the schema already carries parameter meaning. The description adds no parameter-level detail (e.g., which project/batch pairing is valid), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a Work in a Batch') plus the precondition that the Batch must not be ready, and explicitly contrasts the initial state with begin_work. An agent can distinguish it from begin_work, create_batch_plan, and update_work without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use condition ('a Batch that is not ready') and names the alternative that handles the next step ('begin_work starts it'). It stops short of naming exclusions such as when to prefer create_batch_plan or what to do if the Batch is already ready.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_workflowCreate workflowBInspect

Create a named workflow: a route of stages that requests in the project follow. name is unique in the project; description says which requests it handles. Nodes are Start, optional Idea, Batch, optional Delivery, optional Workflow (the step that improves the Workflow itself after a run), and Finish. Start can connect to an existing Batch without Idea, and Batch can connect to Workflow or Finish without Delivery. Start can also connect to Delivery, for a workflow that only creates or adds to a Delivery: Start to Delivery to Finish, or Start to Delivery to Workflow to Finish, with no Batch node. Each edge.data must contain instructionDocumentIds (existing project document IDs) and contextKinds (features and/or dictionary; each adds only a short outline, top-level feature names or dictionary words, that the agent expands with the read tools). edge.data may also carry rulebookIds, an ordered list of the project's ruleset IDs attached to that connection; leaving it out attaches none. Select documents on connections; do not create Instructions, Dictionary, or Features nodes or put text on lines. Projects work without a Workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesWorkflow name, unique in the Project.
edgesYesConnections between stages, each with its Instructions, context kinds, and Rulesets.
nodesYesWorkflow stages.
projectIdYesId of the Project.
requestIdNoOptional unique ID for this create. After a timeout, retry with the same requestId and arguments: the record already created is returned instead of a duplicate.
descriptionYesWhich requests this Workflow handles.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false, so the description must carry most of the behavioral burden. It does disclose the non-obvious constraint that instruction documents must be attached on edges rather than as nodes ('do not create Instructions, Dictionary, or Features nodes or put text on lines'), plus the uniqueness of name. It doesn't cover permissions, what happens on validation failure, or retry semantics beyond what requestId's schema already says.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One long paragraph that mixes scope, node taxonomy, edge requirements, and constraining rules in no clear order. Critical usage constraints ('Each edge.data must contain instructionDocumentIds and contextKinds') are buried mid-paragraph after an exhaustive node enumeration, and the dependency lists (Start→Delivery→Workflow, no Batch) read like schematic notes rather than actionable structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter creation tool with no output schema and untyped nested arrays, the description covers the domain space reasonably but omits the return value and whether the created workflow is immediately active. It is adequate for building a valid graph but leaves the calling agent without confirmation of what success returns or how to detect invalid combinations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents name, projectId, description, and requestId. The description adds real domain meaning to edges (instructionDocumentIds, contextKinds, rulebookIds ordering and defaults) that the schema's generic 'each with its Instructions, context kinds, and Rulesets' does not. However nodes and edges are untyped arrays in the schema, and the description's node/edge semantics are the only docs for their internal shape, which partially compensates but leaves structure under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a named workflow') and defines the domain object concretely: 'a route of stages that requests in the project follow'. It's distinguishable from sibling create_* tools by its domain vocabulary, though it doesn't name alternatives explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage by describing the node graph model in depth, and the closing line 'Projects work without a Workflow' hints at optionality. But it never says when an agent should reach for this tool versus update_workflow or start_workflow, and there are no explicit exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decline_ideaDecline ideaA
Idempotent
Inspect

Decline an idea the vibe coder decided against, with their reason in their own words. It leaves the planning order and is listed with declined ideas. Declining an already declined idea replaces its reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Idea.
reasonYesThe vibe coder's reason, in their own words.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=false and idempotentHint=true, so safety is covered; the description adds genuine state-change context (removal from the planning order, move to the declined list) and the important non-idempotent side effect that re-declining overwrites the existing reason. It doesn't mention recoverability, though restore_idea exists as a sibling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action front-loaded and the edge-case behavior last. No filler, though the middle sentence about planning order could be folded into the first without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description adequately conveys the resulting state (removed from planning order, listed as declined) plus the overwrite semantics. For a two-parameter mutation with annotations covering the safety profile, this is nearly complete; only the reversibility path via restore_idea is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented in the schema, so the schema does the heavy lifting. The description echoes 'reason in their own words' and implies the reason is replaceable on repeat calls, but adds no syntax, length, or id-format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (decline) and resource (idea), and adds the distinguishing outcome: the idea 'leaves the planning order and is listed with declined ideas,' which separates it from delete_idea and archive_idea. It stops short of naming those siblings explicitly, so the differentiation is inferable rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear trigger ('an idea the vibe coder decided against') and covers the repeat-call case ('declining an already declined idea replaces its reason'), but never routes the agent between this and the near-siblings delete_idea, archive_idea, or restore_idea. Usage is implied rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_batchDelete batchA
Destructive
Inspect

Delete a batch. Batches that still have works can't be deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Batch.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered. The description adds a non-obvious behavioral constraint not present in annotations: deletion is blocked for batches that still contain works, which is meaningful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the key constraint. No filler, no repetition, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with rich annotations and no output schema, the description covers purpose and the main failure precondition. It omits details like whether deletion is permanent or what errors are returned, but those are minor given the annotation coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter (id) and schema description coverage is 100%, so the schema fully documents the input. The description adds nothing about the id parameter, so the baseline of 3 applies when structured data does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Delete a batch'), so the operation is unambiguous, but it largely restates the tool name/title and does not explicitly distinguish it from siblings such as revert_batch or delete_work. The resource name carries the differentiation rather than the prose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives a usable precondition ('Batches that still have works can't be deleted'), which implies when the call will fail. However, there is no guidance on alternatives (revert_batch to undo, archive workflows instead) or on when to prefer another sibling, leaving usage mostly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_deliveryDelete deliveryA
Destructive
Inspect

Delete a delivery. Its batches come out of it unassigned, ready to go into another one; they are not deleted. Deleting a shipped delivery rewrites public history.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered structurally. The description goes well beyond that by disclosing the cascade behavior (batches become unassigned rather than deleted) and the public-history side effect of deleting a shipped delivery.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and followed immediately by the two consequences that matter. No filler; each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive single-parameter tool whose safety hints are already in annotations and which has no output schema, the description covers the essential behavioral surprises. It stops just short of routing to a reversible alternative such as archive_delivery, which would make it fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'id' parameter, so the schema already documents it fully. The description adds no meaning about the identifier (format, source, where to obtain it), which is the baseline expectation when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete a delivery') that is unambiguous on its own. It does not, however, differentiate itself from close siblings like delete_delivery_presentation or archive_delivery, so the agent must infer the distinction from names alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The caveat about deleting a shipped delivery rewrites public history implies caution and therefore shapes when it is appropriate to call. But it never names the alternative (e.g., archive_delivery for reversible removal) or states an explicit when-to-use/when-not-to-use condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_delivery_presentationDelete delivery presentationA
Destructive
Inspect

Delete a delivery's current release presentation, so the Public Delivery Page shows none. It stays restorable from the presentation history in the app.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered. The description adds non-obvious behavioral context: the operation is reversible via presentation history and its effect on the public page. It does not mention required permissions, but it goes well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, no filler. The primary action and its effect are front-loaded, followed by the restore caveat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with annotations covering safety and no output schema, the description supplies the essential context (what it deletes, where it's reflected, that it's restorable). Minor gaps remain around prerequisites, but the core is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With one parameter and 100% schema description coverage, the baseline is already high. The description implies the target is a delivery's presentation, adding mild context over the schema's 'Id of the Delivery'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Delete a delivery's current release presentation') and names the observable effect ('Public Delivery Page shows none'). It is easily distinguished from siblings like delete_delivery or update_delivery_presentation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose implies usage, and the mention of the presentation history hints at alternatives, but the description never explicitly states when to use this tool versus update_delivery_presentation or get_delivery_presentation, nor does it lay out preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_dictionary_termDelete dictionary termA
Destructive
Inspect

Permanently delete an archived dictionary term. Active terms must be archived first.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Dictionary term.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered structurally. The description adds genuinely new behavioral facts beyond that: the deletion is permanent and irreversible, and the term must already be archived. It doesn't cover auth requirements or failure modes, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the permanence warning front-loaded before the precondition. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation with no output schema, the description covers what an agent needs: it's permanent and requires prior archival. Missing only recovery/undo behavior and error conditions, which are minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single 'id' parameter, so the schema fully documents it. The description adds no syntax, format, or sourcing guidance for the id, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete) plus resource (dictionary term) and adds the crucial qualifier 'permanently', which distinguishes it from the sibling archive_dictionary_term and unarchive_dictionary_term. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition: 'Active terms must be archived first', which routes the agent to archive_dictionary_term when the term is still active. It stops short of naming the alternative tool outright, so it's clear context rather than a full when/when-not rule set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_ideaDelete ideaA
Destructive
Inspect

Permanently delete an archived or declined idea. Active ideas must be archived or declined first.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Idea.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds meaningful context beyond that: the operation is 'Permanently' irreversible and gated on a specific idea state. It does not address whether repeat calls on a deleted id error (relevant given idempotentHint=false), a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste, and the destructive scope constraint plus precondition are front-loaded where an agent will read them.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers irreversibility and the state precondition, which is what an agent needs to call it safely. It stops short of noting error behavior for repeat/locked calls, keeping it just shy of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'id' parameter, so the baseline is 3. The description adds no format or example details about the id beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('delete') and resource ('idea') with an explicit scope constraint ('archived or declined'). This differentiates it reasonably from archive_idea and decline_idea, though it never names those siblings as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context ('archived or declined idea') and an explicit prerequisite for the active case ('Active ideas must be archived or declined first'), which implicitly points to the archive/decline siblings. No explicit when-not beyond the state precondition, but the routing guidance is solid.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_instructionDelete instructionA
Destructive
Inspect

Permanently delete an archived instruction document, together with its revision history. Active instructions must be archived first.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Instruction.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is structured data. The description still adds real value: it warns that deletion is permanent and cascades to revision history, and that an archive step is a precondition. It does not describe failure behavior or confirm irreversibility semantics beyond 'permanently'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the destructive action and scope front-loaded and the precondition second. No filler, nothing repeated from the title or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive one-parameter mutation with no output schema, the description covers permanence, cascade scope, and the archive precondition, with annotations carrying the rest. Minor gaps (error cases, whether deletion can be undone) remain but nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'id' parameter, so the schema already carries the parameter meaning. The description adds nothing about the identifier format or what happens with an unknown id, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete) plus resource (instruction document) and an explicit scope: only archived instructions, and the revision history goes with it. This distinguishes it cleanly from archive_instruction and unarchive_instruction without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The prerequisite 'Active instructions must be archived first' gives a concrete condition for when this call is legal, implying archive_instruction as the preceding step. It stops short of naming the alternative tool explicitly or stating when not to use deletion at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_projectDelete projectA
Destructive
Inspect

Delete a project and everything under it (batches and works).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Project.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavioral context beyond that: the delete cascades to child batches and works, which tells the agent what collateral damage to expect before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action and immediately qualified by the cascade scope. Nothing redundant, nothing padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool with no output schema and annotations covering the safety flags, the key missing piece an agent needs — that the deletion is cascading — is present. Only the reversibility/undo question (and the archive alternative) is left unstated, which is a minor gap given the destructiveHint.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'id' parameter is documented as 'Id of the Project.' The description adds nothing about identifier format or source, so the baseline 3 for a fully-documented single-param schema is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Delete') plus resource ('project'), with added scope detail that everything beneath it (batches and works) is removed. It does not differentiate itself from the nearby archive_project/unarchive_project or the other delete_* siblings, so an agent gets a clear purpose but no routing signal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the obvious alternative archive_project for a reversible removal. The sibling set is dense with archive/unarchive/delete variants, yet the description offers no condition for picking this one over them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_ruleDelete ruleB
Destructive
Inspect

Permanently delete one rule.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Rule.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description adds 'Permanently,' reinforcing irreversibility, but says nothing about cascading effects, required permissions, or confirmation requirements, so it goes only modestly beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with zero filler and the key qualifier ('Permanently') front-loaded. It is efficiently sized, though its brevity borders on under-specification for a destructive operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent delete with one required id and no output schema, the description conveys the essential irreversibility but omits cascade behavior, error conditions (e.g. missing id), and any routing versus archive/update siblings. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single 'id' parameter with 100% schema description coverage, so the schema already documents it fully. The description adds no meaning about the id (format, source, or lookup), making this the baseline case where the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('delete') and resource ('rule') with a clear scope qualifier ('one'). It does not differentiate itself from siblings like delete_rulebook or the many archive_* tools, but an agent can identify the operation unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives such as archive or update, and no prerequisites stated. The word 'Permanently' hints that it is irreversible versus an archive, but that inference is left to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_rulebookDelete rulesetA
Destructive
Inspect

Permanently delete a ruleset and its rules, and detach it from every workflow connection.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Ruleset.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the deletion is permanent (no undo) and cascades to child rules and workflow connections. It stops short of stating what happens on a non-existent id or that repeats fail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and resource, and every clause (rules removal, workflow detach) carries information. No waste or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive delete with full schema coverage and no output schema, the description covers the essentials: permanence and cascade effects. Minor gaps remain around behavior for an unknown/already-deleted id (relevant given idempotentHint=false).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (the single 'id' param is documented as 'Id of the Ruleset.'), so the schema carries the parameter burden. The description adds no syntax or format detail beyond that, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (delete) plus resource (ruleset), with cascade scope spelled out: it also removes the ruleset's rules and detaches it from every workflow connection. That scope distinguishes it from the sibling delete_rule and create_rulebook/update_rulebook without needing the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance, no mention of a non-destructive alternative (there is no archive_rulebook, but update_rulebook and delete_rule exist as adjacent choices), and no prerequisites such as required permissions. Only the adverb 'Permanently' hints at context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_workDelete workC
Destructive
Inspect

Delete a work item.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Work.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered by structured data. However, the description adds nothing beyond that: it doesn't disclose irreversibility, cascade effects, required permissions, or what happens on a repeat/cross-referenced call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no waste, but also no substance. It is concise to the point of being under-specified for a destructive operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema and no annotations explaining consequences, the description should say more about irreversibility and side effects. It leaves the agent without the context needed to use a delete safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'id' parameter has 100% schema description coverage ('Id of the Work.'), so the schema carries parameter semantics. The description adds no format or sourcing detail, which is the expected baseline when schema coverage is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource ('Delete a work item'), but it merely restates the tool name and title without distinguishing it from the many other delete_* siblings (delete_batch, delete_idea, delete_project, etc.). An agent knows what it does but not what makes it the right 'delete' tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. Notably, archive_* siblings exist for several resources, and delete future-proof guidance (or a note that deletion is permanent versus archiving) is absent, leaving routing entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_workflowDelete workflowA
Destructive
Inspect

Permanently delete a workflow together with all of its runs, step reports, and revisions. Batches, Deliveries, and tool calls it touched stay. confirmName must equal the workflow's exact name. archive_workflow hides a workflow and keeps its history.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesId of the Project.
workflowIdYesId of the Workflow.
confirmNameYesThe Workflow's exact name, to confirm deletion.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With destructiveHint=true already declared, the description goes well beyond the annotation by enumerating exactly what is destroyed (runs, step reports, revisions) and what survives (batches, deliveries, tool calls). This is the difference between the destructive and non-destructive siblings made concrete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the destructive verb and its consequences, then the confirmation rule, then the sibling alternative. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, irreversible mutation with no output schema, the description covers irreversibility, retained entities, the confirmation requirement, and the escape hatch. Nothing an agent needs before calling it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so projectId/workflowId/confirmName are already documented. The description restates the confirmName invariant ('must equal the workflow's exact name') rather than adding new syntax or format detail, so it earns the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Permanently delete a workflow') and immediately scopes the blast radius. It is clearly distinguishable from archive_workflow, cancel_workflow, and the other delete_* siblings by naming the retained vs destroyed entities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the prerequisite for invocation (confirmName must equal the workflow's exact name) and points at archive_workflow as the non-destructive alternative. It does not spell out the explicit 'use this when...' condition, but the routing intent is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

finish_workFinish workA
Destructive
Inspect

Record completed repository implementation in one boundary call. Replaces the Work description with the factual implementedSummary and marks it done. If this closes the batch, batchSummary closes it in the same call; otherwise the response asks only for that missing summary. Calling it again on a done Work with batchSummary closes a Batch left open. A Work left planned or in progress inside a ready Batch is finished here too. Use it once implementation is complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Work.
metricsNoOptional client measurements, stored on the workflow step this call completes.
batchSummaryNoWhat the Batch changed; needed when this closes the Batch.
implementedSummaryYesWhat was actually implemented; replaces the Work description.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply destructiveHint=true; the description goes well beyond that by disclosing that the Work description is replaced (overwritten), that the call can cascade to close a Batch and even finish other Works inside a ready Batch, and how the response behaves when a summary is missing. This is rich, non-obvious side-effect disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and stays dense, but the middle sentences about batchSummary closing and re-calling on a done Work are somewhat tangled and require re-reading. It is appropriately sized overall with no filler, but structure could be cleaner.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still explains the response contract ('the response asks only for that missing summary') and covers the mutation, cascade, and idempotency-like re-call behavior. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds conditional meaning: implementedSummary overwrites the Work description, and batchSummary is required only when this call closes the Batch. That conditional requirement is not expressed in the schema, so it earns above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Record completed repository implementation', marks the Work done) and clearly differentiates itself from siblings like update_work and begin_work by describing a terminal state transition rather than a generic edit. An agent can identify the tool's role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use it once implementation is complete') plus conditional guidance: when batchSummary is needed, what happens on a repeat call against a done Work, and how a Work left planned/in-progress inside a ready Batch is handled. These are the exact conditions an agent needs to select and sequence this call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_batchGet batchC
Read-only
Inspect

Get one batch.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Batch.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safe-read profile is covered structurally. The description adds nothing beyond that: it doesn't say what is returned, whether the batch could be missing, or any error behavior. With a single trivial sentence, the description contributes no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with no filler, so it is concise and front-loaded. However, the brevity reflects under-specification rather than efficient communication, so it earns only a middle score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description bears some burden for explaining what a caller receives, and it says nothing about the returned batch shape or error cases. For a simple read-only getter this is tolerable but still an incomplete definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the single required parameter 'id' documented as "Id of the Batch." Per the baseline rule, full schema coverage yields a 3 even when the description adds no parameter meaning. The description indeed adds nothing here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get one batch" restates the tool name and title almost verbatim, adding no distinguishing detail. It does imply a single-item retrieval versus the sibling list_batches, but that contrast is only inferable, not stated. This is essentially a tautology of the title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as list_batches, get_work, or the many other get_* siblings. No prerequisites, no exclusions, no context are provided. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_deliveryGet deliveryA
Read-only
Inspect

Get one delivery by id, with its shipped date, rationale, summary, assigned batches, work progress, and archived state.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, so the safety profile is covered. The description goes beyond that by inventorying the returned fields (shipped date, rationale, summary, assigned batches, work progress, archived state), which is genuinely useful given there is no output schema; it stops short of stating error behavior for a missing or unauthorized id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and keying mechanism, with every listed clause earning its place as return-value information. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with annotations covering safety, the description supplies the field inventory that an absent output schema would otherwise leave blank, making it largely self-sufficient. A note on not-found/error behavior would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter with 100% schema description coverage ('Id of the Delivery.'), so the schema already carries the semantics. The description adds no format, source, or constraint detail for the id beyond what the schema states, which is the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get one delivery by id') and enumerates the payload (shipped date, rationale, summary, assigned batches, work progress, archived state). It implicitly contrasts with list_deliveries via 'one delivery by id', but never names a sibling such as get_delivery_presentation or get_work to disambiguate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the 'by id' phrasing signals a lookup when a delivery id is already known, but there is no explicit when-to-use, when-not-to-use, or named alternative. An agent can infer the context but gets no routing guidance among the many get_* and list_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_delivery_presentationGet delivery presentationA
Read-only
Inspect

Read a delivery's current release presentation document, or null when it has none.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds genuine behavioral value by disclosing the null-when-none result, which tells the agent how absence is represented rather than just what is fetched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero waste; the resource being fetched is front-loaded and the null case is appended compactly. Nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does mention the null case plus the document type, which is enough for an agent to call and interpret it. It stops short of describing the document's structure or format, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single id parameter is fully documented in the schema itself. The description adds no additional meaning about the id (format, source, or constraints), so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (read) and a specific resource (a delivery's current release presentation document). It is easily distinguishable from siblings like get_delivery, delete_delivery_presentation, and update_delivery_presentation because it names the presentation document as its subject.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to use this tool versus alternatives such as get_delivery or update_delivery_presentation, nor does it state prerequisites or any exclusions. Usage is only inferable from the tool name, and no explicit guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dictionary_termGet dictionary termC
Read-only
Inspect

Get one dictionary term.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Dictionary term.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered structurally. The description adds nothing on top of that: no behavior for a nonexistent or archived id, no note on whether archived terms are returned, no sense of payload size. With annotations doing the only work, the description contributes zero behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is free of padding, but it is not concise so much as empty — it conveys nothing an agent could not read off the tool name. A terse sentence that earns no place is under-specification, not conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is the simplest possible tool shape: one required string id, full schema coverage, and a readOnly annotation. That alone makes it callable. However, with no output schema, the description should have said something about what comes back or what happens when the id is not found, so it is only minimally viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single required parameter, so the baseline of 3 applies. "Id of the Dictionary term." in the schema is as informative as the description is, and the description supplies no format hints (UUID vs slug) or examples. No penalty, but no added value either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get one dictionary term" restates the tool name and title almost verbatim, giving no information beyond the identifier. It never distinguishes this from siblings like list_dictionary_terms, get_idea, or get_project, nor does it say what a 'dictionary term' contains. This is the tautology case the rubric reserves a 2 for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all: nothing says when to call this versus list_dictionary_terms (to enumerate) or update_dictionary_term (to modify). The closest thing to guidance is the word "one", which weakly implies single-item retrieval, but no alternative or condition is named. Not misleading, just absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_featuresGet featuresA
Read-only
Inspect

Get a project's features as a compact tree. The project name and description are the fixed root; each feature carries id, name, description, and children, with handsOff: true on protected features, which also covers every feature under them. Pass featureId to read just that feature and the features under it. Hands-off features cannot be changed. update_features needs the returned revision.

ParametersJSON Schema
NameRequiredDescriptionDefault
featureIdNoRead only this feature and the features under it.
projectIdYesId of the Project.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, so the description carries most of the burden and delivers: it discloses the return structure (id, name, description, children), the handsOff flag semantics including inheritance to descendants, the immutability of protected features, and the revision value that update_features consumes. It stops short of auth, rate limit, or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each carrying distinct payload: return shape, protection semantics, featureId scoping, and the update_features dependency. The scoping and structure information is front-loaded with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain the return value itself — and it does, enumerating the tree's fields and the handsOff/revision semantics. For a two-parameter read tool with readOnlyHint, nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description's featureId guidance ('read just that feature and the features under it') essentially restates the schema text for that parameter. projectId is only implied by 'a project's features' and adds no syntax or format detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get a project's features') plus the shape of the result ('compact tree'), which is more than the name alone conveys. It also implicitly separates itself from the mutating sibling by noting update_features depends on the returned revision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete usage branch: pass featureId to read just that feature and its descendants. It also names the downstream consumer (update_features needs the revision) and the constraint that hands-off features cannot be changed, but it never states an explicit when-not or a fully alternative tool choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_ideaGet ideaB
Read-only
Inspect

Get one idea with its full summary and content.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Idea.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe read. The description adds that the result includes the 'full summary and content', which is useful return context given there is no output schema. Beyond that it discloses no pagination, auth, or error behavior, so a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the verb and resource, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only single-get with full annotation and schema coverage, the definition is minimally adequate. The absence of an output schema means the vague phrase 'full summary and content' is the only hint at return shape, leaving some room for improvement.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema coverage, so the schema already documents 'id'. The description adds no format, source, or lookup guidance for the id, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (idea) with scope ('one idea'), which implicitly distinguishes it from the sibling list_ideas. However, it does not explicitly name or contrast with siblings the way a 5 would.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance or alternatives; it never mentions list_ideas for retrieval of many, nor the create/update/delete/archive siblings. Usage is only weakly implied by 'one idea'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_projectGet projectA
Read-only
Inspect

Get one project's own details: name, description, links, tags, Public Delivery Page settings (deliveryPageSettings, socials, publicDeliveryPage, showcase listing), and archived state. For next steps call get_project_state; for records use list_batches, search_works, list_ideas, and list_deliveries; instruction documents are the pizzadeveloper://projects/{projectId}/instructions/{documentId} resources.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Project.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered; the description adds the substantive behavioral detail that the payload includes Public Delivery Page configuration and archived state, which the agent cannot infer from the schema. It omits auth/rate-limit or error behavior, but for a read-only single-fetch that is a minor omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the return content and puts routing second, with essentially no filler. The parenthetical field enumerations are slightly dense but each item maps to real returned data, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of telling the agent what comes back, and it does so concretely while also pointing to the follow-up tool, sibling list tools, and the related resource URI. Nothing needed to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (id) and schema description coverage is 100%, so the schema already documents it fully. The description adds no syntactic or format guidance about the id beyond what the schema provides, making this the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Get one project's own details') and then enumerates exactly which fields are returned (name, description, links, tags, deliveryPageSettings, archived state). It is trivially distinguishable from siblings like get_project_state, list_projects, and update_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: get_project_state for next steps, list_batches/search_works/list_ideas/list_deliveries for records, and the pizzadeveloper://projects/{projectId}/instructions/{documentId} resource template for instruction documents. This is a rare case of naming both the alternatives and the condition that selects them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_stateGet project stateA
Read-only
Inspect

Get the planning state needed to resume work on one project. nextAction names the tool to call next. When the project has workflows, it lists them with their open runs and nextAction is start_workflow; resume_workflow continues an open run that already covers the request. Otherwise nextAction is begin_work (resume the active Work, claim the next one, or start the next planned Batch), finish_work (close a Batch whose Works are all done, with batchSummary, on developing.finishWorkId), create_work (add a Work to a developing Batch that has none), create_batch_plan (plan from the listed Ideas), or project_complete when nothing is left. It reports only Pizza Developer records, never repository, runtime, Git, deployment, or agent presence state.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesId of the Project.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered; the description adds real context beyond that — it discloses that nextAction names the next tool and, importantly, that it reports only Pizza Developer records and never repository, runtime, Git, deployment, or agent-presence state. That negative scope boundary is a genuinely useful behavioral disclosure, though return-format details are not covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, but the bulk is a single long run-on sentence enumerating nextAction branches. The information is relevant yet would be clearer as a list; the structure works against scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain return values — and it does explain nextAction semantics and the branch logic well. It leaves other returned state fields (workflow/run details beyond 'lists them with their open runs') only loosely described, so it is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single projectId parameter, so the schema already documents it fully. The description adds only the notion of 'one project' and no syntax or format detail, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Get the planning state needed to resume work on one project.' The scope (one project) and the resume-work purpose are clear. It stops short of explicitly contrasting with sibling get_project/get_work, so differentiation is only implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear context ('needed to resume work on one project') and unusually detailed routing guidance via nextAction branches (start_workflow, resume_workflow, begin_work, finish_work, create_work, create_batch_plan, project_complete). However, it frames these as next calls to make, not as explicit when-to-use-this-vs-alternatives for get_project_state itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workGet workB
Read-only
Inspect

Get one work item. reference is the batch-scoped name it is known by outside this app: 1-2 is work 2 of batch 1, the same reference a commit message carries.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Work.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe read. The description adds useful domain context about the batch-scoped reference format, but discloses nothing about the return shape or whether a missing work item errors versus returns empty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action. The second sentence earns its place by explaining the reference format, though it introduces terminology that doesn't match the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with a readOnlyHint and no output schema, this is close to adequate. But the 'reference' vs 'id' naming mismatch leaves ambiguity about exactly what value to pass, which is the one thing this tool's description most needed to resolve.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the single 'id' parameter is already documented, establishing a baseline of 3. The description adds genuine format meaning ('1-2 is work 2 of batch 1'), but confusingly names the concept 'reference' while the schema calls the parameter 'id', leaving the relationship between the two unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get one work item'), which is clear against list_works and get_batch siblings. However it never explicitly differentiates itself from those siblings or notes it returns a single item by id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no alternatives named, despite many siblings (get_batch, list_batches, search_works) that could retrieve overlapping data. The only context is about the reference format, not about when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workflowGet workflowA
Read-only
Inspect

Get one saved workflow with its name, description, nodes, connections, and the instruction and ruleset IDs attached to each connection.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesId of the Project.
workflowIdYesId of the Workflow.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, so safety is covered; the description adds real behavioral value by disclosing exactly what the read returns (nodes, connections, attached instruction/ruleset IDs), which is non-obvious and important with no output schema present. It stops short of stating error/not-found behavior or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the resource is named first and the payload detail follows immediately. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-param read tool with no output schema, the description usefully compensates by enumerating the returned fields, which is the main thing an agent needs. It would be fully complete with a note on not-found behavior, but nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with both required parameters documented, so the schema does the heavy lifting and the baseline is 3. The description adds no format, scoping, or relationship detail about projectId vs workflowId beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get') and resource ('one saved workflow') and enumerates the payload: name, description, nodes, connections, and per-connection instruction/ruleset IDs. It is clearly distinguishable from list_workflows by the singular 'one saved workflow', though it never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Retrieval by id is implied by 'one saved workflow' plus the required projectId/workflowId, but there is no explicit when-to-use guidance, no statement about when to prefer list_workflows, and no mention of what happens if the workflow does not exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_batchesList batchesA
Read-only
Inspect

List a project's batches as bounded summaries (id, number, name, status, and queue position). Call get_batch for rationale, summary, and Work detail. Batches are numbered sequentially per project (1,2,3…).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows to return.
offsetNoRows to skip before the first one returned.
projectIdYesId of the Project.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true is declared in annotations, so safety is covered; the description adds value beyond that by disclosing the exact shape of what comes back (id, number, name, status, queue position) and the sequential per-project numbering rule. It does not mention ordering of the list or pagination semantics, which would complete the picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary action and result shape, followed by a sibling pointer and a domain fact. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing the returned fields, and it explains the sequential batch numbering. Minor gaps remain around list ordering and how limit/offset interact with the numbered batches, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit, offset, and projectId are already documented in the schema. The description only gestures at bounding via 'bounded summaries' and adds nothing about pagination syntax or defaults, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (batches) scoped to a project, and enumerates the returned fields. It explicitly distinguishes itself from the sibling get_batch by framing results as 'bounded summaries' rather than full detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool (get_batch) and the condition that selects it ('for rationale, summary, and Work detail'). There is no explicit when-not guidance or mention of pagination-driven iteration, but the routing is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_deliveriesList deliveriesA
Read-only
Inspect

List a project's deliveries, newest shipped first, as bounded summaries: id, public version, shipped date, batch and Work counts, and whether it has a presentation. Lists active deliveries; pass archived=true for archived ones. Call get_delivery for rationale, summary, and assigned batches.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows to return.
offsetNoRows to skip before the first one returned.
archivedNoTrue to list archived records instead of active ones.
projectIdYesId of the Project.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint already declares this a safe read, but the description adds behavior the annotations do not: the sort order (newest shipped first) and the fact that rows are bounded summaries rather than full records. It does not mention pagination behavior or default limit, which are minor gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the resource and ordering, then the mode switch, then the routing hint. The field enumeration is dense but earns its place since there is no output schema. Slightly long, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-shape burden and does so by enumerating the summary fields (id, version, shipped date, counts, presentation flag). Combined with the active/archived switch and the get_delivery handoff, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (limit, offset, archived, projectId) are already documented in structured data. The description reinforces archived rather than adding syntax or format detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (list a project's deliveries) with scope (per-project), ordering (newest shipped first), and the exact shape of what is returned. An agent can distinguish it from get_delivery or list_batches without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between modes: lists active deliveries by default, archived=true for archived ones, and directs to get_delivery for rationale/summary/batches. Both the condition for a parameter and the alternative tool are named, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_dictionary_termsList dictionary termsA
Read-only
Inspect

List a project's active dictionary terms as bounded summaries (id, word, and aliases) in the vibe coder's own order. Pass archived=true to read archived terms instead. Call get_dictionary_term for the approved definition.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows to return.
offsetNoRows to skip before the first one returned.
archivedNoTrue to list archived records instead of active ones.
projectIdYesId of the Project.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description adds real behavioral context beyond that: the return shape (id, word, aliases) and that results come back in the vibe coder's own order. It doesn't discuss pagination behavior, but limit/offset are schema-documented, so this is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly-written sentences, front-loaded with the core purpose, followed by the mode switch and the alternative. Zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description compensates by naming the returned fields and ordering. For a read-only list tool with fully-covered params, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all four parameters are already documented. The description restates the archived switch (already in the schema) and implies bounding via limit/offset but adds no syntax or format beyond that. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (dictionary terms) with scope (a project's active terms) and even the return shape (bounded summaries of id, word, aliases). It clearly distinguishes itself from get_dictionary_term, which it names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly covers when to use it (active terms), how to switch modes (archived=true), and routes the agent to the right alternative (get_dictionary_term for the approved definition). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_ideasList ideasA
Read-only
Inspect

List a project's ideas in one state as bounded summaries in the vibe coder's own order, front first. state is active by default; archived ideas were completed; declined ideas were decided against and carry the reason. Only active ideas are planning input; their first two positions are the planning ideas. Call get_idea for full content.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows to return.
stateNoWhich Ideas to list; active by default.
offsetNoRows to skip before the first one returned.
projectIdYesId of the Project.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true covers the safety profile, and the description goes well beyond it: it discloses ordering ('vibe coder's own order, front first'), result shape ('bounded summaries'), state semantics, and planning significance ('first two positions are the planning ideas'). It stops short of stating default limit/pagination behavior, so not a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, front-loaded with the core action and scope, then state semantics, then the planning nuance, then the sibling routing. Every clause carries information and none is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully characterizes the return as 'bounded summaries' and points to get_idea for details, and it covers ordering. It does not describe pagination defaults or the limit/offset interplay, a minor gap for a read-only list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real meaning over the schema's terse enum by defining archived ('completed') and declined ('decided against and carry the reason'), which helps the agent interpret results per state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List a project's ideas') plus scope ('in one state as bounded summaries') and ordering. It is clearly separable from get_idea ('Call get_idea for full content') and from the archive_/delete_idea siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains what each state value means (active is default, archived=completed, declined=decided against with a reason) and routes to the alternative, get_idea, for full content. This gives the agent an explicit basis for choosing both state and tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsList projectsB
Read-only
Inspect

List every project the signed-in user owns or that has been shared with them.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows to return.
offsetNoRows to skip before the first one returned.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the agent knows this is a safe read. The description adds the visibility scope (owned or shared), which is useful, but doesn't cover pagination behavior or output characteristics beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise sentence that front-loads the action and resource, with the scoping detail efficiently included. No waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and 100% schema coverage, the description is adequate for a list tool. However, for a tool with many siblings, it could benefit from clarifying the relationship to get_project or how shared projects are represented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents limit and offset. The description adds no parameter information, but the baseline is 3 when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb (List) and resource (project) and specifies the scope (owned by or shared with the signed-in user). It doesn't differentiate itself from other list_* siblings like list_batches, but those are clearly distinct resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus get_project or other project-related tools. The description doesn't indicate alternatives or prerequisites, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_rulebooksList rulesetsA
Read-only
Inspect

List the project's rulesets with their rules and, in attachedTo, the workflows and connections each one is attached to. A ruleset is a name and ordered rules; each rule is one statement at level must (a failure refuses the write) or should (a failure is reported). Rules run only inside workflow runs: a ruleset is checked when it is attached to a workflow connection (rulebookIds in update_workflow) and the write that completes that step happens.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesId of the Project.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already indicate readOnlyHint=true, so the description does not need to restate safety. It adds useful behavioral context: what a ruleset is, the two rule levels (must vs should) and their failure semantics, and the fact that rules only run inside workflow runs when attached to a workflow connection. These details go beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description front-loads the main purpose in its first clause and then uses subsequent sentences to define rulesets, rule levels, and attachment behavior. It is appropriately sized for a list tool without an output schema, though the domain explanation could be slightly tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of explaining the return shape: rulesets contain ordered rules, and attachedTo shows workflows and connections. It also explains the meaning of rule levels. Pagination or ordering of the ruleset list itself is not addressed, but the core contextual needs are covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for its single parameter, projectId. The description does not add any additional syntax, format, or filtering semantics for that parameter, so the schema already carries the full parameter explanation. This meets the baseline of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('List the project's rulesets') and immediately qualifies the returned content: rules and attachment relationships. It clearly distinguishes this from sibling CRUD tools like create_rulebook or update_rulebook by specifying that this is a read operation listing rulesets with their rules and attachments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the domain context of rulesets and when rules execute, but it does not explicitly state when an agent should call this tool versus alternatives such as get_workflow or check_rules. Usage is implied by the list operation, but no direct when-to-use or when-not-to-use guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_workflowsList workflowsA
Read-only
Inspect

List the project's workflows with each one's name, description, entry points (idea, batch, and/or delivery), whether it can run, and its graph. A project can have several workflows, each described by the kind of request it handles. Archived workflows are left out unless includeArchived is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesId of the Project.
includeArchivedNoTrue to include archived Workflows.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true already declaring the safe-read profile, the description adds genuine behavioral context: archived workflows are filtered out by default unless includeArchived is true, and it sketches the per-item payload. It does not mention pagination, ordering, or behavior when the project has none.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and the returned fields, followed by the archiving rule. Every sentence carries information, and the field enumeration is justified because there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully compensates by naming the returned fields (name, description, entry points, runnability, graph). Combined with the read annotation and archived-filter note, an agent has what it needs, with only pagination/ordering left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both projectId and includeArchived are already documented; baseline is 3. The description restates the archived-inclusion rule but adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (the project's workflows) and enumerates what each item carries. It is clearly distinguishable from get_workflow and list_projects, though it never names the singular sibling or another list alternative to route the agent explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the natural reading is 'call this to enumerate a project's workflows,' but there is no statement of when to prefer it over get_workflow or how to find a specific workflow. The includeArchived flag is the only conditional guidance offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reorderReorder recordsA
DestructiveIdempotent
Inspect

Set the manual order of one kind of record, first shown first. orderedIds lists the records; parentId names where they live:

  • project: your own projects on the dashboard; no parentId; a subset keeps every other project in place.

  • idea: every active Idea; parentId is the project; the first two are planning input.

  • batch: every planned Batch; parentId is the project; developing and ready Batches keep their place.

  • work: every Work in the Batch; parentId is that Batch, which must not be ready.

  • instruction: every active instruction; parentId is the project; during a workflow run, accepted only in its self-improvement stage.

  • dictionary_term: every active term; parentId is the project.

  • rulebook: every ruleset; parentId is the project.

  • rule: every rule in one ruleset; parentId is that ruleset.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesKind of record to order.
parentIdNoWhere the records live; the tool description says what each kind takes.
orderedIdsYesRecord ids in the wanted order, first shown first.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=true, so the safety profile is partly covered. The description adds genuinely non-obvious behavior: partial reorders leave untouched records in place ('a subset keeps every other project in place', 'developing and ready Batches keep their place'), plus a hard precondition on the work parent and a workflow-stage restriction on instructions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the opening clause, and the bullet list is the right structure for eight kinds with divergent rules. It is longer than average but each bullet carries distinct, load-bearing constraints, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the effects, preconditions, and per-kind scoping an agent needs. It omits what the call returns or how errors surface, but the core invocation requirements are complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema itself defers parentId semantics to the description ('the tool description says what each kind takes'). The description delivers on that, specifying per-kind parentId targets and the orderedIds ordering rule, adding real meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb+resource ('Set the manual order of one kind of record') and adds the ordering semantics ('first shown first'). No sibling tool performs reordering, so it is unambiguously distinguishable, and the kind list makes scope concrete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bulleted per-kind rules effectively state when each variant applies: which kinds require a parentId, that a 'work' parent batch must not be ready, and that 'instruction' reordering is only accepted during a workflow's self-improvement stage. It gives rich context but does not name alternative tools (there are none), so it stops short of explicit when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_problemReport a problemAInspect

Report a problem with Pizza Developer itself (the app, not a project tracked inside it) on behalf of the signed-in vibe coder. Never send one on your own initiative: first show the vibe coder the exact message and path, and call only after they explicitly confirm that report, passing userConfirmed: true. A report cannot be read back, edited, or withdrawn once sent. Pass path when it is about a particular page (/projects/{id}), and say plainly in message what went wrong and what was expected. Docs: https://pizzadeveloper.com/docs/report-a-problem

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoApp path the report is about.
messageYesThe problem, exactly as shown to the vibe coder.
userConfirmedYesTrue only after the vibe coder approved this exact report.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, idempotentHint=false, destructiveHint=false. The description adds genuinely new, non-obvious behavior: a report cannot be read back, edited, or withdrawn once sent, and it requires a confirmation gate. That is the key trait an agent must know for a one-way write, though it does not discuss failure modes or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the identity of the tool and then the critical confirmation constraint; every sentence carries actionable content. It is somewhat dense with imperatives but nothing is wasted, and the docs link is a reasonable trailing pointer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and annotations covering only the safety profile, the description supplies everything an agent needs: scope, mandatory confirmation flow, irreversibility, and parameter construction guidance. Nothing required for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes further by illustrating the path format (/projects/{id}) and instructing what message should contain ('what went wrong and what was expected'), adding usage meaning beyond the schema's terse field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('report a problem with Pizza Developer itself') and explicitly scopes out the alternative interpretation ('the app, not a project tracked inside it'), which separates it from project/workflow siblings such as report_workflow_step. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when/when-not rules: never send on your own initiative, show the exact message and path first, and call only after explicit user confirmation with userConfirmed: true. It also says when to include path (a particular page like /projects/{id}), so the routing and precondition logic is complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_workflow_stepReport workflow stepAInspect

Report the agent client's outcome and optional measured costs for a workflow step that has no expected lifecycle tool, a failure, the self-improvement step, or the client score. A step with an expected lifecycle tool is completed by that tool, which also accepts the client metrics, so it needs no report; failed blocks a step for another attempt. Optional stages are bypassed only by choosing a configured available edge with resume_workflow before entering them; once entered, every step is required. A completed self-improvement report closes that step and its reason describes what changed and why, or why nothing did. Research content is never stored. When one route follows, the response carries the next step's instructions; when several do, choose one with resume_workflow and its instructions arrive then.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYesId of the Workflow run.
reasonNoWhat happened and why.
statusYescompleted, or failed to block the step for another attempt.
stepIdYesId of the step being reported.
metricsNoOptional client measurements for the step.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=false and destructiveHint=false already declaring the write/safe profile, the description still adds real behavioral context: 'failed blocks a step for another attempt', a completed self-improvement report 'closes that step', and 'Research content is never stored'. These are non-obvious state effects an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the description runs long and dense, packing several distinct rules into multi-clause sentences. Useful content is present, yet the lack of separation between usage, state effects and response behavior makes it harder to scan than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does address it ('When one route follows, the response carries the next step's instructions; when several do, choose one with resume_workflow'). For a mutation tool it also covers the key state effects and the no-research-storage rule, leaving only minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents runId, stepId, status, reason and metrics; the baseline of 3 applies. The description adds only marginal parameter meaning ('failed blocks the step for another attempt', reason as 'what changed and why'), which largely overlaps with behavioral context rather than extending syntax or format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Report ... for a workflow step') and enumerates the outcome types it covers (no lifecycle tool, failure, self-improvement step, client score). An agent can identify the action, though the opening clause is syntactically tangled enough that the scope takes a second read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use conditions ('a workflow step that has no expected lifecycle tool, a failure, the self-improvement step, or the client score') and routes the agent elsewhere for lifecycle steps ('A step with an expected lifecycle tool is completed by that tool'), naming resume_workflow as the alternative for optional stages. Guidance is present but buried in dense prose rather than stated crisply as when/when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_ideaRestore ideaA
Idempotent
Inspect

Restore an archived or declined idea to the back of the active planning order.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Idea.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, idempotent, non-destructive, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: the restored idea lands at the BACK of the active planning order, a positional side effect an agent could not infer from the schema or annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, no waste, and the action verb plus scope is front-loaded. Every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, single-parameter mutation with no output schema, the description covers the action, valid source states, and the resulting position. It omits only edge behavior such as errors when the idea is not archived/declined, which is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required 'id' parameter whose schema description is complete (100% coverage), so the schema does the work. The description adds no syntax, format, or sourcing guidance for the id, making 3 the appropriate baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restore) and resource (idea) plus the source states (archived or declined). It implies the contrast with archive_idea without naming it, so it is clear but not fully sibling-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The condition 'archived or declined' implies when the tool applies, but there is no explicit when-not guidance or reference to the inverse archive_idea / decline_idea siblings. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resume_workflowResume workflow runAInspect

Reconcile and read an open workflow run, then deliver its current step. Each workflow has at most one open run; name it with runId or workflowId when several are open. Pass an available edgeId before reporting to choose a configured branch, including one that names the optional stages it bypasses. Without an edgeId, a step already delivered is re-sent in full, which is how a new session recovers its instructions; other calls return only the step's instruction titles after its first delivery.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdNoWorkflow run to resume, when several are open.
edgeIdNoAvailable connection to choose before reporting.
projectIdYesId of the Project.
workflowIdNoWorkflow whose open run to resume.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare it is non-read-only and non-destructive; the description adds substantial behavioral context beyond that. It discloses that reconcile mutates run state, that at most one run is open per workflow, that edgeId selects a branch (including bypassed optional stages), and crucially that absent edgeId a delivered step is re-sent in full while other calls return only instruction titles — a subtle, high-value detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and free of filler, with each sentence contributing behavior. It is dense and somewhat compressed, requiring careful reading, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-behavior burden and does so well, explaining that the current step is delivered and that later calls return only instruction titles. Combined with the branch and run-selection behavior, it is largely complete for a stateful resume tool, though a brief statement of what the delivered step contains would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description genuinely enriches the parameters: it explains that runId or workflowId disambiguate when several runs are open, and that edgeId chooses a configured branch including bypassed optional stages. This adds meaning the terse schema descriptions do not carry.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives specific verbs (reconcile, read, deliver) and a clear resource (an open workflow run and its current step), distinguishing it conceptually from start_workflow and report_workflow_step. However, it never names those siblings to make the boundary explicit, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'resume... an open workflow run' and the note about naming the run when several are open. It explains parameter-level conditional behavior well, but provides no explicit when-to-use guidance relative to alternatives like start_workflow or report_workflow_step, so the agent must infer the selection context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revert_batchRevert batchAInspect

Move a batch one step back as an undo: developing → planned, and only while every work inside it is still planned. Ready batches are frozen history and can't be reverted.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Batch.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=false and destructiveHint=false already covering the safety profile, the description adds real context: it is an undo, it moves exactly one step backward, it has a precondition (all contained work still planned), and ready batches are immutable. It doesn't state the return shape or failure behavior, but the mutation semantics are well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence opens with the action, then layers the transition, the precondition, and the exclusion — every clause carries new information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter mutation tool with no output schema, the description covers the state change, the precondition, and the irreversible-case exclusion. Only the failure/error behavior when the precondition is unmet is left unstated, a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single id parameter, so the schema already carries the parameter documentation. The description adds nothing about the id format or constraints, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (revert / move one step back) and resource (a batch), and spells out the exact state transition developing → planned. This clearly distinguishes it from update_batch, delete_batch, and get_batch without needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when (only while every work inside is still planned) and an explicit when-not (ready batches are frozen history and can't be reverted). No alternative sibling is named for the case where revert is not allowed, so it falls just short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_worksSearch worksA
Read-only
Inspect

List or find Works, newest-updated first, as trimmed rows (id, reference, title, size, status, batchId, projectId; call get_work for full detail). Pass projectId to list one project's Works; query searches title and description across every project you can reach, or inside that project. status is planned, in-progress, or done. reference is the batch-scoped name a Work is known by outside this app: 1-2 is work 2 of batch 1.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows to return.
queryNoText to find in Work titles and descriptions.
offsetNoRows to skip before the first one returned.
statusNoLimit to one Work status.
batchIdNoLimit to one Batch.
projectIdNoLimit to one Project.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true; the description adds real behavioral context beyond that: newest-updated-first ordering, trimmed result shape, cross-project reach of query, and the distinction between reference and title. It does not mention pagination semantics (limit/offset behavior) or result-count limits, which are relevant for a search tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: purpose and row shape first, then scoping, then field semantics. Every sentence adds information; the only slight clutter is the parenthetical field list, but it earns its place by documenting the return shape in the absence of an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description helpfully enumerates returned fields, ordering, and scoping rules, which covers most of what an agent needs. It omits pagination behavior and result-count/limit interaction, which are the remaining gaps for a search/list tool with limit and offset parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining the interaction between query and projectId, defining the allowed status values in prose, and clarifying what reference means (batch-scoped external name, e.g. '1-2'), which the schema left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('List or find Works') and immediately distinguishes from the sibling get_work by stating it returns trimmed rows and telling the caller to use get_work for full detail. The row fields and ordering are spelled out, so an agent knows exactly what this tool returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains that projectId scopes to one project, that query searches across every reachable project or inside that project, and that get_work is the alternative for full detail. It does not spell out exclusions or edge cases (e.g., query+batchId combination behavior), so it stops short of a full when/when-not guide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_workflowStart workflow runAInspect

Start a run of one workflow from an Idea, a planned existing Batch whose Work has not begun, or a new batch plan (name and ordered works, as in create_batch_plan). A workflow that starts at Delivery takes none of them: the run begins on its Start to Delivery connection, and create_delivery or assign_batch_to_delivery completes it. With batch, the Batch is planned and the run started in one call, its review step is skipped, and edgeId can name the Batch stage route; a retry while that run is open returns the same run. A Batch run governs that Batch's Work queue until it is ready. The run is pinned to the workflow's current revision and returns its current step. Each instruction body is sent once per run; a document already sent earlier in the run comes back as a uri reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
batchNoNew Batch plan to create and start the run from.
edgeIdNoBatch stage route to take when the run starts from a Batch.
ideaIdNoIdea to start the run from.
batchIdNoPlanned Batch to start the run from.
projectIdYesId of the Project.
workflowIdYesId of the Workflow.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say this is a non-read, non-destructive mutation. The description adds substantial behavioral context beyond them: the run is pinned to the workflow's current revision, returns its current step, each instruction body is sent once per run, previously sent documents come back as uri references, the review step is skipped on the batch path, and a retry while the run is open is idempotent. This is unusually rich disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and source modes are front-loaded in the first sentence, and no sentence is filler given the tool's complexity. However, several sentences pack multiple unrelated facts (idempotency, review-step skipping, revision pinning) into dense clauses, which costs readability for a 6-parameter tool with three start modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by stating the return value (the current step). Combined with the source-mode coverage, idempotency, revision pinning, and the Delivery exception, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at 100%, the baseline is 3, but the description adds real meaning: it explains the mutually exclusive source inputs (ideaId/batchId/batch), that edgeId names the Batch stage route for the batch path, and that the batch object is a full plan (name + ordered works). That is value beyond the schema's one-line parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (start a run of one workflow) and enumerates the three possible sources (Idea, planned Batch, new batch plan), plus the Delivery-start exception where none apply. An agent can distinguish it from create_batch_plan, create_delivery, and assign_batch_to_delivery from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names when this tool applies vs. alternatives: it points to create_batch_plan for the plan shape, and states that a Delivery-starting workflow takes none of the sources and is completed by create_delivery or assign_batch_to_delivery. The routing conditions are stated, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_deliveryUnarchive deliveryA
Idempotent
Inspect

Unarchive a delivery, returning it to the Deliveries page and the public surfaces it was listed on.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral context by disclosing the visibility side effect on public surfaces, which annotations cannot convey. It stops short of stating auth/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb and resource first, followed immediately by the effect. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter state-reversal tool with annotations covering idempotency and safety and no output schema, the description supplies enough to call it correctly. Only the absence of any precondition (e.g., that the delivery must already be archived) keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100%, so the schema fully documents the id field. The description adds no format or constraint detail beyond what the schema provides, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Unarchive) and resource (delivery) and goes further by describing the resulting state: return to the Deliveries page and to the public surfaces it was listed on. This distinguishes it clearly from the many sibling unarchive_* tools, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent infers this is the inverse of archive_delivery and applies when a delivery should be made visible again. There is no explicit when-to-use, when-not-to-use, or mention of alternatives such as update_public_delivery_page.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_dictionary_termUnarchive dictionary termA
Idempotent
Inspect

Unarchive a dictionary term, returning it to the back of the active dictionary.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Dictionary term.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-readOnly, idempotent, non-destructive, so the safety profile is covered. The description adds a genuinely useful behavioral detail beyond annotations: the term returns 'to the back of the active dictionary,' disclosing the ordering side effect an agent could not learn from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and appends the meaningful consequence. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter unarchive operation whose annotations carry the safety profile, the description covers the action and its ordering effect. It omits preconditions (e.g., term must already be archived) and failure behavior, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter with 100% schema coverage, so the schema already documents 'id'. The description adds no syntax or format detail beyond the schema, which is the expected baseline for this case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Unarchive') and resource ('dictionary term'), and describes the resulting state. It is clearly distinguishable from archive_dictionary_term and delete_dictionary_term without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage as the inverse of archive_dictionary_term and describes the effect, but never states when to prefer this over delete_dictionary_term or what the precondition is for unarchiving. Usage is inferable but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_instructionUnarchive instructionA
Idempotent
Inspect

Unarchive an instruction document, returning it to the back of the active list.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Instruction.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With destructiveHint=false and idempotentHint=true already declaring the safety profile, the description contributes a genuine behavioral detail: the restored instruction is placed at the back of the active list, not its old position. That ordering effect is not derivable from annotations or schema and is useful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the action first and the consequential side effect second. Nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter state-change tool with full annotation coverage and no output schema, the description covers what an agent needs to call it. It leaves minor gaps around error behavior (e.g., already-active instructions) and required permissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single id parameter, so the schema already carries the parameter documentation. The description adds no syntax, format, or identifier-source detail beyond it, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (unarchive) and resource (instruction document), making the inverse relationship to the sibling archive_instruction obvious from the name. It is clear what the tool does, though it never explicitly names a differentiating sibling or scope boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and the framing 'returning it to the back of the active list' signals the intended state transition. There is no explicit when-to-use/when-not guidance, no mention of alternatives, and no prerequisites for a previously archived instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_projectUnarchive projectAInspect

Unarchive a project, restoring it to the default dashboard view.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Project.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the agent knows this is a non-destructive write. The description adds the side effect ('restoring it to the default dashboard view'), which is useful context beyond what annotations provide. However, it doesn't state whether this requires prior archival state, what permissions are needed, or how it interacts with the dashboard view for other users.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero waste, front-loaded with the action verb and the resource, followed immediately by the side effect. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter state-change tool with annotations covering the safety profile and full schema coverage, the description is adequate. It states the primary side effect (restoring to the dashboard view), which is the main behavioral detail an agent needs. Missing only edge cases like idempotency or required prior state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single parameter (id) is fully documented in the schema as 'Id of the Project.' The description adds no additional parameter semantics. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Unarchive a project') and names the sibling it is not. An agent can distinguish this from archive_project and other unarchive_* siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (restoring an archived project to the default dashboard) but provides no explicit when-to-use or when-not-to-use guidance. No alternative tool is named. Implied usage only, which is the definition of a 3.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_workflowUnarchive workflowAInspect

Unarchive a workflow so it can be edited and run again.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesId of the Project.
workflowIdYesId of the Workflow.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, so the agent already knows this is a non-destructive mutation. The description usefully adds the resulting state (editable and runnable again), but says nothing about idempotency, error behavior when the workflow is not archived, or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the operation and its outcome front-loaded; every word earns its place. It is perhaps slightly too terse to route between siblings, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter state-toggle tool with no output schema and annotations covering the safety profile, the description supplies what an agent needs to call it. Only edge-case behavior (already unarchived, missing workflow) is unspecified, which is a modest gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both projectId and workflowId fully documented in the schema. The description adds no extra meaning about parameter formats or constraints, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (unarchive) and resource (workflow) and adds the state effect: it becomes editable and runnable again. It does not explicitly distinguish itself from the many sibling unarchive_* tools, but the resource scoping makes the target unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the phrase 'so it can be edited and run again' — an agent can infer that this is the inverse of archive_workflow, but there is no explicit when-to-use, when-not, or reference to alternatives such as update_workflow or resume_workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_batchUpdate batchA
Destructive
Inspect

Edit a batch's name, rationale, or change summary; only supplied fields change, and an empty string clears rationale or summary. A name is 1 to 3 lowercase words joined by single hyphens, like dark-mode-toggle. Rationale is the vibe coder's stated reason for building the batch, in one or two plain sentences. Summary is what the batch changed, drafted from its recorded work, in two or three factual sentences. A ready batch refuses rationale and summary edits. Status moves through begin_work, finish_work, and revert_batch, and batch numbers are immutable.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Batch.
nameNoNew Batch name.
summaryNoWhat the Batch changed; an empty string clears it.
rationaleNoThe vibe coder's reason for the Batch; an empty string clears it.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=true; the description supplies the substance behind that: partial-update semantics (only supplied fields change), the destructive clearing behavior of an empty string, the guardrail that a ready batch rejects rationale/summary edits, and the immutability of batch numbers. This is exactly the behavioral detail annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The editable fields and partial-update rule are front-loaded in the first clause, and the field definitions follow. It is dense but slightly overloaded by the trailing sentence about status transitions and immutable batch numbers, which is useful routing context but tangential to editing fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no output schema and only readOnly/destructive annotations, the description covers field semantics, clearing behavior, guardrails, and routing to status tools. It omits any mention of the return value or permission requirements, which are minor gaps given the strong coverage elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so 3 is the floor, but the description goes beyond the terse schema descriptions: it translates the name pattern into '1 to 3 lowercase words joined by single hyphens', restates the empty-string clearing rule for rationale and summary, and defines what each free-text field should contain (one or two plain sentences vs. two or three factual sentences).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Edit) and resource (a batch) plus the exact editable fields (name, rationale, change summary), which cleanly separates it from update_delivery, update_idea, update_work, and the status-transition siblings. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use conditions (only supplied fields change; empty string clears rationale/summary) and a real exclusion ('A ready batch refuses rationale and summary edits'). It also names begin_work, finish_work, and revert_batch as the tools that move status, implicitly routing status changes away from this tool, but it never explicitly says 'use X for status changes'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_deliveryUpdate deliveryA
Destructive
Inspect

Edit a delivery's own record (version, rationale, summary, ship date, or which undelivered batches ride in it). Only supplied fields change; pass an empty string to clear a text field. batchIds, when present, is the full desired set of batches and replaces the current one. shippedAt corrects when it shipped (ISO date or date-time, not in the future). The version is a free-form label and must stay non-empty and unique within its project.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.
summaryNoWhat shipped; an empty string clears it.
versionNoVersion label, such as 2.4.0.
batchIdsNoFull replacement set of Batches.
rationaleNoWhy the Delivery exists; an empty string clears it.
shippedAtNoShip date, as an ISO date or date-time.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=true; the description goes well beyond that by disclosing partial-update semantics ('Only supplied fields change'), empty-string clearing behavior, full-replacement semantics for batchIds, the non-future constraint on shippedAt, and version uniqueness within a project. This is rich behavioral context an agent cannot get from the annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each carrying a distinct rule (scope, partial update, batch replacement, shippedAt/version constraints). Front-loaded with the core purpose and zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with destructiveHint and no output schema, the description covers all six parameters, clearing semantics, and key constraints, which is nearly complete. It omits anything about error/conflict behavior or what the response returns, though the latter is less critical without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real semantics: empty-string clearing, batchIds as a full replacement set, shippedAt format and the non-future constraint, and version uniqueness/non-empty rules. Some of this (clearing, replacement set) overlaps the schema descriptions, so it is a genuine but partial addition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Edit') and resource ('a delivery's own record') and enumerates exactly which fields are editable (version, rationale, summary, ship date, batches). It implicitly distinguishes itself from siblings like update_batch and assign_batch_to_delivery by scoping to the delivery's own record.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it by describing what fields it edits and how partial updates behave, but it never explicitly names alternatives (e.g., assign_batch_to_delivery vs. using batchIds here) or states prerequisites. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_delivery_presentationUpdate delivery presentationA
Destructive
Inspect

Save a delivery's release presentation: scenes built only from the listed elements, played on the Public Delivery Page in the brand's style. The story comes from the delivery summary and real product copy; brand colors, radius, and font come from the product's own logo and styles. The logo uploaded on the project's public page is the only brand mark shown, and the presentation holds no source code. Release notes lead with the changes element for the delivery's headlines and show what moved with before-after or navigation. Every scene lasts at least its reading time, about 1.5 seconds plus one second per three words shown, and at least 3 seconds. Scenes last whole seconds, at most 60 seconds in total. Replaces the current presentation; earlier versions stay restorable in the app. An invalid presentation returns field-level errors to fix.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Delivery.
presentationYesPresentation document: brand and scenes.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the write/destructive profile is covered. The description still adds real value beyond them: it confirms the replacement semantics, notes earlier versions stay restorable in the app, and states that invalid input returns field-level errors. It does not describe auth requirements or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the core action before descending into content and timing constraints, and most sentences carry information an agent needs (timing math, element ordering, validation behavior). It is dense and long, with a few details that would sit better as schema annotations, but little is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deeply nested, destructive tool with no output schema, the description covers content constraints, style sourcing, timing rules, replacement/restore behavior, and error reporting. The only gap is that it does not characterize a successful return, which slightly matters given the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema itself fully enumerates the nested brand/scenes structure, so the baseline is 3. The description adds conceptual meaning (brand derives from logo and product styles, story comes from the delivery summary) but gives no field-level syntax or defaults beyond what the schema already encodes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Save a delivery's release presentation') and immediately characterizes the artifact and its rendering surface (Public Delivery Page, brand's style). It distinguishes itself from the get/delete/update siblings by framing the operation as a save that replaces the current presentation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the artifact's purpose and the note that it 'Replaces the current presentation,' but the definition never states when to choose this over update_delivery, get_delivery_presentation, or delete_delivery_presentation, nor any prerequisites. The agent must infer routing from sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_dictionary_termUpdate dictionary termB
Destructive
Inspect

Update a dictionary term. Only the fields you provide change.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Dictionary term.
wordNoThe Dictionary word.
aliasesNoOther words that mean the same.
definitionNoThe approved meaning of the word.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, destructiveHint=true), so the description does not need to restate mutability. It does add genuinely useful behavior beyond that: 'Only the fields you provide change' discloses partial-update semantics, telling the agent omitted fields are preserved rather than cleared. It stops short of noting permissions or the tension with destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the behavioral constraint is front-loaded right after the purpose statement. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter mutation tool with no output schema, the description is adequate but thin: it never says what a successful update returns, whether id must reference an existing term, or how an empty aliases array is treated. Annotations carry the safety load, but the operational picture is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (id, word, aliases, definition) are already documented. The description's partial-update sentence adds a small amount of semantic value about how optional params behave, but no field-level meaning beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (update) and resource (dictionary term), and the name distinguishes it from create_dictionary_term / delete_dictionary_term / archive_dictionary_term siblings. It does not need to explain those siblings because the naming convention already separates them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when this tool is appropriate versus archive_dictionary_term, delete_dictionary_term, or get_dictionary_term, and no mention of prerequisites such as required id lookup. The only hint is implicit patch semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_featuresUpdate featuresA
Destructive
Inspect

Change a project's features by id: add (a feature, with optional children, under parentId or null for top level), update (name and/or description), move (to parentId, optional position), remove (the feature and every feature under it). Every feature needs a concise name and description. Changes apply in order as one atomic write against baseRevision from get_features; split a large restructure into several calls, each with the revision the previous one returned. Hands-off marks belong to the vibe coder: hands-off features and every feature under them cannot be changed, moved, removed, or added to. Stale writes are refused, as is a change to a feature the vibe coder has open for editing in the browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
changesYesChanges to apply in order: add, update, move, or remove.
projectIdYesId of the Project.
baseRevisionYesRevision returned by get_features.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only flag destructiveHint=true; the description goes far beyond, disclosing atomicity ("one atomic write"), in-order application, optimistic-concurrency via baseRevision, cascade behavior on remove ("the feature and every feature under it"), and the hands-off/editing-lock protections. These are exactly the behavioral traits an agent cannot infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the operation list, then layers constraints; nearly every clause carries non-redundant information. It is a dense single paragraph that would scan faster as a bulleted list, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, multi-op, concurrency-sensitive mutation with no output schema, the description covers the operation grammar, atomicity, revision threading, and all refusal conditions an agent must anticipate. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: parentId null means top level, position controls ordering, update accepts name and/or description, children are nested under add, and remove cascades. It stops short of explaining the revision/position interaction or id format, so it is not a full 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ("Change a project's features by id") and then enumerates all four operations (add/update/move/remove) with their semantics, making it immediately distinguishable from the sibling read tool get_features.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes to get_features for the baseRevision, explains how to split a large restructure into sequential calls each using the revision returned by the previous, and states three concrete refusal conditions (stale writes, hands-off subtrees, features open for browser editing). This is genuine when/when-not guidance with alternatives named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_handoff_noteUpdate handoff noteA
Destructive
Inspect

Replace the handoff note on the Work being implemented, so another session can continue it: what is done, what remains, files touched, and whether changes are uncommitted. begin_work returns it to whoever resumes the Work; finish_work clears it. Pass an empty note to clear it.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteYesThe Handoff note text.
workIdYesWork being implemented.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true; the description is consistent with that, reinforcing 'Replace'. It adds genuinely useful behavior beyond the annotations: the note's expected content categories and that an empty string clears it, plus the interaction with begin_work/finish_work.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core verb+resource, then purpose, then related-tool behavior. Three compact sentences; the only mild redundancy is restating the clearing behavior twice (empty note, finish_work).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but the annotations carry the safety profile and the schema carries parameter constraints (maxLength, required). The description fills the remaining gaps: note content and clearing semantics. Adequate for a two-parameter mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description adds real meaning: it defines what the note should contain (done, remaining, files touched, uncommitted status) and the empty-note-as-clear semantics – behavior not captured by the schema's terse parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Replace) and resource (the handoff note on the Work being implemented), plus the purpose (so another session can continue it). An agent can distinguish this from update_work or finish_work without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the use context (continuation across sessions) and names related tools: begin_work returns the note, finish_work clears it, and passing an empty note clears it directly. No explicit 'when not to use' exclusions, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_ideaUpdate ideaA
Destructive
Inspect

Update an idea's short name of at most 20 characters, optional detail, or impact. Nothing is generated. Only supplied fields change. Order is separate: move an idea with reorder. Archiving, declining, and restoring Ideas don't run rules.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Idea.
titleNoNew short name, at most 20 characters.
impactNoProduct scope of the Idea: minor or major.
contentNoNew full detail.
summaryNoNew summary.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=false and destructiveHint=true already declared, the description adds genuine behavioral context: 'Only supplied fields change' discloses partial-update/PATCH semantics, and 'Nothing is generated' tells the agent no rules or derived output are triggered. It stops short of explaining the destructive implication flagged by the annotation or what the call returns, but it meaningfully enriches the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the primary action and kept to a handful of short sentences with no filler. The phrasing is slightly clipped and telegraphic ('short name of at most 20 characters, optional detail, or impact'), but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter update tool with full schema coverage and no output schema, the description supplies the key missing pieces: partial-update semantics, no side-effect generation, and routing to reorder for ordering. It leaves the destructive hint unexplained, but otherwise covers what an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented, including the 20-character limit on title and the minor/major enum for impact. The description restates those fields but adds no format, constraint, or interaction detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Update) and resource (idea), and enumerates the exact mutable fields (short name, detail, impact). It actively differentiates from siblings by routing order changes to reorder and noting archive/decline/restore as separate operations, so an agent can distinguish it from the surrounding update_* and idea tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description names an explicit alternative for a closely related operation ('Order is separate: move an idea with reorder'), giving clear routing guidance. It does not, however, state when to prefer update_idea over get_idea or how it relates to the other update_* tools, so the guidance is real but partial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_instructionUpdate instructionA
Destructive
Inspect

Rename an instruction document or rewrite its text whole. Only supplied fields change; content replaces the entire text, and an empty string clears it. Earlier text stays in the document's revision history. Read current text from its pizzadeveloper://projects/{projectId}/instructions/{id} resource first. During an active workflow run, instruction changes are accepted only in its self-improvement stage.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Instruction.
titleNoNew title.
contentNoFull replacement text; an empty string clears it.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description discloses rich behavior: only supplied fields change, content is a full replacement, an empty string clears the text, and prior text is retained in revision history. The workflow-stage acceptance rule is a constraint no annotation conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core mutation, then replacement semantics, history behavior, prerequisite read, and the workflow constraint. No sentence is filler and the ordering prioritizes what affects the call.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the prerequisite read, the replacement/clearing behavior, reversibility (revision history), and the environmental timing constraint. Nothing an agent needs in order to call it correctly appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents id, title, and content including the empty-string-clear semantics. The description's 'only supplied fields change' adds partial-update meaning, but overall the schema does the heavy lifting, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb-and-resource pair ('Rename an instruction document or rewrite its text whole') and makes the partial-update semantics explicit. An agent can distinguish this from create_instruction, delete_instruction, and archive_instruction without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear prerequisite ('Read current text ... resource first') and a timing constraint (mutations accepted only during a workflow run's self-improvement stage). It does not, however, explicitly name the read alternative or contrast with sibling update tools, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_projectUpdate projectA
Destructive
Inspect

Update a project's own details. Only the fields you provide change. Owner-only. Instruction documents change through update_instruction.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Project.
nameNoNew Project name.
tagsNoProject tags.
liveUrlNoLive URL of the Project.
repoUrlNoRepository URL of the Project.
descriptionNoNew Project description.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the safety profile is covered. The description adds value beyond that: it discloses patch/partial-update behavior (omitted fields are left untouched) and an authorization requirement (Owner-only). It does not describe the response or side effects of the destructive flag, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each carrying distinct information (scope, patch behavior, permission + alternative), with the core purpose front-loaded. Nothing is repeated from the name, title, or schema, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with full schema coverage, an owner-only constraint, and no output schema, the description covers scope, partial-update behavior, permissions, and the instruction alternative. The main residual gap is any note on what a successful update returns or whether changes are reversible, which the destructiveHint alone does not clarify.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters, making 3 the baseline. The description adds the meaningful patch semantic that only supplied fields are modified, but contributes no per-parameter detail (e.g., URL formats for liveUrl/repoUrl) beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Update a project's own details') and explicitly scopes it against the nearest sibling by routing instruction-document changes to update_instruction. It stops short of naming other update_* siblings, but none of those operate on projects, so the risk of confusion is low.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real selection guidance: partial-update semantics ('Only the fields you provide change'), a permission prerequisite ('Owner-only'), and an explicit alternative for a related case ('Instruction documents change through update_instruction'). No explicit when-not-to-use, but the alternative routing covers the main ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_public_delivery_pageUpdate Public Delivery PageA
Destructive
Inspect

Set whether the project's Public Delivery Page shows dates and configure supported social profiles: YouTube, X, Discord, LinkedIn, Instagram, TikTok, Bluesky, Threads, Twitch, Spotify and Apple Music. Owner-only. Read current values with get_project. Omitted fields stay unchanged; empty strings remove profiles. Custom links are not supported. Upload or remove logo/cover images in the app (PNG/JPG, 4 MB); image references cannot change here. Does not publish the Public Delivery Page.

ParametersJSON Schema
NameRequiredDescriptionDefault
xUrlNoX profile URL.
profileNoOther page settings: images, dates, and social profile URLs.
projectIdYesId of the Project.
youtubeUrlNoYouTube profile URL.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, so the mutation profile is known, but the description adds substantial value beyond that: partial-update semantics ('Omitted fields stay unchanged'), removal behavior ('empty strings remove profiles'), and the fact that image references cannot be changed here. It does not mention reversibility or rate limits, so 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and field scope, then constrained by owner-only, partial-update, and non-publish caveats. It is dense and the 11-platform enumeration is long, but every clause carries operational weight; only the platform list borders on padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers permissions, partial-update and removal semantics, unsupported cases, and where to do related work (images in the app, reads via get_project). Nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage the baseline is 3, but the description adds semantics the schema lacks: the update is partial (omitted fields unchanged), empty strings clear a profile, and logo/cover references are immutable through this tool. These go beyond the terse per-field descriptions in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Set whether the project's Public Delivery Page shows dates and configure ... social profiles') and enumerates exactly what is configurable. It clearly distinguishes this tool from siblings like update_project or update_delivery, which touch different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the prerequisite ('Owner-only'), routes the agent to get_project for reading current values, and explicitly sends image upload/removal to the app. It also scopes the tool with 'Custom links are not supported' and 'Does not publish the Public Delivery Page', giving clear when/when-not boundaries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_ruleUpdate ruleA
Destructive
Inspect

Edit a rule's text or level. Only supplied fields change.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Rule.
textNoThe Rule as one narrow statement.
levelNomust refuses a failing write; should only reports it.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the agent knows this is a mutating, potentially destructive call. The description adds real value by disclosing patch semantics (only supplied fields change), but it does not mention permissions, reversibility, or the effect of flipping a level from 'should' to 'must'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and immediately followed by the most decision-relevant behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a small mutation tool with complete schema and annotations covering the safety profile, the description is adequate but thin; it omits any note on what happens to omitted fields beyond the patch hint, error behavior for an unknown id, and whether the change is reversible.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including a detailed enum explanation for level, so the schema carries the parameter burden and the baseline is 3. The patch-semantics sentence complements the schema by explaining that id is the only required field and others are optional, but adds no syntax or constraint detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Edit) and resource (rule) plus the two mutable fields (text, level), which cleanly separates it from create_rule and delete_rule. It does not explicitly contrast with the other update_* siblings, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Only supplied fields change" is genuinely useful usage context for a partial-update call, but there is no guidance on when to reach for this versus create_rule/delete_rule or on prerequisites such as the rule needing to exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_rulebookUpdate rulesetA
Destructive
Inspect

Rename a ruleset. Attach or detach it with rulebookIds on a workflow connection in update_workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Ruleset.
nameYesNew Ruleset name.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the agent knows this is a mutating, potentially destructive operation. The description adds the useful scope fact that attach/detach is handled elsewhere, but it never explains why a rename is flagged destructive or what happens to workflows already referencing the ruleset.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the core action is front-loaded in the first sentence. The routing information follows immediately and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema needed, two fully documented required parameters, and annotations covering the mutation profile, the description is nearly sufficient for calling the tool correctly. It could say more about name-uniqueness failures or effects on referencing workflows, but nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both required parameters (id, name) are documented in the schema, so the baseline is 3. The description mentions rulebookIds, but that is a parameter of a different tool, adding nothing about the two parameters this tool actually takes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ("Rename a ruleset") and corrects the broader tool name/title by scoping it to renaming only. It also distinguishes itself from create_rulebook, delete_rulebook, and list_rulebooks by naming what it does not do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the other rulebook-related operation (attach/detach) to update_workflow with rulebookIds, which is genuine when-to-use guidance. It stops short of stating prerequisites or when to prefer this over archive/delete, so it is clear context without full exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_workUpdate workA
Destructive
Inspect

Edit a Work's planning details. Only supplied fields change. batchId deliberately moves a planned Work. status accepts only planned, to send a started Work back to the queue; a Work starts only through begin_work and finishes only through finish_work. In a workflow run, a change to the title, description, labels, size, or Batch is checked as part of that Batch's plan against the rulesets on the connection the run entered Batch on, until one of its Works is done; a Must rule that does not pass refuses the write.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesId of the Work.
sizeNoWork size.
titleNoNew title.
labelsNoLabels on the Work.
statusNoOnly planned, to send a started Work back to the queue.
batchIdNoBatch to move a planned Work to.
descriptionNoNew description.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only declaring readOnlyHint=false and destructiveHint=true, the description carries real behavioral weight: it discloses partial-update semantics, the side effect of batchId (deliberately moving a planned Work), the rule-checking pass against Batch rulesets during an active workflow run, and the refusal behavior when a Must rule fails. That is well beyond what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the partial-update rule, then layers constraints. The final sentence about Batch ruleset validation is information-dense and takes parsing, but every clause carries distinct, non-redundant information, so nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no output schema, it covers partial updates, the status constraint, batch movement, and rule-refusal behavior. It omits what a successful call returns and how label/size changes interact with downstream state, but the critical call-time risks are addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, establishing a baseline of 3, but the description adds meaning the schema lacks: the const-constrained status is explained by intent (rewind a started Work to the queue) and batchId is clarified as a deliberate move of a planned Work. It just stops short of describing the title/description/labels/size update nuances.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Edit a Work's planning details') and immediately frames the scope as partial-update ('Only supplied fields change'). It distinguishes itself from the sibling mutators begin_work and finish_work by explicitly declaring that those are the only paths by which a Work starts or finishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: status here only accepts 'planned' to send a started Work back to the queue, while starting and finishing go through begin_work and finish_work respectively. This names the alternatives and the exact conditions that select them rather than leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_workflowUpdate workflowA
Destructive
Inspect

Edit one workflow. Pass nodes and edges together to replace its graph, and name or description to change its details; omitted parts stay as they are. Nodes are Start, optional Idea, Batch, optional Delivery, optional Workflow (the step that improves the Workflow itself after a run), and Finish. Start can connect to an existing Batch without Idea, and Batch can connect to Workflow or Finish without Delivery. Start can also connect to Delivery, for a workflow that only creates or adds to a Delivery: Start to Delivery to Finish, or Start to Delivery to Workflow to Finish, with no Batch node. Each edge.data must contain instructionDocumentIds (existing project document IDs) and contextKinds (features and/or dictionary; each adds only a short outline, top-level feature names or dictionary words, that the agent expands with the read tools). edge.data may also carry rulebookIds, an ordered list of the project's ruleset IDs attached to that connection; leaving it out attaches none. Select documents on connections; do not create Instructions, Dictionary, or Features nodes or put text on lines. Updates made during a run's self-improvement stage are described in the report that closes that step. Open runs keep their frozen revision. Graph writes are refused while a vibe coder has the workflow open for editing in the browser, and archived workflows must be restored first.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoNew Workflow name.
edgesNoReplacement connections; pass together with nodes.
nodesNoReplacement stages; pass together with edges.
projectIdYesId of the Project.
workflowIdYesId of the Workflow.
descriptionNoNew description of the requests it handles.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=true; the description adds substantial context beyond that: partial-update semantics (omitted parts unchanged), concurrency refusal while open in browser, the archived-must-be-restored prerequisite, frozen revisions on open runs, and self-improvement updates surfacing in a report. This is rich behavioral disclosure an agent needs to call it safely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action and update model are front-loaded in the first sentence. The dense paragraph is long but each clause carries non-redundant meaning given the empty item schemas; a bulleted breakdown of node types and edge.data would improve scannability without cutting content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex graph-mutation tool with no output schema and empty item schemas, the description supplies the graph grammar, edge payload contract, valid topologies, and all write-precondition behavior. Nothing essential to invoking it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% but the nodes/edges items are untyped empty objects, so the schema provides almost no structural meaning. The description compensates heavily by defining the node vocabulary (Start, Idea, Batch, Delivery, Workflow, Finish) and the required edge.data fields (instructionDocumentIds, contextKinds, optional ordered rulebookIds), plus the valid connection topologies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ("Edit one workflow") and immediately states the update model: nodes/edges replace the graph, name/description change details, omitted parts stay. An agent can distinguish this from get_workflow, delete_workflow, and archive_workflow without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete when-to-use conditions (pass nodes+edges together to replace the graph; pass name/description to rename) and explicit when-not cases (refused while a vibe coder has it open, archived workflows must be restored first). It also sets exclusions ("do not create Instructions, Dictionary, or Features nodes or put text on lines"), though it never names a sibling tool as an alternative for those cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 77 tool updates
    • First observedarchive_delivery
    • First observedarchive_dictionary_term
    • First observedarchive_idea
    • First observedarchive_instruction
    • First observedarchive_project
    • First observedarchive_workflow
    • First observedassign_batch_to_delivery
    • First observedbegin_work
    • First observedcancel_workflow
    • First observedcheck_rules
    • First observedcreate_batch_plan
    • First observedcreate_delivery
    • First observedcreate_dictionary_term
    • First observedcreate_idea
    • First observedcreate_instruction
    • First observedcreate_project
    • First observedcreate_rule
    • First observedcreate_rulebook
    • First observedcreate_work
    • First observedcreate_workflow
    • First observeddecline_idea
    • First observeddelete_batch
    • First observeddelete_delivery
    • First observeddelete_delivery_presentation
    • First observeddelete_dictionary_term
    • First observeddelete_idea
    • First observeddelete_instruction
    • First observeddelete_project
    • First observeddelete_rule
    • First observeddelete_rulebook
    • First observeddelete_work
    • First observeddelete_workflow
    • First observedfinish_work
    • First observedget_batch
    • First observedget_delivery
    • First observedget_delivery_presentation
    • First observedget_dictionary_term
    • First observedget_features
    • First observedget_idea
    • First observedget_project
    • First observedget_project_state
    • First observedget_work
    • First observedget_workflow
    • First observedlist_batches
    • First observedlist_deliveries
    • First observedlist_dictionary_terms
    • First observedlist_ideas
    • First observedlist_projects
    • First observedlist_rulebooks
    • First observedlist_workflows
    • First observedreorder
    • First observedreport_problem
    • First observedreport_workflow_step
    • First observedrestore_idea
    • First observedresume_workflow
    • First observedrevert_batch
    • First observedsearch_works
    • First observedstart_workflow
    • First observedunarchive_delivery
    • First observedunarchive_dictionary_term
    • First observedunarchive_instruction
    • First observedunarchive_project
    • First observedunarchive_workflow
    • First observedupdate_batch
    • First observedupdate_delivery
    • First observedupdate_delivery_presentation
    • First observedupdate_dictionary_term
    • First observedupdate_features
    • First observedupdate_handoff_note
    • First observedupdate_idea
    • First observedupdate_instruction
    • First observedupdate_project
    • First observedupdate_public_delivery_page
    • First observedupdate_rule
    • First observedupdate_rulebook
    • First observedupdate_work
    • First observedupdate_workflow

Publisher details

Operator
Pizza Developer
Vendor relationship
First-party
Restrictions
Requires a Pizza Developer account. Every account starts with a free trial, with no payment card needed.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    YOUR AGENTS COLLIDE. File a FlightPlan. Run multiple coding agents without them stepping on each other. FlightPlan gives each agent the same preflight picture, lets them coordinate while plans are still cheap to change, and leaves behind what changed and why for whichever agent comes next. Across sessions. Across people. Across providers. Across time. One shared picture.
    6
    5
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Serves instructions and templates that make any AI coding tool plan before building, work through a project checklist in the right order, and hold a consistent standard, while all reading and writing of your code stays on your machine.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Connects Claude Code, Codex, Cursor, and other coding agents to a project's idea, feature topology, and ordered task queue over MCP, so they can answer planning jobs, claim, build, and tick off implementation tasks live in your repository, or assemble a shareable PRD.
    1 npm
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Routes each coding task to the best-fitting spec-driven-development framework with explainable rules, then enforces deterministic lifecycle gate checks between phases so progress can't skip required artifacts or evidence. Assembles per-phase context packs from a company knowledge base with app-scoped memory and cross-app lookup, while the host agent remains the only actor that edits files or runs tests.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources