Skip to main content
Glama

Server Details

Build and publish static sites on Dotsy with draft-first MCP tools, OAuth Connect, and API keys.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

B3.1/5.0

Scored across 89 tools

Disambiguation2/5

The 89-tool set creates many overlapping clusters, especially around content staging/editing (put_content, patch_content, batch_patch_content, bulk_put_content, apply_changeset, stage_content_chunk, edit_page) and route/menu/domain management. Descriptions do clarify thresholds, but the agent must parse dense guidance to choose correctly, leaving several boundaries unclear.

Naming Consistency4/5

Most names follow a readable snake_case action-object pattern (get_site, put_content, delete_route, list_collections). A few prefixed exceptions like adapt_site_analyze, structure_site_apply, and get_structure_status deviate from strict verb_noun style, but the overall convention is predictable.

Tool Count1/5

At 89 tools, the server exposes an entire platform in one MCP surface. This is far beyond a well-scoped toolset and creates heavy selection overhead, context bloat, and risk of misselection.

Completeness4/5

Coverage is extensive: site lifecycle, import/adapt/structure, content CRUD, routes, menus, assets, domains, experiments, analytics, collections, forms, and publishing are all represented. Some subdomains lack full CRUD (e.g., collection create/delete, direct asset update), but the major workflows are covered.

Available Tools

89 tools
adapt_site_analyzeA
Read-onlyIdempotent
Inspect

Propose nav wiring + blog collectionization for an imported mirror (draft, opt-in). Read-only; returns per-adapter reports + proposalEtag. Nav decline may include nextStep (e.g. mega-menu → edit fragments/header.html). Wait for mirror-vs-editable answer if requiredDecision pending.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoReturn the full per-member harvest report (default: compacted for the token budget).
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
optionsNo
adaptersNoWhich adapters to analyze (default: both).
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
refreshProposalNoReplace an existing UNCONSUMED proposal (otherwise re-analyze 409s proposal_exists).

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/non-destructive, so the bar is lower, and the description still adds real behavior: returns per-adapter reports plus a proposalEtag, nav declines may carry nextStep (e.g. mega-menu → edit fragments/header.html), and it must wait on a pending requiredDecision. That is meaningful context beyond the structured hints, though it doesn't cover rate limits or the 409/proposal-exists path beyond the param.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three telegraphic sentences that front-load the core action, then return shape, then the decision-gating constraint. Dense but every clause carries information; the terse style borders on cryptic but wastes little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 7-param tool with a nested options object and no output schema, the description does address both the return payload (per-adapter reports, proposalEtag) and the blocking requiredDecision state. It is largely complete, with the main gap being the relationship to the apply/revert siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents nearly every parameter (full, site, context, options.*, guestToken, refreshProposal). The description adds almost no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (propose/analyze) and resource (nav wiring + blog collectionization) scoped to an imported mirror, and marks it as a draft/opt-in proposal, which distinguishes it from adapt_site_apply. The jargon 'collectionization' is a minor clarity cost, and it doesn't explicitly name its sibling apply/revert tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'draft, opt-in' and 'wait for mirror-vs-editable answer if requiredDecision pending' imply a pre-apply workflow, but the description never names adapt_site_apply or adapt_site_revert as the next/alternative steps, nor states explicit when-not conditions. Usage context is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

adapt_site_applyA
Destructive
Inspect

Apply the adapt proposal (draft): wire nav into the menu store and/or rewrite the blog index to a live . One undo covers the whole apply; a second apply is refused until adapt_site_revert (or force). Queued applies are polled to completion.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
forceNoAn undo record from a prior apply exists (409 adapt_undo_exists): defer-discard it — the prior record is dropped only when this apply SUCCEEDS; a partial-failure rollback reinstates it. Prefer adapt revert first.
dryRunNoFull computation, would-write set + previews, ZERO writes, no undo.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
ifMatchYesThe proposal etag from analyze (or GET …/adapt/proposal / status).
adaptersYesWhich analyzed adapters to apply.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint=true, idempotentHint=false) by disclosing that one undo covers the whole apply, that a repeat apply is refused until revert/force, and that queued applies are polled to completion. These are exactly the behavioral traits an agent needs before mutating.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, action and scope front-loaded, followed by the undo/refusal guardrail and the polling behavior. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive 7-parameter mutation tool with no output schema, the description supplies the missing operational picture: what is changed, the undo model, the refusal-and-force escape hatch, and how queued applies complete. Nothing essential to correct invocation is omitted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema carries the parameter documentation and the baseline is 3. The description nevertheless adds cross-parameter meaning: it explains the force/revert relationship and the undo lifecycle that frames ifMatch and adapters, which the schema only states locally.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Apply) and resource (the adapt proposal/draft), then enumerates the concrete effects: wiring nav into the menu store and/or rewriting the blog index to a live <dotsy-collection>. This clearly differentiates it from adapt_site_analyze (which produces the proposal) and adapt_site_revert (which undoes it).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real when-to-use context: a second apply is refused until adapt_site_revert or force, and the force param description itself says 'Prefer adapt revert first'. It does not explicitly route the agent to dryRun or analyze in prose, but the conditions around refusal and undo are stated concretely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

adapt_site_revertA
Destructive
Inspect

Atomically restore everything the last adapt_site_apply wrote (fragments, the menu doc when the op touched it, index, member configs, routes).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered. The description adds real value beyond that: 'Atomically' discloses all-or-nothing semantics, the parenthetical enumerates the exact scope of mutation, and 'the last' implies only the most recent apply is revertible — a meaningful non-idempotency constraint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the atomic-restore action and the target, with the scope enumeration in a compact parenthetical. No filler or restated title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers what gets restored and that the operation is atomic, and annotations carry the safety signals. It omits the failure mode when no prior apply exists and whether the revert itself can be undone, which keeps it short of full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so site, context, and guestToken are all documented in the schema itself. The description adds no parameter-level detail (e.g., which site scope the revert targets), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restore/revert) and precisely scopes the resource: everything the last adapt_site_apply wrote, with an enumeration of the artifacts (fragments, menu doc, index, member configs, routes). This clearly distinguishes it from sibling adapt_site_apply and from the narrower revert_change.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'the last adapt_site_apply' implies the usage context (undo the most recent apply), but the description never explicitly states when to use this versus revert_change, discard_drafts, or what happens if no prior apply exists. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_custom_domainBInspect

Connect apex and/or www on a paid plan; starts a DNS connect session.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
hostsNo1–2 hostnames: apex (example.com) and/or www (www.example.com).
domainNoBare domain when hosts omitted — apex expands to apex + www; a subdomain connects as-is.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-read-only, non-idempotent, open-world write. The description adds two useful facts beyond them: the paid-plan requirement and that it initiates a DNS connect session (implying a multi-step flow). It does not disclose that the connect is asynchronous or what the completion step is, so the picture remains partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the plan constraint and the session outcome front-loaded; every clause earns its place. It is arguably over-terse for a mutation with follow-up steps, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 100% schema coverage and annotations carrying the safety profile, the definition is callable as-is. However, for a domain-connect flow the agent lacks the crucial next-step and polling context, and no output schema exists to fall back on, leaving the workflow incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: site, hosts, domain, context, and guestToken are each fully documented in the schema, including the hosts-vs-domain fallback and the guestToken lifecycle. The description adds no parameter-level meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (connect apex and/or www) and the initiating effect (starts a DNS connect session), which distinguishes it from read-side siblings like get_custom_domain_connect and list_custom_domains. It stops short of explicitly differentiating from refresh_custom_domain, reapply_custom_domain_oauth, or set_primary_custom_domain, so it is clear but not sibling-aware.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"on a paid plan" gives one precondition, but there is no guidance on when to prefer this over refresh/reapply/set_primary, nor any statement of what the caller must do next (e.g., wait_for_custom_domain). Usage is implied by the action rather than explained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

apply_changesetA
Destructive
Inspect

Atomic writes[] + deletes[]; attach routes via writes[].route. At most 10 writes; each body ≤60KB and the call ≤80KB. Larger new pages: stage_content_chunk then writes[].bodyStaging, or dotsy push / PUT content-raw.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
writesNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
deletesNo
messageNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds meaningful context beyond them: the operation is atomic, it is capped at 10 writes, and each body/call has 60KB/80KB size ceilings. It still omits auth/guestToken requirements and what a failed batch leaves behind, but the constraint disclosure is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the core operation and then the limits and the large-page escape hatch. Nothing is padding, though the telegraphic 'dotsy push / PUT content-raw' shorthand assumes prior knowledge and reads more like a note than an instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent batch mutation with six parameters and deep nested objects, the description covers atomicity and size limits but not permissions, partial-failure behavior, or the required context/guestToken flow. The rich input schema carries much of the load and there is no output schema to explain, but the behavioral envelope is only partially closed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 50% schema description coverage, the description must compensate and it partially does: it explains writes[].route for attaching routes and writes[].bodyStaging for staged bodies. It says nothing about deletes[], message, context, or guestToken, and does not explain the key/body fields, leaving the remaining undocumented half of the schema unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the operation concretely: an atomic set of writes[] plus deletes[], with routes attached via writes[].route. An agent can tell it performs a batched mutation rather than a single put/patch, though it never names the sibling tools (batch_patch_content, bulk_put_content) it overlaps with, so differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real operational guidance for the large-page case (stage_content_chunk then writes[].bodyStaging, or dotsy push / PUT content-raw), which is a genuine routing hint. However there is no explicit statement of when to prefer apply_changeset over batch_patch_content, bulk_put_content, or put_route, nor any when-not condition, so usage is implied rather than specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch_patch_contentA
Destructive
Inspect

patch_content on multiple keys; all-or-nothing validation. Preferred edit path for pages over 60KB. replace-string value/replace ≤60KB; whole call ≤80KB.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
itemsYes
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=false, so safety is covered. The description adds genuinely new behavioral facts: all-or-nothing (atomic) validation and hard size ceilings (replace-string ≤60KB, call ≤80KB), which an agent cannot derive from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact clauses with no filler, and the defining scope ('on multiple keys') is front-loaded. The phrasing is telegraphic, but every clause carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and for a destructive batch mutation the key concerns — atomicity, size limits, scope — are addressed, so an agent can call it correctly. It leaves per-op semantics to the schema enum, which is a reasonable delegation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema carries most parameter meaning. The description adds quantitative constraints tied to the replace-string op and the overall call size, which is useful, but it does not clarify any of the eight op enums or the expect/find/selector fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('patch_content on multiple keys'), which immediately distinguishes it from the sibling patch_content and from bulk_put_content/find_replace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear selection rule for the sibling: 'Preferred edit path for pages over 60KB,' implying plain patch_content handles the smaller cases. No explicit when-not or prerequisite/auth guidance, but the routing signal is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bulk_put_assetsA
DestructiveIdempotent
Inspect

Upload 1–200 CMS assets; key = path after /assets/. Optional bodySha256 verifies submitted bytes before write; response sha256 is stored bytes (JS/CSS match; WebP rasters may differ). Uploaded files are live in the editor and draft preview at once; the published site picks them up at the next publish.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
itemsYes
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: it discloses that bodySha256 is verified before any write with a named failure mode (body_sha256_mismatch), and explains that the response sha256 is the stored blob, which can differ from submitted bytes for WebP rasters after IMAGES optimization. It also states the immediate-liveness rule (editor and draft preview now, published site at next publish), which is exactly the kind of effect an agent needs and annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences that front-load the action and scope, then layer verification and publication semantics. No filler, though the semicolon-chained second sentence packs three distinct ideas and could be split for scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, so the description carries return semantics and it does (sha256 of stored bytes, mismatch error). The notable remaining gap is batch failure behavior: with 1–200 items it never says whether a single bad item aborts the whole upload or writes partially, which an agent must know before committing a large batch.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description supplements it meaningfully: it clarifies the key prefix rule and, more importantly, disambiguates submitted-bytes sha256 from stored-bytes sha256, a distinction the schema only hints at. It adds little for contentType or importFromUrl, keeping it short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Upload 1–200 CMS assets.' The key semantics ('key = path after /assets/') and the batch limit make it instantly distinguishable from get_asset, list_assets, and delete_asset, which are the only nearby siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the 1–200 batch bound and the CMS-asset framing, and the publish-timing sentence tells the agent why it might care. But it never states when to reach for this over another write tool (e.g. put_content), nor any precondition beyond the schema's guestToken field.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bulk_put_contentA
DestructiveIdempotent
Inspect

Stage many content keys (pages, layouts, fragments). Content only — not routes or assets. Each body ≤60KB and the whole call ≤80KB; larger pages use stage_content_chunk, dotsy push, or PUT content-raw. May queue above the inline threshold.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
titleNo
contentYes
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so safety is partly carried. The description adds non-obvious behavior: hard payload limits (60KB per body, 80KB total) and that large calls may queue above an inline threshold. It does not explain what happens to pre-existing content at a staged key.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with what the tool does and its content-only scope, then constraints, then alternatives. Dense and mostly waste-free, though 'May queue above the inline threshold' is vague about where that threshold sits.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters, no output schema, and no enums, the description covers the operation, its scope, its size constraints, and escape hatches for oversized payloads. Missing only the outcome semantics of staging/queuing and overwrite behavior for existing keys.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60% and the nested content item fields (key, body, contentType) have no schema descriptions, so the description compensates by stating the per-body and per-call size limits and the kind of keys accepted. It says nothing about the required 'context' or 'guestToken' fields, but those are documented in the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Stage many content keys') and immediately scopes what is included ('pages, layouts, fragments') versus excluded ('not routes or assets'), which lets an agent separate it from bulk_put_assets and put_route without opening schemas. It does not explicitly distinguish itself from batch_patch_content or put_content, which is the only gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit routing for an important case: bodies over 60KB or calls over 80KB should use stage_content_chunk, dotsy push, or PUT content-raw. It also hints at the inline-threshold/queueing condition. It lacks guidance on when to prefer this over single-key put_content or batch_patch_content.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

claim_siteAInspect

Move the guest draft onto your signed-in Dotsy account. Requires OAuth Connect; pass the guestToken from start_site.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenYesGuest token from start_site (required). Pass the pgg_ value returned by start_site.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-read-only, non-idempotent, non-destructive, closed-world; the description adds the OAuth Connect auth requirement, which is behavioral context the annotations do not convey. It stops short of saying what happens to the guest draft or whether a repeat call fails or re-claims, which matters given idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences: the action is front-loaded and the prerequisite/parameter hint follows. Nothing is redundant or padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema, the description covers purpose, auth prerequisite, and token source. It omits the post-call outcome (what the claimed site looks like, error behavior on re-claim), which is a minor gap given non-idempotent semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are fully documented in the schema, including the 'one sentence, publicly visible' constraint on context and the pgg_ prefix on guestToken. The description's pointer to start_site for the token largely repeats the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb+resource: moving a guest draft onto the signed-in account, which clarifies the otherwise cryptic name 'claim_site'. It implicitly distinguishes itself from siblings like start_site (which creates the guest token) and create_site/graduate_site, though it never names an alternative outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states a clear precondition ('Requires OAuth Connect') and routes the agent to start_site as the source of the guestToken, which is genuinely useful invocation guidance. It does not spell out when NOT to call it or name a competing sibling for the same task.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_siteAInspect

Create empty site; seeds brief. Globally taken slugs auto-suffix to {slug}-{adjective}-{noun} (default); pass onSlugTaken:"fail" to get a 409 instead of a renamed site.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYes
titleNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
purposeNo
workspaceNoWorkspace org id or slug (see get_account_context memberships). Only for multi-workspace users; omit otherwise. Ignored for pga_ keys (fixed workspace).
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
editingModeNoHow the editor's Save behaves. 'live' (default) publishes each save immediately; 'staged' keeps saves as drafts until publish_site. Prefer 'staged' when an agent will hand the editor to a person (see start_edit_session → draftEditorUrl) so keystrokes don't go straight to the public site. API writes always stage regardless of this setting.
onSlugTakenNoauto-suffix
keyConventionNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare it is a non-readOnly, non-idempotent, non-destructive, non-openWorld create operation. The description adds meaningful behavior beyond that: globally taken slugs are auto-suffixed to a slug-adjective-noun pattern, and onSlugTaken:'fail' returns a 409 instead of renaming. It does not cover permissions or the guestToken/claim lifecycle, but the auto-suffix and 409 semantics are real added context on top of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences front-load the core action and then qualify slug-collision behavior. There is some density with the inline code and default annotation, but no wasted filler. It is appropriately sized for a create tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description needn't explain returns. But for a 9-parameter create tool with only 44% schema coverage, the description omits guidance on guestToken/claim_site flow, workspace selection, editingMode, and keyConvention, which are relevant to calling it correctly. It covers slug behavior well but is incomplete overall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 44%, below the 80% threshold, so the description should help compensate. It only adds meaning for one parameter (onSlugTaken) by explaining the fail/409 behavior. The other eight parameters, including required slug and context plus guestToken, editingMode, keyConvention, and workspace, receive no explanation in the description even though several of those schema entries do carry their own descriptions. This leaves a partial gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create empty site') and adds a notable side effect ('seeds brief'). This distinguishes it from sibling create-like tools such as start_site or import_site_from_url, though it never explicitly names those alternatives. The core purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is the entry point for making a site, and the onSlugTaken guidance ('pass onSlugTaken:"fail" to get a 409') gives conditional usage for one parameter. However, it does not say when to prefer this over start_site or import_site_from_url, nor does it state prerequisites like guestToken or claim_site flow. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decompose_analyzeB
Idempotent
Inspect

Deprecated — use structure_site_analyze.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoReturn the full manifest incl. per-page HTML samples (default: compacted — etag/action/counts kept, samples truncated — to fit the token budget).
siteYesSite slug, e.g. acme-studio
scopeNoDefault flat-only; analyze clusters flat keys only.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
persistNoFirst call: false (preview only). true only after user confirms editable/CMS mode.
profileIdNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
sampleLimitNo
refreshManifestNoReplace persisted manifest when true.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=false. The description adds one genuinely non-structured fact — that the tool is deprecated — which is real behavioral context, but discloses nothing about execution, auth, or failure mode.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, and the deprecation flag leads before the redirect. Exactly the right size for a tombstone definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description is minimal, but deprecation justifies brevity. It omits whether calls still succeed, whether results remain valid, and when the tool will be removed — details an agent choosing between this and structure_site_analyze would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, so the schema documents most of the 9 parameters itself. The description contributes zero parameter-level information, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description never states what the tool does; it only declares deprecation status and names a replacement. The purpose is recoverable from the name and title ('Analyze site structure'), but the description text itself supplies no verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'use structure_site_analyze' is an explicit, unambiguous routing instruction that names the alternative, which is the single most important usage fact for a deprecated tool. It stops short of saying whether this tool still functions or errors if called.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_assetA
DestructiveIdempotent
Inspect

Delete a static asset by key. The editor and draft preview drop it at once; the published site drops it at the next publish.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesPath after /assets/ on the site — e.g. photos/hero.webp or style.css. Do not prefix with assets/ or /. Uploaded files are live in the editor and draft preview at once; the published site picks them up at the next publish.
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so safety is covered structurally. The description adds genuinely non-obvious behavior: the editor and draft preview drop the asset immediately while the published site only drops it at the next publish. It omits auth requirements (guestToken is left to the schema) and error semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The deletion action and its scope are front-loaded ahead of the propagation caveat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive 4-parameter mutation tool with rich annotations and a fully documented schema, the description covers the essential non-obvious trait (split-brain propagation between draft and published). It does not mention guestToken auth flow or failure behavior, but those are largely schema- and annotation-territory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (key, site, context, guestToken) are already documented in detail, including examples. The description adds nothing beyond the schema's own key-format and propagation notes, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete a static asset') plus the addressing mode ('by key'). This distinguishes it from delete_content and bulk_put_assets, whose names imply different resources and operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus delete_content, bulk_put_assets, or get/list_assets. The only usage-like signal is the propagation note, which describes effect rather than selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_contentB
DestructiveIdempotent
Inspect

Delete content key (tombstone if published).

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so safety is covered. The description does add real value by disclosing that published content is tombstoned rather than hard-deleted, but says nothing about reversibility of that tombstone or permission/auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the destructive scope front-loaded and no filler. It earns its brevity, though it borders on under-specification for a destructive operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, four-parameter tool with no output schema, the definition covers the key behavioral nuance but omits what happens to unpublished content, whether the tombstone can be restored, and the guestToken auth flow. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, with the 'key' parameter undocumented in the schema, and the description only names it rather than clarifying its format or lookup behavior. site, context, and guestToken are covered by the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delete) and resource (content key), which separates it from the many delete_* siblings like delete_asset, delete_site, and delete_route. The parenthetical adds scope behavior, though 'content key' is jargon an outside agent must infer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives such as discard_drafts, revert_change, or patch_content. The only hint is the tombstone clause, which is behavioral rather than conditional routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_draft_testB
DestructiveIdempotent
Inspect

Remove a draft A/B test (releases variant bodies; does not publish).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
experimentIdYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the safety profile is covered. The description goes beyond them with two useful facts: variant bodies are released, and the operation does not publish anything. It still omits authorization/permission requirements and whether the deletion is recoverable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a parenthetical that earns its place by clarifying the non-publishing semantics. No wasted words, though it is terse enough to need more content elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with annotations covering the safety hints, the description is adequate but thin: no state preconditions, no note on reversibility, no mention of what happens to the experiment record after the variant bodies are released. No output schema exists, so return expectations are also unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% (site, context and guestToken are documented, experimentId is not), so the schema does most of the work. The description adds nothing about parameters, leaving the required experimentId undocumented in both places; a bare 3 is appropriate for this coverage level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Remove a draft A/B test') and the parenthetical narrows the scope to the draft state. It implicitly separates itself from siblings like discard_drafts and promote_experiment, though it never names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no stated condition for choosing this over alternatives such as discard_drafts, delete_content, or promote_experiment, nor any prerequisite (e.g. the test must still be a draft). Usage must be inferred from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_menu_itemB
DestructiveIdempotent
Inspect

Remove one item from a menu (page stays live).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
menuYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the safety profile is covered. The description usefully clarifies the blast radius ('page stays live'), which is genuine added context, but it says nothing about authentication requirements, the fact that a required context sentence is logged publicly, or whether the removal is recoverable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler, and the scope qualifier is placed immediately after the action so the agent reads the constraint first. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, five-parameter tool with 60% schema coverage and no output schema, the description is thin. Safety is carried by annotations, but the agent gets no signal about the required context/guestToken workflow, which are the most likely causes of a malformed call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 60% and the description contributes no parameter information at all. The undocumented id and menu parameters are inferable from their names, but the description does nothing to compensate for the coverage gap or to explain the mandatory audited 'context' field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Remove one item from a menu') and adds a scope qualifier ('page stays live') that separates it from the page-level delete siblings. It does not name an explicit sibling (e.g. delete_route, delete_content) but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(page stays live)' implies the tool is for removing a single menu entry without deleting the underlying page, which is a usage hint. However there is no explicit when-to-use guidance, no mention of when to prefer reorder_menu or put_menu_item, and no statement of prerequisites such as the guest token or the mandatory context field.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_routeC
DestructiveIdempotent
Inspect

Remove a route by path.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered by structured data. The description adds nothing beyond that — it says nothing about what happens to dependent routes, whether removal is reversible, or the activity-log/auth context that the schema alludes to.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the verb and key identifier front-loaded and zero filler. It is efficient, though its brevity borders on under-specification rather than pure concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a destructive mutation tool with no output schema and only a one-line description. It omits consequences of deletion, handling of child/dependent routes, and any interaction with the required site context — a significant gap even accounting for the annotations that cover idempotency and destructiveness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% (site, context, and guestToken are documented; path is not), so the schema does most of the work. The phrase 'by path' marginally clarifies that path is the deletion identifier, which is slightly more than the bare schema offers, but no format or matching semantics are provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (remove) plus resource (route) and the identifier (path), which cleanly distinguishes it from siblings like delete_content, delete_asset, and delete_menu_item. It does not, however, mention the site scoping that the schema requires, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance, and no mention of the obvious related siblings (get_route, put_route, patch_route) that an agent would need to choose between. The agent gets no signal about prerequisites or the effect on routes that reference this path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_siteB
DestructiveIdempotent
Inspect

Permanently delete site; confirmSlug must equal slug.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
confirmSlugYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the destructive nature and repeat-call behavior are covered by structured data; the description's 'Permanently' merely reinforces that. It does add one genuine behavioral fact, the confirmSlug==slug interlock, but says nothing about what is destroyed (content, assets, domains, activity log), whether it is reversible, or what quota/ownership checks apply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the destructive verb front-loaded and the constraint immediately after. Every clause earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an irreversible site-level destruction with three required parameters and no output schema, the description is thin: it never says what happens to the site's pages, assets, domains, or guests, nor whether any recovery window exists. The confirmSlug gate is documented, but an agent about to invoke this cannot tell what the blast radius is from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the one undescribed parameter (confirmSlug, which has no schema description at all) is exactly what the description explains: it must equal the slug. That is meaningful semantics added beyond the schema. It does not clarify the site slug format beyond what the schema's own example provides, so it is not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (permanently delete a site) in the very first words, which cleanly separates it from the many delete_* siblings such as delete_asset, delete_content, and delete_route. It does not name or contrast any sibling explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no stated preconditions (e.g. site must be unclaimed, unpublished, or domain-free), and no mention of when a softer alternative like update_site or rename_site would be preferable. The only usage-adjacent content is the confirmSlug constraint, which is a safety interlock rather than routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discard_draftsA
DestructiveIdempotent
Inspect

Discard unpublished drafts: all (omit key/keys/routes), one content key (key), or a selection (keys, routes). Destructive — check get_draft_status first.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNo
keysNo
siteYesSite slug, e.g. acme-studio
routesNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the adjectival 'Destructive' in the description largely restates structured data. It does add real value by scoping destruction to 'unpublished drafts' and by recommending a status check beforehand, but it omits anything about reversibility, required permissions, or whether discarded drafts can be recovered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence for the selection semantics followed by a short safety warning; the destructive nature and the get_draft_status precondition are front-loaded, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, no-output-schema mutation tool with annotations already covering the safety profile, the description covers the ambiguous selection logic well. The remaining gaps — auth/permission requirements and what a successful discard returns — are minor given the schema documents guestToken and context and the annotations flag destructiveness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%, and the four undocumented parameters (key, keys, routes) are exactly the ones the description explains — including the important default that omitting all three discards everything, and the key-vs-(keys, routes) distinction. This is meaning the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Discard unpublished drafts') and then enumerates the three distinct scopes the tool can operate on: all, one key, or a selection. 'Discard' plus 'unpublished drafts' cleanly separates it from the sibling delete_content/delete_route tools, which target published/committed resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit precondition ('check get_draft_status first') and encodes the selection rules (omit key/keys/routes for all; key for one; keys/routes for a subset), so an agent knows how to invoke each mode. It does not, however, state when *not* to use this versus delete_content or what happens if the drafts were already published.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_pageA
Read-only
Inspect

Open the interactive page canvas to change content: the user clicks an element, then you call patch_content. Use when the intent is editing — not for verify-only checks (use preview_page).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRoute path (default /).
siteYesSite slug.
stageNoPreview stage. Default draft; published is planned but not supported yet.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds real behavioral context beyond that: this tool only opens an interactive canvas and requires the user to click an element, with actual mutations deferred to patch_content – which also reconciles the seemingly write-oriented name with the readOnly annotation. It stops short of describing session lifecycle or what happens if the canvas is abandoned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, dense sentence that front-loads the action, then the flow, then the alternative. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only 3 fully documented params, no output schema, and simple annotations, the description covers what an agent needs: what happens (interactive canvas), who acts (the user), and what to call next (patch_content). Minor gap in not clarifying whether the canvas session persists or how long it stays open.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (site, path, stage all documented, including the caveat that published is not yet supported), so the schema carries the load. The description adds no parameter-level detail, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (open the interactive page canvas) and its purpose (change content), and explicitly distinguishes itself from preview_page (verify-only) and names patch_content as the follow-up mutation step. An agent can place this tool among its many siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when ('intent is editing') and when-not ('not for verify-only checks'), and routes the alternative case to a named sibling, preview_page. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

estimate_testB
Read-onlyIdempotent
Inspect

Estimate days to reach promote readiness for a page (traffic baseline). Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesRoute path, e.g. /pricing
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the appended 'Read-only' repeats structured data. The useful addition is '(traffic baseline)', disclosing the estimation basis, but no return shape or precision (e.g., a day count vs. a range) is given despite no output schema existing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no padding. Minor waste in the second sentence, which restates the readOnlyHint annotation rather than adding new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The safety profile is covered by annotations, but with no output schema the description should say what the estimate returns (a number of days, a confidence band, etc.) and whether it depends on current traffic. That omission keeps it at a minimally viable level for a 4-parameter analysis tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter (site, path, context, guestToken) is already documented, so the baseline is 3. The description adds only an implicit binding of 'a page' to the path parameter and clarifies nothing about context or guestToken semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (estimate) plus resource (days to promote readiness) and scope (for a page), with a parenthetical clarifying the method (traffic baseline). 'Promote readiness' is jargon that only resolves by cross-referencing the sibling promote_experiment, so it is clear but not sibling-differentiating on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use, prerequisites, or alternatives. With siblings like start_test and promote_experiment, the agent is left to infer that this is a pre-promotion estimate; nothing names the sequencing or when NOT to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_replaceA
Destructive
Inspect

Case-insensitive find→replace across pages in one staged commit. Batch all replacements into ONE call — never parallelize find_replace on the same key (concurrent edits surface as per-key conflict). find/replace ≤60KB.

ParametersJSON Schema
NameRequiredDescriptionDefault
findYes
keysNo
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
replaceYes
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
expectPerKeyNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, idempotent=false, readOnly=false; the description still adds real value beyond them by disclosing case-insensitivity, the staged-commit semantics, the per-key concurrency conflict behavior, and the 60KB find/replace limit. It stops short of saying how the staged change is applied or reverted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the operation and scope, followed by the two operational constraints. No filler; every clause carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with seven parameters, no output schema, and 43% schema coverage, the description covers concurrency and size limits well but omits how the staged commit is finalized (apply_changeset/publish), what happens on conflict/partial match, and the meaning of keys and expectPerKey.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43%, and the description compensates for at most one parameter by mentioning the 60KB find/replace size bound. It says nothing about what 'keys' scopes, what 'expectPerKey' validates, or that 'replace' can be empty, leaving four parameters semantically thin in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (case-insensitive find→replace across pages) and adds the notable scope detail that it lands in one staged commit. It does not name or distinguish itself from near siblings such as patch_content, batch_patch_content, or bulk_put_content, so an agent must infer when this is the right mutation tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a strong invocation rule (batch everything into ONE call, never parallelize on the same key because of per-key conflicts) and a size ceiling, but says nothing about when to choose this over the overlapping content-mutation siblings. The guidance is about how to call it safely, not when to prefer it over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_account_contextC
Read-onlyIdempotent
Inspect

Account, user, memberships[]. Pass workspace on writes when memberships.length > 1.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
workspaceNoWorkspace org id or slug (see get_account_context memberships). Only for multi-workspace users; omit otherwise. Ignored for pga_ keys (fixed workspace).
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds cross-tool context (this tool's memberships drive whether workspace must be passed on writes), which is genuinely useful relational behavior. It does not, however, disclose return format or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse fragments with no filler, and the workspace rule is front-loaded effectively. But it is under-specified rather than genuinely concise — the noun-list style borders on cryptic and omits the actual operation, so brevity costs clarity here.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only context tool with no output schema, the description lists the returned entities (account, user, memberships) and 100% param coverage in the schema carries the inputs. It still lacks any statement of when the tool should be invoked, leaving a real gap for what is likely a first-call bootstrap tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description's 'memberships.length > 1' threshold is a slightly more concrete restatement of the schema's 'Only for multi-workspace users' note, adding marginal value but no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Account, user, memberships[]' tells the agent what entities the tool surfaces, which is more than a pure tautology of the name. However, it never states a verb or the operation itself, so the purpose is inferred rather than declared, and nothing differentiates it from context-bearing siblings like get_site or list_my_sites.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Pass workspace on writes when memberships.length > 1' gives a conditional rule, but it governs how to use the workspace parameter in *other* (write) tools, not when to call get_account_context itself. There is no when-to-use, no prerequisite, and no alternative named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activityA
Read-onlyIdempotent
Inspect

List site activity or fetch one change diff (pass id). Filters: who, state, actor (me), since, limit≤20. Notes from context are visible to everyone on the site — no secrets.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoEvent id — returns diff for this change instead of the list.
keyNoContent key when id is set (default first key in the event).
whoNoDefault all.
slugYesSite slug, e.g. acme-studio
actorNoActor id or me.
limitNoPage size (default 20).
sinceNoISO date (default 7 days, plan retention floor).
stateNoneeds_review = agent drafts awaiting a person's publish.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive). The description adds the meaningful disclosure that context notes are publicly visible to everyone on the site with 'no secrets'. Retention floor on since and limit≤20 are also behavioral, though pagination behavior is unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the primary action and mode, then filters, then the public-visibility warning. Dense but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list/diff tool with 100% schema coverage and no output schema, the description covers mode selection, filters, and the safety-relevant visibility of context. Pagination and cursor behavior are the only notable gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds cross-parameter semantics not obvious from the schema alone: id switches the return type to a diff, context must be a one-sentence reason, and 'since' has a 7-day default with a plan retention floor.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names specific verb (List/fetch) and resource (site activity/change diff) with a clear mode distinction keyed on the id parameter. Distinguishes itself from siblings like get_analytics and list_pages, though it doesn't cite a sibling by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States two operating modes and when to use each ('pass id' to get a diff instead of the list), plus enumerates available filters. Clear context but no explicit exclusions or alternative-tool routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_adapt_statusA
Read-onlyIdempotent
Inspect

Read adapt state: proposal summary + proposalEtag (+ applied/consumed flag), undo-record presence, active or queued job progress. Use to resume in a fresh session.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, so safety is covered. With no output schema, the description usefully discloses what state is returned (etag, applied/consumed flag, undo presence, queued vs active jobs), which an agent cannot get from structured fields. It adds real behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the resource and returned data front-loaded ahead of the usage clause. Dense and well-ordered, though the parenthetical flag notation is slightly compressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately covers return content, and annotations cover the safety profile; the schema covers parameters. The main gap is the absence of any mention of continuation behavior between calls, but for a status-read tool this is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so 'site', 'context', and 'guestToken' are fully documented in the schema itself, establishing the baseline of 3. The description contributes nothing further about parameter usage or format, so it neither gains nor loses on this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (adapt state) and enumerates the concrete payload: proposal summary, proposalEtag with applied/consumed flag, undo-record presence, and job progress. This clearly separates it from the mutating siblings adapt_site_analyze/apply/revert, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The single clause 'Use to resume in a fresh session' gives one concrete usage context, which is genuinely helpful. However, there is no guidance on when not to use it or which sibling to prefer for overlapping needs (e.g. get_draft_status, get_site_build_progress), leaving the routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_analyticsA
Read-onlyIdempotent
Inspect

Read-only traffic overview (default) or full breakdowns via include[]; optional bot id for crawler page drill-down.

ParametersJSON Schema
NameRequiredDescriptionDefault
botNoCrawler bot id (e.g. googlebot, gptbot) → that bot's top crawled pages under `crawlerBotPages`.
siteYesSite slug, e.g. acme-studio
rangeNoTime window; default 30d.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
includeNoFull breakdown sections to add under `detail`: pages, sources, countries, devices, browsers, os, utm, goals, series, crawlers, experiments, events — or 'all'. Omit for the overview only.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description's 'Read-only' largely restates the annotation, and it adds only mode-selection behavior rather than new traits like rate limits or response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single densely packed sentence with the default mode front-loaded and secondary capabilities (include[], bot) following. No filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters and no output schema, the description covers invocation modes but does not describe the return structure (overview fields, detail sections, crawlerBotPages). Annotations and full schema coverage compensate for much of the gap, making this adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents site, range, context, bot, include, and guestToken in detail. The description adds marginal gloss on include[] and bot, which is the expected baseline-3 outcome when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource (read-only traffic overview / breakdowns) and distinguishes two modes (default overview vs include[] breakdowns) plus a bot drill-down variant. It is clear what the tool does, though it doesn't explicitly contrast itself against sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states which mode is default ('overview (default)'), how to get more ('full breakdowns via include[]'), and the condition for the bot parameter ('for crawler page drill-down'). There are no explicit when-not-to-use statements or named alternatives, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_assetA
Read-onlyIdempotent
Inspect

Read one asset (metadata + base64 + stored sha256). sha256 is the stored blob; may differ from upload bodySha256 when IMAGES re-encodes rasters. Uploaded files are live in the editor and draft preview at once; the published site picks them up at the next publish.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesPath after /assets/ on the site — e.g. photos/hero.webp or style.css. Do not prefix with assets/ or /. Uploaded files are live in the editor and draft preview at once; the published site picks them up at the next publish.
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds real value beyond them: it specifies the exact return payload and clarifies that the stored sha256 can differ from the upload bodySha256 when IMAGES re-encodes rasters. That hash semantics detail is non-obvious and useful, though it stops short of noting size limits or large-base64 caveats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and return contents, and the sha256 clarification earns its place. The third sentence about files being live in the editor/published site is largely duplicated from the 'key' parameter description and is only loosely relevant to a read tool, slightly diluting focus.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the burden of describing returns (metadata, base64, stored hash) and explains the hash's provenance. It is complete enough for correct invocation, though it omits any hint about response size or how the base64 should be consumed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters (key path format, site slug, context, guestToken). The description adds no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource ('Read one asset') and immediately enumerates what is returned (metadata + base64 + stored sha256). The singular 'one asset' cleanly distinguishes it from the sibling list_assets and from mutators like bulk_put_assets/delete_asset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to reach for this tool versus alternatives such as list_assets, get_site_manifest, or search_content. The upload/publish sentence describes asset lifecycle behavior but does not tell the agent when this read tool is the right choice or what prerequisites apply.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_contentA
Read-onlyIdempotent
Inspect

Read one content key; optional selector (region) or published=true (live body). Bodies over 40KB without a selector return a truncated summary — use selector, batch_patch_content, or dotsy pull.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
selectorNo
publishedNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds genuinely new behavior beyond annotations: the 40KB truncation behavior and the published=true live-body mode. It does not describe auth/guestToken flow, but that lives in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, zero filler, with the truncation caveat and its remedies front-loaded after the core read statement. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only single-key tool with no output schema, the description covers the key gotcha (large-body truncation) and the live/published distinction. Minor gap: it doesn't clarify the guestToken vs signed-in context, but that is documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%. The description does add meaning for two undocumented params: selector => 'region' and published => 'live body'. However 'key' and 'guestToken' remain unexplained, so the description only partially compensates for the coverage gap; baseline 3 is right.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Read one content key.' The scope word 'one' implicitly separates it from list_content and search_content siblings. It does not explicitly name those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete conditional: bodies over 40KB without a selector return a truncated summary, and names alternatives (selector, batch_patch_content, dotsy pull). This tells the agent when to add a selector, though it offers no guidance on when to prefer this over list_content or search_content for discovery.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_custom_domain_connectC
Read-onlyIdempotent
Inspect

Current connect session and nextAction (browser URL or manual DNS guide).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
sessionIdNoConnect session id from add_custom_domain; omit to resume pending.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds the useful detail that the result may be either a browser URL or a manual DNS guide, but omits auth/session resume behavior and any rate or polling notes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero filler. It is efficient, though arguably under-specified rather than maximally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a workflow/state-inspection tool with no output schema, it does indicate the shape of nextAction, which is the key return field. However it leaves the lifecycle relationship with add_custom_domain, wait_for_custom_domain, and reapply_custom_domain_oauth unstated, which an agent needs to sequence calls correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema documents site, context, sessionId (from add_custom_domain, omit to resume pending), and guestToken in full. The description adds no parameter-level meaning beyond that baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a noun phrase describing the returned payload (current connect session, nextAction with browser URL or manual DNS guide) rather than a verb stating what the tool does. A reader can infer it is a read of connect-session state and distinguish it from add/remove/refresh siblings, but the action itself is only implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no named alternatives, despite a dense cluster of sibling domain tools (add_custom_domain, wait_for_custom_domain, reapply_custom_domain_oauth). Nothing tells the agent when to call this versus waiting or re-applying.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_draft_statusA
Read-onlyIdempotent
Inspect

Pending draft changes: content keys and routes with status. May include requiredDecision (publish_confirm) before publish_site.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the safety profile is fully covered. The description adds one genuine behavioral fact beyond that – the response may carry a requiredDecision (publish_confirm) gating publish_site – but says nothing about freshness, scope of 'pending', or how the decision should be handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the payload stated first and the publish-gating note second; nothing is wasted. The telegraphic style ('content keys and routes with status') is compact but slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does list the returned elements plus the possible requiredDecision field, which is the key thing an agent needs to branch on. It stops short of naming the decision's values or the exact field shape, but is adequate for a three-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself explains site, context (including the activity-log warning), and guestToken in detail. The description adds no parameter meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource and payload clearly: pending draft changes with content keys, routes, and status. It implies a read of draft state and is distinguishable from mutations like discard_drafts or publish_site, though it never uses an explicit verb like 'return' or 'list'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied via the trailing phrase 'before publish_site', which hints at a pre-publish checkpoint. There is no explicit when-to-use statement, no mention of alternatives such as get_content or list_content, and no note on when this is unnecessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_form_submissionA
Read-onlyIdempotent
Inspect

Read one form submission with the full field payload (explicit PII access — use only when needed).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesSubmission id from list_form_submissions
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so safety is covered. The description adds meaningful non-structural context: this call exposes PII in full, which is exactly the kind of consequence an agent should weigh before invoking. It does not, however, discuss logging or retention implications of the returned data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the action, the object, the payload scope, and the sensitivity caveat. Nothing is padded and the warning is not buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter read tool with no output schema, the description conveys what comes back ('full field payload') and flags the PII risk, which is the main decision factor. Missing only a pointer to the list tool for the lighter-weight alternative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter (id, site, context, guestToken) carries its own schema description, so the baseline is 3. The description adds no parameter-level detail beyond noting the payload is 'full', which does not disambiguate any argument.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read one form submission') plus the scope ('full field payload'), which distinguishes it from the sibling list_form_submissions by implying single-record retrieval. It stops short of naming that sibling explicitly, so the differentiation is inferential rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical 'explicit PII access — use only when needed' gives a caution and an implicit when-to-use condition, but it never names the cheaper alternative (list_form_submissions) or states a concrete trigger for the full-payload read. Usage is hinted rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_menuA
Read-onlyIdempotent
Inspect

Read a menu's draft items and etag; menu rules → get_skill('nav').

ParametersJSON Schema
NameRequiredDescriptionDefault
menuYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds that the result contains 'draft items and etag', useful return-content context in the absence of an output schema, but says nothing about permissions, guest-token behavior, or what an absent menu returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no wasted words, and the core action ('draft items and etag') is front-loaded ahead of the secondary routing hint. Nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter read tool with no output schema, the description gives only a thumbnail of the return ('draft items and etag') and leaves the relationship to get_menu_doc and the semantics of a draft-vs-published menu unstated. It is minimally viable but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and includes the unusually well-documented 'context' and 'guestToken' parameters, so the schema does the heavy lifting. The description adds no parameter meaning beyond what the schema already states, which is the baseline expectation at this coverage level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('a menu's draft items and etag'), which is precise enough that an agent knows this fetches draft menu state. However, it never distinguishes itself from the sibling get_menu_doc (published doc) or get_menu_doc/put_menu_doc, so the boundary between reading a menu's draft vs. its doc is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'menu rules → get_skill('nav')' is a genuine routing hint for a related need, which is more than most tools offer. But there is no guidance on when to call this versus get_menu_doc or when a menu read is appropriate at all, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_menu_docA
Read-onlyIdempotent
Inspect

Read the whole menu doc (all menus + shell/template), etag for put_menu_doc, and issues (by item id) — page_not_routable, page_missing, page_dead_link, page_is_redirect — computed from the draft routes, so a link that won't show or leads nowhere shows up before publish.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
sourceNoWhich doc to read. Default draft.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safe-read profile (readOnlyHint, idempotentHint, destructiveHint=false, closed world), so the bar is lower; the description adds real value by disclosing the response contents and, crucially, that the issues (page_not_routable, page_missing, page_dead_link, page_is_redirect) are computed from draft routes. It does not restate safety behavior, which is fine given the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence that leads with the core action, followed by the return contents and the rationale. The parenthetical enumerations of issues and doc parts make it dense, but every clause carries information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so reasonably: it lists the menu doc, the etag for the paired write, and the four issue categories. Combined with the fully documented input schema, an agent has enough to call and interpret the result, though a note on scoping source=draft/live in the description would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with 'source', 'context', and 'guestToken' fully documented in the schema, so the baseline is 3. The description mentions draft routes (implying the draft source) but does not add format or selection detail beyond the schema, so it neither compensates nor detracts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Read the whole menu doc'), enumerates what the doc contains (all menus + shell/template), and clarifies scope against the sibling get_menu by emphasizing 'whole' vs a single menu. It also ties itself to put_menu_doc by naming the etag it supplies, so an agent can place it in the read/write pair without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage timing is only implied by 'shows up before publish', which hints this is a pre-publish verification read, and by the etag mention implying it precedes put_menu_doc. There is no explicit when-to-use/when-not statement or named alternative (e.g. get_menu for a single menu), so the agent must infer the routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_routeB
Read-onlyIdempotent
Inspect

Get one route by path (home is '/').

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, fully covering the safety profile. The description adds no behavioral context beyond that — nothing about authentication needs, error behavior for a missing path, or any distinguishing characteristic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the verb and resource, with the path convention tucked in parenthetically. Zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a get tool with no output schema, the description says nothing about what a retrieved route contains (layout, template, redirect target, etc.), leaving the agent unable to reason about the result. The auth/guest-token story is carried by the schema, which mitigates but does not eliminate the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%; site, context, and guestToken are all documented in the schema, including the notable requirement to send guestToken until claim_site. The description contributes only the '/' convention for path, which the schema leaves undocumented, so it adds marginal value over the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: fetch a single route identified by its path, with the home-path convention stated. It does not differentiate itself from near-siblings like get_menu, get_site_outline, or patch_route, so it lands just below the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives (patch_route/put_route/delete_route all operate on the same resource). The only hint is the parenthetical about '/' for home, which is a path convention rather than usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_siteA
Read-onlyIdempotent
Inspect

Site detail, stats, publicUrl, brief head (default-on), and readiness flags. pages[] is a ~25-entry head by default (stats.pages has the true count) — pass pages: true for the full array (capped at 1000), or enumerate everything with list_content/get_site_manifest. Large sites (200+ pages): list_pages is paginated — use cursor or get_site_manifest for the full route set.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
briefNoInclude parsed brief head (default true).
pagesNoReturn the full pages[] array instead of the default ~25-entry head (server-capped at 1000 either way). Default false.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/non-destructive/openWorld=false, so the safety profile is covered. The description adds genuinely useful behavior the annotations cannot: pages[] is a ~25-entry head by default, stats.pages holds the true count, and pages:true is server-capped at 1000. The truncation semantics are the key disclosure an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the return contents come first, then the pages/pagination guidance. Every clause earns its place, though the em-dash-heavy packing makes it slightly harder to scan than a clean sentence would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must carry the return shape, and it does (detail, stats, publicUrl, brief, readiness flags, pages truncation). Auth via guestToken and the required context sentence are documented in the schema, so the description needn't repeat them. Adequate for a 5-param read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all five parameters are already documented, including pages default/cap and brief default. The description reinforces pages semantics (the ~25-entry head, the 1000 cap) and the stats.pages relationship, but this largely restates the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (site) and enumerates what it returns: detail, stats, publicUrl, brief head, readiness flags. It implicitly distinguishes itself from siblings like list_pages and get_site_manifest in the same breath. It never states the verb outright, but the read-of-a-site intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent explicitly: use pages:true for the full array, or enumerate everything with list_content/get_site_manifest, and it flags that list_pages is paginated for large sites. It does not address when to prefer this over get_site_outline, list_my_sites, or get_analytics, but the alternatives it does name are the relevant ones.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_site_build_progressA
Read-only
Inspect

Read-only build progress for humans (build-theater view) during structure/adapt or after import. Agent reads structuredContent.forHumans — optional; do not call in a loop during import_site_from_url (that tool blocks until done). Prefer get_site for agent logic.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, destructiveHint=false, and openWorldHint=false, so safety is covered. The description goes beyond them by flagging that the payload lives in structuredContent.forHumans, that the call is optional, and that polling is wasteful because import_site_from_url blocks until completion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the operation and its timing. The only drag is the parenthetical jargon 'build-theater view', which costs a little clarity without adding much.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully names the return field (structuredContent.forHumans) and the conditions under which calling is worthwhile. A brief note on what that payload contains would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single 'site' parameter at 100% schema description coverage, so the schema already defines it fully. The description adds no syntax, format, or slug-resolution detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read-only build progress') plus the scoping condition ('during structure/adapt or after import') and explicitly names the sibling to prefer for agent logic ('Prefer get_site'). An agent can distinguish this from get_site, get_structure_status, and get_adapt_status without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the when (during structure/adapt or after import), the when-not ('do not call in a loop during import_site_from_url (that tool blocks until done)'), and the alternative ('Prefer get_site for agent logic'). All three selection conditions are explicit rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_site_manifestA
Read-onlyIdempotent
Inspect

Content keys + hashes only (optional routes=true) — static files live in list_assets, not manifest. Diff keys+hashes before apply_changeset; stage only changed bodies.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo
siteYesSite slug, e.g. acme-studio
limitNo
cursorNo
routesNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld=false). The description adds genuine context beyond them by disclosing response content (keys+hashes only) and what is excluded (static file bodies, route detail unless routes=true). It omits pagination behavior despite limit/cursor params, but the added content disclosure is real.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the return content and immediate clarifications. Every clause earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and with 7 parameters and low schema coverage the description leaves gaps around mode, pagination, and auth parameters. It covers purpose and workflow well but is thin for a tool with this much surface area.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, so the description should compensate, yet it addresses just one parameter (routes). It never explains mode (published/draft), limit, cursor, or the guest-token flow, leaving several parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (site manifest) and precisely what it returns (content keys + hashes only), and distinguishes its scope from siblings by noting static files belong to list_assets. The verb 'get/lists manifest' is implied rather than explicit, but an agent can tell what the tool yields.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete workflow ('Diff keys+hashes before apply_changeset; stage only changed bodies') and routes the agent away from this tool when dealing with static files ('list_assets, not manifest'). No explicit when-not conditions, but the alternative is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_site_outlineA
Read-onlyIdempotent
Inspect

Site shape: every layout and section with its page count and staged-change flag, whether a layout is missing, redirects count, menus grouped by where they show, and collections with their page placements. Use for 'is my layout missing', 'what does this section reach', 'where does this menu show' — without listing every page. Read-only; indexing mirrors list_pages (a freshly imported/graduated site may still be indexing).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive, closed-world. The description adds real behavioral context beyond them: the returned facets, and importantly the indexing caveat that a freshly imported/graduated site may still be indexing, mirroring list_pages. It doesn't describe pagination or result-size behavior, but the added indexing nuance is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the return shape, then usage triggers, then the indexing caveat. Dense but every clause earns its place; the enumeration is long but serves the missing output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating what is returned (layouts, sections, counts, flags, redirects, menu groupings, collection placements) and adds the indexing timing caveat. An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so site, context and guestToken are already fully documented in the schema. The description adds no syntax or semantics beyond that, which is the correct baseline when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (site shape: layouts, sections, redirects, menus, collections) with concrete attributes like page count and staged-change flags. It explicitly distinguishes itself from list_pages by promising structure 'without listing every page', so an agent can route without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three concrete usage triggers ('is my layout missing', 'what does this section reach', 'where does this menu show') and implies the list_pages alternative via 'without listing every page'. It stops short of naming the sibling tool or stating when NOT to use it, so it's clear context rather than explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_skillA
Read-onlyIdempotent
Inspect

Skill body + optional questions[]. Accepts an id OR a synonym (e.g. 'list', 'menu') — resolves aliases and suggests near-misses. Start with intake when ambiguous.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesSkill id or synonym from list_skills, e.g. intake, clone, list, menu
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only, idempotent safety profile, so the bar is lower. The description still adds real behavioral context the annotations don't: alias resolution to synonyms like 'list'/'menu', near-miss suggestion on failure, and that the response includes a questions[] payload. This is meaningful disclosure of resolution and return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses, front-loaded with the return shape then the key behavioral fact (synonym acceptance). Slightly telegraphic ('Skill body + optional questions[]' is a fragment) but every sentence carries load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully describes the return (body + questions[]) and the alias/fallback behavior. Given only two fully-documented required params and a read-only annotation set, this is close to complete, though the 'intake' reference assumes cross-tool knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (id and context) are already fully documented, including the privacy warning about context being publicly logged. The description's note that id may be a synonym duplicates the schema rather than extending it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (skill) and what comes back (skill body plus optional questions[]), and the id-or-synonym acceptance makes it distinguishable from list_skills. The verb is implicit ('get') but the scope is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It offers one conditional hint ('Start with intake when ambiguous') and gestures at fallback behavior via near-miss suggestions, but never states when to call get_skill versus list_skills or other siblings. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_structure_resultB
Read-onlyIdempotent
Inspect

Last structure apply job result (structuredKeys, conflicts, partial).

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoReturn the full manifest/detail (default: compacted for the token budget).
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered structurally. The description adds only the shape of the returned result, not behavior such as what happens when no prior apply job exists or whether the result is cached/expiring.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the resource front-loaded and no filler. It is efficiently sized, though the parenthetical field list is a bare enumeration rather than explanatory.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should carry more of the return-value burden; listing the field names helps but says nothing about structure, empty/absent-result states, or the distinction from get_structure_status. Adequate but with clear gaps for a job-status retrieval tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the noteworthy `full` default-compaction flag and the `context` logging warning, so the schema does the heavy lifting. The description adds no parameter meaning beyond that, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (the last structure apply job result) and names the payload fields (structuredKeys, conflicts, partial), so an agent knows what it gets back. It does not differentiate itself from the sibling get_structure_status or structure_site_apply, which keeps it below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to call this versus get_structure_status, nor any prerequisite (e.g., that a structure apply job must have been started by structure_site_apply first). Usage has to be inferred entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_structure_statusB
Read-onlyIdempotent
Inspect

Flat vs structured page counts; informs user — do not auto-run analyze/apply while requiredDecision pending.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint=false, so safety is covered. The description adds genuinely useful operational context (the requiredDecision gate on analyze/apply), but the term is left undefined, limiting its value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely compact — a single line with no filler, and the return-value hint is front-loaded before the behavioral warning. The telegraphic style is terse but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should carry more of the return-value burden; it only hints at 'flat vs structured page counts' and leaves the return shape, the meaning of requiredDecision, and any next steps unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and all three parameters (site, context, guestToken) are documented in the schema, including the guestToken lifecycle. The description adds nothing about parameters, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys what the tool returns — flat vs structured page counts — but uses a noun phrase with no explicit verb or subject (no 'get'/'retrieve'). It gestures at the resource but a reader cannot cleanly distinguish it from the sibling get_structure_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the relevant alternatives (analyze/apply) and gives a when-not rule: don't auto-run them while requiredDecision is pending. However, there is no positive statement of when to call this tool and no explanation of what a 'requiredDecision' is or how it is set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graduate_siteA
Idempotent
Inspect

Enable site search & collections (kv → do-sqlite), one-way. Unlocks search_content, collection indexes, visitor /__paige/search, and moves any R2 form-submission ledger into SiteDb. 409 blocked covers three unrelated gates: an unsupported source backend (no override, ever); strictDrafts — default true, this client sends true unless you pass strictDrafts:false — blocking on unpublished drafts or unrecovered KV content rows; or a KV snapshot that looks incomplete (e.g. a page deleted in the last couple of minutes). force:true overrides ONLY that last, incomplete-snapshot gate — never the backend or strictDrafts gates. The blocked response's reason names which gate fired, and audit (+ audit.warnings) is aggregate COUNTS only (published/route/pending-draft counts) — never row keys — so judge the size of the gap, not a row list, before forcing. The flip is one-way: anything outside the frozen snapshot is unreachable once it lands.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
forceNoProceed past the incomplete-KV-snapshot 409 only — not the unsupported-backend or strictDrafts blocks, which force does nothing for. Read the blocked response's `reason` + `audit` first; the flip is one-way, so rows outside the snapshot are unreachable afterwards. Default false.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
strictDraftsNoBlock on unpublished drafts or unrecovered KV content rows (default true — publish/discard first, or pass false to proceed anyway).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile; the description goes far beyond by decomposing the 409 into three distinct gates, describing exactly which gate force overrides, clarifying that the audit payload is aggregate counts rather than row keys, and hammering home the irreversibility of the flip. This is the kind of context an agent needs to avoid a destructive-by-omission mistake.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then gate details; every sentence is substantive. Minor redundancy: the 'never the backend or strictDrafts gates' constraint is stated in both the body and the force parameter's schema text, adding a little length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by explaining the blocked response's reason/audit fields. For a one-way mutation it is nearly complete, though it omits auth/permission prerequisites (guestToken is only in the schema) and what happens to already-published content post-flip.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter and client-behavior meaning: force's scope is bounded against the other gates and strictDrafts' default is explained with what this client actually sends. It largely restates the schema's own force/strictDrafts text, hence not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Enable') plus the exact resource and scope ('site search & collections (kv → do-sqlite), one-way'), and enumerates the concrete capabilities it unlocks (search_content, collection indexes, /__paige/search). An agent can distinguish this migration/flip tool from siblings like reindex_collection or inspect_collection without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when the call is appropriate and, importantly, explains when force should and should not be used, plus that the operation is one-way. It does not name a sibling alternative (e.g. a non-migrating path) or state prerequisites like timing relative to publish/discard, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_site_from_urlBInspect

Import URL to draft. On complete, response may include requiredDecision — preview then ask mirror vs editable before structure_*.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoDefault auto — per-page mirror with silent render escalation
siteYesSite slug, e.g. acme-studio
pagesNoExplicit page list; omit to import sourceUrl as index.html
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
segmentsNoSplit-origin segments — each expands into pages by fetching its own `sitemap`. A segment without `sitemap` imports only its pathPrefix hub page (with a warning), not the whole prefix.
sourceUrlYes
workspaceNoWorkspace org id or slug (see get_account_context memberships). Only for multi-workspace users; omit otherwise. Ignored for pga_ keys (fixed workspace).
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
onSlugTakenNoauto-suffix
keyConventionNoContent key shape: directory = Astro dist (about/index.html); default flat

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, open-world, non-idempotent, non-destructive. The description adds real value beyond that by disclosing that the response may contain requiredDecision and that a user decision is required before structure_*, which is not encoded anywhere else. It omits permissions, size/crawl limits, and failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, no filler, and the action is front-loaded before the follow-up workflow. The second sentence is dense with opaque tokens (requiredDecision, structure_*) but each word carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, open-world, non-idempotent import with no output schema, the description covers the critical decision handoff but is thin elsewhere: it never clarifies what the resulting 'draft' is, whether prior content is affected, or the auth/guest flow, leaving the schema to carry most of the context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80% and the schema already documents mode, pages, segments, guestToken, workspace, and onSlugTaken in detail. The description contributes no parameter meaning of its own, so the baseline 3 for schema-carried semantics applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Import URL to draft'), which is enough to distinguish it from create_site/start_site and the structure_* siblings. However it never says the import targets a whole site (crawl/sitemap expansion), so the scope of 'import' is left to the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives post-call workflow guidance ('preview then ask mirror vs editable before structure_*'), which is genuinely useful routing. It does not say when to choose this over create_site, start_site, or adapt_site_*, nor any prerequisite (guestToken/claim_site), so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_collectionA
Read-onlyIdempotent
Inspect

Collection state — count, server.mode, sample entries, tag (membershipKind, match=) when discovered; plus name, noun, URL pattern, layout, where it's listed, and post fields used. Pass status/q/cursor/limit (any one, even empty) for a full items[] page instead: status draft = never published; a published post with unpublished edits is status published with hasChanges:true (never draft) and shows its staged title/date/tags/author; scheduled = future visibleAt. Never-published drafts come from staged content (no extra reads, no publish needed to see them); published/scheduled come from the index. limit defaults to 50, clamped 1-200. Omit all four for just the base inspect (no items[]).

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoCase-insensitive title search within items[].
siteYesSite slug, e.g. acme-studio
limitNoItems per page (default 50, max 200).
cursorNoResume from a prior response's nextCursor.
prefixYesCollection id or prefix, e.g. "posts" or "blog/" — use the key exactly as list_collections returned it.
statusNoFilter items[] by status (see the tool description for what each means). Providing this (or q/cursor/limit) is what makes the response include items[].
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, and the description goes well beyond them: status semantics (draft = never published; published+hasChanges:true is never draft; scheduled = future visibleAt), the data source split (staged content vs index), and the limit default of 50 clamped to 1-200. This is meaningful behavioral context on top of the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded in a loose sense — the base inspect output comes first and the items[] trigger second — but it is a single semicolon-chained block mixing outputs, status semantics, data sources and limits, which makes it dense and harder to scan than it needs to be. Little is wasted, but structure suffers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries return-value disclosure and does so adequately: what the base response contains, when items[] appears, pagination via cursor/limit, and status semantics. Auth (guestToken) and the context requirement are covered in the schema, so no major gap remains for calling it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description earns more: the status enum is deliberately deferred to the description ('see the tool description for what each means') and the semantics are spelled out there (draft/published/scheduled, hasChanges, visibleAt), plus the '(any one, even empty)' activation rule and limit clamping.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description enumerates exactly what the tool surfaces (count, server.mode, sample entries, tags, name/noun/URL pattern/layout, post fields) and how it relates to a full items[] page, which clearly separates it from list_collections and list_content. It never states a plain verb like 'inspect', relying instead on the resource name, but the output inventory is specific enough that an agent knows what it gets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete branch: pass status/q/cursor/limit (any one, even empty) for a full items[] page, or omit all four for the base inspect with no items[]. It also ties prefix back to list_collections output ('use the key exactly as list_collections returned it'). It doesn't name a competing sibling to avoid, but the conditional routing is explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_assetsA
Read-onlyIdempotent
Inspect

CMS uploads under assets/ only — not mirrored clone CDN (get_site mirroredAssets or import warmJob). Not content — use list_content. Uploaded files are live in the editor and draft preview at once; the published site picks them up at the next publish.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
limitNo
cursorNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds valuable behavioral context that uploads appear immediately in editor/draft preview but only reach the published site at next publish, though it does not discuss authorization or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each contributing distinct information: scope, content exclusion, and publication timing. The front-loaded scope sentence is useful, though the first sentence is a fragment rather than a fully explicit action statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for distinguishing this list tool from key siblings, and annotations cover safety. However, pagination semantics and return-value shape are absent with no output schema, leaving an agent to infer how limit/cursor behave and what an asset listing contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description contains no parameter guidance at all. With 60% schema description coverage, the schema documents site, context, and guestToken, but limit and cursor remain undescribed in both places, so the description does not compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource as CMS uploads under assets/ and clearly excludes mirrored clone CDN assets and content assets. The action (listing) is implied by the tool name and title but not stated explicitly in the description itself, so it falls short of a full 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly directs the agent to list_content for content and distinguishes get_site mirroredAssets and import warmJob for clone CDN assets. These are concrete when-not and alternative-tool signals that make routing unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_collectionsA
Read-onlyIdempotent
Inspect

Collections on the site — name, noun, item count (or 'not published yet'), URL pattern, and whether any member is staged. Includes a collection only a draft page lists. Returns a bare array (not wrapped in an object) — pass each row's key verbatim to inspect_collection/update_collection.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/non-destructive, but the description adds real behavioral context beyond them: the return is a bare array rather than a wrapped object, and the listing includes a collection that only a draft page lists. It also notes the item-count placeholder 'not published yet', which signals how unpublished collections appear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both front-loaded: the first lists the returned fields and the draft-collection nuance, the second covers the unusual return shape and the downstream handoff. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by describing the returned fields and the bare-array shape. It stops short of mentioning ordering or any size limits, but for a simple read-only list tool it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (site, context, guestToken are all documented in the schema), so the schema carries parameter meaning. The description adds no syntax or format detail for any parameter, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (collections on the site) and enumerates the exact fields returned, so an agent knows precisely what this produces. The verb is implicit rather than stated, and while it names inspect_collection/update_collection as consumers, it does not explicitly contrast itself with them as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the handoff instruction to pass each row's key to inspect_collection/update_collection, which reveals the intended workflow. However, it never states when to call this versus a sibling like inspect_collection or list_content, nor any prerequisite beyond the schema's guestToken field.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_contentA
Read-onlyIdempotent
Inspect

List content keys (HTML/md); not assets — use list_assets. Defaults to 50 keys per page — raise limit (max 1000) or pass cursor (from nextCursor) to page through the rest. Per-row size/missing are measured automatically only on pages of ≤100 rows; past that they're ABSENT (not measured, never null) unless you pass withSize:true, which forces measurement at a cost of one storage read per returned row. Only size: null + missing: true means the page's stored bytes are gone.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
limitNoKeys per page (default 50, max 1000).
cursorNoResume from a prior response's nextCursor.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
withSizeNoForce size/missing measurement on a page bigger than 100 rows (costs one storage read per returned row). Default false — rows in a ≤100-row page are measured either way.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description goes further by disclosing the measurement behavior: size/missing are auto-measured only on pages of ≤100 rows and are ABSENT past that, with withSize forcing measurement at one storage read per row. This is non-obvious operational cost/behavior an agent could not infer from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with purpose and sibling exclusion before paging and measurement caveats. Every clause carries distinct operational information; nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of return semantics, and it does: what a key is, what per-row size/missing mean, when they are absent, and how to page via nextCursor. An agent has everything needed to call and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents default/max for limit, the cursor source, and the withSize cost, so much of the description's parameter text is restated. It does add real meaning beyond the schema, though: the ≤100-row auto-measurement rule that conditions whether size/missing appear, and the precise 'size: null + missing: true' interpretation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List content keys (HTML/md)') and immediately distinguishes itself from the closest sibling by name ('not assets — use list_assets'). An agent can pick this over list_assets or list_pages without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to list_assets for the asset case, and details the operating conditions: raise limit (max 1000) or pass cursor from nextCursor to page. It also states exactly when to pass withSize (page >100 rows) and what it costs, so the paging decision is fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_custom_domainsB
Read-onlyIdempotent
Inspect

List all custom hostnames and connection status for the site.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is fully covered. The description adds the useful fact that the response includes connection status, which is behavioral context a caller needs since there is no output schema. It says nothing about ordering, pagination, or empty results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the resource and key output field front-loaded and zero filler. Being a one-liner for a 3-parameter tool is lean but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with complete annotations and full schema coverage, the definition covers what it does and what comes back (hostnames plus connection status) without needing to document return values it has no output schema for. Only the absence of any usage/routing guidance keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the notably instructive context and guestToken parameters, so the schema carries the parameter burden. The description mentions no parameters, which is acceptable at this coverage level but adds no extra meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (custom hostnames/domains) plus the scope implied by 'all', which distinguishes it from the singular get_custom_domain_connect. It stops short of explicitly contrasting with sibling write tools like add_custom_domain or remove_custom_domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as get_custom_domain_connect (single domain), wait_for_custom_domain (polling), or the custom-domain mutation tools. Usage must be inferred from the word 'List' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_form_submissionsB
Read-onlyIdempotent
Inspect

List form submissions (metadata + bounded field preview — not full PII). Works on kv and do-sqlite backends.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoSearch reply-to email (indexed)
formYesForm id from dotsy-config, e.g. contact
siteYesSite slug, e.g. acme-studio
toMsNo
limitNo
cursorNo
fromMsNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/no-destructive, so the bar is lower, and the description usefully adds that results are a bounded field preview and not full PII, plus the kv/do-sqlite backend support. It stops short of describing pagination behavior via cursor or how 'bounded' the preview actually is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key constraint (not full PII) front-loaded and no filler. It is perhaps overly terse for a nine-parameter tool, losing no words but also offering little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nine-parameter, read-only listing tool with no output schema and only 56% schema coverage, the description omits pagination, time filters, and search semantics that an agent needs to invoke it correctly. The PII note helps but does not close the larger parameter gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 56%, so the description should compensate, yet it mentions none of the nine parameters (q, fromMs, toMs, limit, cursor). Nothing explains how time-range filtering or cursor paging works, leaving the gaps entirely in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (list form submissions) and adds meaningful scope: metadata plus a bounded field preview rather than full PII. It implicitly contrasts with the singular get_form_submission sibling but never names it, so the differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance and no mention of the obvious alternative, get_form_submission, for retrieving a single record. The agent must infer the list-vs-single distinction from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_goalsA
Read-onlyIdempotent
Inspect

List site conversion goals (pageview paths and form submits). Required before start_test.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is fully covered structurally. The description's contribution beyond that is workflow context (prerequisite to start_test) rather than behavior; it says nothing about output shape or empty-result handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler; the purpose and the critical prerequisite are both front-loaded within a very short string.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with full schema coverage and no output schema, the definition covers purpose and the required ordering with start_test. It could mention whether goals must exist before start_test or what an empty list implies, but it is essentially sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented in the schema (including the context-logging caveat and guestToken lifecycle). The description adds no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (site conversion goals), and clarifies what a goal is with parenthetical examples (pageview paths, form submits). It doesn't explicitly distinguish itself from the sibling put_goals, but the read verb and the examples make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Required before start_test" gives an explicit precondition that tells the agent when this call is mandatory. There is no stated when-not guidance or alternative route to the same data, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_my_sitesA
Read-only
Inspect

List your Dotsy sites. Renders a site-picker for humans; agent reads structuredContent. Use to choose a site — not for editing content (use edit_page) or verify-only checks (use preview_page). Defaults to your active workspace; pass workspace (org id or slug from get_account_context memberships) to list another workspace's sites without switching the dashboard session.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace org id or slug (see get_account_context memberships). Only for multi-workspace users; omit otherwise. Ignored for pga_ keys (fixed workspace).

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false). The description adds non-obvious behavior beyond them: it renders a site-picker for humans while agents read structuredContent, and it defaults to the active workspace. It does not discuss pagination or result size limits, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but tightly packed with routing and behavioral facts; the purpose and exclusions are front-loaded. It is slightly overloaded with parenthetical cross-references, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Single optional parameter, no output schema, annotations present. The description covers purpose, usage routing, rendering behavior, and the parameter's default and source, so nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single optional parameter, so the baseline is 3. The description goes further by explaining where the org id/slug comes from (get_account_context memberships) and that omitting it uses the active workspace, adding practical meaning beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List your Dotsy sites') and explicitly distinguishes itself from the sibling tools that could be confused with it (edit_page, preview_page). An agent can identify the tool's role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('Use to choose a site') and when not to, naming the alternative for each excluded case (edit_page for editing, preview_page for verify-only). It also explains the workspace parameter's purpose: listing another workspace without switching the dashboard session.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pagesA
Read-only
Inspect

List pages on a site — route settings (path/title/navLabel/purpose) by default, one row per route. Renders a page-picker for humans (redirect and worker routes left out there — nothing to open in edit_page); agent reads structuredContent. Use after picking a site — not for editing (use edit_page) or verify-only checks (use preview_page). Pass view:'pages' to group by layout instead — adds layout/noLayout/q filters, what each page places, and indexing progress; use that for 'what pages use this layout' or 'is this page indexed yet', not for route settings. Paginated — pass cursor (from nextCursor) for more; partial:true means this page is a window, not the full route set.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoOnly with view:'pages'. Case-insensitive substring match on path or title.
siteYesSite slug.
viewNopages: group by layout — enables layout/noLayout/q below, adds page-structure fields and indexing progress. Omit for the default per-route listing.
limitNoPages per page (server default applies when omitted).
cursorNoResume from a prior response's nextCursor.
layoutNoOnly with view:'pages'. Layout file key (e.g. layouts/store.html, or bare 'store') — only pages using it. Exclusive with noLayout.
noLayoutNoOnly with view:'pages'. Only pages with no layout (routeless keys included). Exclusive with layout.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false). The description adds real behavioral context beyond that: pagination via cursor/nextCursor, the meaning of partial:true as a window rather than the full route set, and the fact that redirect and worker routes are omitted from the human page-picker. It stops short of describing auth or rate limits, but the added disclosure is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded, and the dense em-dash clauses each carry information (exclusions, alternatives, mode switch, pagination). It is long but not padded; a slightly tighter phrasing could improve readability, but every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, two-mode list tool with no output schema, the description covers the default output shape, the alternate view and its added fields, the human-facing rendering caveat, and pagination/partial semantics. An agent has enough to call it correctly in either mode without inspecting further.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents each parameter's meaning and baseline would be 3. The description goes beyond that by explaining the mode interaction — that view:'pages' enables layout/noLayout/q and adds page-structure and indexing fields — giving semantic grouping that individual parameter descriptions don't.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List pages on a site') and immediately distinguishes the tool from siblings by naming edit_page and preview_page. It also crisply contrasts its two modes (default route settings vs view:'pages' layout grouping), so an agent can tell exactly what it does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('after picking a site'), when not to ('not for editing — use edit_page; not for verify-only checks — use preview_page'), and routes the alternative mode to specific questions ('what pages use this layout' or 'is this page indexed yet'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_skillsA
Read-onlyIdempotent
Inspect

Paige build skills (id + whenToUse). Pull get_skill on demand — do not duplicate in tool text.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds that each entry carries id and whenToUse, but says nothing about list size, filtering, or whether skills are site-specific.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what is returned and followed by the follow-up action. Dense and waste-free, though the opening phrase is cryptic enough to cost a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A simple read-only list tool with no output schema; the description compensates by stating the returned fields and pointing to get_skill for elaboration. Only pagination/size behavior is unaddressed, which is minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is a single required 'context' parameter whose purpose, format, and privacy constraint are fully documented in the schema. The description adds no parameter meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (build skills) and names the return shape '(id + whenToUse)'. It implicitly distinguishes itself from get_skill, which it tells the agent to call for detail, though the term 'Paige build skills' is domain jargon not defined here.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear workflow rule: enumerate with this tool, then 'Pull get_skill on demand'. The instruction not to duplicate skill text in tool output is a concrete usage constraint. No explicit when-not-to-use, but the alternative is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

new_post_templateA
Read-onlyIdempotent
Inspect

Read-only scaffold for a new post — suggested key/path, membership meta, blank or duplicate body. Then put_post (+ optional put_route).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoblank (default) or duplicate full body
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
collectionYesCollection id or prefix, e.g. "posts" or "blog/"
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
templateKeyNoExemplar content key; default = most recent member

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds useful behavioral content beyond that: it tells the agent what the call yields (a scaffold with suggested key/path, membership meta, and a body that is either blank or a duplicate). No rate limits or error behavior, but the return-shape hint is valuable for a tool with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with the read-only nature and outputs front-loaded before the follow-up tool. It is efficient, though the em-dash list is slightly clipped and a phrase like 'membership meta' leans on jargon the schema must disambiguate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only scaffold with six fully documented parameters and no output schema, the description covers purpose, produced values, and the next step in the workflow. It stops short of describing pagination or error cases, but nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including the mode enum and templateKey exemplar is already documented in the schema. The description's mention of 'blank or duplicate body' and 'membership meta' loosely echoes mode/collection but adds no syntax, defaults, or format guidance beyond the schema. Baseline 3 for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read-only scaffold for a new post') and enumerates what it produces: suggested key/path, membership meta, and blank or duplicate body. It also distinguishes itself from the write siblings by naming put_post and put_route as the follow-up step, so an agent can place it in the content-creation chain without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Then put_post (+ optional put_route)' clause makes the intended workflow position explicit: this is the preparatory read step before writing a post. It gives clear context on when to reach for it, though it states no exclusions (e.g. when not to use it versus inspect_collection or get_content).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patch_contentB
Destructive
Inspect

Edit one content key via ops[] (selector ops or replace-string). Stages draft; op shapes in schema.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
opsYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, readOnlyHint=false, so the grammar is known. The description adds one valuable behavioral fact beyond annotations: it 'Stages draft' rather than publishing immediately. It does not say whether staged drafts are reversible, how they are committed, or what effect a failed op has.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, scoping information front-loaded, zero filler. It is arguably too terse ('op shapes in schema') but nothing is wasted and the reader reaches the key fact immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with 5 params, no output schema, and 60% schema coverage, the description does cover the single-key scope and the draft-staging behavior. However it omits the draft lifecycle (how the staged change reaches the live site) and any failure/partial-application behavior, leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%; get_context's 'context' and guestToken are already documented in the schema. The description adds grouping meaning by telling the agent ops[] splits into selector ops vs replace-string families, and defers detail to the schema. Useful but shallow – it names no field like 'expect' or 'selector' semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Edit one content key via ops[]', which distinguishes it from batch_patch_content (batch scope) and put_content (whole-key write). The parenthetical '(selector ops or replace-string)' clarifies the two editing modes. It doesn't name siblings explicitly, so not a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative named. The word 'one' implicitly contrasts with batch_patch_content and find_replace, but the agent must infer that. No prerequisites or follow-up guidance (e.g. how the staged draft gets applied) are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patch_routeA
DestructiveIdempotent
Inspect

Partial route update (only sent fields change; an OMITTED field is left alone). On a field that accepts null (contentKey, redirectTo, redirectStatus, preserveQuery), sending null CLEARS it back to its default — different from omitting it. To set redirectTo/redirectStatus, include purpose: "redirect" in the SAME call — the check looks only at what's sent, not the route's stored purpose, so a PATCH that sets redirectTo without repeating purpose: "redirect" (even on a route that's already a redirect) 400s invalid_body.

ParametersJSON Schema
NameRequiredDescriptionDefault
navNoDEPRECATED — use put_menu_item / menu API. Menu memberships on the route row. Still accepted during transition and synced to menu items. See get_skill('nav').
pathYes
siteYesSite slug, e.g. acme-studio
varsNo
titleNo
layoutNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
noindexNoExclude from sitemap/RSS; inject noindex meta (0.69).
purposeNo
navLabelNoDEPRECATED — use put_menu_item / menu API. Menu link text for <dotsy-nav> (separate from title). Still accepted during transition and synced to menu items.
contentKeyNoExplicit content key (defaults to a path-derived key when omitted). null clears an explicit key back to that path default — same as omitting it, but usable on a PATCH to undo a prior set.
experimentNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
redirectToNoRedirect target, a root-relative path (/…). Required when purpose is "redirect". null clears it — only valid once purpose stops being "redirect" (a redirect route always needs a target).
preserveQueryNoAppend the visitor's original query string to the redirect target (default true on a redirect). null clears an explicit override.
redirectStatusNo301 permanent / 302 or 308 temporary (default 301 on a redirect). null clears an explicit override.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=true, but the description adds genuinely non-obvious behavior beyond them: null-clears differ from omission, and a PATCH that sets redirectTo without repeating purpose returns 400 invalid_body even on an already-redirect route. That failure mode is exactly the kind of context an agent cannot derive from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the partial-update contract first, then the null semantics, then the redirect trap; every sentence carries operational weight. The final sentence is long with nested parentheses, but the density is justified by a real 400 error an agent would otherwise hit.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter mutation tool with no output schema and 63% schema coverage, the description covers the highest-risk semantics (null vs omitted, redirect invariant). It omits nothing critical for calling correctly, though it says nothing about auth/guestToken flow or the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 63% across 16 params, so the description must compensate, and it does for the riskiest ones: it names which fields accept null (contentKey, redirectTo, redirectStatus, preserveQuery) and spells out the purpose/redirectTo coupling. It does not touch nav, navLabel, experiment, or vars, which remain schema-documented only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource and scopes it: 'Partial route update (only sent fields change; an OMITTED field is left alone).' The word 'partial' implicitly distinguishes it from put_route, but the sibling is never named, so the agent must infer the alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating conditions: omit a field to leave it alone, send null to clear, repeat purpose: "redirect" in the same call to set redirect fields. It does not state when to prefer put_route or delete_route instead, so alternatives are left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preflight_siteB
Read-onlyIdempotent
Inspect

Plan check before large build (counts, route soft limit). Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
requiredAssetsNoHow many assets you plan to upload
requiredRoutesNoHow many routes you plan to add

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the trailing 'Read-only' is redundant. The one piece of new behavioral context is the existence of a route soft limit, but the description does not say what happens if the plan exceeds it, whether an error or warning is returned, or what a failure looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded fragment with zero waste; the purpose precedes the scope qualifiers. It is terse to the point of being slightly cryptic ('plan check'), but nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 params, no output schema, and safety fully covered by annotations, the remaining burden is describing what the check yields. '(counts, route soft limit)' hints at it but leaves the result semantics and limit-exceeded behavior unstated, which is a meaningful gap for a preflight/validation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (site, context, guestToken, requiredAssets, requiredRoutes) are already explained in the schema. The parenthetical '(counts, route soft limit)' loosely gestures at requiredAssets/requiredRoutes but adds no format or semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and scope: a plan/preflight check performed before a large build, with the things it examines (counts, route soft limit). An agent can grasp the intent without opening the schema, though it does not name or contrast any sibling (e.g. get_site_build_progress, estimate_test).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'before large build' is a real timing cue, but it is implied rather than explicit, and there is no guidance on when NOT to call it or which sibling to prefer for a smaller build. No thresholds or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_pageC
Read-onlyIdempotent
Inspect

Verify-only composed HTML check; home preview may include requiredDecision — ask mirror vs editable before structure_*.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
variantNoA/B variant arm to compose (e.g. "b") — same as ?dotsy-variant= on the public URL.
containsNoAssert substrings present in composed HTML.
selectorNoReturn matched elements only (tag/attrs/text).
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
notContainsNoAssert absent — e.g. `<dotsy-`, `<paige-` for unresolved template tags.

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds one genuinely useful behavioral fact — that a home preview may return a requiredDecision forcing a mirror-vs-editable choice — but leaves the response shape otherwise opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single compact sentence with the key constraint front-loaded, which is good. However, the density comes from unexplained internal jargon rather than tight phrasing, so brevity is achieved at the cost of legibility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema, the description omits what the check actually returns (aside from the requiredDecision aside) and doesn't clarify the mirror-vs-editable decision it references. An agent has to guess at the return contract and the workflow trigger.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 88%, so the schema already documents nearly all eight parameters (context, guestToken, contains/notContains, etc.). The description adds no syntax or meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

It signals a non-mutating verification of 'composed HTML', which is a recognizable verb+resource, but the phrasing is jargon-heavy ('composed HTML', 'Verify-only'). It does not clearly distinguish itself from the sibling verify_page, so an agent could confuse the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is the cryptic fragment 'ask mirror vs editable before structure_*', which hints at a conditional workflow but names no clear when-to-use condition or alternative tool. There is no explicit routing versus verify_page or edit_page.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

promote_experimentA
Destructive
Inspect

Promote a winning A/B variant to live — explicit user instruction only; see get_skill('verify') § A/B.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
variantIdYesWinning variant arm id (usually "b").
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
experimentIdYes
loserDispositionNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is known. The description adds genuinely new context beyond the annotations: the call requires explicit user consent and is governed by a verification skill section. It still does not say what 'live' does to the losing variant or whether the change is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence, front-loaded with the action, then the consent constraint, then the reference. Every clause carries information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema and 67% parameter coverage, the description covers the consent requirement and points at the procedural doc, which is enough to act safely. It leaves the post-promotion state (what happens to the loser, reversibility) unstated, and does not compensate for the undocumented parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description contributes no parameter meaning at all, while schema coverage is only 67%: experimentId has no description, and loserDisposition is an enum ('discard'/'keep_draft') with no explanation of the trade-off — exactly the semantics a sentence here could supply. With moderate coverage and zero compensating text, this falls below the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('promote a winning A/B variant to live'), which cleanly separates it from siblings like start_test, estimate_test, and set_experiment_note. It stops short of explicitly naming those alternatives, so differentiation is inferred from the phrase 'winning variant' rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Explicit user instruction only' is a concrete when-to-use gate that tells the agent not to invoke this autonomously, and the pointer to get_skill('verify') § A/B routes it to the governing procedure. It does not name a counterpart for the not-yet-decided case (e.g. use start_test), so the guidance is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

publish_siteB
Destructive
Inspect

Publish draft to live. feedback and handoff required (use dummy feedback + one-line handoff if clean).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
pagesNoPublish only these pages — use when a publish returns rules_not_fully_checked; publish the rest in further batches (≤50 route paths, e.g. /pricing).
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
handoffYesOne sentence: decisions, preferences, or open threads for the next session (appended to the brief).
feedbackYesRequired. Canonical dummy OK on friction-free sessions; tool intents attached automatically.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
acknowledgeSeoWarningsNoRequired when brief.priorities.seo is must and verify_page reported SEO errors (not warnings). User must have explicitly chosen to publish despite errors.
confirmedDraftFingerprintNoRequired when get_draft_status returned requiredDecision (publish_confirm). Must match the current draft fingerprint from get_draft_status.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the agent knows this mutates live state. The description adds the required-payload workflow (dummy feedback acceptable), but never states that publishing overwrites live content or what the draft-to-live transition does to the existing site, leaving a meaningful gap for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, with the core action front-loaded before the payload requirement. Nothing is wasted, though the parenthetical phrasing is slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter destructive tool with nested objects and no output schema, the description covers the mandatory payloads but omits the conditional branches (SEO acknowledgement, fingerprint confirmation) and any mention that sub-batching via 'pages' exists for partial publishes. The schema compensates, but the description alone is thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all eight parameters, including the conditional requirements for acknowledgeSeoWarnings and confirmedDraftFingerprint. The description only echoes the feedback/handoff requirement already spelled out in the schema, adding no syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource: 'Publish draft to live.' An agent can distinguish this from read/inspect siblings like get_draft_status or preflight_site. It doesn't explicitly name which sibling to use first, but the action itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states that feedback and handoff are required and offers a workaround for clean sessions ('use dummy feedback + one-line handoff if clean'), which is useful invocation guidance. However, it never says when to publish versus when to run preflight_site, verify_page, or get_draft_status first, so sequencing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

put_contentA
DestructiveIdempotent
Inspect

Stage one content key. Body ≤60KB. Larger pages: stage_content_chunk + apply_changeset bodyStaging, dotsy push, or PUT content-raw. Edits to an existing large page: batch_patch_content.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
bodyYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
contentTypeNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the safety profile is covered. The description adds the operationally critical 60KB body limit and the escalation paths when it is exceeded, which prevents failed calls, but it never says what staging does to existing content at that key or that a publish step is still required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four telegraphic fragments, zero waste, with the core action and size limit front-loaded ahead of the routing alternatives. The terseness is efficient but leans on unexplained references ('dotsy push', 'bodyStaging') that cost some clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Safety is carried by annotations and there is no output schema to explain, so the description mostly needs to cover staging semantics. It covers limits and alternatives well but omits what 'staging' implies operationally (draft vs live, required follow-up publish) and leaves contentType unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: site, context and guestToken are documented in the schema, while key, body and contentType are not. The description only compensates for 'body' by stating its 60KB ceiling; it adds nothing about the required 'key' format or 'contentType'. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening 'Stage one content key' gives a specific verb (stage) and resource (a content key), and the body-size limit follows immediately. It is clearly differentiable from siblings like stage_content_chunk, apply_changeset and batch_patch_content, though 'stage' is domain jargon that the description never defines.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states an explicit threshold ('Body ≤60KB') and then names concrete alternatives for the over-limit case (stage_content_chunk + apply_changeset, dotsy push, PUT content-raw) and for edits to an existing large page (batch_patch_content). An agent knows when to pick this tool versus a sibling without guessing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

put_goalsA
DestructiveIdempotent
Inspect

Create or replace all site conversion goals (≤50). Required before start_test when no goal exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
goalsYes
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, and the description's "replace all" is consistent with that. It adds genuine context beyond the annotations: the ≤50 ceiling and the ordering prerequisite relative to start_test. It does not explain the fate of pre-existing goals or impact on running tests.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the primary action and immediately followed by the prerequisite. No filler, no repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, idempotent replacement tool with no output schema, annotations cover the safety profile and the description covers scope plus the start_test prerequisite. Missing only what happens to existing goals and any auth nuance, which is minor given the rich schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema itself documents site, context (including the activity-log/no-secrets warning), guestToken, and the goal variants in detail. The description only adds the ≤50 item cap, which the schema already enforces via maxItems. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with clear scope: "Create or replace all site conversion goals (≤50)." An agent can immediately distinguish this from the sibling list_goals (read) and understand it is a full-replacement write.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the dependency explicitly — "Required before start_test when no goal exists" — giving the agent a concrete trigger and referencing a sibling tool. It stops short of naming alternatives (e.g. list_goals to check first) or stating when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

put_menu_docA
DestructiveIdempotent
Inspect

Replace the whole menu doc (ifMatch from get_menu_doc required).

ParametersJSON Schema
NameRequiredDescriptionDefault
docYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
ifMatchYesWhole-doc etag from get_menu_doc. Required (sent as If-Match).
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so safety is covered. The description adds the optimistic-concurrency requirement (ifMatch sourced from get_menu_doc), which is genuine behavioral context not present in the annotations. It stops short of saying existing menus are discarded on replace.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the operation and the key precondition with zero filler. Nothing could be cut without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive whole-document replace with a deeply nested object and no output schema, the description covers the operation and the concurrency token but omits what happens to prior menus, failure behavior on an ifMatch mismatch, and how the doc payload maps to the existing document. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents site, context, ifMatch, and guestToken. The description's only parameter note (ifMatch from get_menu_doc) restates the schema. The large nested `doc` structure is undocumented in both places, but the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ("Replace the whole menu doc") and the word "whole" implicitly distinguishes it from single-item siblings like put_menu_item. It's clear but doesn't explicitly name the alternatives it differs from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical "ifMatch from get_menu_doc required" implies the prerequisite call sequence but never states when to choose this whole-doc replace over put_menu_item, delete_menu_item, or reorder_menu. Usage is only inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

put_menu_itemA
DestructiveIdempotent
Inspect

Partial upsert one menu item: provided fields overwrite; omitted html/childTemplate/childShell/attrs/data are preserved. To remove those keys, use put_menu_doc. A brand-new id dedupes by (menu, ref) — reuses an existing item linking the same target instead of creating a duplicate. An id that already names an item always updates that exact item, even if another item links the same page.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
refYes
dataNo
htmlNo
kindYes
menuYes
siteYesSite slug, e.g. acme-studio
sortYes
attrsNo
labelYes
newTabNo
parentYes
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
childShellNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
childTemplateNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=true. The description adds non-obvious behavior: partial update preserves omitted fields, a new id dedupes by (menu, ref), an existing id forces update of the exact item. This goes beyond the annotations by explaining the dedup and targeting semantics, though it doesn't discuss auth or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences with no filler. The partial-upsert rule is front-loaded, then the delete alternative, then the dedup/targeting edge cases. Every sentence adds decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter mutation tool with no output schema and low schema coverage, the description covers the critical semantics: partial update behavior, dedup rule, and id targeting. It omits some parameter details and auth requirements, but given annotations already flag destructive and idempotent behavior, it is largely complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 19% (only site, context, guestToken documented), so the description must compensate. It does by explaining which fields (html/childTemplate/childShell/attrs/data) are preserved when omitted and the significance of the id and ref fields. It does not explain other parameters like sort, parent, kind, newTab, but the key upsert fields are covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (partial upsert) and resource (menu item), and precisely distinguishes the semantics: provided fields overwrite, omitted fields preserved. Explicitly names the sibling put_menu_doc for key removal and contrasts behavior, so an agent can distinguish it from siblings like put_menu_doc and reorder_menu.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to use it (partial upsert of one item) and names an alternative (put_menu_doc) for the delete-keys case. It does not, however, cover when not to use it versus other menu operations, though no stronger sibling ambiguity seems present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

put_postB
DestructiveIdempotent
Inspect

Stage a blog post via structured meta+body (?format=post). Key e.g. blog/my-post.html. Declared membership: meta.collections; schedule: meta.visibleAt; SEO title override: meta.seoTitle (card always shows meta.title); custom fields: meta.custom.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
bodyYes
headNo
metaYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=true, and the description's use of 'Stage' usefully signals the content is not immediately live — context beyond the annotations. However, it never warns that staging a post at an existing key overwrites/updates previous content, nor explains the staging-to-publish lifecycle, which is the most important behavioral fact for a destructive write.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then a compact semicolon list of meta-field meanings; every clause carries information and there is no filler. Slightly dense/telegraphic for an agent without prior Dotsy context, but the structure is sound.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and this is a complex tool with a nested meta object and 5 required params, so the description should do more. It documents several meta fields but omits the body/head shape, the guestToken/claim flow, and — critically — the lifecycle step of publishing staged content, leaving the agent with an incomplete mental model.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, so the description must compensate, and it partly does: it explains meta.collections as 'declared membership', meta.visibleAt as 'schedule', meta.seoTitle as a title/og:title override with 'card always shows meta.title', meta.custom as custom fields, and gives a key format example (blog/my-post.html). It leaves body format, head, tags, image, excerpt, author, date, url and dateDisplay unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Stage a blog post') and even flags the wire format ('via structured meta+body (?format=post)'), which lets an agent distinguish it from the generic put_content/patch_content siblings. It stops short of explicitly contrasting with new_post_template or put_content, so it's clear but not fully sibling-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies 'staging' as the mode, but gives no explicit when-to-use guidance, no mention of when this is preferable to new_post_template, put_content, or edit_page, and no prerequisites (e.g. auth/guestToken flow, or that publishing happens later via publish_site).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

put_routeB
DestructiveIdempotent
Inspect

Create/replace route; contentKey derived from path if omitted. On a field that accepts null (contentKey, redirectTo, redirectStatus, preserveQuery), sending null clears it back to its default.

ParametersJSON Schema
NameRequiredDescriptionDefault
navNoDEPRECATED — use put_menu_item / menu API. Menu memberships on the route row. Still accepted during transition and synced to menu items. See get_skill('nav').
pathYes
siteYesSite slug, e.g. acme-studio
varsNo
titleNo
layoutNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
noindexNoExclude from sitemap/RSS; inject noindex meta (0.69).
purposeNo
navLabelNoDEPRECATED — use put_menu_item / menu API. Menu link text for <dotsy-nav> (separate from title). Still accepted during transition and synced to menu items.
contentKeyNoExplicit content key (defaults to a path-derived key when omitted). null clears an explicit key back to that path default — same as omitting it, but usable on a PATCH to undo a prior set.
experimentNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
redirectToNoRedirect target, a root-relative path (/…). Required when purpose is "redirect". null clears it — only valid once purpose stops being "redirect" (a redirect route always needs a target).
preserveQueryNoAppend the visitor's original query string to the redirect target (default true on a redirect). null clears an explicit override.
redirectStatusNo301 permanent / 302 or 308 temporary (default 301 on a redirect). null clears an explicit override.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true and readOnlyHint=false, so the overwrite/idempotency profile is covered. The description's null-clearing rule is useful but largely restates what the contentKey, redirectTo, redirectStatus and preserveQuery schema descriptions already say; it adds no new behavioral context such as auth requirements or what a replace destroys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficiently packed sentences with the core action front-loaded, and the null-clearing caveat placed after the primary purpose. No wasted words, though the second sentence is dense and reads as a reference note rather than usage guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter destructive mutation with no output schema and rich nested/experiment structure, the description is thin. It omits the deprecated nav/navLabel transition, the experiment round-trip caveats, the required context/guestToken fields, and the required-when-redirect rule, all of which live only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 63%, and the description only touches contentKey and the four nullable fields, both already documented in the schema. Many parameters (title, layout, vars, nav, experiment, guestToken, purpose) are given no added meaning beyond the schema, so the description only marginally compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource: 'Create/replace route', which clearly signals PUT-style upsert semantics. It does not, however, distinguish itself from the sibling patch_route, leaving the agent to infer the create-vs-patch distinction from the names alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives such as patch_route or delete_route. The only usage-like content is a narrow rule about null clearing, which does not help the agent choose this tool over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reapply_custom_domain_oauthA
Idempotent
Inspect

Re-apply DNS with stored Cloudflare OAuth grant (no browser).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (not read-only, open-world, idempotent, non-destructive), so the bar is lower. The description usefully adds the auth path detail that no browser interaction is needed, but omits what happens on failure, whether an existing grant/domain is required, and what state it mutates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action and the key constraint front-loaded; every word earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an idempotent, open-world mutation with no output schema, the description is minimally adequate but leaves key operational questions unanswered: required prior state (connected domain/grant), failure behavior, and how it relates to the several other custom-domain tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents site, context, and guestToken fully. The description mentions no parameters at all and adds nothing beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Re-apply DNS') plus the mechanism ('stored Cloudflare OAuth grant'), which is enough for an agent to tell it apart from generic domain siblings like add_custom_domain or refresh_custom_domain. It stops short of explicitly naming or contrasting those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(no browser)' implicitly signals the use case: re-establishing DNS without an interactive OAuth flow. But it never states when to prefer this over refresh_custom_domain, wait_for_custom_domain, or get_custom_domain_connect, nor any prerequisites such as an existing connected domain or grant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refresh_custom_domainC
Idempotent
Inspect

Poll one hostname until DNS/cert propagates.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostYesHostname to refresh, e.g. www.example.com
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true, destructiveHint=false, openWorldHint=true, and notably readOnlyHint=false. The description's 'poll until propagates' hints at repeated/blocking behavior but does not explain duration, blocking semantics, retry cadence, or why a polling tool is flagged non-read-only. It adds little beyond the annotation profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the verb and target, with no filler. It is efficiently sized, though almost too terse given the tool's polling complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-read-only, open-world polling tool with no output schema, the description omits how long it polls, what it returns on success/failure, and how it relates to wait_for_custom_domain. An agent cannot predict behavior or outcome from this one-liner.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (host, site, context, guestToken) are already documented, including the context-logging caution. The description adds nothing about parameter meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (poll) and its target (one hostname's DNS/cert propagation), which is more than a tautology. However, it does not distinguish this tool from the very similarly-named sibling wait_for_custom_domain, nor from get_custom_domain_connect, leaving an agent unable to tell them apart from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus wait_for_custom_domain or get_custom_domain_connect, and no prerequisites or exclusions. The near-duplicate sibling makes the absence of routing guidance a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reindex_collectionA
Idempotent
Inspect

Build/finish collection indexes after publish returns collectionDataPending or collectionSqlPending.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds the important operational condition for when the write is needed, which goes beyond the annotations, though it does not discuss permissions, rate limits, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that contains no filler and states the core action and the precise trigger condition. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with rich annotations and no output schema, the description gives the essential trigger and purpose. It could say more about error handling or expected duration, but it is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (site, context, guestToken) are fully documented in the schema itself. The description adds no parameter-level details, yielding the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (build/finish) and resource (collection indexes) with no ambiguity. The trigger condition (after publish returns collectionDataPending or collectionSqlPending) further distinguishes it from all siblings, none of which are named reindex.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to invoke this tool: after publish returns collectionDataPending or collectionSqlPending. It does not name alternatives or state when not to use it, but the trigger is precise and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_custom_domainC
DestructiveIdempotent
Inspect

Remove one custom hostname from the site.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostYesHostname to remove
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say what removal destroys (verification records, certificates), whether it is reversible, or what downstream state changes. With zero added behavioral context, only the annotation floor holds this up.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single eight-word sentence with the action front-loaded and no wasted clauses. It is efficient, though its brevity is partly under-specification rather than disciplined economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a destructive, open-world mutation with no output schema, and the description supplies none of the consequence information an agent needs — what the removal breaks, whether it is recoverable, or what happens to related DNS/verification state. Annotations cover the safety flag, but the broader context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters including the important context and guestToken fields are documented in the schema itself. The description's 'one custom hostname' loosely maps to the host parameter but adds no syntax or format detail. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Remove') and resource ('custom hostname') with a clear scope limit of one hostname. It does not name the sibling add_custom_domain or otherwise differentiate against the many domain-related siblings (refresh_custom_domain, set_primary_custom_domain), but the purpose itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no alternatives. The agent is not told whether the domain must first be disconnected, whether this requires a guest token separately (that lives only in the schema), or what to use instead if the intent is to change rather than drop a domain.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_contentA
Destructive
Inspect

Re-key a page by slug (draft or published static). Server-side hash copy — no body round-trip. Default deletes static routes (obscure-URL); movePath:true repoints the route, rewrites menu refs and other redirects aimed at the old path, and — only for an already-PUBLISHED page — leaves a 301 there too (leaveRedirect, default true; no effect on a never-published page, or when movePath is false). Retry-safe: a part-way failure (409 rename_failed) rolls everything back, so resend the same call. Then publish_site.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesCurrent content key, e.g. dashboard.html
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
newSlugYesNew filename stem; directory prefix preserved (blog/foo → blog/bar)
movePathNofalse (default): delete routes bound to old key. true: move static route path to match slug.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
leaveRedirectNoWith movePath:true on an already-published page, leave a 301 at the old path (default true). false removes the old path outright — old links break. No effect when movePath is false, or on a page that was never published.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing the default destructive behavior (deletes static routes / obscure-URL), the full scope of movePath (rewrites menu refs, redirects, and leaves a 301 only for published pages), and atomic rollback semantics on partial failure. This is exactly the authorization/rollback/failure detail the destructiveHint=true annotation cannot carry on its own.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core operation and keeps everything in one dense but purposeful paragraph; every clause carries information (defaults, rollback, follow-up). It is heavily parenthetical, which slightly hurts scannability, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with 7 parameters and no output schema, the description covers defaults, conditional behavior, failure mode and recovery, and the required next step. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters, including movePath and leaveRedirect semantics. The description restates those defaults and effects clearly but adds no syntax or constraint detail that is not already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: "Re-key a page by slug," backed by concrete mechanics (server-side hash copy, no body round-trip) and a concrete example key format. It does not explicitly differentiate itself from the sibling rename_page, so an agent must infer that this operates on the content key rather than the page object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly spells out when each mode applies: default deletes static routes, movePath:true repoints the route, and leaveRedirect only has effect with movePath:true on an already-published page ("no effect on a never-published page, or when movePath is false"). It also names the required follow-up ("Then publish_site") and tells the agent to resend the identical call after a 409 rename_failed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_pageA
Destructive
Inspect

Change a page's URL: moves its route from from to to. Menu links and other redirects that pointed at the old path retarget automatically; a 301 is left at the old path by default (leaveRedirect). Retry-safe (resend the same call after any error). Errors: 404 route_not_found (from isn't a route); 400 reserved_route (to is under /__dotsy/) or invalid_body (missing/malformed from or to); 422 rename_same_path (from equals to already), rename_pattern_route (from or to contains a :param segment — pattern routes can't be renamed this way), rename_redirect_route (from is itself a different redirect — edit its redirectTo instead); 409 rename_experiment_running (an A/B test is running on from — end it first) or rename_target_exists (to is already a different route — unless to already serves the exact content from would, which is treated as already done and returns 200). Then publish_site.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesNew route path, e.g. /new-page
fromYesCurrent route path, e.g. /old-page
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
leaveRedirectNoLeave a 301 from the old path to the new one (default true). false removes the old path outright — old links break.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: menu links and redirects retarget automatically, a 301 is left at the old path by default, and the 'retry-safe' guarantee plus the full error-code catalogue (404/400/422/409 with distinct causes) tell the agent what will happen and how to recover. This is unusually rich behavioral disclosure for a mutation tool whose annotations only give destructive/idempotent hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and semantics, then error cases. Dense but every clause carries information; the long semicolon-chained error list is somewhat heavy, though it maps to distinct recovery paths an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no output schema, the description covers outcome (redirects, 301), retry behavior, all failure modes mapped to specific codes, and the required next step (publish_site). Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: it defines from/to as route paths and explains leaveRedirect's default (true) and the consequence of false ('old links break'), which the schema only hints at.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Change a page's URL: moves its route from `from` to `to`'), and by naming the route-level mechanism it distinguishes itself from siblings like rename_content or patch_route. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when this applies (renaming a page's route) and closes with 'Then publish_site', implying the follow-up step. It does not explicitly contrast against patch_route/put_route, so an agent must infer the routing-vs-content split from the error list rather than from a direct alternative statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_siteA
DestructiveIdempotent
Inspect

Rename the site slug (public hostname {slug}.dotsy.site). Owner-only. After success: use the returned slug in every later tool. Published site: the old slug redirects on the API (301 for reads, 308 for writes) while it stays reserved; the old host 301s and stays reserved until delete or rename-back. Never published: the old slug 404s on the API and the old hostname is released (no 301; create_site can reuse it). Analytics history stays on the old slug. HTML absolute URLs are not rewritten — publish_site refreshes sitemap; find_replace if stored HTML must change.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
newSlugYesNew site slug (lowercase hostname stem)
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations (destructiveHint/idempotentHint/readOnlyHint): it details redirect behavior for published sites (301 reads, 308 writes, host retained), 404/release behavior for never-published sites, analytics history retention, and that absolute HTML URLs are not rewritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and owner constraint are front-loaded, and the dense redirect/analytics/HTML clauses each carry distinct operational value. It is longer than typical but nearly every sentence earns its place, with only slight density cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, yet the description covers the key return contract ('use the returned slug') plus the divergent post-rename states an agent must reason about. Complete enough to call correctly and act on the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description conveys the returned-slug contract and slug/hostname semantics but adds no syntax or format detail for site, newSlug, context, or guestToken beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource plus the concrete effect: 'Rename the site slug (public hostname {slug}.dotsy.site)'. The 'Owner-only' qualifier and slug focus clearly distinguish it from rename_page, rename_content, and update_site.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes to alternatives with conditions: use the returned slug in later tools, use publish_site to refresh the sitemap, use find_replace if stored HTML must change. It also names the owner-only prerequisite, giving both when-to-use and when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reorder_menuA
Idempotent
Inspect

Set final menu order. Pass ifMatch from get_menu; stale → 409 with current items.

ParametersJSON Schema
NameRequiredDescriptionDefault
menuYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
ifMatchNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
orderedIdsYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (idempotentHint=true, destructiveHint=false), it discloses the optimistic-concurrency contract: a stale ifMatch returns 409 with current items. That is a genuinely useful behavioral trait not derivable from structured fields, though nothing is said about partial reorder or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short clauses, action first, failure mode second. No filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema and 6 parameters, it covers the operation and its failure mode but omits key semantics: whether orderedIds must be exhaustive, whether nesting is affected, and any permission/auth prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (site, context, guestToken documented). The description fills the most important gap by explaining ifMatch's provenance, but orderedIds and menu remain undocumented in both places despite orderedIds being the core payload.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

'Set final menu order' is a precise verb+resource and clearly distinct from put_menu_item, delete_menu_item, and get_menu. It does not explicitly name how it relates to the sibling mutators (put_menu_doc, put_menu_item), but the scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It tells the agent to source ifMatch from get_menu and what happens when it is stale, which is real usage guidance. It never states when to reach for this instead of put_menu_item / put_menu_doc, so the alternatives are left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revert_changeAInspect

Undo one activity event by staging the prior version as draft only — call publish_site to go live. Respects AI access.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesActivity event id from get_activity.
keysNoUndo only these keys of the change. Default: every key.
slugYesSite slug, e.g. acme-studio
reasonNoWhy — shown next to 'Reverted by …'. Visible to everyone on the site.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds real behavioral value beyond annotations: it discloses that the operation does not go live but only stages a draft, and that publishing is a separate step. 'Respects AI access' gestures at a permission gate, though it is too vague to be actionable. Annotations already cover readOnly=false and destructive=false, so the staging detail is genuine added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and tightly compressed, with no filler sentences. The trailing fragment 'Respects AI access' is terse to the point of ambiguity, but the overall structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations and no output schema, the description covers the essential non-obvious behavior (draft staging, publish follow-up). It leaves the permission semantics and any failure modes underexplained, but the core contract is complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all six parameters (including id, keys, slug, reason, context, guestToken) are fully documented in the schema itself. The description adds no parameter-specific syntax or defaults beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: undo one activity event, with the crucial scope that it stages the prior version as a draft rather than reverting live content. An agent can tell what it does, though it never explicitly distinguishes itself from the sibling adapt_site_revert, which a reader must disambiguate on their own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied through the chained instruction 'call publish_site to go live', which tells the agent what to do next but not when to prefer this over adapt_site_revert, apply_changeset, or discard_drafts. No explicit when-to-use or exclusion conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_contentA
Read-onlyIdempotent
Inspect

Find pages containing a phrase (grep without reading bodies). Requires do-sqlite — graduate_site first if search_unavailable.

ParametersJSON Schema
NameRequiredDescriptionDefault
qYes
siteYesSite slug, e.g. acme-studio
limitNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world behavior. The description adds useful non-annotation context: it operates like grep without reading bodies, depends on do-sqlite, and may be unavailable unless graduate_site is run first. It does not cover result format or pagination, but the annotation bar is met well.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the purpose front-loaded and the operational prerequisite immediately after. Every clause carries information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with annotations covering safety and a schema covering most auth/context fields, this is nearly complete: it states purpose, dependency, and fallback. It is slightly thin on return behavior because there is no output schema and no mention of result limits or pagination.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 5 parameters and only 60% schema description coverage, the description should compensate for undocumented parameters. It clarifies that 'q' is a phrase search, but says nothing about required 'site', required 'context', 'limit', or 'guestToken', leaving important parameter semantics unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find pages containing a phrase') and adds a concise analogy ('grep without reading bodies') that distinguishes it from content-reading siblings such as get_content. An agent can identify the tool from the sentence alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear prerequisite ('Requires do-sqlite') and a fallback path ('graduate_site first if search_unavailable'), which is meaningful context for invoking it correctly. It does not explicitly compare against alternatives like find_replace or list_content, so it stops short of full usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_experiment_noteA
Idempotent
Inspect

Record what was learned from an ended A/B test (≤500 chars; null clears).

ParametersJSON Schema
NameRequiredDescriptionDefault
noteYesThe learning, in plain words (≤500 chars, multiline ok) — null clears it.
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
experimentIdYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds real behavioral detail beyond that: a 500-character cap and that passing null clears the existing note, which tells the agent this is an idempotent overwrite rather than an append.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action, the scope, and the two key behavioral constraints in one breath. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 5-parameter setter with no output schema and rich schema-level descriptions for context and guestToken, the description covers the essential nuance (length limit, clearing behavior, test must have ended). It omits any note about authorization/guest-token flow, but the schema carries that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80% and the schema already documents note (max length, multiline, null clears), site, context, and guestToken. The description's '≤500 chars; null clears' merely restates the note parameter's schema text and adds nothing about site, experimentId, or context. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (record) and resource (learning/note on an ended A/B test), which cleanly separates it from sibling experiment tools like start_test, estimate_test, and promote_experiment. It does not explicitly name a sibling, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'from an ended A/B test' implies the precondition (only after the test concludes) but gives no explicit when-not guidance and names no alternative tool for capturing experiment learnings. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_primary_custom_domainA
Idempotent
Inspect

Set the canonical host (301 target) when multiple hostnames are connected.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostYesConnected hostname to make primary
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare mutation (readOnlyHint=false), idempotence, and non-destructiveness. The description adds meaningful semantic context by identifying the effect as setting the '301 target', but does not explain what happens to other hostnames, auth requirements, or side effects, so it adds only moderate value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single focused sentence with no wasted words. The purpose and the key precondition are front-loaded, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich annotations and full schema descriptions, the description supplies enough to select and invoke the tool correctly. It could mention verification steps or related tools (e.g., wait_for_custom_domain), but for a simple idempotent set operation the core context is covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (host, site, context, guestToken) are fully documented in the schema. The description adds no parameter-level guidance, which is acceptable at baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Set') and resource ('canonical host') with the technical scope ('301 target'). It does not explicitly name any sibling tool, so it lacks the explicit sibling differentiation required for a 5, but the action is clearly distinct from add/remove/list domain operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear usage condition: 'when multiple hostnames are connected', which tells the agent this tool is only relevant in a multi-host scenario. It does not mention when-not-to-use or point to an alternative tool, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stage_content_chunkAInspect

Append one ≤60KB chunk of a large new page. Omit stagingId on seq 0; pass the returned stagingId on later seq. Set finalize:true on the last chunk, then apply_changeset with writes[].bodyStaging.

ParametersJSON Schema
NameRequiredDescriptionDefault
seqYes
bodyYes
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
finalizeNo
stagingIdNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
contentTypeNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the safety profile (readOnly=false, destructive=false, idempotent=false), and the description adds real context beyond that: staging is stateful across calls, the stagingId is returned and must be reused, and the chunks are inert until apply_changeset is called. It does not mention auth/guest-token requirements or rate limits, but the key non-obvious behavior (staging ≠ publishing) is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core action and size constraint, then the cross-call protocol. Zero filler; every clause carries operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful multi-call staging tool with no output schema, the description covers the full lifecycle and the follow-up apply_changeset with writes[].bodyStaging. It leaves the return payload shape and error/validation behavior (e.g. what happens on a failed chunk or bad seq) unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 38%, so the description has to compensate, and it does for the critical params: seq 0 vs later semantics for stagingId, the meaning of finalize, and the 60KB body limit (which the schema's minLength:1 does not convey). contentType and the guestToken/context semantics are left to the schema, so it is not fully compensating.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+constraint: 'Append one ≤60KB chunk of a large new page.' Combined with the stagingId/finalize protocol, an agent can tell this apart from put_content, bulk_put_content, and patch_content without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit call-sequence guidance: omit stagingId on seq 0, pass the returned stagingId on later seqs, set finalize on the last chunk, then call apply_changeset. It never names an alternative tool (e.g. put_content) for the small-page case, so the boundary condition is implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_edit_sessionCInspect

Open a draft session. Returns preview/edit URLs for the person (see linkGuidance). draftEditorUrl only when the site gives AI agents Full access. path defaults to '/'.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly=false, destructive=false, and idempotent=false, so the safety profile is covered. The description adds genuine context beyond that: the return content (preview/edit URLs) and a permission gate (draftEditorUrl only when the site grants AI agents Full access). It still doesn't say what the session mutates, how long it lives, or whether it must be discarded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and the return value. The only waste is the unresolvable "(see linkGuidance)" pointer, which costs a clause without conveying information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and four parameters, the description should carry more: the shape of the returned URLs is hand-waved via an external, undefined reference, and session lifecycle (creation, reuse, discard via discard_drafts) is absent. For a session-establishing tool this leaves the agent guessing about next steps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with site, context, and guestToken already documented in-schema; the description's contribution is limited to clarifying that path defaults to '/'. It adds no format or constraint detail for path or for how site/path interact in a session.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Open a draft session" states a verb and resource, but "draft session" is never defined, so an agent cannot tell what state change this creates or how it relates to siblings like edit_page, apply_changeset, or discard_drafts. It does concretely say it returns preview/edit URLs, which adds some specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to open a session versus editing directly with edit_page/put_content, nor when to end one. The only routing hints are incidental: the guestToken parameter description points at start_site and claim_site, and the parenthetical "see linkGuidance" is not resolvable from anything in the tool definition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_siteAInspect

Start one draft site with no Dotsy account. Returns slug + guestToken — pass guestToken on every later tool call until claim_site.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
purposeNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-readOnly, non-idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuinely useful operational context beyond that: it returns slug + guestToken and the guestToken must be passed on every subsequent call until claim_site – an auth/credential handling detail the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with zero filler; the creation scope comes first and the credential requirement second, which matches the order an agent needs them.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description covers the essential return contract (slug, guestToken) and the critical lifecycle rule for the token. The only real gap is the undocumented title/purpose inputs, which leaves the definition slightly short for a 3-parameter creation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% – 'title' and 'purpose' are undocumented in both schema and description, and the description adds no meaning for any of the three inputs. It does describe the return values (slug + guestToken), but that is output, not parameter semantics, so it cannot compensate for the input gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start one draft site') plus the distinguishing constraint 'with no Dotsy account', which separates it from create_site/graduate_site. The mention of claim_site as the follow-on step anchors the guest flow, so an agent can identify this tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'no Dotsy account' and by the instruction to pass guestToken until claim_site, but the description never explicitly contrasts this with create_site or states when a guest site is preferred over a real one. Adequate as a hint, not as routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_testCInspect

Start a same-URL A/B test (draft or live source). Rules → get_skill('verify') § A/B.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
pathYes
siteYesSite slug, e.g. acme-studio
goalIdYes
sourceYes
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
weightsNo
handsOffNoHands-off mode (0.56e M4): Dotsy finishes the test WITHOUT further user instruction — the LIVE SITE may change later on its own (a ≥95% winner is auto-promoted, old page kept as a draft; or after 30 days with no clear winner the test auto-ends keeping the control). Only set when the user explicitly asked for hands-off / auto-finish.
changeNoteNoOne line on what the variant changes, shown on the test card (≤120 chars).
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
hypothesisNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the agent knows this is a non-idempotent write. The description adds the useful "same-URL" scope and the draft/live source dimension, but says nothing about what starting a test changes on the site or the guest-token flow — though the schema's handsOff and guestToken descriptions carry that load.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the action and the rules pointer. No waste, though for an 11-parameter mutation tool the brevity borders on under-specification rather than tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 11-param, 5-required creation tool with no output schema, the description omits the draft-vs-live consequence, lifecycle, and any return/next-step context, instead outsourcing everything to an external skill document the agent may not have.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 11 parameters and only 45% schema description coverage, the description must compensate and largely does not. It only gestures at the "source" enum; required params (site, path, goalId, context) and weight/hypothesis semantics get no explanation beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: "Start a same-URL A/B test," and narrows scope with "draft or live source." It is distinguishable from siblings like estimate_test, promote_experiment, and delete_draft_test, though it doesn't name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no alternative named (e.g., estimate_test to size the test first). The only direction is an outbound pointer to get_skill('verify') § A/B, which defers the actual rules rather than stating them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

structure_site_analyzeC
Idempotent
Inspect

Propose shared layout + fragments; persist:false first. Wait for mirror-vs-editable answer if requiredDecision pending.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoReturn the full manifest incl. per-page HTML samples (default: compacted — etag/action/counts kept, samples truncated — to fit the token budget).
siteYesSite slug, e.g. acme-studio
scopeNoDefault flat-only; analyze clusters flat keys only.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
persistNoFirst call: false (preview only). true only after user confirms editable/CMS mode.
profileIdNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
sampleLimitNo
refreshManifestNoReplace persisted manifest when true.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, destructiveHint=false and idempotentHint=true, so the safety profile is already provided. The description usefully adds the two-phase preview-then-persist workflow, which is behavioral context beyond the annotations, but the cryptic 'mirror-vs-editable' decision language adds confusion rather than clarity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and the core action is front-loaded, so it is not padded. But two of the three clauses are sentence fragments heavy with undefined internal vocabulary, which harms usability at this length rather than helping.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter analysis tool with no output schema, the description leaves major gaps: what a proposal contains, what happens after preview, and what 'requiredDecision' or 'mirror-vs-editable' actually mean. The schema carries most of the burden, and the description does not compensate where it is silent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, so the schema already documents most parameters, including persist's 'false first' semantics that the description duplicates. The description adds no syntax or format detail beyond what the schema provides, making baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'Propose' with 'shared layout + fragments' gestures at generating a structural proposal, but the phrasing is vague and never plainly says it analyzes an existing site's structure. It does not clearly distinguish itself from the sibling structure_site_apply, which an agent would need to choose between.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It does give workflow context ('persist:false first') and a conditional gate ('Wait for mirror-vs-editable answer if requiredDecision pending'), which is more than nothing. However, the conditions are written in terse jargon that an agent cannot reliably act on without guessing what requiredDecision means.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

structure_site_applyA
Destructive
Inspect

Rewrite flat pages to layout + fragments (draft). Do not call while requiredDecision pending — user must pick editable first.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysNoRequired when scope is keys.
modeNoDefault apply. Neither value is a dry run — review-only persists the same content/layout write as apply. Only pass review-only after a review_required error (low analyze confidence); it labels the result outcome "review" and defers primary-nav auto-wiring to a later apply call.
siteYesSite slug, e.g. acme-studio
scopeNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
ifMatchYesManifest etag from structure_site_analyze or get_structure_status.
overridesNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
restructureNoRequired true when scope all and site has structured pages.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so safety is covered structurally. The description adds genuine behavioral context beyond that: the write lands in draft state and is blocked by a pending requiredDecision gate. It stops short of saying what existing flat page content is lost on rewrite.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, with the action front-loaded and the precondition second. No filler. It is arguably too terse for a 9-parameter destructive tool, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-param, destructive, non-idempotent tool with no output schema, the description covers the operation and one gating rule but omits scope selection guidance and the meaning of review-only vs apply. The rich schema partly compensates, but an agent choosing scope/restructure still has to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, so the schema already documents most parameters including the important mode/ifMatch/scope semantics. The description adds no parameter-level meaning (e.g. scope values, overrides, restructure) beyond the word 'draft'. Baseline 3 is appropriate when the schema carries this much detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+transformation ('Rewrite flat pages to layout + fragments') and marks it as a draft operation. It implicitly separates itself from structure_site_analyze, though it never names the analyze sibling explicitly, so the differentiation is derivable rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-not condition: 'Do not call while requiredDecision pending — user must pick editable first.' That is a real gating rule an agent can act on. It does not, however, point to the alternative tool to consult status or resolve the decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_feedbackCInspect

Optional mid-session feedback (required feedback rides with publish_site).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
titleYes
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
categoryYesmissing_tool | blocking_error | extra_round_trips | confusing_response | workaround | other
findingsNo
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
descriptionNo
impactPercentYes
suggestedImprovementNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=false, so the safety/mutation profile is covered. The description adds only the optional-vs-required distinction and says nothing about what submitting does (activity-log write, public visibility, non-idempotent retry behavior), so it contributes little behavioral context beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler and the key routing fact (required feedback rides with publish_site) placed up front. It is efficient, though for a nine-parameter mutation tool the extreme brevity borders on under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 params, 5 required, no output schema, and low schema coverage, the one-line description leaves the agent without the information needed to fill required fields correctly (what category to pick, what impactPercent means, whether guestToken is needed). It is not misleading, but it is far from complete for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Nine parameters with only 44% schema description coverage, and the description mentions none of them — not site, category, impactPercent, title, context, nor guestToken. With required fields like impactPercent (1-100) and a sensitive public 'context' field, the description does nothing to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says it is 'mid-session feedback,' which gives a verb-ish action and a rough resource, and it explicitly separates itself from publish_site. But it never states what the feedback is about (agent/tooling experience) or what it does with the submission, so an agent must infer the domain from the schema's category enum.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear routing context: this tool is for optional mid-session feedback, while required feedback goes with publish_site. That is a genuine when-to-use/when-to-use-something-else signal, though it does not state any prerequisites (e.g. the guestToken requirement noted in the schema).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_briefC
DestructiveIdempotent
Inspect

Update _intent/site.md (site brief) or a page brief when page is set. See field descriptions for rules, notes, and head enums.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoReplace narrative body (site ≤50k chars; page ≤10k).
goalNoCurrent goal (site ≤4k chars; page ≤2k).
headNoHead fields to merge. Site: purpose, audience, kind (exploration|production), priorities ({seo,analytics,design,mobile}: must|nice|skip), editingModel (ai|simple-editor|creator), sections, design.{palette,type,layout}. Page: purpose, design.{direction,hero}.
pageNoContent key — targets that page's brief (`_intent/pages/<key>.md`). Omit for the site brief.
siteYesSite slug, e.g. acme-studio
notesNoSite: free-form ## Notes string (≤4k) OR per-section notes object for Design system / Content & CMS / SEO (null clears one). Page brief: string only (≤2k).
rulesNoRules & guidelines lines to add or change (upsert by id — send only fields you change; omitted fields keep their stored value; check: null turns a Rule into a Guideline). A supported check (title-max, title-suffix, description-length, one-h1, banned-words, image-alt, allowed-colors, allowed-fonts) makes a Rule — use level block for hard musts (banned words, title suffix), warn for softer checks. locked can only be set by a person; locked: false unlocks.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
removeRulesNoIds of rules or guidelines to remove (site brief only).
appendHandoffNoAppend a hand-off log line — site brief only; ignored when page is set.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, so the safety profile is covered structurally. The description adds only 'see field descriptions for rules, notes, and head enums' — a deferral that discloses nothing new about what gets destroyed, auth needs, or the activity-log side effect of `context`.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core action front-loaded. The second sentence is a deferral rather than added information, but the overall footprint is appropriately small for a tool whose schema carries the detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive 11-parameter mutation tool with no output schema, the description leans entirely on annotations for safety and on the schema for parameter meaning. Nothing critical is missing, but the guestToken auth flow and the destructive replacement/merge semantics of the update are not surfaced in the description itself.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters and their enums. The description's pointer to field descriptions adds no syntax or semantics beyond what the schema provides, making 3 the appropriate baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: updating `_intent/site.md` (the site brief) or a page brief when `page` is set. The target is named precisely enough to distinguish it from update_site or update_collection, though it never explicitly routes away from those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is 'when `page` is set' for the page-brief branch. There is no when-to-use/when-not guidance, no mention of prerequisites, and no named alternative tools such as update_site or update_collection to steer the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_collectionA
Idempotent
Inspect

Rename a collection's display name and/or noun (e.g. the generic 'posts' key shown as 'Recipes' / 'recipe') — cosmetic only, stored separately from the tag that defines membership; doesn't need publish. Provide at least one of name/noun. null (or "") resets that field back to its derived default — different from omitting it, which leaves the current value alone.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesCollection key exactly as list_collections returned it, e.g. "posts" or "blog/" (a hand-typed key may 404 or get silently normalized).
nameNoDisplay name (≤80 chars). null resets to the derived name.
nounNoSingular noun for one item, e.g. "post" (≤40 chars). null resets to the derived noun.
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (idempotent, non-destructive, non-readOnly) by disclosing the critical null-vs-omit distinction, the reset-to-derived-default behavior, and the fact that the change is purely presentational and needs no publish. These are exactly the traits an agent needs to avoid clobbering a field accidentally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence that front-loads the action and packs constraints efficiently; the parenthetical examples of name/noun are illustrative rather than filler. It is information-dense to the point of being slightly crowded, which keeps it just short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no output schema, the description covers the update semantics, the reset behavior, the cosmetic scoping, and the minimum-argument constraint, while the schema itself documents site, key, context, and guestToken. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds the semantics the schema omits: that null/"" resets a field while omitting it preserves the current value, and that at least one of name/noun is required (a constraint not enforced by the schema's required list). That is genuine added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Rename) and resource (a collection's display name and/or noun), and immediately clarifies the resource is the cosmetic label rather than the membership tag. This cleanly separates it from rename_content, rename_page, rename_site, and reindex_collection in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real usage context: cosmetic-only, no publish required, at least one of name/noun must be supplied. It doesn't explicitly route to a named alternative, but the scope statement ('stored separately from the tag that defines membership') implicitly tells the agent this is not the tool for changing membership.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_siteA
Idempotent
Inspect

Update site settings: title, keyConvention, or editingMode ('live' = each editor save publishes immediately; 'staged' = saves stay drafts until publish_site). Slug is the hostname — change it with rename_site, not this tool. Set 'staged' before handing the editor to a person via start_edit_session's draftEditorUrl. signature toggles the "Made with Dotsy" badge ({enabled, mark: logo|dot, reveal: pill|seal}); it is live immediately.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
titleNo
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
signatureNo"Made with Dotsy" signature badge. Takes effect on the live site immediately, no publish needed.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
editingModeNoHow the editor's Save behaves. 'live' (default) publishes each save immediately; 'staged' keeps saves as drafts until publish_site. Prefer 'staged' when an agent will hand the editor to a person (see start_edit_session → draftEditorUrl) so keystrokes don't go straight to the public site. API writes always stage regardless of this setting.
keyConventionNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, idempotent, non-destructive, so the safety bar is covered; the description adds genuinely non-obvious behavior: 'API writes always stage regardless of this setting', signature changes 'take effect on the live site immediately, no publish needed', and the staging implication for human handoff. These are traits an agent could not infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the field list comes first, then per-field clarifications and the sibling redirect. Every sentence carries behavioral or routing information; the signature sentence is the wordiest but still informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with a nested object and no output schema, the description covers the consequential fields and side effects well. Minor gaps remain (keyConvention values, no mention of the context/guestToken obligations), but those are documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, so the baseline is 3, but the description earns more by explaining editingMode ('live' publishes each save, 'staged' keeps drafts until publish_site) and the signature badge semantics (enabled/mark/reveal with hover behavior). keyConvention is named but its 'flat' vs 'directory' values remain unexplained, leaving one gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Update site settings') and enumerates the fields it touches (title, keyConvention, editingMode, signature). It explicitly distinguishes itself from the sibling rename_site ('Slug is the hostname — change it with rename_site, not this tool'), so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete conditions for choosing values: set 'staged' before handing the editor to a person via start_edit_session's draftEditorUrl, and route slug changes to rename_site. There is no general when-not guidance (e.g., when to use a more specific tool like update_collection or adapt_site_apply instead), but the actionable context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_pageA
Read-onlyIdempotent
Inspect

Pre-publish lint (tags, assets, layout). Pair with preview_page — see get_skill('verify').

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYesSite slug, e.g. acme-studio
pathsNoRoutes to check; omit to verify every route on the site.
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered structurally. The description adds the lint scope (tags, assets, layout) and points to get_skill('verify'), but says nothing about output format, blocking behavior, or failure reporting.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse fragments that front-load the purpose before the pairing hint. Nothing is wasted, though the fragmentary style ('see get_skill(verify)') is slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only lint tool with full schema coverage and annotations, the description covers purpose and pairing. It does not indicate what the lint returns or how failures surface, though the get_skill('verify') referral partially compensates and there is no output schema to imply structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all four parameters (including the notable 'context' logging constraint and guestToken flow) are already documented in the schema. The description adds no parameter-level detail, so it lands at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation (pre-publish lint) and enumerates what it checks (tags, assets, layout), which lets an agent distinguish it from siblings like preview_page. It never says 'page' explicitly, relying on the tool name, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Pre-publish' sets a clear triggering context, and 'Pair with preview_page' names the related sibling it is meant to accompany. No explicit when-not condition is given, but the pairing guidance is actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_custom_domainA
Idempotent
Inspect

Poll one hostname until it is connected or the timeout is reached.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostYesHostname to wait for, e.g. www.example.com
siteYesSite slug, e.g. acme-studio
contextYesOne sentence: why you are calling this tool and what you are trying to accomplish. Stored on the site's activity log and visible to everyone on the site — never put secrets or personal data here.
guestTokenNoGuest token from start_site. Send it on every call until claim_site. Omit when this chat is already signed in to Dotsy.
timeoutSecNoMax seconds to wait (default 300, max 600).
intervalSecNoSeconds between polls (default 30, min 10, max 60).

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is largely covered. The description adds the polling-until-connected-or-timeout behavior, but doesn't say what happens on timeout (error vs. status return) or what gets logged. Note the mild tension between 'Poll' (read-like) and readOnlyHint=false, though logging to the activity log plausibly explains the non-read-only hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb, scope, and termination condition are all in the first clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a polling tool with no output schema and full schema descriptions, the core intent is covered, but the description omits what a timeout yields (error or result), whether it blocks, and how guestToken/site auth factor in. Enough to call it correctly, but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so host, site, context, guestToken, timeoutSec, and intervalSec are all documented in the schema. The description adds no parameter-level detail, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (poll) and resource (one hostname) plus the termination condition, which clearly separates it from siblings like list_custom_domains or get_custom_domain_connect. It does not explicitly name an alternative, but the polling semantics are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the polling behavior: call it to wait for a hostname to become connected. There is no explicit when-to-use, when-not-to-use, or reference to a sibling alternative such as get_custom_domain_connect, leaving the agent to infer the context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 89 tool updates
    • First observedadapt_site_analyze
    • First observedadapt_site_apply
    • First observedadapt_site_revert
    • First observedadd_custom_domain
    • First observedapply_changeset
    • First observedbatch_patch_content
    • First observedbulk_put_assets
    • First observedbulk_put_content
    • First observedclaim_site
    • First observedcreate_site
    • First observeddecompose_analyze
    • First observeddelete_asset
    • First observeddelete_content
    • First observeddelete_draft_test
    • First observeddelete_menu_item
    • First observeddelete_route
    • First observeddelete_site
    • First observeddiscard_drafts
    • First observededit_page
    • First observedestimate_test
    • First observedfind_replace
    • First observedget_account_context
    • First observedget_activity
    • First observedget_adapt_status
    • First observedget_analytics
    • First observedget_asset
    • First observedget_content
    • First observedget_custom_domain_connect
    • First observedget_draft_status
    • First observedget_form_submission
    • First observedget_menu
    • First observedget_menu_doc
    • First observedget_route
    • First observedget_site
    • First observedget_site_build_progress
    • First observedget_site_manifest
    • First observedget_site_outline
    • First observedget_skill
    • First observedget_structure_result
    • First observedget_structure_status
    • First observedgraduate_site
    • First observedimport_site_from_url
    • First observedinspect_collection
    • First observedlist_assets
    • First observedlist_collections
    • First observedlist_content
    • First observedlist_custom_domains
    • First observedlist_form_submissions
    • First observedlist_goals
    • First observedlist_my_sites
    • First observedlist_pages
    • First observedlist_skills
    • First observednew_post_template
    • First observedpatch_content
    • First observedpatch_route
    • First observedpreflight_site
    • First observedpreview_page
    • First observedpromote_experiment
    • First observedpublish_site
    • First observedput_content
    • First observedput_goals
    • First observedput_menu_doc
    • First observedput_menu_item
    • First observedput_post
    • First observedput_route
    • First observedreapply_custom_domain_oauth
    • First observedrefresh_custom_domain
    • First observedreindex_collection
    • First observedremove_custom_domain
    • First observedrename_content
    • First observedrename_page
    • First observedrename_site
    • First observedreorder_menu
    • First observedrevert_change
    • First observedsearch_content
    • First observedset_experiment_note
    • First observedset_primary_custom_domain
    • First observedstage_content_chunk
    • First observedstart_edit_session
    • First observedstart_site
    • First observedstart_test
    • First observedstructure_site_analyze
    • First observedstructure_site_apply
    • First observedsubmit_feedback
    • First observedupdate_brief
    • First observedupdate_collection
    • First observedupdate_site
    • First observedverify_page
    • First observedwait_for_custom_domain

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to publish static sites and build folders as live websites, push updates, provision and manage custom domains with DNS, and send transactional email. Ships as a local pip/uvx server that reads build folders directly, or as a hosted HTTP endpoint for Claude Code, Claude Desktop, and other MCP clients.
    11
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for self-hosted static site publishing on Cloudflare Workers. Enables AI coding agents to deploy pages with a single 'publish' tool and get live URLs, with support for atomic updates, versioning, and per-site passwords.
    5 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables agents to manage markdown-based landing pages across many domains over MCP, with tools to list, read, write, and delete sites. Writes publish immediately and are atomic and versioned.
    -
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources