AffiliateSpy
Server Details
Find the TikTok, YouTube and Instagram creators and the blogs, review sites and roundups already promoting your competitors, with the proof, then recruit them: 32 tools for competitor rosters, graded creators, roundup placements, verified contact reveals, outreach from your own inbox, pipeline and Autopilot. OAuth 2.1 (dynamic client registration) or API key; free quick scan and sandbox keys for building agents.
- Status
- Healthy
- Uptime
- 97.4% over 22 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 87 tools
Most tools target distinct resources and actions, and descriptions explicitly cross-reference adjacent tools (e.g., creator vs website campaign drafts, autopilot mode vs floor). However, with 87 tools there are several close pairs, such as add_competitor vs add_suggested_competitor and multiple get/set autopilot tools, where an agent could still misselect without careful reading.
Tool names consistently use snake_case with a predictable verb_noun structure: add_*, get_*, list_*, create_*, update_*, set_*, remove_*, start_*, etc. There is no mixing of camelCase or conflicting conventions.
87 tools is an extreme mismatch for a single MCP server; even a broad affiliate-marketing domain does not justify this many discrete endpoints in one surface. The count alone makes discovery, selection, and maintenance unwieldy for an agent.
The surface covers creators, websites, campaigns, inbox replies, autopilot, deliverability, competitors, keywords, scans, deals, billing, and project settings very thoroughly. Minor lifecycle gaps remain, such as no explicit delete for deals, tracked apps, or individual creators, but these are not core dead ends.
Available Tools
87 toolsadd_competitorAdd competitorAInspect
Add a competitor to a project by domain or URL; the next scan mines its affiliates deep. Same rules as the Competitors page: the plan's competitor cap (too_many) and a brand already on the list (already_tracked). Undo with remove_competitor.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | How the brand is written, e.g. "AI Allure"; used when its homepage gives no fitting name. | |
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| domain | Yes | The competitor's domain or URL, e.g. rival.com. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare write, non-destructive, non-idempotent, and closed-world behavior. The description adds specific behavioral context beyond that: the next scan mines affiliates deep, the exact error conditions too_many and already_tracked, and the undo tool remove_competitor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with purpose, then constraints, then undo. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with full schema coverage and annotations covering its safety profile, the description supplies error codes, scan impact, and undo path. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are fully documented in the schema. The description adds only 'by domain or URL' for the required domain, which matches schema text and provides no additional syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the verb (add), resource (competitor), and method (by domain or URL). It does not explicitly contrast with the sibling add_suggested_competitor, so sibling differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use: same rules as the Competitors page, plan cap (too_many), duplicate brand (already_tracked), and the undo path via remove_competitor. It does not explicitly say when to choose this over add_suggested_competitor, but the constraints and reversibility guidance are strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_creators_to_campaignAdd creators to campaignADestructiveInspect
Add creators to a campaign, as Add to campaign on the Creators tab does. A draft only grows its audience (nothing sends until launch). An active or paused campaign ENROLS them now: only creators with a REVEALED email that never opted out or bounced; the rest come back as excluded (dm_only have no email at all). Enrolled creators get the campaign's first email on the engine's schedule, window and daily cap. Campaigns written for websites refuse creators. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | A campaign id from list_campaigns. | |
| creator_ids | Yes | Creator ids from list_creators, in the campaign's project. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only declaring destructive/not-idempotent/non-read-only, the description carries strong extra context: nothing sends until launch in draft, active/paused campaigns enrol immediately, only REVEALED-email creators qualify, others return as excluded, dm_only have no email, and the approval_token guard returns an approval_required error. This is exactly the behavioral depth annotations cannot provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded with the verb and resource, then ordered by behavioural importance (draft, enrolment, eligibility, guard). The closing 'dashboard guardrail' sentence earns its place as a scope statement, though the prose is heavier than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, non-idempotent, destructive mutation with no output schema, the description covers the mutation semantics, eligibility filtering, error/approval flow, and downstream scheduling effects. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds real meaning: creator_ids eligibility rules (revealed email, no opt-out/bounce) and the approval_token round-trip (error response, then re-call with token). It goes beyond the schema's one-line parameter docs without restating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add creators to a campaign') and immediately distinguishes the operation from a plain write by explaining draft-growth vs active/paused enrolment. An agent can tell this apart from siblings like add_websites_to_campaign or launch_campaign without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly anchors the tool to the dashboard 'Add to campaign' action, spells out when it applies (draft vs active/paused), and gives a real exclusion: 'Campaigns written for websites refuse creators.' Alternatives and preconditions are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_keywordAdd keywordAInspect
Track a search keyword on a project (30 at most). A keyword you add is searched first by every sweep. To undo remove_keyword, pass its undo.restore so the keyword goes back where it was.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| keyword | Yes | ||
| restore | No | Undo of remove_keyword: its undo.restore (position, and whether it was one you added). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds real context beyond them: a hard cap of 30 keywords on a project, priority ordering in sweeps, and the restore path from remove_keyword. It stops short of covering duplicate handling or how the 'confirm' guard interacts with adding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, followed by behavioral detail and the undo routing. No filler, though the undo sentence could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation tool with no output schema and annotations covering safety, the description supplies the missing behavioral facts (cap, sweep priority, undo path). Gaps remain around duplicate keywords and the confirm-guard flow, but the essentials for correct invocation are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 75% schema coverage the schema carries most of the load, but the description adds genuinely new meaning: the 30-keyword cap constrains the required keyword parameter, and it explains that the nested restore object replays remove_keyword's undo (position plus whether it was custom) — details the schema descriptions don't fully spell out.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: tracking a search keyword on a project. The limit (30 at most) and the routing note to remove_keyword's undo.restore help distinguish it from list_keywords/remove_keyword, though it doesn't explicitly name those siblings as alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one concrete usage rule — pass undo.restore to reverse remove_keyword — and implies the tool's role ('searched first by every sweep'). But it never states when to add a keyword versus using add_suggested_competitor or list_keywords, and gives no prerequisites beyond the implicit 30-keyword cap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_seed_addressAdd seed addressAInspect
Add a seed inbox for inbox placement tests whose results the owner checks by hand (Outlook, Yahoo). Seeds belong to the account, not one project. For a Gmail seed whose results are read automatically, use get_seed_connect_link instead.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| address | Yes | The seed's email address. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, so the write/safety profile is covered. The description adds useful scope behavior ('seeds belong to the account, not one project'), but says nothing about reversibility, the confirm guard referenced in the schema, or what happens on a duplicate seed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with zero waste: purpose first, then the account-scope constraint, then the sibling routing. Nothing is buried or repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must stand alone, and it covers purpose, scope, and alternative adequately for a simple add operation. It stops short of describing duplicate handling or when the confirm flag is expected, which are the only meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (app_id, address, confirm) are already documented in the schema itself. The description adds no parameter-level detail beyond the Gmail-vs-manual distinction, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Add a seed inbox') and immediately scopes it to manual-check inbox placement tests (Outlook, Yahoo), which cleanly separates it from the Gmail/automatic sibling. An agent can identify this tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes to the alternative: 'For a Gmail seed whose results are read automatically, use get_seed_connect_link instead.' The condition that selects this tool (results the owner checks by hand) and the condition that selects the sibling are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_suggested_competitorAdd suggested competitorAInspect
Add a suggestion from get_competitor_suggestions to the list (mined deep on the next scan), as the rail's Add does. At the cap, pass swap_out_id to take out a competitor already mined deep in the same write, as the rail's Swap does. Undo an add with remove_competitor; undo a swap with restore_competitor (competitor: swap_out_id, replacing: the new id).
| Name | Required | Description | Default |
|---|---|---|---|
| root | Yes | A suggestion's root from get_competitor_suggestions. | |
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| swap_out_id | No | At the cap: the competitor to take out, usually swap_out.id from get_competitor_suggestions. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds genuinely non-obvious behavior: the suggestion is 'mined deep on the next scan' (a deferred effect, not immediate), the cap triggers a swap mode, and both operations are reversible. It also explains confirm's role via the secondary path. Minor gap: 'as the rail's Add/Swap does' is UI jargon an agent cannot resolve, and the return/receipt format is unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, reasonably front-loaded with the primary action before the cap/swap and undo branches. The 'rail's Add/Swap' analogies and 'mined deep' phrasing consume space without adding agent-resolvable meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and full schema coverage, the description supplies the missing pieces an agent needs: source of input, cap handling, and undo paths. Only the post-call result/pagination behavior is not covered, which is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description adds meaning by tying swap_out_id to the 'at the cap' condition and clarifying it is usually swap_out.id from get_competitor_suggestions. root/app_id/confirm are left to the schema, which already documents them well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Add a suggestion ... to the list') and identifies the source tool (get_competitor_suggestions) versus the sibling add_competitor, which adds competitors directly. An agent can tell what this does and where the input comes from without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit conditions: use swap_out_id when at the cap, undo an add with remove_competitor, undo a swap with restore_competitor. It names the alternatives and the trigger for each, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_websites_to_campaignAdd websites to campaignADestructiveInspect
Add sites to a campaign, as Add to campaign on the Websites tab does. A draft only grows its audience (nothing sends until launch). An active or paused campaign ENROLS them now through the enrol gate: a site needs a contact email, a clean review, status new or prospect, no earlier contact by either product, no opt-out or bounce, and must not be a competitor's or your own site; enrolled sites get the campaign's first email on the engine's schedule, window and daily cap (a paused campaign once resumed). Counts say why sites were skipped. Campaigns written for creators refuse websites. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| site_ids | Yes | Site ids from list_websites, in the campaign's project. | |
| campaign_id | Yes | A campaign id from list_campaigns. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Far beyond the annotations (destructive/openWorld/idempotent hints), it discloses the full enrol-gate eligibility rules, the skip-reason counting behavior, the approval_required two-call flow with the exact error payload shape, and that dashboard guardrails still apply. This is unusually rich behavioral disclosure for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the enrol-gate mechanics are front-loaded, and most sentences carry distinct facts. It is a dense paragraph that could be broken up, but there is little outright waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies return-side context (counts explaining skips, the approval_required error object), the gating semantics, and scope limits, so an agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3; the description adds genuine value by explaining the approval_token retry flow that the schema only describes statically ('call again with the token'). It adds little about site_ids/campaign_id beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add sites to a campaign') and anchors it to a familiar UI action ('Add to campaign on the Websites tab'). It is clearly distinguishable from the sibling add_creators_to_campaign, which targets a different entity type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong conditional context: draft campaigns only grow the audience while active/paused campaigns enrol immediately through the enrol gate, and creator-type campaigns refuse websites. It does not explicitly name an alternative tool or a when-not-to-use condition, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
approve_autopilot_actionApprove autopilot actionADestructiveInspect
Approve a queued autopilot action and EXECUTE it. Depending on its type: a website batch enrols its sites and their cold first emails go out on the lane campaign's schedule; a sister-program intro, a welcome or a check-in is sent to that partner; a follow-up or a reply is sent into its thread; a campaign proposal reveals missing contacts (metered) and launches. The approval summary names what this one does, and an edit after it was shown needs a fresh approval. Same gated paths as the dashboard's approve button, including its edits: body replaces the drafted text, remove_ids drops sites from a website batch (they are flagged so autopilot never proposes them again). Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | Your edited text for a drafted email, sent instead of the draft. | |
| action_id | Yes | An id from get_autopilot's queue. | |
| remove_ids | No | Sites to drop from a website batch (ids from get_autopilot_action). | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint=true, idempotentHint=false) by disclosing the metered contact reveal, that edits require a fresh approval, that remove_ids permanently flags sites so autopilot never re-proposes them, and that missing approval_token yields a structured error plus token. It also ties the guardrails to the dashboard's approve button, defining the blast radius.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Approve a queued autopilot action and EXECUTE it') and then details side effects, edits, and the token guard in a logical order. It is dense and long, but nearly every clause carries operational information; the enumerated action-type list is the only mildly expendable part.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with no output schema, the description covers the trigger (queue ids), the guard flow (token), the side effects per action type, and the edit semantics. An agent has enough to call it correctly without opening another tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: body replaces the drafted text, remove_ids drops sites and flags them, and approval_token is explicitly framed as coming from a prior approval_required response after user confirmation. That is value beyond the field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Approve ... and EXECUTE') plus the resource (a queued autopilot action) and enumerates the concrete action types it can trigger, so the agent immediately understands what happens. The name alone would be ambiguous against hold/reject/get siblings, but the description removes that ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context (action ids come from get_autopilot's queue) and the exact two-step token flow to use when the tool returns approval_required. It does not explicitly contrast with hold_autopilot_action or reject_autopilot_action, so the when-not guidance is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_draftCheck draftARead-onlyIdempotentInspect
Quality hints for an email draft, the same check the dashboard's editor and Inbox composer show: judged against the campaign's offer (or, without campaign_id, the project's program and the autopilot offer). Hints, up to 3: Reads as a template, Promises something not in your offer, Mentions a fee, Not safe to send without edits, Nothing specific about this creator. Pass the rendered email (preview_campaign_step gives it), not {{tokens}}. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | The email body as it would be sent. | |
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| subject | No | Omit for a reply in an existing thread. | |
| campaign_id | No | A campaign id from list_campaigns. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, and the description adds genuinely new behavioral content: the hint vocabulary and the up-to-3 cap, which tells the agent what to expect back. The standalone 'Read-only' sentence is redundant with annotations rather than additive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph, but front-loaded with the purpose and the hint list, then the input-format caveat. The redundant 'Read-only' tail is the only dispensable element, so it is efficient without being minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by enumerating the possible hints and the cap, plus the judgment basis and required input form. An agent can call this and interpret the result without guessing, though the exact response shape is still only described informally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: campaign_id selects the campaign offer as the judgment basis and its absence falls back to program + autopilot offer, and body must be the rendered email rather than raw tokens.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (quality hints / check) on a specific resource (an email draft) and grounds it in the familiar dashboard editor behavior. It is clearly distinguishable from siblings like preview_campaign_step, which it names as the upstream supplier of the input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage context: pass the rendered email from preview_campaign_step, not {{tokens}}, and explains the fallback judgment basis when campaign_id is absent. It lacks an explicit 'when not to use' clause, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_campaign_draftCreate campaign draftAInspect
Create a DRAFT outreach campaign from creator ids (or your saved list) with the default 4-step sequence. DRAFT ONLY, this tool cannot launch, enroll, or send. Review and launch in the dashboard at https://affiliatespy.io/dashboard/campaigns.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| offer | No | The offer line, e.g. "30% recurring". Default empty. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| use_saved | No | Use your saved creators as the audience. | |
| creator_ids | No | Audience creator ids from list_creators. Omit with use_saved=true to use your saved list. | |
| followup_gaps_days | No | Follow-up day offsets, default [3,7,12]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is not read-only or destructive, but the description adds the critical behavioral guarantee 'DRAFT ONLY, this tool cannot launch, enroll, or send'. This clarifies that no side effects beyond draft creation occur, which is valuable context beyond the annotations. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences: purpose, limitation, and follow-up action. The draft-only constraint is front-loaded immediately after the purpose, and the dashboard link is provided without extra fluff. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a draft-creation tool with no output schema, the description covers the essential purpose, input methods, and the critical limitation. However, it does not indicate what the tool returns (e.g., campaign id or confirmation), which an agent might need. The guard behavior for the confirm parameter is only in the schema, not the description. These are minor gaps given the otherwise clear usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, and the schema itself documents most parameters (offer, confirm, use_saved, creator_ids, followup_gaps_days). The description adds little beyond schema, mostly reinforcing the use_saved/creator_ids relationship and mentioning the default 4-step sequence (not a parameter). The name parameter lacks schema description but is self-explanatory. With high schema coverage, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'create' and the resource 'DRAFT outreach campaign', specifies the input source ('creator ids or your saved list') and the default 4-step sequence. It explicitly distinguishes this from launching by stating 'DRAFT ONLY, this tool cannot launch, enroll, or send', which sets it apart from sibling tools like launch_campaign.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-not guidance: it cannot launch, enroll, or send, and points to the dashboard for review and launch. It implies that launching is done elsewhere, but does not explicitly name the sibling tool launch_campaign. The audience selection guidance (creator ids vs saved list) is also provided, adding practical usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_dealCreate dealAInspect
Put a creator in the pipeline, from a creator id or an inbox thread. One deal per creator: on an existing deal it moves the stage. Website threads have no deals yet (they arrive with the partner sync); track a site by its status instead.
| Name | Required | Description | Default |
|---|---|---|---|
| stage | No | Default contacted. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| creator_id | No | A creator id from list_creators or list_deals' candidates. | |
| enrollment_id | No | An enrollment_id from list_inbox_threads, instead of creator_id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, non-destructive, non-idempotent. The description adds genuinely useful behavior beyond them: one-deal-per-creator uniqueness and the upsert-like stage-move on an existing deal. It omits return shape and error behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the uniqueness rule, then the exclusion. Dense but no wasted words; slight compression makes the second sentence need a re-read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description could say what a successful call returns, but it fully covers the input routes and the create-vs-update behavior for an agent to act correctly. Minor gap in return-value disclosure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents stage, confirm, creator_id and enrollment_id with sources for each id. The description adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Put a creator in the pipeline') and clarifies the two input routes (creator id or inbox thread). The clarification that an existing deal causes a stage move slightly blurs the create-vs-update boundary, but it is still distinguishable from siblings like update_deal and move_deal_stage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool and an explicit exclusion: website threads have no deals yet, so track a site by its status instead. It does not explicitly name update_deal or move_deal_stage as alternatives, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_website_campaign_draftCreate website campaign draftAInspect
Create a DRAFT campaign for sites from list_websites (one project), with the website sequence autopilot's affiliate lane sends unless you pass steps, and an optional schedule. The copy must pass the same check enrolment runs (rendered for a sample site: no links, bare domains, stock phrases, unknown tokens or empty offer), else nothing is created. Website copy personalises with tokens such as {{greeting}}, {{opener}}, {{tease_pitch}}, {{offer_pitch}} and {{signoff}}: get_campaign shows a campaign's templates and preview_campaign_step renders one for a site. DRAFT ONLY: nothing is enrolled or sent; launch_campaign (guarded) starts it.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Sending days. Default Mon to Fri. | |
| name | No | ||
| offer | No | The commission the emails promise, e.g. "30% recurring". Default: the program commission in App settings. | |
| steps | No | Your own sequence instead of the default website one. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| hour_to | No | Window end hour, exclusive. Default 17. | |
| site_ids | Yes | Site ids from list_websites, all in one project. | |
| timezone | No | IANA timezone, e.g. "Europe/London". Default America/New_York. | |
| hour_from | No | Window start hour, 24h, in timezone. Default 9. | |
| daily_limit | No | Most emails this campaign sends a day. The mailbox's own daily cap still applies. | |
| followup_gaps_days | No | Day offset of each follow-up from step 1, e.g. [4,14,28,42]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the basic non-read-only/non-destructive/non-idempotent profile; the description adds substantial behavioral detail beyond them: the copy must pass the same validation enrolment runs (no links, bare domains, stock phrases, unknown tokens, or empty offer) or nothing is created, and it explicitly states DRAFT ONLY with nothing enrolled or sent. This is meaningful, non-obvious behavior disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then layers validation, tokens, and the draft/launch boundary in information-dense sentences with little waste. It is slightly long and packs several distinct topics, but each sentence carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter write tool with no output schema, the description covers the critical unknowns: draft-only semantics, validation gating, personalisation tokens, and the launch transition. It defers scheduling detail to the 91%-covered schema, which is reasonable, leaving only minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 91%, so the schema already documents nearly every parameter, making 3 the baseline. The description adds semantic value on tokens ({{greeting}}, {{opener}}, {{offer_pitch}}, etc.) and on the default-sequence-vs-steps behavior, but it does not explain the scheduling parameters (days, hour_from/hour_to, timezone, daily_limit, followup_gaps_days) beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a DRAFT campaign for sites from list_websites (one project)'. It scopes the resource (website campaigns, single project) and distinguishes itself from the generic draft creator by tying inputs to list_websites, so an agent can tell it apart from create_campaign_draft without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: use list_websites for site_ids, get_campaign and preview_campaign_step to inspect/render templates, and launch_campaign to actually start sending. It also clarifies the draft-vs-launch boundary. It does not explicitly contrast against the sibling create_campaign_draft, so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_campaignDelete campaignADestructiveInspect
Delete a campaign, as the dashboard does: it leaves every list and its queued emails stop. Enrolments and messages stay as contact history, so nobody in it can be emailed again from scratch, and late replies still reach the Inbox. The dashboard has no undo. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | A campaign id from list_campaigns. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, but the description adds substantial context beyond them: what is destroyed (lists, queued emails), what is preserved as contact history (enrolments, messages), that recipients cannot be re-emailed from scratch, that late replies still reach the Inbox, and that there is no undo. It also discloses the approval_token gate and the exact shape of the approval_required response.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and its consequences, and the approval mechanism is stated before the closing scope note. Slightly redundant phrasing ('as the dashboard does' / 'Every dashboard guardrail still applies') costs a little, but nearly every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the full picture: irreversibility, data retention, the approval handshake, and the error payload an agent must parse. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema by explaining the approval_token lifecycle: the first call without it returns {error, approval_token, summary}, and the second call supplies the token after user confirmation. That is operational semantics the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a campaign') and immediately distinguishes the operation from sibling lifecycle tools by describing its effects: it leaves every list and queued emails stop. An agent can tell this apart from pause_campaign/resume_campaign/duplicate_campaign without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when this is appropriate ('exactly what a user clicking the dashboard could do, no more') and explains the guarded two-step confirmation flow. It stops short of explicitly routing the agent to an alternative such as pause_campaign for a reversible stop, so the when-not guidance is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dismiss_competitor_suggestionDismiss competitor suggestionAInspect
Hide a suggestion for good (later scans keep the choice). undo: true puts a dismissed suggestion back on the rail.
| Name | Required | Description | Default |
|---|---|---|---|
| root | Yes | A suggestion's root from get_competitor_suggestions. | |
| undo | No | true to bring a dismissed suggestion back. | |
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false and idempotentHint=false, so the safety profile is largely covered. The description adds genuine extra context the annotations lack: the dismissal is persistent ('for good' — later scans keep the choice) and is reversible via undo, which is the key behavior an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler. The persistence trait comes first and the reversal path second, which is the right ordering for an agent deciding whether to call it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter mutation with full schema coverage, no output schema and annotations covering safety, the description covers what the tool does and its persistence/reversal semantics. It omits the confirm-guard interaction, but that parameter is fully explained in the schema, so the gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (root, undo, app_id, confirm) are already documented in the schema. The description restates the undo semantics ('puts a dismissed suggestion back on the rail') but adds no syntax or format detail beyond what the schema provides, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource ('Hide a suggestion for good'), making the mutation unambiguous. It implicitly distinguishes itself from the read-side siblings like get_competitor_suggestions, but never names an alternative, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the undo branch ('undo: true puts a dismissed suggestion back on the rail'), which tells the agent when to use that mode. However, it gives no guidance on when to prefer dismiss vs. add_suggested_competitor or get_competitor_suggestions, leaving the primary usage context implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_campaignDuplicate campaignAInspect
Copy a campaign as a new DRAFT: same audience, sequence and settings, no enrolments. Nothing is sent until the copy is launched.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| campaign_id | Yes | A campaign id from list_campaigns. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=false and readOnly=false. The description adds real behavioral context beyond them: what is copied (audience, sequence, settings), what is omitted (no enrolments), and the safety property that nothing sends until launch. It does not explain the confirm guard or idempotency, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the verb and result, with the no-send guarantee as supporting context. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema and full annotation coverage, the description covers what is created and the safety profile adequately. It leaves the return value (presumably the new draft's id) unstated, a minor gap since no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both params (campaign_id, confirm) are self-documented, so the baseline is 3. The description adds no parameter syntax or behavior beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (copy/duplicate) and resource (campaign) and immediately scopes the result as a new DRAFT with same audience/sequence/settings. An agent can distinguish it from create_campaign_draft, launch_campaign, and update_campaign_settings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Nothing is sent until the copy is launched' clause implies the draft-state context and contrasts implicitly with launch_campaign, but no sibling is named and there is no explicit when-to-use / when-not-to-use guidance. Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_creators_csvExport creators CSVARead-onlyIdempotentInspect
CSV of your creator list (same filters as list_creators). Paid feature, requires an active subscription, same gate as the dashboard export. Emails appear ONLY for creators you have revealed.
| Name | Required | Description | Default |
|---|---|---|---|
| saved | No | Only creators you saved. | |
| app_id | No | App to list creators for. Omit to use your most recently added app (see list_apps). | |
| search | No | Substring match on handle, name, or bio. | |
| platform | No | ||
| min_score | No | Minimum quality score, 0-100. | |
| has_contact | No | Only creators with an email on file (value hidden until revealed). | |
| proven_partner | No | Only creators with proof of a paid competitor partnership. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only and idempotent behavior, and the description adds valuable non-obvious details: the paid feature/subscription gate and the email privacy rule ('Emails appear ONLY for creators you have revealed'). These are behaviorally significant and not inferable from the annotations. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct information: output format and filter equivalence, subscription gating, and email redaction. It is front-loaded with the core purpose and avoids filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an export tool with no required parameters and a high-coverage schema, the description covers the important non-schema aspects: subscription requirement and data visibility restrictions. The main gap is that it does not describe how the CSV is delivered (e.g., file download vs URL), and there is no output schema to fill that in, but this is not critical for selecting and invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so individual parameters are already documented. The description's main added semantic is 'same filters as list_creators', which helps conceptually but does not explain any individual parameter. It clarifies that 'has_contact' values are hidden until revealed, but that is behavioral rather than semantic detail. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the deliverable clearly: 'CSV of your creator list', which distinguishes it from list_creators by output format. It also references the sibling list_creators to clarify that the same filtering applies caeteris. The operation is slightly implicit, relying on the title for the 'Export' verb, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says 'same filters as list_creators', connecting it to a known sibling and making the output format the key differentiating factor. It also provides important selection context: 'requires an active subscription, same gate as the dashboard export'. It lacks an explicit 'when not to use' statement, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_website_contactFind website contactsADestructiveInspect
Find a contact email for up to 25 sites, as the Websites Find email button does: the free chain first (cache, the site's contact pages, policy and deliverability checks), then ONE paid contact-finder lookup per site when that finds nothing and the project's monthly finder budget has room. Found addresses (and the contact's name when known) are saved on the site. Sites that already show a contact path are skipped, as in the dashboard. Runs one site at a time for up to about four minutes; sites not reached come back as not_started. Nothing is sent. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| site_ids | Yes | Up to 25 site ids from list_websites. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses the free-then-one-paid-lookup flow, the monthly finder budget gate, one-site-at-a-time ~4-minute execution, 'not_started' for unreached sites, that nothing is sent, and the exact guarded approval flow including the {error:'approval_required', approval_token, summary} shape and the re-call requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded, with the core action stated first and guardrails after. Given the tool's complexity the length is largely justified, though the closing 'Every dashboard guardrail still applies... no more' restates earlier points and is mildly redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden and does so: it explains results (addresses and contact name saved on the site), partial-completion semantics (not_started), spend limits, and the approval handshake error payload. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning about the approval_token round-trip and the site_ids skip/budget behavior that the schema alone doesn't convey. It doesn't add new detail on site_ids beyond the schema's 'up to 25'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource+scope: 'Find a contact email for up to 25 sites,' and anchors it to a known UI action ('as the Websites Find email button does'). It is clearly distinguishable from siblings like set_website_contact (manual entry), reveal_contact, and list_websites.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives strong context: when contacts already exist the site is skipped, the free chain runs before the paid lookup, and budget must have room. It stops short of naming sibling alternatives (e.g., reveal_contact or set_website_contact) for when the agent should prefer manual entry instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_accountGet accountARead-onlyIdempotentInspect
Your AffiliateSpy plan and usage: plan name, caps, contact reveals used/left this month (resets on the 1st, UTC), and tracked apps used vs cap.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it readOnly, idempotent, and non-destructive. The description adds valuable behavioral context beyond that: it specifies the return content (plan name, caps, contact reveal usage with UTC reset on the 1st, tracked app usage vs cap). This tells the agent exactly what to expect without any side effects. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the purpose ('Your AffiliateSpy plan and usage') and packs in all relevant details without redundancy. Every clause earns its place, and it is immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only tool with no output schema, the description fully covers what the agent needs: it enumerates the returned fields and even notes the reset schedule. There is no missing information that would hinder correct invocation or interpretation of the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4 per the rubric. There is nothing to explain beyond the schema, which is already empty. The description doesn't need to add parameter meaning, and it doesn't. This is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the resource (AffiliateSpy account) and the specific data returned: plan name, caps, contact reveal usage with reset schedule, and tracked app usage. This distinguishes it from sibling get_* tools, which all target different resources (campaign, creator, thread, etc.). No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While no explicit 'use when' guidance is provided, the description makes it self-evident that this is the tool to check plan limits and quota usage. It names the exact metrics (contact reveals used/left, tracked apps used vs cap), which implies it should be consulted before quota-consuming operations. It doesn't state exclusions or alternatives, but there is no competing sibling for account-level info, so this is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_autopilotGet autopilot stateARead-onlyIdempotentInspect
The Discovery Autopilot's state for one project (default: the project it runs for): status on/off for this project, as the Autopilot page shows it (autopilot_status and running_for say whether it runs for another project), the commission offer it may promise, today's send count vs cap, what it is waiting on before it can recruit, the approval queue (campaign launches, website batches, reply drafts and check-ins waiting for a human), and recent approved/rejected activity. held_by_breaker means autopilot is bound here but the bounce breaker is holding every send and proposal until set_autopilot on (or choosing a mailbox, per paused_reason) clears it. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id. Omit for the project autopilot runs for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the redundant 'Read-only' adds little. However, the description goes well beyond structured data by decoding domain semantics: what held_by_breaker means, how autopilot_status/running_for relate, and what paused_reason governs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and free of filler, but it is one very long run-on sentence whose parenthetical asides make it dense. Every clause carries field meaning, so size is justified even if readability suffers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining returns and does so for the major fields (status, cap, waiting-on, queue, recent activity), giving an agent enough to interpret results. Minor return fields likely remain unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single optional parameter with 100% schema description coverage, so the schema already documents app_id and its 'omit for the running project' default. The description's 'default: the project it runs for' merely echoes the schema, adding no syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (get the Discovery Autopilot state for a project) and enumerates what 'state' comprises (status, cap usage, waiting-on, approval queue, recent activity). It is clearly a getter contrasted with the set_autopilot sibling, though it never explicitly names that contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated. The clause about the default project and the note that set_autopilot clears held_by_breaker give context, but there is no explicit 'use this when...' guidance or named alternative for retrieving different state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_autopilot_actionGet autopilot actionARead-onlyIdempotentInspect
One autopilot queue or activity item in full, as the dashboard's approval card shows it: type, the lane (mode key) that governs it, status, error, the drafted email (to, subject, body, the original draft when edited), the thread it answers and their message, reply triage and QA flags, and for a website batch every site with its fit score, the competitors its page links to, its opener and any sensitive content label (selected:false means dropped). Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| action_id | Yes | An id from get_autopilot's queue. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered and the trailing 'Read-only' is partly redundant. The description still adds genuine value by disclosing the response payload shape and the non-obvious flag semantic 'selected:false means dropped'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in the first clause, then a dense but information-rich enumeration of payload fields. It is one long run-on sentence, but nearly every clause earns its place by describing a distinct part of the returned record.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the burden of describing return values and does so thoroughly (type, lane, status, error, draft, thread, triage/QA flags, website batch details). The only gap is usage routing relative to approve/hold/reject siblings that share the same id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single action_id parameter, and the schema already notes the id comes from get_autopilot's queue. The description adds nothing beyond the schema for parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — fetching ONE autopilot action in full — and explicitly distinguishes it from the queue-level sibling by naming the 'approval card' view and enumerating the fields returned. An agent can tell this apart from get_autopilot without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: it retrieves a single item, and the schema notes the id comes from get_autopilot's queue. However, there is no explicit when-to-use/when-not guidance and no mention of the approve/hold/reject siblings that also act on an action_id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_campaignGet campaignARead-onlyIdempotentInspect
One campaign's settings, sequence templates, enrollment rows and latest reply previews.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds value by specifying the exact content of the response, which goes beyond the annotations. It does not mention auth requirements or rate limits, but these are not critical for a read-only get.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that lists the key data components. It is front-loaded with the resource type and immediately states what is included. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only get with one parameter and no output schema, the description adequately conveys what is returned. The annotations cover safety, and the listed components give a clear picture of the response. It could optionally describe error cases or format, but that is not essential for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. However, the description does not mention campaign_id at all. While the parameter name and tool name make its purpose obvious, the description adds no meaning beyond the schema. It only describes the output, not the input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the resource ('One campaign') and enumerates the specific data returned: settings, sequence templates, enrollment rows, and latest reply previews. This distinguishes it from list_campaigns (which lists many) and other get_* tools like get_creator or get_thread.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: use this when you have a campaign_id and need detailed information about that single campaign. It does not explicitly mention alternatives or exclusions, but the sibling list makes the distinction obvious. No exclusions are needed for a straightforward get operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_checkout_linkGet checkout linkAInspect
Stripe Checkout URL for a plan (starter $99, growth $299, agency $799 a month; yearly = two months free). The USER opens it and pays in the browser; never enter payment details yourself. On an active Stripe subscription the plan is switched in place (prorated) and a settings URL is returned instead. Access is granted by the payment webhook, so call get_account afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | Default starter. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| interval | No | Default monthly. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations, it discloses that the user performs payment in the browser, that an active subscription is switched in place and prorated, and that access is only granted via the payment webhook. This materially changes how an agent should handle the returned URL and follow-up.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with the primary purpose front-loaded. Every sentence carries operational information: what is returned, how the user pays, what happens on an existing subscription, and what to call next.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The return variants are described explicitly, the payment flow is specified, and the follow-up call to get_account is provided. With no output schema, this description gives an agent enough behavioral context to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers all three parameters, so the baseline is 3. The description adds useful pricing and billing-period context for plan and interval, although confirm is not mentioned in the description; the schema covers its meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: it returns a Stripe Checkout URL for a named plan. It also gives concrete pricing and the active-subscription alternative (settings URL), so there is no ambiguity about what this tool produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: the user opens the URL and pays, the agent must never enter payment details, and after payment the agent should call get_account. It does not explicitly state when not to use the tool or compare it to alternatives, but the workflow guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_competitor_suggestionsGet competitor suggestionsBRead-onlyIdempotentInspect
The Competitors page's Suggested rail: up to 12 brands the scans found that are not on the list, best first, with why (sites that co-promote them, bumps, a program, paid ads, App Store similar, shared searches), and swap_out: at the cap, the mined competitor with the fewest receipt sites that a swap would take out. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/destructive, but the description adds real behavioral context: the 12-item cap, ranking order, the evidence categories behind each suggestion, and the precise semantics of swap_out (the mined competitor with fewest receipt sites). It does not mention pagination or auth, but the cap and swap semantics are non-obvious and usefully disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded sentence that starts with the resource, which is good. However it is a run-on packed with six enumerated 'why' examples and parenthetical detail that is heavier than needed for tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so well, explaining the suggested brands, their rationale fields, and swap_out. It is largely complete for a read-only, single-param tool, missing only explicit routing to the add/dismiss sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (app_id) exists and the schema already documents it at 100% coverage, including the 'omit to use most recent project' behavior. The description adds nothing about the parameter, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description concretely states what the tool returns: up to 12 mined brands not on the competitors list, ranked best-first, with the 'why' signals and a swap_out target. This distinguishes it from list_competitors and get_partners, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit statement of when to call this versus list_competitors, add_suggested_competitor, or dismiss_competitor_suggestion. The 'at the cap' swap_out detail implies a scenario, but the agent must infer the workflow and follow-up actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_creatorGet creator detailARead-onlyIdempotentInspect
Full detail for one creator: grade reasoning, evidence receipts, recent posts metadata, cross-platform profiles, monetization signals. Contact value stays hidden unless already revealed.
| Name | Required | Description | Default |
|---|---|---|---|
| creator_id | Yes | An id from list_creators. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering safety. The description adds meaningful behavioral nuance by stating that contact value stays hidden unless already revealed, which is a non-obvious access-control behavior an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one tight sentence that front-loads the main purpose, uses a colon to introduce the content categories, and ends with the contact-visibility caveat. Every clause earns its place and there is no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description names the major response categories and the one conditional behavior that affects returned data. For a simple read-only detail tool with a single required parameter, this is enough for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, creator_id, is fully documented in the schema with a description pointing to list_creators as the source. The tool description adds no further parameter-level meaning, so the schema already carries this burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Full detail for one creator' and then enumerates the specific data categories returned, making the tool's purpose explicit. It is clearly distinguishable from list_creators and other sibling tools because it targets a single creator's full record rather than a collection or action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies this is the detail-fetch tool for one creator, and the schema's parameter description ('An id from list_creators') reinforces the expected workflow of listing creators first. It does not explicitly name alternative tools or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_deliverabilityGet deliverabilityBRead-onlyIdempotentInspect
The project's sending health (App settings, Sending): the From domain check (SPF covers Google, DKIM, DMARC policy, MX, MTA-STS, Spamhaus blocklist, with errors and warnings in plain words), the separate outreach domain, the mailbox warm-up (week 1 to 3 caps 10, 20, 30 a day; established skips it), seed inboxes and how their results are read, and the latest inbox placement test: which tab each seed got each step in (primary, promotions, updates, spam, missing), whether it covers the current template, and the gate verdict.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuine behavioral detail beyond that: warm-up caps of 10/20/30 a day for weeks 1-3 with 'established skips it', errors/warnings surfaced 'in plain words', whether the placement test covers the current template, and a gate verdict. It does not mention error conditions (e.g. unconfigured domain) or how stale the results may be.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded ('The project's sending health'), but the entire body is one long run-on sentence with a nested parenthetical list that is hard to scan. Most of the bulk is return-value enumeration, which is legitimate compensation for the missing output schema, yet it is packaged in a form that is denser than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the burden of describing what comes back, and it covers the major components (domain checks, warm-up, seeds, placement test, gate verdict) thoroughly. It omits how to force a fresh test or what happens when deliverability data does not yet exist, but for a single-optional-param read tool it is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional app_id parameter at 100% schema description coverage, so the schema already explains that it comes from list_apps and defaults to the most recently added project. The description adds nothing about the parameter. Per the rubric, high coverage with no added param info lands at baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource (the project's sending health) and enumerates the exact scope: SPF/DKIM/DMARC/MX/MTA-STS checks, outreach domain, warm-up caps, seed inboxes, and inbox placement results. This lets an agent distinguish it from generic siblings like get_overview or get_project_settings. It never names a sibling explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. The only contextual anchor is '(App settings, Sending)', which is a UI location, not a routing cue. Given siblings like run_inbox_test, refresh_inbox_test_results, and recheck_from_domain that mutate/refresh this same data, the absence of any 'use X to trigger a new test' guidance is a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mailbox_connect_linkGet mailbox connect linkARead-onlyIdempotentInspect
The link a human opens in a browser, signed in to AffiliateSpy, to connect a Gmail mailbox through Google's consent screen; it is linked to the project afterwards. With reconnect, it re-consents that connected mailbox instead (for example to grant read access). Do not open it yourself.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| reconnect | No | Address of a connected mailbox to re-consent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotent, non-destructive, so the safety profile is covered; the description adds real context beyond that: the link is consumed by a human in a browser, requires an AffiliateSpy session, routes through Google's consent screen, and the mailbox is linked to the project afterwards. It omits link expiry/reuse behavior, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but waste-free: the core function and its human-in-the-browser constraint come first, the reconnect variant second, and the imperative warning last. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-optional-param link generator with no output schema, the description covers what the link is, who uses it, and the reconnect path. Only the lifetime/reuse behavior of the returned link is left unstated, a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics for 'reconnect' — it means re-consenting an existing connected mailbox, with a concrete use case (granting read access) — and frames app_id implicitly as the target project.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (getting a mailbox connect link) and precisely what the artifact is: a browser link a signed-in AffiliateSpy human opens to run Gmail through Google's consent screen. This clearly distinguishes it from siblings like get_seed_connect_link and get_checkout_link, which produce different link types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the operating condition well: a human must open it in a browser, use reconnect to re-consent an already-connected mailbox (e.g., to grant read access), and explicitly warns 'Do not open it yourself.' It does not name competing tools, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mailboxesGet mailboxesBRead-onlyIdempotentInspect
Every connected Gmail mailbox: read access (needed to see replies and bounces), free mailbox flag, over the plan's mailbox limit, warm-up, the Send mail as aliases Gmail lists and whether it accepted each, and the projects sending from it with their alias. project: the given (or newest) project's mailbox and From. Connecting, disconnecting and revoking happen in the dashboard.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is fully covered. The description adds useful context: the JSON-style scoping paragraph indicates the return shape and that 'project' changes the view. No auth/permission requirements or rate limits are mentioned, so it's adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single-property description is a run-on sentence with comma-dangled list items and a sentence fragment ('project: the given ...'). It mixes return-value enumeration, parameter semantics and scope notes without clear sentence boundaries, making it effortful to parse. Not front-loaded on the primary purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotations cover safety and the input schema fully documents app_id, so no safety or parameter details are missing. However, with no output schema and a verbose, loosely structured description, an agent must parse the paragraph to infer the return structure. It's usable but not complete or crisp for a tool enumerating many mailbox fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter: the input schema already documents app_id fully (UUID format, 'Project id from list_apps. Omit to use your most recently added project.'). The description adds the concept that 'project' returns the given or newest project's mailbox and From, but doesn't give syntax beyond what the schema says. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (Gmail mailboxes) and enumerates the exact data returned: read access status, free mailbox flag, plan limit status, warm-up, aliases, and sending projects. It's a clear list/retrieve verb with a detailed resource. However, it doesn't explicitly distinguish itself from close siblings like get_mailbox_connect_link or list_apps, so it doesn't reach 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a read/list use case and notes that connecting, disconnecting and revoking happen in the dashboard, which softly brackets its scope. But there is no explicit when-to-use, when-not-to-use, or routing to alternatives (e.g. get_mailbox_connect_link for authorization). The paragraph about 'project:' hints at an alternate view but doesn't state selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_offer_benchmarkGet offer benchmarkARead-onlyIdempotentInspect
How the project's commission offer compares with its competitors' published programs: their rates and cookie windows, the market median, how many open prospect sites promote a program paying more, a terms chip per competitor domain with its source, and the affiliate program directories the competitors are listed on. benchmark null when no competitor program terms were found yet. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered and 'Read-only' at the end is largely redundant. The description nonetheless adds real behavioral context beyond the annotations: the null/empty-state behavior when no competitor program terms have been collected, which is the only place that is documented given there is no output schema. No auth or rate-limit detail, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause, and the following enumeration is dense but every item is a distinct return field rather than filler. The trailing 'Read-only' restates the readOnlyHint annotation and is the one line that does not earn its place, keeping it out of 5 territory.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of describing what comes back, and it does so field by field, plus it documents the empty/null case and the precondition that competitor terms must have been discovered. Nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is exactly one parameter and schema description coverage is 100%, with the schema itself explaining 'Project id from list_apps. Omit to use your most recently added project.' The description adds no further meaning about app_id semantics or defaulting, so the baseline of 3 for high-coverage schemas applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause names a specific computed resource — the project's commission offer versus competitors' published programs — which is clearly distinct from sibling readers like list_competitors (raw competitor list) or get_program (own program). It then enumerates the return contents (rates, cookie windows, market median, higher-paying open prospect sites, terms chips, affiliate directories), so an agent can identify the tool without opening a schema. It stops short of explicitly naming the sibling it should be chosen over.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'benchmark null when no competitor program terms were found yet' hints that competitor data must exist first, which is a useful precondition signal. But there is no explicit when-to-use/when-not, no mention of how this differs from list_competitors or get_program, and no guidance on cadence or freshness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_overviewGet overviewARead-onlyIdempotentInspect
The project's Overview metrics (live partners, signed up, recruited by outreach, converting, in outreach, contacted, replied, reply rate, prospects, with contact, lost, competitors tracked), each one's change over 7, 30 and 90 days (null until a daily snapshot that old exists), and the sites each competitor gained in the last 7 days. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so safety is covered. The description adds real behavioral context beyond that: the 7/30/90-day deltas are null until a daily snapshot that old exists, and it discloses the 7-day competitor site-gain data. 'Read-only' merely repeats the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the resource and its contents, ending with the read-only marker. The long parenthetical metric list is dense but every item is informative, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by naming all metrics plus the null-snapshot caveat. It is largely complete, though the absence of any usage routing is the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single app_id parameter is fully documented in the schema (including the 'omit to use most recent project' fallback). The description adds no parameter-level detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('the project's Overview metrics') and enumerates the exact metrics returned, so an agent knows precisely what payload to expect. This clearly distinguishes it from sibling readers like get_account, get_program, or get_deliverability.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to prefer this over the many other read siblings (get_account, get_program, get_competitor_suggestions, etc.). The 'Read-only' tag and app_id behavior imply it's a safe lookup, but nothing routes the agent here versus an alternative view.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_partnersGet partnersARead-onlyIdempotentInspect
Everyone already promoting the project, as the Partners page shows it: creators in a live or paid deal, sites featuring you, and your affiliate program's roster with per-affiliate clicks, conversions, revenue and commission (cents). Tabs: active, pending, lost, nofollow (sites linking without passing authority), missing_email. counts covers every tab. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| tab | No | Default active. | |
| limit | No | Default 100. | |
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, so the bar is lower. The description adds useful behavioral detail beyond them: the returned fields (clicks, conversions, revenue, commission in cents) and 'counts covers every tab', which tells the agent counts are tab-independent. Doesn't address pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: the scope definition comes first, then tab semantics, then the read-only note. Every clause carries information, though the run-on construction packs three partner categories into one sentence, slightly at the cost of scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by naming the return fields and their units (cents) plus tab semantics. It covers the enum and the general shape of the response. Missing only pagination/limit behavior and the app_id defaulting nuance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so tab, limit, and app_id already carry descriptions. The description adds real meaning on the enum: it glosses 'nofollow' as sites linking without passing authority and clarifies 'counts covers every tab'. It says nothing about limit/offset, but those are schema-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (partners), and crucially defines the composite scope: creators in live/paid deals, featuring sites, and the affiliate roster. This effectively distinguishes it from list_creators/list_deals/get_program without naming them. The generic verb 'get' and lack of sibling naming keep it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied via 'as the Partners page shows it' and the tab breakdown, letting an agent infer this is the aggregate partners view. However, there is no explicit when-to-use statement, no exclusion, and no pointer to a sibling for a narrower need (e.g. per-creator or per-site lookups).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_programGet affiliate programARead-onlyIdempotentInspect
The project's affiliate program connection, as the Autopilot page's card shows it: provider, program id, status, affiliates (approved, pending, inactive), revenue and commission, last sync, last error and signup approval mode; and the program's recent sales stats (conversions, average payment, commission, top affiliate's commission, window hours) when a sync stored them. Connecting or disconnecting a program takes an API key, so it happens in the dashboard (https://affiliatespy.io/dashboard/autopilot). Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds genuine behavior beyond them: sales stats are only present 'when a sync stored them', the last-error field is exposed, and connect/disconnect requires an API key and lives in the dashboard rather than here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the identity of the resource before the long field list, and every clause names a returned field rather than padding. The semicolon-chained enumeration is dense but each element earns its place since there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full return-shape burden and does so comprehensively, including the conditional nature of the sales stats. The redirect for connect/disconnect closes the remaining gap, leaving nothing an agent needs to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100% – the schema already documents app_id including the 'omit to use your most recently added project' default. The description adds no additional parameter meaning, so this is the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (get the project's affiliate program connection) and enumerates exactly what it surfaces: provider, program id, status, affiliate counts, revenue/commission, last sync/error, signup approval mode, and sales stats. An agent can distinguish it from sync_program_now and set_signup_approval_mode without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the write-side operations (connecting/disconnecting a program) away to the dashboard URL because they need an API key, which tells the agent when NOT to use this tool. It does not, however, state when to prefer it over read siblings like get_partners or get_autopilot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_project_settingsGet project settingsBRead-onlyIdempotentInspect
Everything App settings shows for a project. offer: the stored offer fields update_offer edits (absent: unset, the legacy program fills it; empty or null: cleared). program_facts: what outreach actually uses. follow_up_schedule: the day each email goes out and the rate it offers (rate null: the commission line as written, or a reminder). sending: effective daily cap, reply reserve and morning contact checks (max is the platform ceiling), today's cap under the mailbox warm-up, paid finder budget, window, timezone and days. Also notifications and re-scan cadence, product summary, market, marketing domain, Play link and the autopilot fit floor.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so the safety profile is covered. The description goes further by disclosing semantics the annotations cannot: absent vs empty/null meaning for offer fields, the distinction between stored offer fields and program_facts ('what outreach actually uses'), and the interpretation of a null rate. That is genuine added behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a good front-load, but the remainder is a dense run-on catalogue of returned fields with nested parentheticals, and the trailing 'Also notifications and re-scan cadence...' clause reads as an afterthought. The field semantics do earn some of their length, but the structure is list-like rather than prioritized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of describing return values, and it does enumerate most of the settings groups plus their edge-case semantics. A few entries stay vague ('product summary, market, marketing domain, Play link and the autopilot fit floor'), so it is strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (app_id) with 100% schema description coverage, including 'from list_apps' and the omit-to-use-most-recent default. The description adds nothing about app_id, but the schema already carries the full burden, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states the resource clearly ('everything App settings shows for a project') and enumerates the setting groups returned (offer, program_facts, follow_up_schedule, sending), so an agent knows this is the read-side counterpart of update_offer/set_autopilot. However it never names the verb explicitly (retrieve/read) and does not differentiate itself from siblings like get_program, get_autopilot, get_overview or get_deliverability, some of which likely overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no routing to alternatives. Given siblings such as get_account, get_overview, get_program and get_autopilot that plausibly expose overlapping settings, the agent is left to guess which read tool to call. The description only implies usage through its enumeration of contents.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_quick_scan_previewGet quick scan previewARead-onlyIdempotentInspect
Progress and locked preview rows for a quick scan (platform, followers, a hint like '★ partner', masked handle; no real handles or contacts). Long-polls: waits server-side (default 45 seconds) until the scan status is done or failed, so one call is usually enough; call again while status is queued or running. Use it to show the user what a paid plan unlocks, then get_checkout_link.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | run_id from start_quick_scan or get_scan_status. | |
| wait_seconds | No | How long to wait for the scan to finish before returning. Default 45. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only and idempotent. The description adds critical behavioral details: long-polling server-side up to 45 seconds, waiting for done/failed status, returning masked handles (no real handles/contacts), and the hint format. This goes well beyond the annotations and provides meaningful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the purpose and content, then explains the long-poll behavior and usage. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description sufficiently describes the return content (progress, locked preview rows, fields) and the long-poll semantics. It also gives the usage sequence and retry guidance, making it fully actionable for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with run_id and wait_seconds fully described including defaults and sources. The description adds no parameter-specific meaning beyond what the schema already provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('get quick scan preview'), enumerates the content (progress and locked preview rows with platform, followers, hint, masked handle), and explicitly distinguishes its purpose from status polling by mentioning the paid-plan unlock use case. It clearly separates from get_scan_status and start_quick_scan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Use it to show the user what a paid plan unlocks, then get_checkout_link.' It also explains the long-poll behavior and advises calling again while status is queued or running. It does not explicitly name alternatives or exclusions, but the context is clear enough for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scan_statusGet scan statusARead-onlyIdempotentInspect
Latest scan run for an app (stages, status, creators found). Starting scans: start_quick_scan (free preview, no plan needed) or start_scan (guarded, plan required).
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Omit to use your most recently added app. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that the tool returns stages, status, and creators found, which is useful context beyond the annotations. It does not mention any edge cases like empty results or permission requirements, but given annotation coverage, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The first sentence states exactly what the tool returns, and the second provides actionable sibling context. The core purpose is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only status tool with one optional parameter, the description is complete. It explains the return values (stages, status, creators found) and even points to the tools used to start scans, giving agents enough context to call it correctly without needing an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with the app_id parameter described ('Omit to use your most recently added app.'). The description adds no additional parameter-level semantics beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action and resource: it retrieves the latest scan run for an app, including stages, status, and creators found. It implicitly distinguishes itself from sibling tools like start_scan and start_quick_scan, which are for initiating scans, and get_quick_scan_preview, which implies preview functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides direct pointers to sibling tools for starting scans (start_quick_scan, start_scan) with conditions (free preview vs. guarded/plan required), but it does not explicitly state when to use this tool versus alternatives like get_quick_scan_preview. The purpose is clear enough to infer usage, and the alternatives are named, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_seed_connect_linkGet seed connect linkARead-onlyIdempotentInspect
The link a human opens in a browser, signed in to AffiliateSpy, to connect a Gmail account as a seed inbox through Google's consent screen. The seed only receives tests and is read for where they land; it never sends and takes no mailbox slot. Do not open it yourself.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, openWorldHint=false), so the bar is lower. The description nonetheless adds real behavioral context: the link is consumed by a human in a browser, and the seed only receives tests, never sends, and takes no mailbox slot. No rate limits or link-expiry behavior is mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what the link is, what the seed does, and the operator-facing warning. The caution is front-loaded appropriately and every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must describe what comes back, and it does — a consent-screen link for a human. It is nearly complete for a zero-param read tool, missing only practical details such as link reuse/expiry or what to do after the human completes consent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema leaves nothing to document and the baseline is 4. The description correctly focuses on output semantics rather than inventing parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource: returns a browser link that connects a Gmail account as a seed inbox via Google's consent screen. The 'seed inbox' qualifier is specific. It does not, however, explicitly distinguish itself from the sibling get_mailbox_connect_link, so an agent must infer the difference between seed and regular mailbox linking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context (a human opens this while signed in to AffiliateSpy to connect a seed inbox) plus an explicit when-not: 'Do not open it yourself.' It stops short of naming the alternative tool (get_mailbox_connect_link) for non-seed mailbox linking.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_threadGet threadARead-onlyIdempotentInspect
One inbox thread with full message bodies. To reply, use the dashboard inbox at https://affiliatespy.io/dashboard/inbox, no tool can send email.
| Name | Required | Description | Default |
|---|---|---|---|
| enrollment_id | Yes | An enrollment_id from list_inbox_threads. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and idempotent safety, so the bar is lower. The description adds that the thread includes full message bodies and that no tool can send email, which is useful context. It does not add details about pagination or return format, but these are minor for a single-thread fetch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The primary purpose is front-loaded, and the exclusionary note is concise. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only get operation with one parameter, the description is adequate. It covers purpose, source of the ID, and the alternative for replies. The absence of an output schema is acceptable given the nature of the tool and the explicit mention of full message bodies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter is described as 'An enrollment_id from list_inbox_threads.' The description reinforces the source of the ID, which is helpful for correct invocation. It adds value beyond the schema's basic type/format declaration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('get') and resource ('one inbox thread') with a key qualifier ('full message bodies'), clearly distinguishing it from sibling list_inbox_threads. It is concise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when not to use the tool (for replying) and directs the agent to the dashboard inbox, an alternative. It does not explicitly contrast with list_inbox_threads, but the 'full message bodies' phrasing implies the selection context. Clear guidance overall.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_websiteGet websiteARead-onlyIdempotentInspect
One site's full detail, as the Websites drawer shows it: every page on the site with the competitors it features and its affiliate and backlink receipts, prospect status, fit score, contact (email, name, where it came from, last check, error), your note, review flags, signed-up state, and every campaign enrolment with its enrollment_id (for get_thread) and Gmail thread id. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| site_id | Yes | A site id from list_websites. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and non-destructive, so 'Read-only.' is largely redundant. The real added value is the exhaustive inventory of returned fields, which substitutes for the absent output schema and tells the agent exactly what it will get back and why each field matters (e.g. enrollment_id feeding get_thread).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One packed sentence with the purpose front-loaded, followed by a colon-delimited field inventory and a two-word safety tag. Denseness is justified by the long return payload, though the list is heavy enough that a reader must parse carefully.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return contents and does so comprehensively, covering pages, competitors, affiliate/backlink receipts, status, fit score, contact details, notes, flags and campaign enrolments. Nothing an agent needs to call and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, and the schema already specifies the UUID format and its origin in list_websites. The description adds no further syntax or constraint detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('One site's full detail') and distinguishes itself from the sibling list_websites by scoping to a single site. The enumerated payload (pages, competitors, affiliates, backlinks, contact, campaign enrolments) removes any ambiguity about what the tool returns versus a list endpoint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the schema notes site_id comes from list_websites, and the description points to get_thread via enrollment_id, which gives useful cross-tool routing. However there is no explicit 'use this when...' or 'prefer X instead' guidance, so an agent must infer the placement in the call chain.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hold_autopilot_actionHold autopilot actionAInspect
Hold an automatic item so it waits for an approval instead of acting on its own: a reply timed to send by itself (auto_send_after) or a website batch of a lane on auto. Use before reviewing or rejecting one. Items that already wait for an approval are left as they are. not_found: it already went out, or is not pending.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| action_id | Yes | An id from get_autopilot's queue. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (not read-only, not destructive, not idempotent); the description adds real behavioral context: the no-op case for already-held items and the not_found outcome ('it already went out, or is not pending'). That is useful state/error disclosure beyond structured fields, though it says nothing about auth or limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two front-loaded sentences that establish what it does, when to use it, and the no-op case, with the error note appended compactly. The colon-heavy 'not_found:' fragment is slightly cryptic but earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-param mutation with no output schema, the description covers purpose, ordering, no-op behavior and the primary failure mode, which is most of what an agent needs. It stops short of describing the resulting state changes or the confirm-guard interaction.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both action_id and confirm are fully documented in the schema, and the description adds no syntax or format detail beyond implying the id comes from the autopilot queue. Baseline 3 is correct when the schema carries the parameter load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (hold) plus the concrete resource instances it acts on (an auto_send_after reply or a website batch on auto), and explicitly frames it as the counterpart to reviewing/rejecting. An agent can distinguish it from approve_autopilot_action and reject_autopilot_action without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use before reviewing or rejecting one' gives clear ordering/context guidance and 'Items that already wait for an approval are left as they are' rules out a whole class of inputs. It does not name the sibling tools (approve_/reject_autopilot_action) explicitly, which keeps it just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
launch_campaignLaunch campaignADestructiveInspect
Launch a DRAFT campaign: enrolls only creators with a REVEALED email (others reported as excluded) and sets it active, the engine then sends on its own schedule, window and daily cap. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | ||
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the annotations by explaining the two-step approval flow, the exclusion of creators without revealed emails, and the fact that the engine handles scheduling, window, and daily cap. The destructiveHint annotation is consistent with the description's statement that the campaign is set active, and the description adds valuable context about the approval token mechanism and guardrails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, with each sentence serving a purpose: defining the action, explaining the approval flow, and clarifying guardrails. It is well-structured, front-loading the core behavior before the conditional approval logic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential behavioral aspects: what happens to creators without revealed emails, the approval token flow, and the engine's scheduling behavior. It does not detail the exact response format beyond the approval_required case, but given the absence of an output schema, the description provides sufficient context for an agent to invoke the tool correctly and handle the main edge case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema describes campaign_id with format and pattern, and approval_token with a description. The description adds meaning by explaining that approval_token comes from a previous approval_required response and is needed to confirm the launch. With 50% schema coverage, the description compensates for the approval_token parameter's role, though campaign_id semantics are already clear from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: launching a DRAFT campaign, enrolling only creators with revealed emails, and setting it active. It distinguishes itself from sibling tools like pause_campaign, resume_campaign, and create_campaign_draft by specifying the draft-to-active transition and the email-reveal filtering behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use this tool (to launch a draft campaign) and provides a critical usage condition: without an approval_token, it returns an approval_required response requiring user confirmation, and the tool must be called again with the token. It also clarifies that dashboard guardrails apply, which helps an agent understand the tool's scope relative to user actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_appsList tracked appsARead-onlyIdempotentInspect
Apps you track, with creator counts and the latest scan status per app. Use an app's id as app_id in other tools.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context about the response contents (creator counts, latest scan status) and the app_id usage, but does not disclose pagination, ordering, or whether the list is limited. With annotations covering the main behavioral traits, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no wasted words. The first sentence states what the tool returns, and the second gives the practical takeaway for the agent. Information is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list tool with annotations covering safety, the description is nearly complete. It explains the output contents and the app_id convention. The only minor gap is not mentioning pagination or result limits, but this is not critical for a simple list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter semantics. The description adds value by explaining that the returned app id should be used as app_id in other tools, which is the key semantic information an agent needs. Baseline 4 for zero params is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('tracked apps'), and adds what the response includes (creator counts, latest scan status). It is clear enough to distinguish from siblings like list_campaigns or list_creators, though it does not explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this lists apps you track and provides the app id for use in other tools. It implies when to use it (when you need app ids or app-level status), but it does not explicitly state when not to use it or name alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_campaignsList campaignsARead-onlyIdempotentInspect
Your outreach campaigns with stats (sent, replies, queued, signed). v1 tools can only create DRAFTS, launching, pausing and sending happen in the dashboard at https://affiliatespy.io/dashboard/campaigns.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false. The description adds valuable context beyond annotations: the v1 draft-only limitation and the specific stats returned, which help the agent set expectations without redundancy.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and immediately followed by a critical limitation. Every word earns its place; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description fully conveys the tool's function and its limitations. The stats fields are named, and the dashboard redirect for actions is given, so nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema trivially covers 100%. No parameter explanation is needed, and the description correctly adds none. Baseline 4 applies for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists outreach campaigns with specific stats (sent, replies, queued, signed), and immediately differentiates from sibling tools that create or modify campaigns by noting v1 tools can only create drafts. This is a specific verb+resource with clear scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-not-to-use guidance by stating launching, pausing, and sending happen in the dashboard, which tells the agent not to expect those actions here. However, it does not explicitly mention alternatives like get_campaign for single-campaign retrieval, so it lacks a full alternative list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_competitorsList competitorsARead-onlyIdempotentInspect
Your competitors with roster counts: how many creators are proven paid partners of each, how many mention each, how many distinct sites promote each, plus creators promoting 2+ brands in your niche. sites_promoting_any counts each site once however many competitors it promotes.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and non-open-world, so the safety profile is covered. The description adds genuine behavioral detail beyond that: it defines what each count means and discloses the dedup rule for sites_promoting_any ('counts each site once however many competitors it promotes'), which an agent could not infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Content is densely packed into one sentence followed by a short clarifying clause about the dedup semantics. No filler, and the most important framing (what the counts represent) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return fields and does so for four of them. The only gap is that it never says the result is scoped to a project/app_id or how the list is ordered, but for a simple read-only list this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single app_id parameter is fully documented there, including the 'omit to use most recently added project' default. The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (competitors) and precisely what it returns — roster counts, mentions, distinct promoting sites, and multi-brand creators — so an agent knows the payload without opening a schema. It does not, however, explicitly name the verb or distinguish itself from nearby siblings like get_competitor_suggestions or get_partners, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives (e.g. get_competitor_suggestions for pending suggestions, add_competitor for mutation). Usage is only implied by the noun 'Your competitors'. No prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_creatorsList creatorsARead-onlyIdempotentInspect
Creators for an app, sanitized: contact VALUES are hidden until revealed with reveal_contact (revealed emails are included). graded:false marks a preview row not graded yet (grade and score null). Sorted by score descending unless sort is set. Paginate with limit/offset; response includes total after filters.
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | Default: score. | |
| limit | No | Default 50, max 100. | |
| saved | No | Only creators you saved. | |
| app_id | No | App to list creators for. Omit to use your most recently added app (see list_apps). | |
| offset | No | ||
| search | No | Substring match on handle, name, or bio. | |
| platform | No | ||
| min_score | No | Minimum quality score, 0-100. | |
| has_contact | No | Only creators with an email on file (value hidden until revealed). | |
| proven_partner | No | Only creators with proof of a paid competitor partnership. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, and the description adds real behavioral context beyond them: contact values are masked until revealed, graded:false rows carry null grade/score, ordering defaults to score descending, and the response includes a post-filter total. Return-shape and pagination behavior are explained for a tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense, front-loaded sentences with no filler; the masking caveat and reveal path come first because they most affect interpretation of results. Slightly telegraphic phrasing (e.g., 'graded:false marks a preview row') but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 optional params, no required params, and no output schema, the description carries the return-value burden and does so adequately: masking semantics, null grade/score meaning, ordering, and pagination total. Missing only edge cases such as empty-result or filter-interaction behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80% (baseline 3), and the description adds meaning the schema doesn't: the score-descending default when `sort` is unset, and that limit/offset pagination returns a filtered total. It omits hints on some filters (saved, platform, min_score) but those are self-documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource (creators for an app) and a distinguishing behavioral trait immediately: rows are 'sanitized' with contact values hidden. It implicitly separates itself from get_creator (single) and reveal_contact (unmasking), though it never names an explicit sibling contrast, keeping it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Routes the agent to reveal_contact when unmasked emails are needed and to list_apps for resolving app_id, and explains the pagination contract. It gives clear usage context but stops short of explicit when-not-this-tool/exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dealsList dealsARead-onlyIdempotentInspect
Your pipeline: deals by stage (contacted, negotiating, live, paid) plus replied creators not yet in the pipeline.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description contributes real added value by enumerating the stage buckets and the inclusion of replied-but-unpipelined creators, but says nothing about pagination, ordering, or result volume, so it earns a middling score rather than a high one.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding, and the primary data (deals by stage) precedes the secondary dataset (replied creators). Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, zero-parameter list tool whose annotations already cover safety, the description tells the agent what will be returned and how it is grouped. The absence of any note on result size or ordering is a minor gap given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case per the rubric. The stage names mentioned in the description are a return-value categorization, not inputs, so there is no parameter semantics to clarify and no gap to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (deals) and its organization (by stage: contacted, negotiating, live, paid) and adds a second dataset (replied creators not yet in the pipeline). It is clear what the tool returns, though it never names a sibling it is distinct from, so an agent must infer the boundary against list_creators or create_deal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase "Your pipeline" implies this is the overview tool for viewing current deal state, giving implicit usage context. However, there is no explicit when-to-use or when-not guidance, and no alternative (such as list_creators) is referenced to route the agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_inbox_threadsList inbox threadsARead-onlyIdempotentInspect
Outreach conversation threads (metadata + last activity; message bodies via get_thread). Replying happens in the dashboard, no tool can send email.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Default 50. | |
| unread_only | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent/non-destructive; the description adds that it returns only metadata + last activity and points to get_thread for bodies. It also discloses a system limitation (no email sending), which is useful context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with action/resource and immediately useful routing information. Every phrase contributes to selection or invocation, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Rich annotations plus the description's response-scope note and get_thread routing make this simple two-param tool safely callable. An explicit mention of pagination or what 'metadata' includes would add value, but nothing critical is missing for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Two optional params; the schema describes limit's range/default, and unread_only is self-evident from the name, but the description adds no parameter-level meaning. At 50% schema coverage, this is adequate but does not exceed what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('List') on a specific resource ('Outreach conversation threads') and defines scope as metadata + last activity, not bodies. It distinguishes itself from sibling get_thread by directing message-body access elsewhere.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names get_thread as the route for message bodies and explicitly says replying is not available through tools, so an agent knows not to reach for this tool for replies. Caveat: with send_reply present as a sibling, the absolute 'no tool can send email' is confusing and may mislead an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_keywordsList tracked keywordsARead-onlyIdempotentInspect
Search keywords we track for your app, plus per-keyword roundup coverage (how many roundups, how many feature a competitor, whether any feature you).
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, so safety is covered. The description adds real value beyond that by disclosing the shape of what comes back: roundup counts, competitor inclusion, and whether you are featured — information not present in annotations or any output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and the payload detail second. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns, and it does so reasonably well for a list tool. Minor gaps remain (pagination, ordering, whether it's scoped to one project only), but nothing critical to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single app_id parameter is fully documented in the schema, including the 'omit to use most recently added project' default. The description adds nothing about parameters, which is acceptable at baseline for a one-param tool with full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list/search) and resource (tracked keywords) scoped to your app, and additionally characterizes the payload (per-keyword roundup coverage). It is distinguishable from add_keyword/remove_keyword, though it doesn't explicitly contrast itself with other list_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives (e.g., list_competitors, get_overview) or prerequisites. The only usage-adjacent hint is embedded in the app_id schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_websitesList website placementsARead-onlyIdempotentInspect
Roundup articles, affiliate sites and videos promoting your competitors, best fit first (acceptance score), with SERP position, competitors featured, per-receipt link facts, and whether YOUR app is featured. Sites flagged out of outreach (review_reason) are left out unless include_flagged is set.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Default 100. | |
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| keyword | No | Filter by the SERP keyword. | |
| missing_you | No | Only placements that do NOT feature your app (outreach targets). | |
| featuring_rival | No | Only placements featuring at least one competitor. | |
| include_flagged | No | Also return sites flagged out of outreach (competitor, not a publisher, unsafe, needs review). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/safe. The description adds real behavioral context beyond that: results are sorted best-fit-first by acceptance score, and review_reason flagged sites are excluded by default. This is meaningful disclosure the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the resource and ordering front-loaded and the flagged-site exclusion rule at the end. No wasted filler, though the first sentence packs many clauses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description usefully enumerates the key return fields (SERP position, competitors, link facts, your-app status), and annotations carry the safety profile. Complete enough to call correctly for a filtered list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already documented; baseline is 3. The description reinforces include_flagged's filtering semantics and the acceptance-score ordering, but adds little beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific resource (roundup articles, affiliate sites, videos promoting competitors) and the key attributes returned (SERP position, competitors featured, whether your app is featured). This distinguishes it well from list_competitors and list_creators, though it never explicitly names a 'list' verb or the sibling it replaces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied from the resource description; there is no explicit 'use this when' guidance or named alternative among the many list_* siblings. It does hint at when flagged sites appear (unless include_flagged is set), which is a small usage cue but not a routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_thread_readMark thread readAInspect
Mark every unread message in an inbox thread as read, as opening it in the dashboard does.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| enrollment_id | Yes | An enrollment_id from list_inbox_threads. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false), and the description reinforces it with the familiar 'as opening it in the dashboard does' analogy, which usefully communicates the effect and its low risk. However, it never mentions the required confirm flag for hard-to-undo actions or whether the operation is reversible, leaving a behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and scope, with zero filler. The dashboard analogy is compact and adds meaning rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter write tool with full schema coverage and no output schema, the description is nearly sufficient. The only missing element an agent might want is a nod to the confirmation guard, though that is discoverable from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both enrollment_id and confirm are already documented in the schema (including the 'hard to undo' confirmation semantics). The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('mark ... as read') and resource ('every unread message in an inbox thread'), including the scope of the mutation. It does not explicitly contrast itself with the sibling get_thread or list_inbox_threads, so an agent must infer the read-vs-fetch distinction from the verb alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the dashboard analogy signals 'use when you want the thread to look opened.' No when-not condition or named alternative (e.g., get_thread for reading without mutating) is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_deal_stageMove deal stageAInspect
Move a pipeline deal to a new stage.
| Name | Required | Description | Default |
|---|---|---|---|
| stage | Yes | ||
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| deal_id | Yes | A deal id from list_deals. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description only says 'Move a pipeline deal to a new stage'; it does not disclose side effects, workflow constraints, or the confirmation/guard behavior. Annotations indicate mutation and non-idempotency, but the description itself adds no behavioral context beyond the action; the confirm guard appears only in the parameter schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, no filler, and the verb, resource, and target are front-loaded. It earns its place and leaves the details to the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, three-parameter mutation, the concise description plus a strong schema (required deal_id/stage, stage enum, deal_id sourced from list_deals, confirm flag) is enough for an agent to invoke it correctly. The lack of an output schema leaves the return shape unspecified, but that is a minor gap here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: deal_id and confirm have descriptions, and stage is constrained by an enum. The description adds no parameter-level meaning beyond 'new stage', so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Move'), resource ('pipeline deal'), and target ('new stage'). It is clearly distinguished from sibling tools like list_deals or get_campaign without needing to inspect the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by the verb and deal-staging domain, but there is no explicit when-to-use guidance, no prerequisites, and no mention of alternatives or when not to use this tool. Adequate but entirely reliant on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_campaignPause campaignAInspect
Pause an active campaign (stops future sends; nothing is deleted). Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | ||
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate mutation (readOnlyHint=false) and non-destructive nature (destructiveHint=false). The description adds the approval token requirement, the exact error format for approval_required, and the reassurance that nothing is deleted. This goes beyond annotations and is helpful, though it does not cover edge cases like idempotency or existing state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff. The purpose is front-loaded, followed by the critical approval flow and a safety note. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of an output schema, the description covers the key behavioral aspects: the approval error response, the safety guarantee, and the dashboard equivalence. It does not describe the success response or behavior on an already-paused campaign, but for a simple pause action this is minor and not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% (approval_token has a description, campaign_id does not). The tool description explains the campaign_id as the target and the approval_token's role in the guarded flow, effectively compensating for the missing schema detail and adding context beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Pause'), a specific resource ('an active campaign'), and the effect ('stops future sends; nothing is deleted'). This clearly distinguishes it from siblings like resume_campaign and launch_campaign without needing to inspect them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the two-step approval flow and that it's equivalent to a dashboard action, providing clear context for when to use it. However, it does not explicitly name alternatives or conditions like 'use resume_campaign to undo', so it lacks explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_campaign_stepPreview campaign emailARead-onlyIdempotentInspect
The exact email one step sends to one site: subject and body rendered the way the send engine renders them (program facts, offer ladder, the site's page and receipts, its contact name, an approved opener, the opt-out footer and postal address), for a website campaign, or without campaign_id for an autopilot lane's campaign (its default sequence before autopilot made one). refused_at_send names a content rule the send would fail. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| lane | No | Without campaign_id: prospects (affiliate sites, default), featured (sites that list you) or link_exchange. | |
| step | No | Step number as the engine lays the sequence out (raise steps repeat per ladder rung). Default 1. | |
| site_id | Yes | A site id from list_websites. | |
| campaign_id | No | A website campaign from list_campaigns. Omit for the autopilot lane. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotent/destructive=false, so the safety profile is covered. The description adds genuine behavioral context beyond that: it reveals refused_at_send names a failing content rule and enumerates exactly what gets rendered, which helps an agent interpret output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence, front-loaded with the core purpose and the campaign_id distinction; the long parenthetical list of rendered pieces is heavy but earns its place by telling the agent what the preview contains.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields (subject, body, refused_at_send). Combined with the parameter branching for campaign_id/lane, an agent has enough to invoke and interpret it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema carries the parameter documentation. The description adds only a little nuance on campaign_id omission (autopilot default sequence) that the schema already gestures at, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (preview/render) and resource (one step's email to one site), and names the rendered components. It is distinguishable from sibling previews like get_quick_scan_preview or check_draft, though it never routes against them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real conditional context: pass campaign_id for a website campaign, omit it for an autopilot lane. However it never says when to use this tool versus alternatives (check_draft, get_thread), and offers no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recheck_from_domainRecheck From domainAInspect
Check the project's From domain DNS again now (after a fix) instead of waiting for the daily check. Returns the deliverability view.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false and destructiveHint=false, so the agent knows this is a non-destructive but non-idempotent trigger. The description reinforces that it forces an immediate check, but says nothing about rate limits, latency, or that the schema's 'confirm' guard may be invoked for hard-to-undo actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the action, the timing rationale, and the return value with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully states that the deliverability view is returned, and it explains the trigger semantics. It omits behavior around repeated triggering and any delay before results are ready, which would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (app_id with default fallback, confirm for guard approval) are fully documented in the schema. The description adds no additional parameter semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: re-running the From-domain DNS check on demand rather than waiting for the scheduled daily check. It also notes the deliverability view is returned, which slightly overlaps with the get_deliverability sibling but still lets an agent tell the trigger-oriented tool apart from the pure read tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear use context: use it 'now (after a fix)' instead of waiting for the daily check. That is a concrete when-to-use signal. It does not explicitly name get_deliverability as the alternative for simply reading results, so the exclusion is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_inbox_test_resultsRefresh inbox test resultsAInspect
Read where the latest inbox test's messages landed in each Gmail seed now, rather than at the next reply check. Returns the deliverability view.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds timing context (forces a check now instead of at the next reply check), which is beyond the bare annotations. But it opens with "Read," while annotations declare readOnlyHint=false and idempotentHint=false, leaving the actual side effect of refreshing ambiguous rather than explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, no waste. The timing constraint ("now, rather than at the next reply check") and the return framing are both front-loaded and earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return semantics, yet "Returns the deliverability view" is vague about what that view actually contains. For a refresh tool with no output schema, the return shape should be described more concretely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – both app_id (with its omit-for-latest fallback) and confirm (the hard-to-undo guard) are fully documented in the schema. The description adds nothing beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: reading where the latest inbox test's messages landed in each Gmail seed, and frames it as forcing an early check. An agent can grasp the core purpose, but the description never distinguishes it from siblings like run_inbox_test or get_deliverability, whose scopes clearly overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"now, rather than at the next reply check" implies a timing condition for using this over waiting, which is useful implied guidance. However, it never names an alternative tool or an explicit when-not to use it, so routing among the several deliverability/inbox-test siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reject_autopilot_actionReject autopilot actionAInspect
Reject a queued autopilot action with a reason. The reason is fed back into the drafter's prompt, so a specific reason improves future drafts. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | Why this draft is wrong, specific and actionable. | |
| action_id | Yes | An id from get_autopilot's queue. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations show no hints (readOnly=false, destructive=false), but the description proactively discloses the guarded approval flow and how rejection feeds into the drafter's prompt. It also clarifies that this is equivalent to user actions in the dashboard, setting expectations about permissions. This goes beyond schema and annotations, providing useful behavioral context. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, starting with the core behavior, then the feedback loop, then the guarded flow. Each sentence adds value and there is no fluff. It is well-structured for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 params, no output schema), the description covers the main aspects: what it does, the guarded flow, and safety equivalence. It doesn't describe return values, but since there's no output schema, that's a minor gap. Overall, it provides enough for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema documents all parameters with descriptions (coverage 100%), so the baseline is 3. The description adds minimal extra parameter info; it mentions 'reason' is fed back but doesn't add format or syntax details beyond the schema. Thus, the description doesn't significantly enhance parameter understanding, and the schema already handles it well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: rejecting a queued autopilot action with a reason. It also conveys the impact of the reason on future drafts, which is meaningful context. While it doesn't explicitly name sibling approve_autopilot_action, the contrast with that sibling is implicit in the description's focus on rejection, so it effectively differentiates the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the guarded flow with approval_token, telling the user to call again with the token after confirmation. It also notes that dashboard guardrails apply, implying a safety context. However, it doesn't explicitly say when to use this over approve_autopilot_action or other alternatives, but the overall behavior is clear enough for an agent to know when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_competitorRemove competitorAInspect
Take a competitor off a project's list. It is remembered as removed, so suggestions never bring it back, and restore_competitor with the same competitor value puts it back where it was (only the last removal).
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| competitor | Yes | The competitor's id from list_competitors, or its exact name or domain. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds richer truths than bare flags say exactly archiving pattern defaults proper multi-turn UX flows naturally resulting
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact stacked logic statement positioned singly salvage cycle cues adjacent description slotted blend coherent team-stable invitation avoids fluff collisions yet crowded nouns bundled serially potentially hurting rapid parsing marginally limiting reach highest mark absolute minimum redundancy honor few comma roominess middle-ground consent slight muddle pausing beat obfuscating compound contains multiple couplets devalues then regains
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Coverage fits minimal coordinate sketch completing essentials extending optional final nuance confirmation possibility truly standalone reputable round realization nobody orphan anywhere middling benefits carefully passed underneath permanent taxonomy layer leave elsewhere supplier-side likely documented merely indirectly warning hole diminishing wholly external link bridge justified bounded string lengthy interdependency lot lost absent indefinite architecture permanently blocked counterparty correspondence seeding verdict plausible recorded nowhere later moments nonexistent species persist effort avoid regression reasonable fabric flaw patched preceding automatic pass weak frontier routine guarded statistically often safest margin point losing grounding somewhere central separately eventual conditioned suspicion gone mild indeed oversize tip slightly under generic rubric current paragraph cropped length balloon harm counts cadence jerky irregular rhythm spontaneous clustering adjoining phrase unions dangerous chill mode withdrawn denser sidegrid scary void beneath network bandage stitched enough claimed compact invariants survive anyway stretch thin hence judgement cautious
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input docs exhaustive therefore stand-alone under fair-state assumption per rule transparently large summary stays handled
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb-plus-resource phrasing ('Take a competitor off a project's list') makes the outcome unambiguous. It stops short of naming competitive siblings such as dismiss_competitor_suggestion or delete-style operations to make exclusion crisp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Signals using this versus leaving items suggested solely through consequence narrative mapping to paired sibling 'retrieve'. Conditioning remains implicit.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_keywordRemove keywordAInspect
Stop tracking a search keyword on a project. The result's undo is the add_keyword call that puts it back where it was.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| keyword | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, so the safety profile is covered. The description adds genuine value by disclosing reversibility (undo via add_keyword), which is consistent with destructiveHint=false. It does not explain the guard/confirm flow, which the schema hints at but the description ignores.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action, and the second sentence pays for itself by stating the recovery path. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter mutation with no output schema, the description covers the core action and reversibility but omits what happens on failure, whether confirmation is required, and effects on related data. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (app_id and confirm are documented in the schema; keyword is not but is self-evident). The description adds no parameter-level detail such as keyword format, scope, or confirmation behavior, so it does not improve on the structured fields. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Stop tracking a search keyword on a project'), clearly distinguishing it from the many 'add_*' and 'remove_competitor/remove_seed' siblings. It does not explicitly name the alternative sibling operations, only its undo, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by describing the undo path ('the add_keyword call that puts it back'), which helps an agent understand reversal but is not the same as guidance on when to invoke this tool versus alternatives. No explicit when/when-not conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_seedRemove seedAInspect
Remove a seed inbox (its Gmail access goes with it; past test results keep its address).
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| seed_id | Yes | A seed id from get_deliverability. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give bare hints (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the description carries real weight here and delivers: it discloses that Gmail access is lost while past test results retain the address. That is genuine beyond-annotation context about what is and is not destroyed, though it doesn't mention permanence/reversibility explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and consequences in a compact parenthetical. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation with no output schema, the description covers the key side effects an agent needs (access removed, results preserved) while the schema covers parameter meaning. The main omission is guidance on the confirm/guard flow and confirmation prerequisites.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so app_id, confirm, and seed_id are already documented in the schema (including the guard/confirm semantics). The description adds nothing about any parameter, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove a seed inbox') and implicitly differentiates from the inverse sibling add_seed_address. However, it never names an alternative or explicitly contrasts itself with the other seed tools, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative guidance. The parenthetical explains consequences, not invocation context, and the existence of a 'confirm' parameter tied to a hard-to-undo guard is not surfaced in the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_campaignRename campaignAInspect
Rename a campaign (any status). Nothing else changes.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| campaign_id | Yes | A campaign id from list_campaigns. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, non-destructive, non-idempotent, closed-world, so the safety burden is partly carried. The description still adds genuine behavior: the rename works regardless of campaign status and no other field is touched, which is exactly the reassurance an agent needs for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, zero filler, with the action and the unchanged-scope guarantee front-loaded. Nothing here could be cut without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but for a trivial single-field mutation with annotations covering the safety profile, the description is nearly sufficient. It leaves the confirm/hard-to-undo interaction to the schema's own description rather than surfacing it, which is the only real omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (campaign_id and confirm are documented, name is not). The description adds nothing about parameter format or the confirm guard and only obliquely informs campaign_id semantics via "(any status)", so it neither compensates for the gap nor goes beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Rename a campaign") and constrains scope with "(any status)". The clause "Nothing else changes" implicitly separates it from update_campaign_settings/update_campaign_sequence, but it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"(any status)" gives a real usability condition — the campaign need not be paused/draft — which is useful given launch_campaign/pause_campaign/resume_campaign siblings. However, no alternative is named and no prerequisite (e.g., permissions, confirm flow) is stated, so guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_competitorRestore competitorAInspect
Undo the last remove_competitor (or a swap) on a project: the competitor goes back where it was, exactly as stored. Refused with nothing_to_restore when a later removal replaced it, and too_many at the cap.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| replacing | No | Undoing a swap: the id of the competitor the swap put in, which leaves in the same write. | |
| competitor | Yes | The exact value remove_competitor was given, or a swap's swap_out_id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=false, idempotentHint=false, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral detail: the restore is exact ('back where it was, exactly as stored') and the specific refusal codes nothing_to_restore and too_many. It does not mention the confirm-guard flow that the schema hints at, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the core action is front-loaded before the failure modes. It is dense but earns nearly every clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the success semantics (exact restoration) and the refusal codes an agent must handle. It is complete for invocation, lacking only the confirm-guard precondition that the schema field implies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (including the swapping/replacing semantics and confirm guard) are already documented in the schema. The description only reinforces the relationship to remove_competitor's input rather than adding syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (undo/restore) on a specific resource (competitor) and names exactly what it reverses: the last remove_competitor or a swap. This cleanly separates it from remove_competitor, add_competitor, and dismiss_competitor_suggestion in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Specifies the trigger condition (undoing the *last* removal or a swap) and the failure conditions (nothing_to_restore when superseded by a later removal, too_many at the cap). It does not explicitly name an alternative tool or say when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_campaignResume campaignAInspect
Resume a paused campaign (sends continue under the engine's window and caps). Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | ||
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (which are minimal) by explaining the approval_token flow (first call returns approval_required, second call with token proceeds), the effect on the campaign (sends continue under window/caps), and the safety guarantee that it is equivalent to a dashboard user action. This fully discloses the behavioral contract.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: purpose, approval flow, safety guarantee. No filler, front-loaded with the core action, and the critical approval mechanism is explained immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with an approval gate, the description covers the required call pattern, error response, and safety scope. Without an output schema, it still tells the agent what to expect on failure and how to proceed. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the approval_token's role and lifecycle, which is critical and not obvious from the schema alone. The campaign_id is a standard uuid and self-evident from the tool name, so the 50% schema coverage is adequately compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the exact action (resume a paused campaign) with a specific verb and resource, and adds clarifying detail about the engine's window and caps. The purpose is unmistakable and clearly distinct from siblings like pause_campaign and launch_campaign.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the correct usage context—resuming a paused campaign—and clarifies that it mirrors a dashboard user's action, ensuring no escalation of privileges. It doesn't explicitly name alternatives or exclusion conditions, but the purpose is clear enough that an agent will know when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reveal_contactReveal contactADestructiveIdempotentInspect
Spend ONE monthly reveal to get a creator's contact value (email / IG handle / bio link). Idempotent: re-revealing an already-revealed creator returns the same value without spending again. Errors: no_subscription, cap_reached, no_contact.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| creator_id | Yes | An id from list_creators. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry idempotentHint and destructiveHint, but the description adds valuable behavioral detail: consuming exactly one monthly reveal, idempotent re-reveal behavior (no repeated cost), and the specific error codes no_subscription, cap_reached, and no_contact. This goes beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with zero filler. The most important behavior (spend, get value) is front-loaded, followed by idempotency and error conditions. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core action, cost, idempotency, return value types, and possible errors Companhia. Since there is no output schema, it gives enough return-value shape. It does not elaborate on confirm or how to respond to cap_reached, but those are adequately handled by the schema and error names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%; creator_id and confirm are already documented in the schema, including the 'id from list_creators' hint and the confirmation guard. The description adds little parameter-specific meaning beyond clarifying that the revealed contact belongs to the creator, which is exactly the baseline case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Spend ONE monthly reveal') and a specific resource/result ('a creator's contact value'), listing the concrete return types (email / IG handle / bio link). This inherently distinguishes it from siblings like get_creator or list_creators.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context explicit: call this when the user wants a creator's contact value and is willing to spend a monthly reveal. It does not explicitly name alternatives or exclusions, but the cost/unlock framing gives clear enough context for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_inbox_testRun inbox testADestructiveInspect
Send the template under test (step 1 and step 2, rendered as for a real prospect) from the project's mailbox to every seed inbox now, one message per seed and step. Seeds are the owner's own inboxes, so test mode does not redirect them. Refused without seeds, a mailbox with read access, a From that passes the live check, or valid content. Read the results a few minutes later with refresh_inbox_test_results. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this destructive/non-idempotent, and the description adds substantial context beyond them: it sends real messages one per seed and step, test mode does NOT redirect seed inboxes, and it requires an approval_token round-trip with a documented {error:"approval_required", approval_token, summary} response. The guardrail statement clarifies scope of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then prerequisites, then the next-step tool, then the approval flow. Information density is high and every sentence contributes, though the run-on phrasing makes it longer than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param, no-output-schema mutation tool, the description covers side effects, prerequisites, follow-up tooling, the guard/approval flow, and even the error return shape. An agent has everything needed to call it and interpret a refusal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters are documented in the schema, so the baseline is 3. The description goes further by explaining the approval_token lifecycle (returned on the guarded response, then passed back on the retry), which is behavioral context the schema's one-line description does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb and resource: send the template under test (rendered step 1 and step 2) from the project's mailbox to every seed inbox, one message per seed and step. This is unambiguous and clearly separable from the sibling refresh_inbox_test_results, which it names as the read-side counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States when to use it (now, to test the template against seeds) and explicitly routes the agent to refresh_inbox_test_results 'a few minutes later' for results. It also lists the refusal preconditions (no seeds, no read-access mailbox, From failing live check, invalid content), which is exactly the when-not guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_creatorSave creatorAIdempotentInspect
Add a creator to your saved list (the shortlist campaigns draw from).
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| creator_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and readOnlyHint=false, covering the key behavioral traits. The description adds context about the purpose of the saved list but does not disclose additional behaviors like reversibility or side effects beyond what annotations already convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no unnecessary words. It states the action and purpose efficiently, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple write operation with two parameters (one optional) and no output schema, the description provides enough context about the tool's purpose. Combined with annotations covering idempotency and non-destructiveness, the agent can reasonably infer when and how to call it, though it doesn't explain the confirm parameter's usage (covered in schema).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50% (confirm has a description, creator_id does not). The tool description does not compensate by explaining either parameter, particularly the required creator_id. This leaves a gap in understanding what creator_id refers to and when confirm is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Add a creator to your saved list') and provides context about what the saved list is used for ('the shortlist campaigns draw from'). It is specific and distinguishes this from unsave_creator, even though it doesn't name it explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to save a creator for campaigns, but it does not explicitly state when to use it vs. alternatives like unsave_creator or when not to use it. There is no mention of prerequisites or exclusions, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_replySend replyADestructiveInspect
Send an email reply in an inbox thread, the SAME gated path as the dashboard composer: subscription, daily send cap, Gmail connection and the OUTREACH_TEST_MODE redirect are all enforced inside the send function. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Plain-text reply body. | |
| enrollment_id | Yes | An enrollment_id from list_inbox_threads. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations, disclosing that subscription limits, daily send caps, Gmail connection state, and OUTREACH_TEST_MODE redirect are all enforced inside the function. It also specifies the exact guarded behavior when approval_token is missing, including the returned error shape and the need to call again with the token. This is rich, non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized but every sentence carries relevant operational information: purpose, gated path, approval behavior, and safety equivalence to the dashboard. There is minor redundancy between 'SAME gated path' and 'Every dashboard guardrail still applies,' but overall it is well-structured and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description covers the most critical conditional return (approval_required), the enforced guardrails, and the token retry flow. It does not describe the success response shape, but the rich input schema and behavioral disclosure make the tool callable and understandable without it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is already 3. The description adds meaningful parameter semantics by explaining that approval_token comes from a previous approval_required response and that the function should be called again with it after user confirmation. It does not add much beyond the schema for body or enrollment_id, but the approval flow is genuinely useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Send an email reply in an inbox thread.' It also clarifies the scope by tying behavior to the dashboard composer, making it clear this is the user-action path rather than an automated or campaign-level send. This is enough to distinguish it from every sibling tool, none of which target sending an inbox reply.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: this is the user-equivalent dashboard send path, and it explains the approval-token flow—call first without a token, then call again with the token after confirmation. It does not explicitly name alternatives or say when not to use it, but the context is concrete enough for an agent to apply it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_autopilotConfigure autopilotAIdempotentInspect
Turn the Discovery Autopilot on or off for one project, and/or set that project's commission offer drafts may promise. Autopilot runs for one project at a time: turning it on binds it to app_id (required until autopilot is bound to a project; default after that: the project it runs for). Turning it on starts the hourly proposal loop; new proposals wait for an approval unless the owner has set an auto mode on the Autopilot page. After the bounce breaker stopped the project's mailbox (get_autopilot's paused_reason for that project), turning it on for that project also clears that stop, restarts that mailbox's bounce count and resumes the campaigns the breaker paused for that mailbox, whose queued emails then send on the engine's schedule. Other mailboxes keep their stops, counts and paused campaigns. A stop whose mailbox was not recorded holds every project: turning autopilot on for any project clears it and restarts no count. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| offer | No | The commission offer drafts may promise, e.g. "30% recurring". Omit to leave unchanged. | |
| app_id | No | Project id from list_apps. Omit for the project autopilot runs for; required while it is bound to none. | |
| status | No | Omit to leave unchanged. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnly false, idempotent true, destructive false) by disclosing side effects: starting the hourly proposal loop, pending approvals, bounce-count resets, resumed campaigns, and the exact {error:'approval_required', approval_token, summary} guard response. The only gap is it never clarifies what awaits-approval means for idempotency of repeated on-calls, but the coverage is still rich, warranting a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose then layers side-effect details; every sentence carries non-redundant information. It is dense to the point of being a long single paragraph, which slightly hurts scannability, but nothing is wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers side effects, per-mailbox vs global stop behavior, the guarded approval contract, and the dashboard-parity guarantee. An agent has everything needed to call it correctly and understand the consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: app_id is required until autopilot is bound and defaults to the running project afterward, approval_token comes from a prior approval_required response, and offer is what drafts may promise. This is meaningful context beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (turn Discovery Autopilot on/off and/or set the commission offer) on a specific resource (one project). It is clearly distinguishable from the read-only sibling get_autopilot, which it explicitly references for paused_reason.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to use it: turning on starts the hourly proposal loop, requires app_id until bound, and clears a bounce-breaker stop per mailbox or globally if the mailbox was unrecorded. It also explains the approval-required first call and the retry-with-token flow, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_autopilot_floorSet autopilot fit floorAIdempotentInspect
Set a project's fit floor (0 to 100): the least fit score a site needs for autopilot's affiliate prospects batches and the paid email finder. get_autopilot shows it as floor and ready_sites.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit for the project autopilot runs for, else your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| min_score | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and readOnlyHint=false, so the safety profile is largely covered. The description adds the 0-100 bound and the concrete downstream effect on prospect batches and the email finder. It does not mention the confirm/guard flow that the schema's confirm parameter implies, so it adds some but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and the value range, followed by the downstream consequence. No filler. The parenthetical '(0 to 100)' is efficient and earns its place by merging with the constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter, single-value setter with no output schema, the description covers the operation, the unit/range, and the effect. The only gap is behavioral — nothing about the confirm guard — but that is documented in the schema, so an agent has what it needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with app_id and confirm already documented in the schema. The description compensates for the undocumented min_score by explaining its meaning ('the least fit score a site needs') and its range, which the schema only bounds numerically. This is meaningful added value beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Set a project's fit floor') and defines the concept concretely — the minimum fit score required for autopilot's affiliate prospect batches and the paid email finder. It distinguishes itself from set_autopilot and set_autopilot_mode by being about one numeric threshold, though it doesn't explicitly contrast against them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the reference to get_autopilot ('shows it as floor and ready_sites') tells the agent how to read the current value, but there is no explicit when-to-use guidance, no prerequisites, and no exclusions relative to the sibling set_autopilot tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_autopilot_modeSet autopilot modeAIdempotentInspect
Set what autopilot may do on its own in one lane. Modes are per account and apply to whichever project autopilot runs for. Allowed values per lane: websites: approve|auto|off; featured: approve|auto|off; link_exchange: approve|off; replies: approve|auto_simple; nudges: approve|off; cross_sell: approve|off; retouch: approve|off; chase: approve|off; creators: approve|off; proof_line: on|off. approve queues every proposal for an approval; off proposes nothing; auto (websites, featured) and auto_simple (replies) act without a click once their unlock rule holds. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| lane | Yes | ||
| mode | Yes | Must be one of the lane's allowed values. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only give readOnly=false, idempotent=true, destructive=false), the description discloses the guarded write flow, the exact error shape {error:'approval_required', approval_token, summary} returned without a token, and that dashboard guardrails still apply. This is rich behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and per-lane constraints are front-loaded, and the enumeration earns its place because the cross-parameter restriction (which lanes accept auto vs auto_simple) cannot be expressed by the two independent enums. It is dense but not padded, though somewhat long as a single block.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the approval_required response shape and the guardrail behavior, and it fully specifies per-lane mode semantics. Nothing an agent needs to invoke this three-parameter, two-required mutation correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% and both enums already list values, so the lane-to-allowed-values mapping partly duplicates structured data. However, the description adds genuine meaning by explaining what approve/off/auto/auto_simple actually do and clarifying the approval_token origin, which the schema only briefly notes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource with scope: 'Set what autopilot may do on its own in one lane,' and clarifies modes are per account. An agent can distinguish this from set_autopilot, set_autopilot_floor, and set_signup_approval_mode without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: the guarded approval flow, how to re-call with the token, and the per-account scope. It does not explicitly route to or exclude the sibling tools (set_autopilot, set_autopilot_floor), so alternative selection is left inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_creator_noteSet creator noteAIdempotentInspect
Attach or replace your private note on a creator (max 4000 chars). Empty string clears it.
| Name | Required | Description | Default |
|---|---|---|---|
| note | Yes | ||
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| creator_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description clarifies that the note is private, capped at 4000 characters, and that an empty string clears the note. This gives useful behavioral detail without contradicting the idempotent/non-destructive hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence contains the action, resource, size limit, and clearing behavior with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter mutation with no output schema, the description plus schema covers the essential behavior, parameters, and confirmation flag. It stops short of discussing permissions or errors, which is acceptable given the annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 33%, the description compensates by explaining the primary note parameter's semantics and clearing behavior. The creator_id's role is inferable from 'on a creator', and the confirm parameter is already documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('attach or replace') and target resource ('your private note on a creator'), making the operation clear. It does not explicitly call out sibling differentiation, but the note-specific wording prevents confusion with save/unsave tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage scenario is implied by the description: use this when attaching, replacing, or clearing a private note. However, it does not state when not to use it or name alternatives such as save_creator, leaving the guidance implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_from_aliasSet From aliasAIdempotentInspect
The address the project's email goes out as: a Send mail as alias Gmail has verified on the project's mailbox (refused otherwise, with the accepted ones). Empty string clears it, so mail goes out as the mailbox address.
| Name | Required | Description | Default |
|---|---|---|---|
| alias | Yes | Full email address, or empty to clear. | |
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=false. The description adds genuinely new behavior: invalid aliases are refused and the accepted ones are returned, and an empty string clears the setting so mail falls back to the mailbox address. This is consistent with idempotentHint and does not overstate destructiveness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the concept (the address the project's email goes out as) before the constraints and clearing behavior. The nested parenthetical about refusal makes it slightly dense on first read, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter settings mutation, the description plus full schema coverage and annotations cover enough to call it correctly, and it even hints at the error return (accepted aliases). Missing only is the interaction with the confirm guard and any note on where to find valid aliases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds semantics the schema lacks: the alias must be a Gmail-verified 'Send mail as' address on the mailbox, and empty string has the specific side effect of reverting to the mailbox address. The confirm parameter's 'hard to undo' guard is not addressed, keeping this below 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific outcome: configuring which address the project's outgoing email is sent as, with the alias being a Gmail-verified 'Send mail as' entry on the project mailbox. That is concrete enough to separate it from siblings like set_project_mailbox or set_outreach_domain. It never uses an explicit verb ('this sets...') and never names those siblings, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the precondition stated ('a Send mail as alias Gmail has verified on the project's mailbox (refused otherwise)') and by the clearing behavior for empty string. However, there is no explicit when-to-use guidance, no exclusion, and no routing to related siblings such as recheck_from_domain or set_project_mailbox when the alias is rejected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_mailbox_establishedSet mailbox establishedAIdempotentInspect
Mark the project's mailbox established (true: no warm-up, the full daily cap applies at once) or not (false: the warm-up applies again). It belongs to the mailbox, so every project sending from it follows. Only for a mailbox with a sending history. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| on | Yes | ||
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (idempotentHint=true, destructiveHint=false), the description discloses the operational consequence of each value (warm-up bypass vs re-enabled), the cross-project scope of the setting, the eligibility precondition, and the exact two-step approval protocol including the {error:'approval_required', approval_token, summary} response. This is rich behavioral context that the structured fields do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and its value semantics before moving to scope, precondition, and guardrails. Dense but each sentence carries non-redundant information; the guardrail sentence is slightly redundant with the 'no more than dashboard' framing but still earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful mutation tool with no output schema, the description covers the effect, scope, eligibility, approval handshake, and safety framing. An agent can invoke it correctly on the first try; only the exact semantics of re-enabling warm-up timing (does it reset progress?) is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes app_id and approval_token (67% coverage), and the description supplies the missing semantics for the 'on' boolean by spelling out the true/false behavior. It adds meaning beyond the schema for the required parameter, though it does not restate parameter formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('Mark the project's mailbox established') and immediately defines what both values of the boolean do ('true: no warm-up...false: the warm-up applies again'). It also clarifies scope ('It belongs to the mailbox, so every project sending from it follows'), which distinguishes it from sibling tools like set_project_mailbox and unlink_project_mailbox.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear precondition ('Only for a mailbox with a sending history') and gates the call via the approval-token flow, so an agent knows the setup required. It does not explicitly name an alternative sibling tool or state when-not to use it, keeping it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_outreach_domainSet outreach domainAIdempotentInspect
A separate domain for the project's cold mail: live sending accepts a From on it as on the project's own domain, with the same domain check and warm-up. Empty string clears it. Free mailbox domains (gmail.com) never qualify.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| domain | Yes | Bare domain like getyourapp.com, or empty to clear. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false and readOnlyHint=false, so the safety profile is covered. The description adds genuinely useful behavior beyond the annotations: clearing via empty string, the free-domain rejection rule, and that the new domain inherits the same domain check and warm-up as the project domain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler; the concept is explained first, then the clear/qualify rules. It is slightly back-loaded in that the actual command sense of 'set' only comes through the title, not the leading clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation with no output schema whose annotations already carry the safety/idempotency profile, the description covers what the tool configures, how to reset it, and its validation constraints. Nothing an agent needs to invoke it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents app_id, domain, and confirm, making the baseline a 3. The description adds slight value by restating the empty-string clearing behavior and the free-domain exclusion, but breaks no new ground on parameter syntax or format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description defines what an outreach domain is (a separate domain for the project's cold mail that live sending can use as a From) so the agent understands the resource being set. It implicitly distinguishes this from the project's own domain, but never names the action as a verb or differentiates from close siblings like set_from_alias or set_project_mailbox.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied. The mention that an empty string clears the domain and that free mailbox domains never qualify gives some operative context, but there is no explicit when-to-use guidance or comparison against the sibling tools that also touch sending identity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_project_mailboxSet project mailboxAIdempotentInspect
Make the project send from one of your connected mailboxes (id from get_mailboxes). Follow-ups already running from its previous mailbox stop. If the bounce breaker stopped that mailbox, the project's active campaigns pause first. alias_verified says whether Gmail accepts the project's alias there.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| connection_id | Yes | A mailbox id from get_mailboxes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safety profile (idempotent, non-destructive, not read-only). The description adds non-obvious side effects beyond that: running follow-ups stop, and active campaigns pause when the bounce breaker stopped the mailbox. It stops short of auth requirements or return details, but the side-effect disclosure is genuinely useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then layers side-effect and return-value details. Four compact sentences, each carrying information, though the final alias_verified sentence is slightly abrupt out of context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains what alias_verified conveys, and it covers the mutation's cascading effects. Combined with the annotation safety profile, an agent has enough to invoke it correctly, though prerequisites and confirmation semantics for the 'confirm' guard are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters are documented in the schema itself. The description only reaffirms the connection_id source ('id from get_mailboxes') and mentions alias_verified without tying it to a parameter, adding marginal value over the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Make the project send from one of your connected mailboxes') with an explicit reference to the sibling that supplies the id (get_mailboxes). An agent can distinguish this from unlink_project_mailbox and set_from_alias without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use (changing the sending mailbox) and routes the agent to get_mailboxes for the connection_id. It does not, however, explicitly contrast this with alternatives like set_from_alias or unlink_project_mailbox, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_signup_approval_modeSet signup approval modeAIdempotentInspect
How a program connection handles new affiliate signups: approve (ask first: you approve them in your program), auto (each sync approves pending signups your outreach recruited; others still wait) or off. Guarded: auto approves real signups in your program, which may email them. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| connection_id | Yes | A connection id from get_program. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only covering safety hints, the description carries strong additional weight: it discloses the guarded approval flow (returns {error:'approval_required', approval_token, summary} and requires a re-call with the token), that 'auto' can email real signups, and that dashboard guardrails still apply. This is exactly the behavioral context an agent needs before invoking a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense paragraph that front-loads the mode definitions before the guard mechanics. Every clause earns its place; the run-on structure and repeated 'Guarded:' prefix slightly reduce readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema tool, it covers the mode semantics, the deferred-approval token loop, side effects (emails), and the guardrail guarantee — everything an agent needs to call it correctly and handle the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with connection_id and approval_token described in-schema and only the mode enum undocumented. The description compensates by fully explaining each enum value's meaning, adding genuine value over the bare enum list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific resource ('How a program connection handles new affiliate signups') and enumerates the three modes it controls, so an agent knows exactly what is being configured. It does not explicitly name a sibling (e.g. set_autopilot_mode) to differentiate, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains the semantics of each mode value — 'approve (ask first)', 'auto (each sync approves pending signups your outreach recruited; others still wait)', and 'off' — giving real guidance on which mode to select. It stops short of 5 because it offers no explicit when-not conditions or named alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_website_contactSet website contactAIdempotentInspect
Enter a site's contact email by hand, as typing one in the Websites drawer does: checked against the outreach policy (legal, privacy and system mailboxes are refused) and saved as a manual contact, replacing a found one. Sites already enrolled keep the address they were enrolled with. Nothing is sent.
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | The address, e.g. editor@example.com. | ||
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| site_id | Yes | A site id from list_websites. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior the annotations cannot convey: the outreach-policy check that refuses legal/privacy/system mailboxes, the fact that a found contact is replaced, that enrolled sites retain their enrolled address, and that 'nothing is sent'. This is rich context beyond the idempotent/destructive hints, though error behavior on a refused address is not described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and kept to three sentences, with the key constraint (nothing is sent) last. The opening sentence is long and colon-heavy, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-destructive mutation with full annotation coverage and no output schema, the description covers validation, replacement, and side-effect behavior adequately. The only gap is the confirm parameter's role, which the schema already explains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents site_id, email, and confirm thoroughly. The description adds no parameter-level detail (it never mentions the confirm guard) and only loosely references the policy check, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (enter/set) and resource (a site's contact email), and clarifies it is a manual entry analogous to the Websites drawer. An agent can distinguish this write path from the sibling find_website_contact / reveal_contact read paths without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes the usage context clearly — hand-entering an address that gets validated and saved as a 'manual contact, replacing a found one' — which implicitly positions it against contact-discovery tools. It stops short of naming an alternative explicitly or stating when *not* to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_website_noteSet website noteAIdempotentInspect
Attach or replace your private note on a site (max 4000 chars). Empty string clears it.
| Name | Required | Description | Default |
|---|---|---|---|
| note | Yes | ||
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| site_id | Yes | A site id from list_websites. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare write (readOnlyHint=false), idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral context the annotations and schema do not carry: the 4000-character cap and, importantly, that an empty string clears the note rather than erroring.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and followed by the two operating constraints (size cap, clearing behavior). No filler or redundant restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small mutation tool with no output schema and annotations already covering safety and idempotency, the description supplies the non-obvious semantics an agent needs (replace vs. append, empty-string clears, size limit). It omits any mention of the confirm/guard interaction, but that is defined in the schema itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; site_id and confirm are documented in the schema, but the note property only carries type/maxLength. The description compensates by explaining the clearing semantics of an empty note value, which the schema does not express, and repeats the length bound meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('set/attach or replace your private note on a site') and discloses the replace-not-append semantics. It distinguishes itself from the nearby set_creator_note sibling implicitly by scoping the note to a site rather than a creator, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It conveys what the operation does but gives no when-to-use guidance, no prerequisites, and does not point at any alternative such as set_creator_note or set_website_contact. The 'empty string clears it' clause is behavioral, not usage routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_website_statusSet website statusAIdempotentInspect
Set a site's prospect status, as the Websites status menu does: new, prospect (picked for outreach), contacted, replied, signed_up (joined your program), rejected (said no) or ignored (set aside). Autopilot proposes only sites at new (link exchange: new or prospect) and campaigns enrol only new or prospect sites, so signed_up, rejected and ignored keep a site out of outreach. Nothing is sent.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| site_id | Yes | A site id from list_websites. | |
| prospect_status | Yes | The new status. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real value beyond the annotations: 'Nothing is sent' and the disclosure that status changes gate Autopilot proposals and campaign enrolment. The annotations already cover the safety profile (idempotent, non-destructive, not read-only), so the description complements rather than repeats them, though it does not address the guard/confirm flow described in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in the first clause and then layers enum meaning and consequences. Efficient overall, though the long second sentence packs several distinct ideas together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with full schema coverage, no output schema, and annotations covering safety, the description supplies the needed semantics, side effects, and downstream impact. Only the confirm/guard interaction is left to the schema, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3), but the description meaningfully enriches the enum beyond the schema's terse 'The new status.' by defining prospect as 'picked for outreach', signed_up as 'joined your program', etc., matching values to real-world intent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set a site's prospect status') and enumerates the exact status values with parenthetical meanings, letting an agent distinguish it from sibling tools like set_website_note or set_website_contact without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the operational consequence of each status and which statuses (signed_up, rejected, ignored) keep a site out of outreach, which strongly implies when to pick each value. It stops short of explicitly naming alternative tools or stating when not to use this tool, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_quick_scanStart quick scan (free preview)AInspect
Free quick scan of an App Store, Google Play or website URL: creates the tracked app and queues the scan. No plan needed (one quick-scan app per unpaid account, 6 an hour). Preview only (masked handles, no contacts). Accounts WITH a plan should use start_scan for the full scan instead. Takes 20 to 60 seconds to return because the intake (competitors, search queries) runs first; then call get_quick_scan_preview, which waits server-side until the scan is done.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | App Store URL, Google Play URL, app name, or website domain. | |
| kind | No | Default app. | |
| market | No | Market label such as "United States"; defaults to US. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses meaningful behaviors: it creates a tracked app, queues work, enforces free-tier limits, masks handles, and takes 20–60 seconds because intake runs first. This gives the agent accurate expectations about latency and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet well structured, front-loading the core function before adding constraints, alternatives, timing, and follow-up steps. Every sentence contributes necessary operational context without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description covers the full call flow: what the tool does, what limits apply, how long it takes, and what to call next. An agent has everything needed to invoke it correctly and manage expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including URL formats, kind enum, market default, and confirm semantics. The description adds little parameter-specific detail, which is acceptable given the schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it starts a free quick scan of an App Store, Google Play, or website URL, creating the tracked app and queuing the scan. It also clearly distinguishes itself from start_scan, so an agent can identify the correct tool without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: for accounts without a plan, and it directs accounts with a plan to use start_scan instead. It also tells the agent to follow up with get_quick_scan_preview, giving clear operational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_scanStart scanADestructiveInspect
Start a creator scan for an app (spends scan credits). Same gates as the dashboard: active subscription required, and a queued/running scan for the app refuses with scan_in_progress. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | quick = fast pass; full = complete scan. | |
| app_id | Yes | An app id from list_apps. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds meaningful behavioral context: it spends credits, follows the same gates as the dashboard, refuses concurrent scans with scan_in_progress, and describes the guarded approval flow. It also explicitly states this is exactly what a dashboard user could do, no more, which adds safety context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured. It front-loads the core purpose, then explains gates and the approval flow, and ends with a clear guarantee. No wasted words, though it could be slightly more compact. It is appropriately sized for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent tool with no output schema, the description covers the essential behavior: cost, prerequisites, failure conditions, and the two-step approval process. It does not state what a successful call returns, but given the approval flow and the fact that the user is mimicking a dashboard action, the missing success response is a minor gap. Overall, the description is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all parameters are documented in the schema. The description adds value by explaining the approval_token flow ('call again with the token') and the overall process, which complements the schema's bare parameter descriptions. It does not introduce new parameter semantics beyond what the schema provides, but the flow context is useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Start a creator scan for an app' with the side effect of spending scan credits. It is specific about the resource and verb, but does not explicitly differentiate from the sibling start_quick_scan, though the 'type' parameter suggests a distinction. Still, the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the approval flow and dashboard gates, but gives no guidance on when to use this tool versus the sibling start_quick_scan. It mentions the general conditions (subscription, no concurrent scan) but does not point to alternatives or exclude cases. There is no 'when not to use' guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_program_nowSync affiliate program nowAInspect
Queue an immediate sync of an affiliate program connection, as the card's Sync now does: roster, signup matching and sales stats. It runs in the background; read get_program's last_synced_at a minute or two later.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| connection_id | Yes | A connection id from get_program. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false and destructiveHint=false, so the description's real added value is that the sync 'runs in the background' and is not immediately complete. That async/latency disclosure plus the verification step is genuinely beyond what annotations provide, though it does not describe failure modes or how long results take to appear beyond 'a minute or two'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: purpose and scope first, async behavior and verification second. Nothing is wasted and the most important action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and no return-value documentation, the description supplies the essential async model and how to check success, which is what an agent needs. The only notable gap is not mentioning that a flagged 'hard to undo' action may require confirm=true, though the schema covers that case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, including connection_id's provenance from get_program. The description adds no syntax or format detail beyond the schema, and the 'confirm' parameter's interaction with a destructive-ish guard is left entirely to the schema, giving the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Queue an immediate sync of an affiliate program connection') and names the exact data surfaces touched (roster, signup matching, sales stats). It is clearly distinguishable from the sibling read tool get_program, which is referenced as the verification path rather than as the actor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for this tool (a manual 'Sync now' trigger) and how to confirm the result via get_program's last_synced_at. It stops short of explicit when-not-to-use guidance, e.g. avoiding repeated calls while a sync is already in flight, so it does not reach 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
undo_opt_outUndo opt-outAInspect
Undo an opt-out the reply reading recorded on a thread (a misread, such as "remove me as the contact, write to partners@"). Only opt-outs this account's own reply reading listed are removed; a click on the unsubscribe link or another sender's opt-out stays. The thread goes back to replied with no more sequence steps, and the contact can be emailed again. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| enrollment_id | Yes | An enrollment_id from list_inbox_threads. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Far exceeds the annotations (readOnlyHint=false, destructiveHint=false, idempotentHint=false). It discloses concrete state changes (thread returns to replied, sequence steps stop, contact becomes emailable again), the guard behavior including the exact error payload shape, and the dashboard-equivalent guardrail limit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and the sentence that follows adds genuine constraints rather than filler. The parenthetical examples and guardrail sentence make it dense, but every sentence carries operative information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no output schema, it explains side effects, scoping limits, and the approval error response. The success return payload is not described, but the state consequences are, which is the more important gap to close here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented, and the description goes further by explaining the two-step approval_token contract (first call returns approval_required with a token, second call uses it after user confirmation). That adds real meaning beyond the format string.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with unusual precision: "Undo an opt-out the reply reading recorded on a thread," including the motivating case (a misread such as "remove me as the contact"). No sibling tool does anything comparable, so the agent can distinguish it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly bounds when the tool applies and when it does not: only opt-outs this account's reply reading listed are removed, while an unsubscribe-link click or another sender's opt-out stays. It also outlines the approval flow, though it names no alternative tool for the excluded cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlink_project_mailboxUnlink project mailboxAInspect
Remove the project's mailbox link. Nothing sends for the project, and its inbox threads cannot be answered, until a mailbox is set again. The Google account stays connected (disconnecting happens in the dashboard).
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds meaningful behavioral context beyond the annotations: the exact functional consequences of unlinking (outbound sending halts, inbox threads become unanswerable) and, crucially, that the Google account stays connected. This is consistent with destructiveHint=false since the unlink is reversible by setting a mailbox again — no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and followed by consequences and scope exclusion. No filler; every sentence adds information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description plus annotations cover the essentials: what happens, what does not happen, and where the excluded behavior lives. It does not discuss reversibility explicitly or the confirm guard's role, but the schema handles the latter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (app_id, confirm) are fully documented in the schema, including the app_id default behavior. The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove the project's mailbox link') and implicitly distinguishes itself from the sibling set_project_mailbox, its inverse. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Conveys the operational context (nothing sends, inbox threads cannot be answered until re-set) and carves out an explicit exclusion: disconnecting the Google account happens in the dashboard, not here. It stops short of stating prerequisites or when this is preferable to alternatives, but the scope boundary is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unsave_creatorUnsave creatorAIdempotentInspect
Remove a creator from your saved list.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| creator_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=true. The description adds that this removes from a saved list, which is a state-changing but non-destructive action. It doesn't add context about confirmation requirements or side effects beyond what the confirm parameter suggests.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, zero waste, front-loaded with the action. The description is minimal but complete for what it states.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple removal tool with annotations covering idempotency and non-destructiveness, the description is mostly adequate. However, it doesn't mention the confirm parameter's role or any prerequisites (e.g., creator must be saved), and there's no output schema. The confirm parameter hints at a guard flow that the description doesn't explain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: creator_id is described by name/format but not in the description text; confirm has a description in the schema. The tool description doesn't explain the parameters, but the schema covers confirm's purpose and creator_id's format. Baseline 3 is appropriate since the schema does partial heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Remove a creator from your saved list' clearly states the verb (remove), resource (creator), and target (saved list). It distinguishes from the sibling save_creator by being the inverse operation, though it doesn't explicitly name the sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: use when a user wants to unsave a creator. It doesn't explicitly state when not to use it or mention alternatives, but the inverse relationship with save_creator is clear from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_campaign_sequenceUpdate campaign sequenceAInspect
Edit one step of a campaign's email sequence (subject, body, delay), as the dashboard's email editor does; on a live campaign the change applies to every enrolled recipient's next unsent step. Follow-ups always reply in step 1's thread as Re: plus its subject, so subject is for step 1 only. Website copy must pass the check enrolment runs (no links, bare domains, stock phrases, unknown tokens or empty offer), else nothing is saved. Website copy personalises with tokens such as {{greeting}}, {{opener}}, {{tease_pitch}}, {{offer_pitch}} and {{signoff}}: get_campaign shows a campaign's templates and preview_campaign_step renders one for a site.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | The body template, with {{tokens}}. | |
| step | Yes | The step to edit, 1 for the opener. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| subject | No | Step 1 only. | |
| delay_days | No | Days after step 1 this step is due (step 1 goes at enrolment). | |
| campaign_id | Yes | A campaign id from list_campaigns. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the safety profile (readOnly false, idempotent false, destructive false). The description goes well beyond by disclosing the key side effect that on a live campaign the change applies to every enrolled recipient's next unsent step, the validation gate that must pass or nothing is saved, and the confirm-flag guard for hard-to-undo actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Core purpose is front-loaded in the first clause, followed by side effects, then token/validation detail. It is dense and slightly run-on across clauses, but nearly every sentence carries actionable constraint information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 6 parameters at full schema coverage, no output schema, and no nested objects, the description covers what is missing from structured data: live-campaign propagation, validation failure behavior, token syntax, and where to preview. An agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: subject is step-1-only, body personalises via named {{tokens}}, and website copy must pass specific checks (no links, bare domains, stock phrases, unknown tokens, empty offer). Those constraints are not in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (edit) and resource (one step of a campaign's email sequence) plus the fields it mutates (subject, body, delay). It is clearly separable from update_campaign_settings and rename_campaign by naming the step-level granularity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (editing a step's subject/body/delay, with an explicit note that subject applies to step 1 only) and routes to get_campaign and preview_campaign_step for tokens/preview. It does not explicitly exclude or contrast with update_campaign_settings for non-step edits, so it falls short of a full when/when-not statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_campaign_settingsUpdate campaign settingsAInspect
Change a DRAFT campaign's setup, as Finish setup saves it: name, offer, daily limit, sending days and hours, timezone and follow-up gaps. Fields you omit keep their value. Live campaigns refuse (not_a_draft): rename them with rename_campaign and edit copy with update_campaign_sequence. Nothing is sent.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Sending days. Default Mon to Fri. | |
| name | No | ||
| offer | No | The commission the emails promise, e.g. "30% recurring". | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| hour_to | No | Window end hour, exclusive. Default 17. | |
| timezone | No | IANA timezone, e.g. "Europe/London". Default America/New_York. | |
| hour_from | No | Window start hour, 24h, in timezone. Default 9. | |
| campaign_id | Yes | A campaign id from list_campaigns. | |
| daily_limit | No | Most emails this campaign sends a day. The mailbox's own daily cap still applies. | |
| followup_gaps_days | No | Day offset of each follow-up from step 1, e.g. [4,14,28,42]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavior beyond the annotations: partial-update semantics ('Fields you omit keep their value'), the not_a_draft error condition, and the side-effect disclaimer 'Nothing is sent.' Annotations only tell the agent this is a non-idempotent, non-destructive write. The one gap is that the confirm parameter's hard-to-undo guard is never explained in prose, so an agent may not know when that prompt appears.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what it edits, the patch semantics, then the exclusion and alternatives. The precondition and routing information is front-loaded and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation with no output schema, the description covers the essential decision inputs: draft restriction, partial update behavior, error path, and side-effect absence. Only the confirm-guard interaction and the shape of the success response are unaddressed, which is acceptable given the rich schema descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 90%, so the baseline is 3; the description adds genuine meaning by grouping the fields and clarifying that this is a patch-style update where omitted fields are preserved. It still says nothing about the confirm flag or the relationship between hour_from/hour_to defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Change) plus resource (a DRAFT campaign's setup) and enumerates the editable surface: name, offer, daily limit, sending days/hours, timezone and follow-up gaps. This distinguishes it cleanly from neighbors like update_campaign_sequence, rename_campaign and update_offer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition (DRAFT only), the failure mode for the wrong state (live campaigns refuse with not_a_draft), and names the two alternatives to use instead for live campaigns. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_dealUpdate dealAInspect
Set a deal's value, type, notes or follow-up date. Omitted fields stay; null clears value or follow-up date. Stage moves: move_deal_stage.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | Replaces the notes; empty string clears them. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| deal_id | Yes | A deal id from list_deals. | |
| deal_type | No | ||
| value_cents | No | Deal value in cents, e.g. 50000 for $500. | |
| follow_up_at | No | ISO date or datetime to follow up, e.g. 2026-10-15. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the mutation safety profile (readOnlyHint=false, non-idempotent, non-destructive). The description adds genuine behavior beyond that: partial-update semantics ('Omitted fields stay') and null-as-clear semantics. It omits permission/auth requirements and the meaning of the guard/confirm flow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses, no filler, with the mutation scope front-loaded before the clearing semantics and the sibling hand-off. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation tool with no output schema, it covers the important unknowns: partial update, clearing, and where stage changes go. Minor gaps remain around the confirm guard and what the call returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the schema does most of the work, but the description adds meaningful semantics the schema lacks: that value_cents and follow_up_at are cleared with null while other fields persist. It doesn't explain the confirm parameter or deal_type enum choices.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Set') plus the exact resource fields it mutates (value, type, notes, follow-up date) and explicitly routes stage changes to move_deal_stage, so the agent can separate it from that sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real usage conditions: omitted fields are preserved, null clears value/follow-up date, and stage moves belong to move_deal_stage. It does not spell out when to prefer this over create_deal, but the alternative routing is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_offerUpdate offerAInspect
Change any of the project's offer fields (App settings, Your offer); fields you leave out keep their value. Strings: empty clears. Numbers: null clears. Saved through the same validation as the dashboard. Applies to the next email. Returns the stored offer.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| assets_url | No | Banners and promo material, shared in replies. | |
| boost_days | No | How long the launch boost lasts. | |
| brand_name | No | How the brand is named in every email. | |
| commission | No | One clause, e.g. "30% recurring commission with a 45-day cookie". | |
| ladder_max | No | Highest rate (%); the last raise always offers it. | |
| nudge_days | No | Days after the first email that the offer follows, in the same thread. Default 4. | |
| signup_url | No | Affiliate signup link, sent in replies to sites that say yes, never in a first email. | |
| cookie_days | No | How long a referred click stays yours. | |
| ladder_step | No | Added (%) at each raise while a site does not reply, in the same thread. | |
| market_rate | No | A re-pitch may say other programs pay up to this (%), never naming one. | |
| sender_name | No | Signs the emails. A person, not a mailbox name like affiliate. | |
| boost_points | No | Launch boost (%) added for a new affiliate's first boost_days days. | |
| ladder_start | No | Rate (%) the first offer makes. null: the commission line as written. | |
| payout_notes | No | Payout timing, minimums and details the AI may quote in replies. | |
| boost_everyone | No | Every affiliate who joins from now on gets the launch boost, not only sites you pitched. The server records when it was first turned on. | |
| lead_with_rate | No | false: step 1 of the website lanes does not state the rate. | |
| postal_address | No | Business postal address, required before live cold email. | |
| placement_deals | No | Accept paid placements as a last resort: the sequence's last step asks what a top spot costs. | |
| raise_every_days | No | Days between raises, counted from the first email. Default 14. | |
| stats_window_hours | No | Hours of recent conversions the program figures cover. Default 72. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations declaring only readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=false, the description carries the real burden and delivers: it discloses partial-update semantics, the non-obvious clearing rules ('Strings: empty clears. Numbers: null clears'), dashboard-equivalent validation, that the effect lands on the next email, and that it returns the stored offer. The clearing behavior is a side effect that the annotations do not flag, and disclosing it is exactly the value-add expected here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each load-bearing: scope, partial-update rule, clearing rules, validation/timing, and return value. No filler, and the scope statement is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 22-parameter, zero-required update tool with no output schema, the description covers scope, patch semantics, clearing rules, validation parity, effect timing, and the return value. Remaining gaps are minor (permission/auth requirements, whether confirm is ever required for this tool, and any rate limits), but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds semantics the schema does not encode: empty string clears a string field and null clears a numeric field, and omitted fields are preserved. That distinction (omit vs. null vs. empty) is the main ambiguity across 22 optional parameters and the description resolves it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (change/update) and a precise resource (the project's offer fields, split into App settings and Your offer), and the partial-update semantics ('fields you leave out keep their value') separate it from a full-replacement tool. It is clearly distinguishable from siblings like update_project or update_campaign_settings, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied rather than stated: the description says changes apply to the next email and are saved through the same validation as the dashboard, which tells the agent this is a configuration edit that takes effect downstream. It never states when to use this versus update_project / update_outreach_settings, nor any prerequisites such as permissions or the confirm flag.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_outreach_settingsUpdate sending settingsAInspect
Sending limits and defaults (App settings, Sending); fields you leave out keep their value. The daily cap and reply reserve belong to the mailbox, so every project sending from it gets them. Values are clamped as in the dashboard; the response has what now applies. Window, timezone and days are the defaults every new campaign starts with. Guarded, since the cap and the paid finder budget change how much mail goes out and what it costs. Guarded: without approval_token it returns {error:"approval_required", approval_token, summary} for the user to confirm; call again with the token. Every dashboard guardrail still applies: this is exactly what a user clicking the dashboard could do, no more.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| hour_to | No | ||
| timezone | No | ||
| daily_cap | No | Emails a day from the mailbox, replies included, up to the platform ceiling. null: the default. | |
| hour_from | No | ||
| sending_days | No | Days campaigns go out on, e.g. ["Mon","Tue","Wed"]. Replaces the whole set. | |
| reply_reserve | No | Part of the daily cap campaigns leave for replies, chases and partner emails; kept below the cap. null: the default. | |
| approval_token | No | Approval token from a previous approval_required response, after the user said yes. | |
| finder_budget_usd | No | USD a month for the paid email finder on sites the free scraper left empty, at or above the autopilot fit floor. 0 keeps it off. | |
| contact_checks_per_day | No | Website prospects the morning sweep checks for an email, best fit first. 0 skips it. null: the default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without any structured help beyond readOnlyHint=false and destructiveHint=false, the description discloses the approval guard, the exact approval_required error payload and how to retry with the token, value clamping behavior, that omitted fields retain their value, and that the response reports what now applies. It also frames the operation's blast radius (mail volume and paid finder spend) and dashboard parity. This is unusually rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the field semantics and scoping before the guard details, and every clause carries information. However, "Guarded" is used twice in adjacent sentences and the closing dashboard-parity sentence is closer to reassurance than instruction, so it is slightly inflated for its content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 all-optional parameters, no output schema and a non-trivial approval flow, the description covers the mechanics an agent needs: partial updates, mailbox vs project scope, clamping, and the approval round-trip. The return value is only described vaguely as "what now applies", which is the one gap for a tool whose success response matters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 70%, and the description adds genuinely new meaning on top: partial-update semantics ("fields you leave out keep their value"), cross-parameter scoping (cap and reply reserve are mailbox-level, not project-level), and the fact that hour_from/hour_to/timezone/sending_days seed new campaigns. It does not explain the individual enum parameters or the per-field behavior of contact_checks_per_day, so it stops short of 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and scope: sending limits and defaults, located as (App settings, Sending), covering the daily cap, reply reserve, window/timezone/days defaults, finder budget and contact checks. It also distinguishes mailbox-level fields from per-campaign defaults, which separates it from update_campaign_settings and update_project_settings. It never states the action as a clean verb first, so it lands just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear routing context: the daily cap and reply reserve belong to the mailbox and therefore affect every project sending from it, while window, timezone and days are the defaults every new campaign starts with. That implicitly tells the agent when to use this rather than a campaign-scoped updater. No sibling tool is named explicitly as an alternative, so 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_projectUpdate projectAInspect
Product summary, market, marketing domain, Google Play link, re-scan cadence and notification emails; fields you leave out keep their value. A changed summary re-grades every graded creator against it (no scan, nothing spent), as saving it in the dashboard does; the response says how many. The market applies from the next scan.
| Name | Required | Description | Default |
|---|---|---|---|
| app_id | No | Project id from list_apps. Omit to use your most recently added project. | |
| market | No | The country every creator and web search runs in. | |
| rescan | No | Automated re-scans (Sundays, early UTC). | |
| confirm | No | Set true after the user confirmed an action the guard flagged as hard to undo. | |
| summary | No | What the product does and who it is for. | |
| digest_day | No | 0 = Sunday. | |
| digest_hour | No | Hour of the day in digest_timezone. | |
| weekly_digest | No | The weekly digest email. | |
| play_store_url | No | Google Play listing. Empty clears it. | |
| digest_timezone | No | ||
| marketing_domain | No | The product's own website, like yourapp.com; live email goes out from an address on it. null unlinks. | |
| scan_results_email | No | One email when a full scan finishes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, idempotentHint=false), it discloses a genuine side effect: changing the summary re-grades every graded creator at no cost, matching dashboard behavior, and 'the response says how many' describes the return payload. The market-timing note ('applies from the next scan') is also useful. It stops short of auth/permission or rate-limit context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the affected fields followed by the impact notes; no filler. The semicolon-separated field list is a little heavy to parse but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter, zero-required mutation tool with no output schema, the description supplies the missing behavioral context (partial update, re-grade side effect, response hint, market timing). It leaves some smaller cadence/email fields to the schema, which is acceptable given 92% coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 92%, so the baseline is 3, but the description adds real semantics: partial-update behavior, the delayed effect of market, and the downstream re-grade triggered by summary. It does not cover the one undescribed schema field (digest_timezone), which holds it below 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates exactly which project settings are mutable (summary, market, marketing domain, Play link, re-scan cadence, notification emails), which clearly separates it from siblings like update_offer or update_campaign_settings. The verb is only implied by the field list rather than stated, but an agent can still tell it edits project settings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Fields you leave out keep their value' communicates partial-update usage, and the market/summary timing notes imply when the change takes effect. However, no alternative or sibling (e.g. get_project_settings for reading, list_apps for the id) is named, and there is no explicit when-to-use statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
set_autopilot_mode1 field changed- changed
Input schema / properties / lane / enumPrevious value: -[ - "websites", - "featured", - "link_exchange", - "replies", - "nudges", - "cross_sell", - "retouch", - "chase", - "proof_line" -]New value: +[ + "websites", + "featured", + "link_exchange", + "replies", + "nudges", + "cross_sell", + "retouch", + "chase", + "creators", + "proof_line" +]
59 tool updates
- Added
add_competitor - Added
add_creators_to_campaign - Added
add_keyword - Added
add_seed_address - Added
add_suggested_competitor - Added
add_websites_to_campaign - Changed
approve_autopilot_action2 fields changed- added
Input schema / properties / bodyAdded value: +{ + "description": "Your edited text for a drafted email, sent instead of the draft.", + "maxLength": 5000, + "minLength": 1, + "type": "string" +} - added
Input schema / properties / remove_idsAdded value: +{ + "description": "Sites to drop from a website batch (ids from get_autopilot_action).", + "items": { + "format": "uuid", + "pattern": "^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$", + "type": "string" + }, + "maxItems": 200, + "type": "array" +}
- Added
check_draft - Added
create_deal - Added
create_website_campaign_draft - Added
delete_campaign - Added
dismiss_competitor_suggestion - Added
duplicate_campaign - Added
find_website_contact - Added
get_autopilot_action - Added
get_competitor_suggestions - Added
get_deliverability - Added
get_mailbox_connect_link - Added
get_mailboxes - Added
get_offer_benchmark - Added
get_overview - Added
get_partners - Added
get_program - Added
get_project_settings - Added
get_seed_connect_link - Added
get_website - Added
hold_autopilot_action - Changed
list_competitors1 field changed- added
Input schema / properties / app_idAdded value: +{ + "description": "Project id from list_apps. Omit to use your most recently added project.", + "format": "uuid", + "pattern": "^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$", + "type": "string" +}
- Changed
list_keywords1 field changed- added
Input schema / properties / app_idAdded value: +{ + "description": "Project id from list_apps. Omit to use your most recently added project.", + "format": "uuid", + "pattern": "^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$", + "type": "string" +}
- Changed
list_websites1 field changed- added
Input schema / properties / app_idAdded value: +{ + "description": "Project id from list_apps. Omit to use your most recently added project.", + "format": "uuid", + "pattern": "^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$", + "type": "string" +}
- Added
mark_thread_read - Added
preview_campaign_step - Added
recheck_from_domain - Added
refresh_inbox_test_results - Added
remove_competitor - Added
remove_keyword - Added
remove_seed - Added
rename_campaign - Added
restore_competitor - Added
run_inbox_test - Added
set_autopilot_floor - Added
set_autopilot_mode - Added
set_from_alias - Added
set_mailbox_established - Added
set_outreach_domain - Added
set_project_mailbox - Added
set_signup_approval_mode - Added
set_website_contact - Added
set_website_note - Added
set_website_status - Added
sync_program_now - Added
undo_opt_out - Added
unlink_project_mailbox - Added
update_campaign_sequence - Added
update_campaign_settings - Added
update_deal - Added
update_offer - Added
update_outreach_settings - Added
update_project
1 tool update
- Changed
list_websites1 field changed- added
Input schema / properties / include_flaggedAdded value: +{ + "description": "Also return sites flagged out of outreach (competitor, not a publisher, unsafe, needs review).", + "type": "boolean" +}
2 tool updates
- Changed
get_autopilot1 field changed- added
Input schema / properties / app_idAdded value: +{ + "description": "Project id. Omit for the project autopilot runs for.", + "format": "uuid", + "pattern": "^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$", + "type": "string" +}
- Changed
set_autopilot2 fields changed- added
Input schema / properties / app_idAdded value: +{ + "description": "Project id from list_apps. Omit for the project autopilot runs for; required while it is bound to none.", + "format": "uuid", + "pattern": "^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$", + "type": "string" +} - changed
Input schema / properties / offer / maxLengthPrevious value: -120New value: +200
1 tool update
- Changed
move_deal_stage1 field changed- changed
Input schema / properties / stage / enumPrevious value: -[ - "contacted", - "negotiating", - "contract", - "live", - "paid" -]New value: +[ + "contacted", + "negotiating", + "live", + "paid" +]
8 tool updates
- Changed
create_campaign_draft1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
- Changed
get_checkout_link1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
- Changed
move_deal_stage1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
- Changed
reveal_contact1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
- Changed
save_creator1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
- Changed
set_creator_note1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
- Changed
start_quick_scan1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
- Changed
unsave_creator1 field changed- added
Input schema / properties / confirmAdded value: +{ + "description": "Set true after the user confirmed an action the guard flagged as hard to undo.", + "type": "boolean" +}
1 tool update
- Changed
get_quick_scan_preview1 field changed- added
Input schema / properties / wait_secondsAdded value: +{ + "description": "How long to wait for the scan to finish before returning. Default 45.", + "maximum": 50, + "minimum": 0, + "type": "integer" +}
32 tool updates
- First observed
approve_autopilot_action - First observed
create_campaign_draft - First observed
export_creators_csv - First observed
get_account - First observed
get_autopilot - First observed
get_campaign - First observed
get_checkout_link - First observed
get_creator - First observed
get_quick_scan_preview - First observed
get_scan_status - First observed
get_thread - First observed
launch_campaign - First observed
list_apps - First observed
list_campaigns - First observed
list_competitors - First observed
list_creators - First observed
list_deals - First observed
list_inbox_threads - First observed
list_keywords - First observed
list_websites - First observed
move_deal_stage - First observed
pause_campaign - First observed
reject_autopilot_action - First observed
resume_campaign - First observed
reveal_contact - First observed
save_creator - First observed
send_reply - First observed
set_autopilot - First observed
set_creator_note - First observed
start_quick_scan - First observed
start_scan - First observed
unsave_creator
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables brand visibility monitoring across major AI platforms like ChatGPT, Claude, Gemini, and Perplexity. It allows users to track visibility scores, analyze competitor data, and receive actionable insights to improve AI-generated brand recommendations.1624 npm1MIT
- AlicenseCqualityBmaintenanceCompetitor Monitor AI - MCP server providing AI-powered tools and automation by MEOK AI Labs1114 npm40 PyPIMIT
- AlicenseAqualityCmaintenanceRevnuvo Company Intelligence tells AI agents what changed at a company, with evidence. It observes company websites, technologies, and DNS over time and returns timestamped, confidence-aware changes, signals, and monitoring.9MIT

industrylens-mcpofficial
AlicenseNot gradedqualityBmaintenanceBrowse IndustryLens's published competitive-intelligence reports and head-to-head competitor comparisons from any AI agent — real, source-backed data.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.