Emailchaser
Server Details
Run cold email outbound from an AI agent: campaigns, leads, replies, sender accounts, autopilot.
- Status
- Healthy
- Uptime
- 92.0% over 23 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 101 tools
Several tools overlap in purpose: add_leads, add_lead_finder_prospects, and source_prospects all add prospects/leads to campaigns; copilot_launch and create_campaign both create campaigns; update_lead_category can duplicate mark_lead_meeting. The detailed descriptions help, but with 101 tools the risk of misselection is real.
Most tools follow a clear snake_case verb_noun pattern (list_campaigns, get_campaign, update_campaign, delete_campaign). Minor deviations like copilot_launch and copilot_plan use noun-first ordering, but the set is otherwise predictable.
101 tools is an extreme mismatch for a single MCP server, far beyond the typical 3-15 range and the 50+ threshold for a score of 1. The surface is enormous and difficult for an agent to navigate effectively.
The tool surface covers the full cold-email lifecycle: campaigns, sequences, leads, senders, verification, autopilot, ICPs, blocklists, webhooks, inbox placement, billing, credits, workspaces, and done-for-you orders. No obvious gaps or dead ends are apparent for the stated domain.
Available Tools
101 toolsadd_blocklist_entriesAdd blocklist entriesAIdempotentInspect
Adds up to 1,000 suppression entries in one call so no campaign ever emails them. Each value is either a full email address (contains @) or a bare domain; the type is inferred per value. Existing entries and invalid values are skipped, so the call is safe to retry and suited to carrying over a suppression list from another sending platform. Scope defaults to "workspace"; "global" blocks across every workspace on the account and only works from the main workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | "workspace" (default) blocks for this workspace only; "global" blocks for every workspace on the account | |
| values | Yes | Email addresses and/or domains to block, e.g. ["competitor.com", "jane@acme.com"] (1 to 1,000 per request) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover idempotency, non-destructiveness and closed-world scope, and the description adds genuinely new behavior: skipped existing/invalid values, per-value type inference, and the main-workspace-only constraint on global scope. It does not say what the call returns (counts of added vs skipped), which is the main remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and batch ceiling, then layers constraints in a compact block with no filler. The final sentences are slightly dense and chain multiple clauses, but every sentence carries load-bearing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter write tool with no output schema, the description covers input rules, scope semantics, skip behavior and retry safety well. The one omission is the shape of the response (e.g., how many were added vs skipped), which an agent would want to know for verification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: how each value's type is inferred from the presence of "@", and that scope defaults to workspace while global requires the main workspace. This exceeds what the enum description conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (adds) and resource (suppression/blocklist entries) with the batch size ceiling, and is instantly distinguishable from siblings like remove_blocklist_entries, get_blocklist_entry and add_leads. The per-value type inference rule further pins down exactly what the tool accepts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete use case (carrying over a suppression list from another sending platform) and a retry-safety condition, plus the scope default and the restriction that "global" only works from the main workspace. It stops short of naming the sibling tools (remove_blocklist_entries, update_blocklist_entry) an agent might confuse it with, so not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_lead_finder_prospectsLead Finder: add people to a campaign (spends credits)AInspect
SPENDS CREDITS, AND BY DEFAULT METERED EMAIL VERIFICATION. Adds people from the Lead Finder contact database to a campaign as leads, revealing their contact details, in the background. Pass refs (from search_lead_finder results; up to 1,000 different people by default) to add exactly those people, or pass filters and count (1 to 5,000 by default) to add count people matching the filters who are not in the workspace yet. A filters add always walks the matches from the top and skips anyone already a lead, free and not counted, so repeating the same add adds the NEXT count people and charges again. Never repeat an add to retry: after an error or a timeout, call list_lead_finder_imports to see whether it started. A filters add stops early when the audience runs out or once it has looked at five people for every one requested (at least 1,000). People already added count toward that limit even though skipping them is free, so after an add of more than about 1,000 people a small follow-up can stop with nobody added. Every person a filters add looks at, added or skipped, also uses the account's daily allowance for filters adds (25,000 people by default). Each person added costs 1 credit ($0.033 at list), reported as creditsPerProspect; the response reports estimatedCredits, the most the add can cost in credits, and the wallet must hold that much for it to start. People already in the workspace, blocklisted, without a usable address, on a personal mailbox when the campaign only takes business addresses, or marked invalid by verification are skipped and cost no credits. verifyEmails (on by default) checks every screened address with the paid email verification waterfall, including the ones it marks invalid: one or two checks per address, billed as metered usage on the next invoice (price per check: GET /email-verification/rates in the REST API), within the account's monthly verification spend limit. Catch-all and unconfirmed addresses are added and charged, and so is every address once that limit is reached; verification.limitReached and verification.unknown on get_lead_finder_import count them. Leads added to a running or paused campaign get their emails at once (a paused campaign sends them when resumed); adding to a completed campaign resumes it, so it starts sending again; in a draft campaign the emails are created at launch. At most 3 adds run at once in a workspace. A refusal names its reason, and a 429 also says when to retry (reason running_imports, daily_import_limit, daily_budget, monthly_budget or fair_use_floor). Poll get_lead_finder_import with the importId until finishedAt is set. Check get_credit_balance first; buy credits with purchase_credits.
| Name | Required | Description | Default |
|---|---|---|---|
| refs | No | Refs of the people to add, from search_lead_finder results. Duplicates and blanks are dropped before sending; up to 1,000 different people per add by default (get_lead_finder_filters has the limit that applies). Use either refs, or filters with count | |
| count | No | How many people not yet in the workspace to add with filters (1 to 5,000 by default; get_lead_finder_filters has the limit that applies) | |
| filters | No | Add people matching these filters (same shape as search_lead_finder). Needs count. The matches are walked from the top every time and people already in the workspace are skipped, so a repeat adds the next count people | |
| campaignId | Yes | The numeric campaign ID (from list_campaigns) | |
| verifyEmails | No | Check each address with the paid email verification waterfall before adding it (default true): one or two checks per address, billed on the next invoice, including addresses that turn out invalid. Invalid addresses are skipped and cost no credits; catch-all and unconfirmed ones are still added and charged, and so is every address once the monthly verification spend limit is reached. false skips verification |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses far more than annotations: costs per person, metered email verification, non-idempotent repeat behavior ('repeating the same add adds the NEXT count people and charges again'), concurrency limits, and response fields like estimatedCredits and importId. It also explains skip conditions and edge cases such as catch-all addresses.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely long and dense, covering many edge cases without visual structure (no headings or bullets). While each sentence adds value, the sheer length makes it harder for an agent to parse quickly. Some details, such as verifyEmails behavior, are repeated from the schema, and the opening warning is front-loaded but the rest is a wall of text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully covers return values, errors, prerequisites, and follow-up actions. It names related tools for polling (get_lead_finder_import), filter limits (get_lead_finder_filters), and credit purchases (purchase_credits), making it complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds critical behavioral semantics beyond the schema: mutual exclusivity of refs and filters, how repeat filters adds walk from the top and skip already-leads, and verification details for verifyEmails. It also clarifies that duplicates are dropped, which is not fully explicit in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Adds people from the Lead Finder contact database to a campaign as leads, revealing their contact details, in the background.' It clearly differentiates this from sibling tools like add_leads by tying it to Lead Finder and credits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly describes two usage modes and when to use each: 'Pass refs ... to add exactly those people, or pass filters and count ... to add count people matching the filters.' It also provides strong anti-retry guidance ('Never repeat an add to retry: after an error or a timeout, call list_lead_finder_imports') and prerequisites ('Check get_credit_balance first').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_leadsAdd leadsAIdempotentInspect
Creates or updates up to 1,000 leads in one request, optionally adding them to a campaign via campaignId. Existing leads (matched by email) are updated instead of duplicated. Returns the id of every lead written, so they can be fetched or updated afterwards. Attributes outside the named fields can be passed in customVariables and become merge tags in email copy.
| Name | Required | Description | Default |
|---|---|---|---|
| leads | Yes | The leads to create or update (1 to 1,000 per request) | |
| campaignId | No | Campaign to add the leads to (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered; the description adds substantive behavior beyond that: dedup-by-email upsert semantics, the return payload ('the id of every lead written'), and how customVariables become merge tags. It doesn't discuss permissions, rate limits, or partial-failure behavior for a 1,000-item batch, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four front-loaded sentences, each carrying distinct information (capacity, upsert rule, return value, custom attributes) with no filler. The closing customVariables sentence mildly overlaps with the schema, so not a perfect 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by stating the return value (lead ids for subsequent fetch/update). Batch limits, dedup behavior, optional campaign attachment, and custom attribute handling are all present, leaving no material gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents email, the named fields, customVariables, and campaignId. The description restates the customVariables merge-tag concept and explains campaignId's purpose, but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates or updates ... leads') with scope ('up to 1,000 in one request') and an optional side effect ('adding them to a campaign via campaignId'). This cleanly separates it from single-record siblings like update_lead and read-only siblings like list_leads/get_lead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the upsert rule clearly ('Existing leads (matched by email) are updated instead of duplicated'), which tells the agent when this tool is the right choice for existing contacts, and notes campaignId is optional. It stops short of explicitly naming alternatives (update_lead for one record, list_leads to discover leads), so it's clear context without explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
approve_autopilot_runAutopilot: approve a runAInspect
Approves an autopilot run that is awaiting approval — the one human gate in the flow. CAN SPEND REAL MONEY: approval first reveals the run's prospects at 1 credit ($0.033 at list) each, then, if the workspace has no connected sending account, places one done-for-you order sized from the plan snapshotted at start, at most 10 .com domains and at most six mailboxes per domain, never more mailboxes than the plan, charged at the prices in effect when the order is placed. So the plan's mailbox and domain counts are a ceiling and its dollar figures an estimate, not an exact quote. No credits are bought, but the reveal and later top-ups spend the wallet's credits. If a sending account is already connected, no infrastructure is bought; a run started without budgetUsd and with no connected account fails. Show the user the plan from the start_autopilot_run response before calling this. Nothing is revealed, bought or sent before this call; after it, sending starts without further confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric autopilot run ID (from list_autopilot_runs) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (openWorldHint, destructiveHint) and do not mention the real-money impact. The description fully compensates by detailing financial implications: 'CAN SPEND REAL MONEY,' cost per credit, order placement conditions, and limits on domains/mailboxes. It also specifies that nothing happens before the call and that sending starts after. This is exemplary transparency beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but every sentence carries critical information: purpose, financial risk, conditions, and prerequisites. It front-loads the core purpose and then layers constraints. While not terse, the density is justified given the high-stakes nature of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (monetary transactions, conditional behavior, caps), the description covers all necessary details: what happens, what fails, the need to show the plan first, and the irreversible nature after the call. There is no output schema, but the description does not need to explain return values for a confirmation action. It is complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'id' is fully documented in the schema with description and constraints. The tool description does not add new semantics about the parameter, but the schema covers it 100%. Since the schema does the heavy lifting, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb, resource, and context: 'approves an autopilot run that is awaiting approval — the one human gate in the flow.' It clearly distinguishes this from siblings like start, pause, resume, and kill by focusing on the approval action. The purpose is immediately clear without confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: it is the human-gate step, and it instructs to 'Show the user the plan from the start_autopilot_run response before calling this.' It also clarifies the consequences (sending starts without further confirmation). It does not explicitly name alternatives but implies its unique role in the flow, making it distinguishable from related tools like kill_autopilot_run.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_campaign_sender_emailsAttach sender emails to campaignAIdempotentInspect
Attaches connected sender email accounts to a campaign. A campaign only sends through the senders attached to it, so this call decides which mailboxes it uses. Idempotent: already-attached senders are left untouched and only missing ones are added. The call is refused whole when any requested sender is not a connected sender email in the workspace. Attaching to a running campaign reschedules it so the new senders are picked up.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) | |
| senderEmailIds | Yes | Sender email IDs to attach (from list_sender_emails). Every ID must be a connected sender email in the workspace |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (idempotentHint, destructiveHint=false), the description discloses three substantive traits: the idempotent merge behavior (existing senders untouched, only missing ones added), the all-or-nothing refusal when any sender is invalid, and the side effect of rescheduling a running campaign. These are exactly the behaviors an agent needs before invoking a mutating attachment.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action, then escalating to idempotency, atomicity, and side effects. No filler, and each sentence carries distinct decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with annotations covering the safety profile and no output schema, the description covers purpose, idempotency, failure mode, and side effects. An agent has everything required to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented there, so the baseline is 3. The description reinforces the constraint that every sender must be a connected workspace sender, but that validation rule is already in the schema; the added value (whole-call refusal) is more behavioral than parameter-specific.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Attaches connected sender email accounts to a campaign') and immediately clarifies the effect ('this call decides which mailboxes it uses'). It is trivially distinguishable from the sibling detach_campaign_sender_email by direction of the operation. Nothing about the purpose is left ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operational context ('a campaign only sends through the senders attached to it') that tells the agent why and when this call matters. It does not, however, explicitly name alternatives or state when-not to use it (e.g., how it relates to connect_sender_email or create_campaign), so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_email_verification_jobCancel an email verification jobAIdempotentInspect
Stops a running verification job. Addresses already verified stay in the results and stay billed; addresses that never ran are never charged, so cancelling costs nothing further.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | The job ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotent and non-destructive hints. The description adds valuable behavioral context about billing and result retention: already-verified addresses remain billed and in results, while unprocessed ones are never charged. This goes beyond what annotations convey and helps the agent understand side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. The primary purpose is front-loaded, and the second sentence concisely explains the billing and results consequences. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema and annotations covering idempotency and non-destructiveness, the description provides the necessary purpose and side effects. It does not mention response format, but that is not critical for a cancellation action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the only parameter (jobId) with a minimal description 'The job ID', so baseline is 3. The tool description adds no extra information about the parameter itself, but the parameter is self-explanatory given the tool name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Stops') and resource ('a running verification job'), clearly distinguishing it from creation, listing, or retrieval of verification jobs. The billing note further clarifies the scope of action, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for stopping a running verification job, which is clear context for when to use it. It does not explicitly mention alternatives or exclusions, but the sibling context and action verb make the use case obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_lead_finder_importLead Finder: stop an addADestructiveIdempotentInspect
Stops a running Lead Finder add. It reads canceled at once. A batch the worker is already writing is still added and charged; a batch still being screened or verified is dropped, though verification checks already made are billed. Credits held for the rest are released. Until finishedAt is set, added, creditsSpent and even status can still change: if that batch was the add's last, it ends completed or failed instead. People already added stay in the campaign. A stopped add cannot be resumed: start a new add_lead_finder_prospects instead. Stopping one that already finished changes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The import ID (importId from add_lead_finder_prospects or list_lead_finder_imports) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though annotations already mark this as destructive and idempotent, the description adds crucial behavioral detail: written batches are still added and charged, screened/verified batches may still incur billing, credits are released, people already added stay, and status fields can change until finishedAt is set. This goes well beyond the annotations and warns about real side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds necessary operational nuance: cancellation timing, billing implications, credit handling, status uncertainty, permanence, and the no-op case. It is front-loaded with the core action and then layers detail in a logical order without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an asynchronous destructive operation with billing side effects and no output schema, the description covers all decision-relevant behavior: what is dropped, what is charged, what happens to credits, eventual status changes, permanence, and alternative next steps. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, id, is fully documented in the schema with 100% coverage, including its source (importId from add_lead_finder_prospects or list_lead_finder_imports). The description does not need to repeat parameter semantics, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: "Stops a running Lead Finder add." It is clearly distinct from sibling tools like add_lead_finder_prospects (which starts imports) and get/list_lead_finder_imports (which monitor them), so an agent can tell exactly what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use context (only a running add can meaningfully be stopped), a when-not-to-use signal ("Stopping one that already finished changes nothing"), and an alternative ("start a new add_lead_finder_prospects instead" since a stopped add cannot be resumed). This fully routes the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_connectionConfirm the connectionAIdempotentInspect
Confirms the API key works and tells Emailchaser which AI assistant connected. Call it as your first action after connecting, and whenever the user asks you to confirm or check the Emailchaser connection: it records the assistant and time on the key, which the app's API & MCP page watches to show the user you are in, and completes the 'Connect your AI' onboarding step. Idempotent: repeat calls simply refresh the record. Works with read-only keys as well as read & write keys.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | No | Your name in any form, e.g. claude, chatgpt, cursor, gemini. Normalized to a short lowercase identifier and echoed back |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover idempotency and non-destructiveness, but the description adds real behavioral context: it writes the assistant name and timestamp onto the key, that record is what the app's API & MCP page surfaces to the user, it completes an onboarding step, and it works with read-only keys. That permission and side-effect disclosure goes well beyond the annotation set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core effect, then trigger conditions, then side effects and the idempotency caveat. It runs slightly long and the idempotency clause partially duplicates the annotation, but every sentence carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter, no-output-schema tool, the description covers purpose, timing, side effects and permission compatibility. It does not describe the shape of the confirmation response (e.g. what 'confirms the API key works' returns on failure), which is the only meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'agent' parameter is thoroughly documented in the schema (free-form name, normalization, echo behavior). The description does not mention the parameter at all, so it adds no value over the structured field; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and effect: confirms the API key works and registers which AI assistant connected. No sibling tool does anything comparable, so the agent can distinguish it immediately from the read/write tools in the list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers: 'call it as your first action after connecting' and 'whenever the user asks you to confirm or check the Emailchaser connection'. It also defines the state it completes ('the Connect your AI onboarding step'), removing any ambiguity about timing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connect_sender_emailConnect sender emailAInspect
Connects an SMTP/IMAP mailbox to the workspace using an app password, with no browser step. Google and Microsoft mailboxes are NOT supported here because both require an interactive consent screen that cannot be automated; connect those in the app. The address must be on a business domain, so gmail.com and similar are rejected. The returned account's healthScore is null until warm-up has checked 20 of its emails (see get_sender_email).
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | The mailbox address, on a business domain | ||
| lastName | No | Sender last name | |
| password | Yes | App password for the mailbox, not the account login password | |
| firstName | No | Sender first name | |
| loginString | No | Login username if it differs from the email address | |
| imapServerUrl | Yes | IMAP host, optionally with a port, e.g. imap.fastmail.com:993 | |
| smtpServerUrl | Yes | SMTP host, optionally with a port, e.g. smtp.fastmail.com:587 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint and destructiveHint=false, and the description adds substantial context: the app-password requirement, provider exclusions, domain rejection behavior, and that the returned account's healthScore stays null until warm-up scans 20 emails. It stops short of covering auth prerequisites for the API or failure/error modes, so a 4 rather than a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four front-loaded sentences that each carry a distinct constraint (mechanism, provider exclusion, domain rule, return semantics). No filler, though it is denser than strictly necessary for a short description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter creation tool with no output schema, the description covers the key behavioral surprises: eligibility rules, the non-automatable providers, and the null healthScore return value. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters including app-password and business-domain semantics are already documented in the schema. The description reinforces the business-domain constraint but adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (connects) and resource (SMTP/IMAP mailbox to workspace) with the mechanism (app password, no browser). It explicitly distinguishes itself from siblings by naming get_sender_email and contrasting with the in-app flow for Google/Microsoft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-not-to-use (Google and Microsoft mailboxes require interactive consent and must be connected in the app) and a hard eligibility rule (business domain only, gmail.com rejected). The agent can route correctly without further inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
copilot_launchCopilot: launch a draft campaignAInspect
Creates a DRAFT campaign from a sequence. With a LinkedIn Sales Navigator search it also starts finding + verifying leads for it; without one the draft holds the emails and leads can be added from Lead Finder or a CSV upload. Nothing sends — the campaign stays a draft until you review and launch it. Returns the new campaign id, plus a lead-finding job id when lead finding started.
| Name | Required | Description | Default |
|---|---|---|---|
| icp | No | Optional ICP object as returned by copilot_plan, stored with the campaign for reference only | |
| name | Yes | Campaign name | |
| sequence | Yes | The email sequence: the first step is the initial email, the rest are follow-ups | |
| salesNavData | No | Optional base64 LinkedIn session payload from the browser extension; required for lead-finding to actually run | |
| salesNavSearchUrl | No | Optional LinkedIn Sales Navigator people-search URL to source leads from. Leave it out to create the draft with its emails only. Required when salesNavData is given. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (destructiveHint=false, idempotentHint=false, openWorldHint=true), so the bar is lower. The description adds real behavior beyond that: nothing is sent, the campaign remains a draft until reviewed, and it returns a campaign id plus a lead-finding job id when finding starts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences, front-loaded with the core action and constraint ('Nothing sends'). Every sentence adds information: mode branches, the safety guarantee, and return values.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description documents the return shape (campaign id, optional lead-finding job id). Combined with 100% schema coverage of the nested sequence object and the annotation safety profile, an agent has what it needs to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes further by explaining the conditional relationship between the LinkedIn search and lead-finding: salesNavSearchUrl's presence/absence changes behavior, and lead finding only runs when the session payload is supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Creates a DRAFT campaign from a sequence') and immediately scopes what it does not do ('Nothing sends — the campaign stays a draft'). This disambiguates it from the sibling launch_campaign, which presumably does send, despite the 'copilot_launch' name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains the two operating modes: with a Sales Navigator search it starts lead finding; without one the draft holds only emails, and leads come from Lead Finder or CSV. It does not explicitly say when to prefer this over sibling create_campaign or when to call copilot_plan first, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
copilot_planCopilot: plan a campaignAInspect
Given a company website, the honest AI SDR builds an ideal-customer-profile (ICP) and a suggested cold-email sequence. Read-only — creates nothing. Returns the ICP (personas, titles, industries, suggested Sales Navigator keywords) and a draft sequence you can review or pass to copilot_launch.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | Optional extra context, e.g. 'we sell to dental clinics in the US' | |
| website | Yes | The company website to analyze, e.g. https://acme.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover openWorldHint and destructiveHint, but the description adds behavioral context: 'Read-only — creates nothing' clarifies persistence semantics that readOnlyHint alone might not convey, and it discloses the return payload (ICP fields plus a draft sequence). It stops short of stating latency or cost, which is fine given the annotation baseline is met.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the input and output, then a behavioral note and a routing pointer. Every sentence carries information an agent needs; nothing is repetitive with the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param planning tool with annotations covering safety and no output schema, the description specifies the input, the produced artifacts, and the downstream handoff. What's missing is whether the output is persisted or ephemeral and whether re-running overwrites anything, but overall an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'website' and 'context' are documented in the schema with examples. The description adds no syntax or format details beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('builds an ideal-customer-profile and a suggested cold-email sequence') with the input that drives it ('given a company website'). It names the sibling consumers ('pass to copilot_launch'), so it is distinguishable from create_icp, plan_autopilot, and copilot_launch despite those siblings' similar names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies when to use it — planning before launch — and names copilot_launch as the downstream step, giving context. It doesn't explicitly state exclusions (e.g., 'use create_icp instead if you want a persisted ICP'), but the planning-to-launch flow is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_campaignCreate campaignAInspect
Creates a campaign in the workspace. The campaign starts as a DRAFT and sends nothing. A new campaign has no emails in it, so it cannot be launched until a sequence is written with replace_campaign_sequence and a sender email is attached. Sending limits and deliverability settings can be set here or changed later with update_campaign.
| Name | Required | Description | Default |
|---|---|---|---|
| flow | No | Campaign flow. Defaults to multiple_leads_scheduled, which behaves exactly like a campaign created in the app. Only pass api if you specifically want lead validation skipped at launch | |
| name | Yes | Campaign name | |
| emoji | No | Campaign emoji shown in the app. Defaults to a generic one | |
| timezone | No | IANA timezone, e.g. America/New_York | |
| dailyLimit | No | Cap on the total emails (initial + follow-ups) the campaign may schedule per calendar day in its timezone. Omit for no campaign-level cap; per-mailbox limits still apply | |
| isEnabledLlm | No | Enable AI (LLM) features for this campaign | |
| minimumHealthScore | No | Minimum email account health score for this campaign, 1-100 (see healthScore in list_sender_emails). An account below it, or with no score yet, sends nothing in this campaign, first emails and follow-ups alike, until its score is back at or above it; its conversations wait for it and never move to another account. 0 means no minimum; omit it to leave the setting as it is | |
| allowNonBusinessEmails | No | Allow sending to free mailbox providers (gmail.com, etc.) | |
| isEnabledEmailVerifier | No | Verify lead emails before sending | |
| ignoreOutOfOfficeReplies | No | Do not stop follow-ups on out-of-office replies | |
| maximumTimeBetweenEmails | No | Maximum gap between two sends, in minutes | |
| minimumTimeBetweenEmails | No | Minimum gap between two sends, in minutes | |
| isEnabledCatchallValidated | No | Send to catch-all validated addresses | |
| isEnabledStopFollowUpsOnReply | No | Stop follow-ups to a lead once they reply | |
| isEnabledIgnoreHardBouncedLeads | No | Skip leads that previously hard-bounced | |
| isEnabledSkipLeadIfAlreadyExists | No | Skip leads that already exist in the workspace | |
| maximumSendingLimitPerSenderEmail | No | Daily sending limit per sender email account | |
| isEnabledStopFollowUpsForSameCompany | No | Company reply stop: once a lead replies (out-of-office and other automatic replies don't count), stop emailing the other leads at the same company in this campaign. Subdomains count as the same company; personal addresses like @gmail.com never do. On by default | |
| isEnabledIgnoreLeadsWhoAlreadyResponded | No | Skip leads who already responded in another campaign | |
| maximumSendingLimitPerSenderEmailVariation | No | Random daily variation applied to the sending limit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint=false and destructiveHint=false; the description adds real behavioral context beyond them — the new campaign is a non-sending DRAFT with no emails, and launch is blocked until sequence and sender are added. It does not describe return values, but the state/lifecycle disclosure is meaningful added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, then the resulting state, then the prerequisite chain, then where settings live. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 20-parameter creation tool with full schema coverage and no output schema, the description supplies the key lifecycle and prerequisite information an agent needs to sequence these calls correctly. It does not mention what the call returns, but with no output schema that is a minor gap against otherwise strong coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 20 parameters in detail. The description only groups them loosely ('sending limits and deliverability settings can be set here'), adding no per-parameter meaning beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Creates a campaign in the workspace') and distinguishes it from siblings by naming replace_campaign_sequence and update_campaign in context. The DRAFT lifecycle statement further clarifies what this tool specifically does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear prerequisites and lifecycle context: the campaign starts as a DRAFT, sends nothing, and cannot be launched until a sequence is written and a sender email attached. It names the follow-up tools (replace_campaign_sequence, update_campaign) but does not phrase explicit when-to-use/when-not conditions, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_dfy_orderCreate done-for-you order (charges real money)AInspect
SPENDS REAL MONEY. Places a real order for sending domains and pre-warmed mailboxes: domains are registered at their listed price (about $13.99/year for a .com — search_dfy_domains shows the exact price per domain) and each mailbox costs about $3 setup plus $6/month. The order is accepted immediately and provisioned in the background; poll get_dfy_order for progress. Purchased domains redirect visitors to forwardingDomain. Use search_dfy_domains first to confirm availability and price. Every domain in domains needs at least one mailbox in mailboxes whose domainName is that domain, or the order is refused before anything is bought; mailboxes can also go on a domain from an earlier order.
| Name | Required | Description | Default |
|---|---|---|---|
| domains | No | Domains to purchase. Each one needs at least one mailbox in mailboxes with the same domainName. | |
| mailboxes | No | Mailboxes to provision: at least one on every domain in domains. A mailbox can also go on a domain from an earlier order. | |
| forwardingDomain | No | Where the purchased domains redirect visitors, e.g. acme.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses the significant financial impact ('SPENDS REAL MONEY'), pricing details, immediate acceptance with background provisioning, and the redirection behavior of purchased domains. It also explains the refusal condition. This is substantial behavioral context the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than typical but every sentence carries critical information: the money warning is front-loaded, then pricing, then workflow, then validation rule. The structure is logical and scannable, with the most important warning in caps first. It is not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that spends money and has 3 parameters, the description covers prerequisites, costs, validation rules, and post-creation polling. The only minor gap is the exact return shape, but since the description points to get_dfy_order for progress, an agent can proceed safely. Given the absence of an output schema, this is complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides descriptions for all three parameters (100% coverage), so the baseline is 3. The description adds value by quantifying costs for domains and mailboxes, and by re-emphasizing the domain-mailbox relationship and the option to attach mailboxes to earlier order domains. While some info duplicates the schema, the cost and process details go beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Places a real order'), the resources (domains and mailboxes), and distinguishes itself from siblings by directing to search_dfy_domains first and get_dfy_order after. It leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs to use search_dfy_domains first to confirm availability and price, and to poll get_dfy_order for progress. It also details the hard precondition (every domain must have a mailbox) and the consequence (order refused). This gives the agent concrete when-to and when-not-to guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_email_verification_jobVerify a list of email addresses (spends money)AInspect
SPENDS MONEY: every address checked adds metered usage to the workspace's next invoice. This does NOT use prospect credits and it does NOT create or send a campaign — it verifies a list on its own and gives back a per-address verdict you can download as CSV. Call get_email_verification_rates first for the current price: an address costs one verification credit when the first provider rejects it outright and two when both providers answer, so multiply your list size by perAddressCeilingUsd to budget. An address that gets no verdict is NOT billed and comes back with result unknown. Duplicates are removed and unparseable entries are returned in invalid_emails before anything is charged, so the total in the response is what will be billed against, not what you submitted. Needs an active subscription: a trialing or past_due workspace is refused with subscription_not_active. If the workspace's monthly verification allowance runs out mid-job the job does not fail — the remaining addresses come back unknown, marked, still in the download, and unbilled. Poll get_email_verification_job until state is done, then read the results with list_email_verification_results.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | What to call this job in the app, e.g. 'Q3 conference list' | |
| emails | Yes | The addresses to verify, up to 50,000. Duplicates and unparseable entries are dropped before billing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description extensively discloses behaviors beyond the annotations: it spends money and adds metered usage to the invoice, does not use prospect credits, bills only valid verdicts, removes duplicates, returns unparseable entries separately, requires an active subscription, and handles quota exhaustion gracefully. This fully compensates for the lack of a readOnlyHint and clarifies the openWorldHint/idempotentHint semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries material operational or financial information. The 'SPENDS MONEY' warning is front-loaded, and the structure flows from cost, to behavior, to subscription requirements, to job lifecycle, making it easy to scan despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a money-spending async job with no output schema, the description covers prerequisites, pricing, edge cases, failure modes, and the full polling/retrieval flow. An agent has everything needed to decide whether to call it, budget for it, and retrieve the results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema already documents both parameters, so the baseline is 3. The description adds meaningful value by explaining the billing implications of the emails parameter, noting that duplicates and unparseable entries are removed before charges, and that the response total is what gets billed. This goes beyond the schema's wording and helps the agent reason about cost.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: it verifies a list of email addresses and returns a per-address verdict downloadable as CSV. It also explicitly distinguishes this from a campaign and from using prospect credits, so an agent can tell it apart from sibling tools like create_campaign and get_email_verification_rates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit workflow: call get_email_verification_rates first to budget, then create the job, poll get_email_verification_job until state is done, and finally read results with list_email_verification_results. It also states what the tool does not do (no campaign, no prospect credits), providing clear routing among alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_icpCreate ICPAInspect
Stores a manually-authored Ideal Customer Profile (ICP): who the workspace sells to, as targeting criteria (titles, seniorities, industries, company sizes, locations, keywords). Set makePrimary to promote it to the workspace's active profile. Use copilot_plan instead when you want the AI to build the ICP from a website.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Profile name, e.g. "Mid-market SaaS RevOps leaders" | |
| titles | No | Job titles to target, e.g. ["CEO", "Head of Sales"] | |
| summary | No | One to three sentences describing who the workspace sells to | |
| keywords | No | Free-text keywords, e.g. ["b2b saas"] | |
| locations | No | Locations, e.g. ["United States"] | |
| industries | No | Industries, e.g. ["Software"] | |
| makePrimary | No | Promote this profile to the workspace's active one, demoting any existing primary | |
| seniorities | No | Seniority levels, e.g. ["owner", "director"] | |
| companySizes | No | Company headcount ranges, e.g. ["11-50", "51-200"] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint=false and destructiveHint=false, so they carry little behavioral weight. The description adds real context: the profile is persisted to the workspace and makePrimary demotes any existing primary. It does not cover permissions, limits on profile count, or what happens to the profile's computed audience, leaving a modest gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The what-it-stores content leads and the routing/alternate-tool guidance follows, which is the correct front-loading for selection decisions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters all fully described in the schema and no output schema to explain, the description's job is selection and behavioral framing, which it covers well. The one meaningful omission is guidance on the relationship to update_icp, get_icp, and set_primary_icp, which a sibling-aware agent would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter, including makePrimary, is documented in the schema itself. The description restates the criteria categories and the makePrimary effect but adds no format or syntax detail beyond what is already structured, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Stores a manually-authored Ideal Customer Profile") and immediately defines what that resource contains (titles, seniorities, industries, sizes, locations, keywords). The word "manually-authored" pre-emptively distinguishes it from the AI-driven sibling, so an agent can tell the two apart without reading either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to copilot_plan "when you want the AI to build the ICP from a website," which is a clear alternative-plus-condition pair. It also notes that makePrimary should be set to activate the profile. It stops short of stating when not to create one (e.g. duplicates of an existing ICP), so it falls just short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_webhookCreate webhookBInspect
Registers a webhook endpoint that receives a POST request whenever the selected event happens (email sent, reply received, bounce, lead created, lead category updated, or campaign status changed).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | HTTPS URL that will receive the events | |
| name | No | Display name for the webhook, shown in the app | |
| type | Yes | The event type to subscribe to |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already state openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds useful delivery semantics (a POST request fires on each event) but omits auth/permission requirements, rate limits, whether a secret/signature is issued, or what happens on duplicate URLs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and the event enumeration trailing. It is efficient and waste-free, though the parenthetical list is long enough to slightly blunt the opening statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter creation tool with a fully documented schema and no output schema, the description conveys purpose and event scope adequately and the schema carries the parameters. It is only slightly thin on create-time specifics such as response contents and required-field behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents url, name and type. The description's event list mirrors the type enum but adds no format or constraint detail (e.g., HTTPS-only enforcement, name length) beyond structured data, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Registers a webhook endpoint') and even enumerates the subscribable events, so an agent immediately knows this creates an event subscription. It is distinct from update_webhook/delete_webhook/list_webhooks by verb, but it never explicitly names or contrasts those siblings, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of alternatives or exclusions (e.g., use update_webhook to modify an existing one, or list_webhooks to inspect). The behavior of the created webhook is described, but nothing routes the agent to or away from this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_workspaceCreate workspaceAInspect
Creates a sub-workspace under the account's main workspace (for example one per client, for an agency) and by default mints a read & write API key bound to it. The key's fullKey is returned exactly ONCE and cannot be retrieved again, so hand it to the user straight away; it never outlives the calling key. Must be called with the main workspace's key: sub-workspace keys are refused. Workspaces are a Professional-plan feature.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name of the new workspace (1 to 60 characters) | |
| generateApiKey | No | Also mint a read & write API key for the new workspace. Defaults to true |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: the API key's fullKey is returned exactly ONCE and cannot be retrieved again, the key never outlives the calling key, sub-workspace keys are rejected, and a plan requirement applies. These are operationally critical facts an agent cannot infer from openWorldHint/idempotentHint/destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct load: what is created, the one-time key disclosure hazard, and the authorization/plan constraints. Front-loaded with the core action and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description properly carries the return-value burden by explaining fullKey's one-time disclosure. For a two-parameter creation tool with full schema coverage, nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning by stating the key is minted by default and is read & write bound to the new workspace, clarifying the generateApiKey side effect beyond the schema's brief wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (creates) and resource (sub-workspace under the account's main workspace) with a concrete use case ('one per client, for an agency'). It also implicitly distinguishes itself from the sibling create_workspace_api_key by explaining what the creation produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context and prerequisites: must be called with the main workspace's key, sub-workspace keys are refused, and it is a Professional-plan feature. It does not explicitly route to sibling tools like create_workspace_api_key for separate key minting, so it stops short of full alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_workspace_api_keyCreate API key for a workspaceAInspect
Mints a new API key bound to a workspace the calling key already owns: its own workspace, or one of its sub-workspaces when called with the main workspace's key. This is how a workspace whose key was lost gets a new one without a UI step. fullKey is returned exactly ONCE and cannot be retrieved again, so hand it to the user straight away. The new key authorizes the named workspace only and never outlives the calling key. Keys can be revoked in the app under API & MCP.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric workspace ID (from list_workspaces) | |
| name | No | Name shown in the app's API key list. Defaults to "<workspace name> key" | |
| readOnly | No | Mint a read-only key. It can use the read (list and get) tools plus confirm_connection, get_audience_size and search_lead_finder, so it can read, size audiences and browse Lead Finder (which still uses the account's daily browsing allowance), but it cannot spend credits or change anything else. Defaults to false, a read & write key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate openWorld=false, idempotent=false, destructive=false, leaving behavioral disclosure to the description. It openly warns that fullKey is returned exactly once and cannot be recovered, notes that the new key never outlives the calling key, and states authorization scope plus revocation path. This is exactly the kind of critical behavioral context the agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact yet information-dense: core action, recovery use case, one-time retrieval warning, authorization scope, and revocation location all fit into four purposeful sentences. Warnings are front-loaded after the main verb, making the critical 'only once' instruction immediately visible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter tool with no output schemaebur, the description covers the full behavioral picture: what the tool does, when to use it, critical irreversibility, scope limitations, lifecycle relative to the calling key, and how to revoke. No essential contextual gap remains for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with id, name, and readOnly each already well documented, including defaults and behavior. The description does not add parameter-level details beyond what the schema provides, so it meets the baseline without compensating for any schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific action, 'Mints a new API key bound to a workspace', and clearly differentiates from siblings like create_workspace by emphasizing it is tied to an already-owned workspace and specifically designed for key recovery. The relation to sub-workspaces and main workspace keys adds precise scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the concrete scenario: recovering a lost workspace key without UI intervention wheelchair. It may not explicitly list alternatives, but the scenario strongly implies this is the specialized tool for that case, while other create_* siblings serve different purposes. Clear context, though no explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_campaignDelete campaignADestructiveIdempotentInspect
Permanently deletes a campaign and all associated data, including its emails, sequences and scheduled tasks. This cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is partly covered. The description adds genuinely new context beyond the annotations: the cascade scope (emails, sequences, scheduled tasks) and the irreversibility warning, which is exactly what an agent needs before invoking a destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero redundancy; the cascade scope and the irreversibility warning are both front-loaded rather than buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with no output schema, the description covers the essentials: what is removed, that it is permanent, and that it cascades. It omits any note on permissions or whether related resources (e.g. leads) survive, which is a minor gap rather than a critical one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'id' parameter is fully documented in the schema, including its provenance ('from list_campaigns'). The description adds no parameter detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Permanently deletes a campaign') and immediately scopes the blast radius ('all associated data, including its emails, sequences and scheduled tasks'). This clearly distinguishes it from siblings like pause_campaign, update_campaign, or delete_icp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb and the irreversibility warning, but the description never says when to choose this over the non-destructive alternative pause_campaign, nor does it state prerequisites such as permissions or whether a campaign must be stopped first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_icpDelete ICPADestructiveIdempotentInspect
Deletes an Ideal Customer Profile. Deleting the primary leaves the workspace without an active profile. This cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric ICP ID (from list_icps) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the safety profile is known, but the description adds substantive context beyond them: irreversibility ("This cannot be undone") and the workspace-level side effect that deleting the primary leaves no active profile. It omits what happens to campaigns or other entities referencing the ICP, but for a single-ID delete this is solid added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, with the core action and the most important warning (irreversibility) front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive single-parameter tool with full annotation coverage and a fully documented schema, the description covers action, irreversibility, and the primary-profile edge case. Only unresolved downstream effects (e.g., impact on campaigns using the ICP) keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the single required id parameter, including its source (list_icps). The description adds nothing about the identifier, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (deletes) and resource (Ideal Customer Profile) in the first sentence, which cleanly separates it from sibling deleters like delete_campaign, delete_lead, and delete_webhook. An agent can identify the target entity without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains a consequence of deleting the primary ICP but gives no when-to-use guidance, no prerequisites, and never points to alternatives such as update_icp (to modify instead) or set_primary_icp (to reassign primacy first). The agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_leadDelete leadADestructiveIdempotentInspect
Permanently deletes a lead by ID. This cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric lead ID (returned when the lead was created) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=false, so the destructive profile is structured data. The description's contribution is 'Permanently ... cannot be undone,' which clarifies that there is no soft-delete or restore path — a real addition, but thin given how much the annotations already carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler: the action and its permanence are front-loaded. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with annotations covering the safety profile and a fully documented schema, the description is nearly sufficient. It could add a note on behavior for a nonexistent ID or whether related records (e.g., conversation history) are affected, which keeps it short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'id' parameter is fully documented, including its origin ('returned when the lead was created'). The description adds only the phrase 'by ID' and no format or validation detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Permanently deletes') and resource ('a lead'), with the identifier form ('by ID'). It is trivially distinguishable from siblings like delete_campaign, delete_icp, and update_lead, so an agent can select it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives (e.g., update_lead to change status, or whether a soft-delete/archive path exists). The irreversibility warning hints at caution but is not usage routing, and no prerequisites or preconditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_webhookDelete webhookADestructiveIdempotentInspect
Deletes a registered webhook endpoint by ID so it stops receiving events.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric webhook ID (from list_webhooks) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds the consequential effect ('stops receiving events'), which tells the agent what is actually disrupted, though it does not address reversibility or whether delivery history is retained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and resource, with the purpose clause appended for value. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter delete with annotations covering destructiveness and idempotency, the description is nearly complete. The only residual gap is whether the deletion is permanent and what happens to already-queued events, which is minor given the annotation coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%; the id parameter is fully documented in the schema including its origin (from list_webhooks). The description only restates 'by ID', adding no syntax or constraint detail beyond the structured field, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Deletes) and resource (registered webhook endpoint) with the selection key (by ID), and the purpose clause distinguishes it from list_webhooks and update_webhook without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and the 'stops receiving events' rationale, but there is no explicit when-to-use/when-not guidance and no mention of the alternative update_webhook for disabling rather than removing a webhook.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detach_campaign_sender_emailDetach sender email from campaignADestructiveIdempotentInspect
Detaches a sender email from a campaign so the campaign stops sending through that mailbox. Unsent follow-ups to emails this sender already sent are canceled and do not come back on re-attach; a running campaign is rescheduled onto its remaining senders; already-sent emails are untouched. Idempotent: detaching a sender that is not attached is a no-op.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) | |
| senderEmailId | Yes | The numeric sender email ID (from list_sender_emails) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=true, but the description goes well beyond them by enumerating exactly what is destroyed (unsent follow-ups are canceled and do not return on re-attach), what is preserved (already-sent emails untouched), and what happens to a live campaign (rescheduled onto remaining senders). This is precisely the operational detail an agent cannot infer from the hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and consequence; the destructive side effects follow in a tight clause chain, and the idempotency note closes it out. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive mutation with no output schema, the description covers side effects, reversibility (do not come back on re-attach), and idempotency, which is everything an agent needs to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (campaign id, senderEmailId) carry source pointers in the schema itself. The description adds no additional parameter-level meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (detaches) plus resource (sender email from a campaign) with a clear outcome clause ('so the campaign stops sending through that mailbox'). The counterpart sibling attach_campaign_sender_emails is implied by 're-attach', making the direction of the operation unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied by describing the effect of detaching, and the idempotent no-op clause tells the agent it is safe to call blindly. However, it never explicitly states when to choose this over the sibling attach_campaign_sender_emails or what an agent should do instead when it wants to swap mailboxes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_audience_sizeGet audience sizeARead-onlyInspect
Counts how many prospects match ad-hoc targeting criteria (titles, seniorities, industries, company sizes, locations, keywords) without saving anything. Free — searching never spends credits; only revealing contact details is metered. Use it to validate targeting before creating an ICP or sourcing leads. A read-only key may call it: a read-only key can use the read (list and get) tools plus confirm_connection, get_audience_size and search_lead_finder.
| Name | Required | Description | Default |
|---|---|---|---|
| titles | No | Job titles to target, e.g. ["CEO", "Head of Sales"] | |
| keywords | No | Free-text keywords, e.g. ["b2b saas"] | |
| locations | No | Locations, e.g. ["United States"] | |
| industries | No | Industries, e.g. ["Software"] | |
| seniorities | No | Seniority levels, e.g. ["owner", "director"] | |
| companySizes | No | Company headcount ranges, e.g. ["11-50", "51-200"] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, and the description reinforces this by stating 'without saving anything' and 'Free — searching never spends credits; only revealing contact details is metered.' This adds valuable behavioral context about credit usage and side-effect-free operation beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, then adds the credit/read-only context and usage guidance. The read-only key sentence is slightly tangential to the tool's core function but still useful for access decisions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only counting tool with no output schema, the description covers purpose, cost behavior, and usage context. It doesn't describe the return format (e.g., whether it returns a single number or a breakdown), but the tool's simplicity and the annotations reduce the need for that detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters with examples. The description adds the overall semantic that these are ad-hoc targeting criteria, but doesn't add per-parameter meaning beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool counts matching prospects for ad-hoc targeting criteria without saving anything, distinguishing it from lead sourcing and ICP creation tools. It names the specific resource (audience size) and the verb (counts), making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use it to validate targeting before creating an ICP or sourcing leads, and clarifies when not to use it (not for revealing contact details). It also notes that a read-only key may call it, which is a clear access guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_autopilot_runAutopilot: get a runARead-onlyInspect
Returns one autopilot run including its full audit trail. It does not return the budget plan: that comes back only from start_autopilot_run (when budgetUsd was given) and from plan_autopilot. Use it to check progress.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric autopilot run ID (from list_autopilot_runs) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, covering the safety profile. The description adds that it returns the full audit trail and that the budget plan is absent, which is useful behavioral context. However, it doesn't describe pagination, size, or other runtime traits; with annotations covering safety, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences with zero waste. It front-loads the core return value, then clarifies exclusions, then states the use case. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with a single fully documented parameter and no output schema, the description is nearly complete: it covers the return (full audit trail) and an important omission (budget plan). It could mention return format nuances, but the annotations and schema cover the rest adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the 'id' parameter fully documented as the numeric run ID from list_autopilot_runs. The description adds no additional parameter meaning beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb and resource: 'Returns one autopilot run including its full audit trail.' It also distinguishes the tool from siblings by explicitly scoping out the budget plan, which comes from start_autopilot_run and plan_autopilot. An agent can identify what this tool does and what it doesn't without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: 'Use it to check progress.' It also distinguishes behavior from start_autopilot_run and plan_autopilot regarding the budget plan. However, it does not explicitly state when not to use it versus list_autopilot_runs or other status-related siblings like get_inbox_placement_run.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_billing_profileGet workspace billing profileARead-onlyInspect
Returns the registrant and postal contact details held for the workspace: what is filed with the registrar when a done-for-you order buys a domain, and the postal address a CAN-SPAM footer must carry. Errors with 404 until set_billing_profile has been called.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: the 404-until-configured failure mode and the semantic purpose of the data (registrar filing vs CAN-SPAM footer). No return format/pagination detail, but not needed for a fixed-profile read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the return content and ending with the error precondition. Dense but every clause carries meaning; the em-dash gloss is slightly elaborate but informative rather than wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return values, and it does so by naming both data categories and their purpose. Combined with the failure-mode note, an agent has enough to call it correctly; only a precise field list/typing is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description correctly implies no input is needed by describing a fixed per-workspace lookup, adding nothing misleading about parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('Returns the registrant and postal contact details held for the workspace') and enumerates the returned fields, which distinguishes it clearly from the sibling set_billing_profile and from generic workspace lookups like get_workspace_details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit precondition — 'Errors with 404 until set_billing_profile has been called' — which routes the agent to the sibling setter before invoking this. It does not, however, state when to prefer this over other workspace-info tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_blocklist_entryGet blocklist entryARead-onlyInspect
Returns one suppression entry by ID: the blocked domain or email address, when it was added and which list it belongs to. Entries inherited from the account-wide list are visible from a sub-workspace too and report scope "global".
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric blocklist entry ID (from list_blocklist_entries) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuinely non-obvious semantics: entries inherited from the account-wide list are visible from a sub-workspace and report scope "global", which an agent could not infer from schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and return payload, then a scoping caveat. Minimal waste, though the clause structure of the first sentence is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields and the scope field's possible value, which is enough for a single-ID lookup tool. Error behavior for a nonexistent ID is the only notable omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and schema description coverage is 100%, and the schema already explains the ID's numeric nature and its source tool. The description adds no syntax or format detail beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (returns) and resource (one suppression/blocklist entry) scoped by ID, and enumerates the fields returned (blocked domain or email, added date, list). This clearly separates it from sibling list_blocklist_entries, which returns many entries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the schema notes the ID comes from list_blocklist_entries, which hints at the list-then-fetch workflow, but the description never says when to use this versus list_blocklist_entries or update_blocklist_entry, nor states preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_campaignGet campaignARead-onlyInspect
Retrieves a campaign by ID, including settings, sending schedule, stats and per-status email counts. minimumHealth lists each connected email account on the campaign with its healthScore and whether the campaign's minimum health score (settings.minimumHealthScore) holds it back, with the reason; when every account is held back, allHeldBack is true and message says the campaign is sending nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real behavioral value by explaining the minimumHealth semantics and the allHeldBack/message signal for a stalled campaign, which is derivable behavior an agent could not otherwise predict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core retrieval statement, then uses the second sentence to explain the one non-obvious return field. It is dense but every clause earns its place given there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns and does so for settings, schedule, stats, counts, and the minimumHealth block. It could still be more complete about naming the sibling stats/sequence tools, but for a single-param read it is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (id), documented at 100% coverage in the schema, including its source (from list_campaigns). The description adds nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (retrieves) and resource (campaign by ID) and enumerates what is returned: settings, schedule, stats, per-status email counts, and minimumHealth details. An agent can distinguish this from get_campaign_stats or get_campaign_sequence without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named, even though siblings like list_campaigns, get_campaign_stats, and get_campaign_sequence overlap. The description tells the agent what comes back but not when to pick this tool over those.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_campaign_sequenceGet campaign sequenceARead-onlyInspect
Returns the ordered emails that make up a campaign's sequence, including any A/B variants. delayDays on each step is counted from the previous step, and the first step is always 0.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description goes further with domain semantics that structured fields do not carry: delayDays is relative to the previous step and the first step is always 0. It omits return format details beyond the basics, but with annotations doing the safety work this is a solid addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler. The primary return description is front-loaded and the delayDays semantics follow immediately as supporting detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return contract, and it does communicate the shape (ordered emails, A/B variants included) plus the delayDays convention. It leaves gaps around empty-sequence behavior and whether variant ordering is guaranteed, but it is sufficiently complete for a simple single-parameter read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100% -- the schema already documents id as 'The numeric campaign ID (from list_campaigns)'. The description adds nothing about the id parameter, so the baseline 3 applies; the schema is fully doing the work here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Returns the ordered emails that make up a campaign's sequence') and clarifies scope with 'including any A/B variants,' which distinguishes it from get_campaign and replace_campaign_sequence. It stops short of naming those siblings explicitly, so an agent must infer the boundary from the resource noun alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the read-only, inspection-oriented phrasing, but there is no explicit when-to-use statement, no mention of when to prefer get_campaign over this, and no note on prerequisites. The reference to a numeric campaign ID from list_campaigns hints at the retrieval flow without spelling out the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_campaign_statsGet campaign statsARead-onlyInspect
Returns a campaign's performance in one call: totals (sent, replied, positive, bounced, meetings), per-A/B-variant results, per-sequence-step results and a daily time series (UTC days). Use it to judge how a campaign is doing or to compare variants. There is no opens metric anywhere in the response: Emailchaser does not track opens, so its absence is deliberate, not an omission.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuinely useful non-obvious context: the response is aggregated in one call, days are UTC, and the absent opens metric is a deliberate product limitation rather than a bug. It still says nothing about result size, pagination or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each doing work: first enumerates the return payload, second gives the use case, third preempts a likely confusion about the missing opens metric. Front-loaded and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden and does it well by listing every section of the response. Combined with annotations covering the read-only profile, an agent has everything needed to call this correctly with a single campaign ID.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100% — the schema already says it is the numeric campaign ID from list_campaigns. The description adds no syntax, format or edge-case guidance beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource and enumerates the exact payload: totals, per-A/B-variant results, per-sequence-step results and a daily UTC time series. This is far more specific than a generic 'get stats', though it never names a sibling (e.g. get_campaign vs get_outcomes_report) to explicitly draw the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use context: 'Use it to judge how a campaign is doing or to compare variants.' That tells the agent the intent, but there are no exclusions or named alternatives for when a different report tool would be the better choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_credit_balanceGet credit balanceARead-onlyInspect
Returns the workspace's credit wallet: available, reserved, lifetime granted and used, and the monthly grant. Credits pay for prospect reveals and AI work (list price $0.033 each). A new workspace sees zeros, not an error. unlimited is true when Emailchaser has made this workspace's credits free: every credit action then goes through without using available and is never refused for lack of credits, so there is no need to check the balance or buy credits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral detail beyond the readOnlyHint annotation: a new workspace sees zeros rather than an error, and the unlimited flag changes whether credit checks or purchases are necessary. These are exactly the edge cases that a caller needs to know and that are not present in the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the return concept and fields, then adds pricing, the zero-state behavior, and the unlimited caveat. Every sentence earns its place, and none of the detail is redundant with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the full explanatory burden. It covers what is returned, the credit consumption model, the new-workspace zero behavior, and the unlimited case, making the tool safely callable and interpretable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters and the input schema has no properties, so the baseline is 4. The description adds useful context about what credits are used for and pricing, but there are no parameter meanings to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Returns the workspace's credit wallet' and enumerates the exact fields returned (available, reserved, lifetime granted/used, monthly grant). It is immediately distinguishable from siblings like purchase_credits and list_credit_transactions because it is framed as a balance/status read, not a purchase or transaction history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context about when checking or buying credits matters: credits pay for prospect reveals and AI work, and if unlimited is true there is 'no need to check the balance or buy credits.' It does not explicitly name alternative tools or list exclusions, but the usage context is strong enough to guide an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_deliverability_insightsGet deliverability insightsARead-onlyInspect
What is wrong with deliverability across the workspace right now, worst first, each with the action that fixes it. Combines the inbox placement measured over the window with the latest health check of every connected mailbox: authentication records, blacklist listings and measured placement per mailbox. The best first call when asked 'why are we landing in spam'.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | How many days of runs to summarise (default 30) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: it aggregates the inbox placement window together with the latest health check of every connected mailbox, and returns results ordered worst first with remediation actions attached.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core question the tool answers, then the data it combines, then the usage trigger. No filler and nothing repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only aggregate with one optional param and no output schema, the description covers the sources (placement window, auth records, blacklist listings, per-mailbox placement) and the output shape at a high level (ranked issues plus fixes). Precise return fields are not enumerated, but no critical calling information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional parameter with 100% schema description coverage, so the schema already explains 'days' and its default. The description refers only obliquely to 'the window' without adding format or semantics beyond what the schema provides, matching the baseline for fully documented params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific diagnostic verb+resource: it surfaces what is wrong with deliverability across the workspace, ranked worst first, each item paired with a fixing action. It is clearly distinguishable from siblings like get_sender_reputation or get_inbox_placement_stats_by_date, which cover only one mailbox or one metric.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: 'The best first call when asked why are we landing in spam,' which tells the agent to prefer this over narrower siblings. It stops short of naming those alternatives or stating when-not-to-use it, so it is strong context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dfy_orderGet done-for-you orderARead-onlyInspect
Returns one done-for-you order: status (created, pending_approval, processing, completed, failed, partially_completed, canceled), cost breakdown and any failure reason. Poll it to track provisioning after create_dfy_order.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric order ID (from create_dfy_order or list_dfy_orders) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety bar is covered. The description adds value beyond annotations by disclosing return content (the full status enum, cost breakdown, failure reason) and noting it is meant for repeated polling, which matters since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first front-loads what is returned, the second gives the usage trigger. The enum listing is dense but earns its place given the absence of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and mostly does so, describing status values, cost breakdown, and failure reason. It omits pagination/format details, but for a single-record read tool it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the id parameter is already documented as coming from create_dfy_order or list_dfy_orders. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns one done-for-you order') and enumerates the returned fields (status enum, cost breakdown, failure reason). The word 'one' implicitly contrasts with the sibling list_dfy_orders, letting an agent distinguish single-fetch from list without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Poll it to track provisioning after create_dfy_order' gives a clear when-to-use context and ties it to the sibling that produces the ID. It lacks an explicit when-not or comparison to list_dfy_orders, but the usage trigger is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_email_verification_jobGet an email verification jobARead-onlyInspect
Returns one verification job: its state, per-verdict counts and what it has billed. Poll until state is done. An errorCode of monthly_limit_reached means the allowance ran out and the remaining addresses came back unknown rather than missing — they are still in the results, marked, and were not billed.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | The job ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond readOnlyHint=true, the description discloses the monthly_limit_reached error semantics, what happens to remaining addresses, and that they are not billed. This is extra behavioral detail not present in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: returns/payload, polling instruction, and a critical edge-case explanation. The most important content is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description covers what the response contains, how to use it (poll), and a non-obvious error case. For a one-parameter read tool, this is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and jobId is already documented as 'The job ID'; the description adds no new parameter-level detail. The baseline of 3 is appropriate since the schema carries the parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description begins 'Returns one verification job' with a specific verb and resource, and immediately enumerates the payload (state, per-verdict counts, billed). The singular 'one' distinguishes it from list_email_verification_jobs and from the cancel/create siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction 'Poll until state is done' tells the agent when to use the tool and how. It does not explicitly name alternatives such as list_email_verification_jobs, but the usage context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_email_verification_ratesGet the email verification rateARead-onlyInspect
Returns what one email verification costs right now, read from the billing plan the meter actually charges against. Call this before create_email_verification_job rather than assuming a price. Every check runs a waterfall of up to two providers and each one that returns a verdict is one billable verification credit, so an address costs ONE credit when the first provider rejects it outright and TWO when both answer. Budget against perAddressCeilingUsd, which is the common case, not perAddressFloorUsd.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations readOnlyHint=true, the description reveals the multi-provider waterfall, the credit-counting rule (one vs two credits), and which billing fields to trust for budgeting. No contradiction with annotations; it adds substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core result, then usage timing, then the pricing rule. Every sentence adds necessary information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even with no output schema, the description explains the return substance (perAddressCeilingUsd vs perAddressFloorUsd) and the billing-credit model, and there are no input arguments to document. An agent has everything it needs to invoke the call and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero parameters, so the parameter-semantics burden is minimal and the baseline is 4. The description correctly focuses on output semantics rather than input details, which are absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object ('Returns what one email verification costs right now') and grounds it in the actual billing plan the meter charges against. It clearly distinguishes this rate-lookup tool from the create_email_verification_job sibling by positioning itself as the pre-budget call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs the agent to call this before create_email_verification_job rather than assuming a price, and gives concrete budgeting guidance (treat perAddressCeilingUsd as the common case, not perAddressFloorUsd). That is clear when-to-use context with an alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_icpGet ICPARead-onlyInspect
Retrieves one Ideal Customer Profile by ID, including its targeting criteria, whether it is AI-generated or human-edited, primary status and estimated audience size.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric ICP ID (from list_icps) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=true, openWorldHint=false), so the bar is lower, and the description adds real value by enumerating what the returned record contains: targeting criteria, AI/human-edited flag, primary status and audience size. That partially compensates for the absent output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One well-formed sentence, front-loaded with the action and the identifier. It is efficient, though the trailing enumeration of returned fields makes it slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-resource read with no output schema, the description supplies the key returned fields and the annotations carry the read-only safety profile. Only the routing decision against get_primary_icp/list_icps is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'id' parameter is documented as 'The numeric ICP ID (from list_icps)'. The description's 'by ID' phrasing adds nothing beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Retrieves') and resource ('one Ideal Customer Profile'), and the 'by ID' scope cleanly separates it from the plural list_icps. It does not name or differentiate against get_primary_icp, which is the closest sibling and a plausible source of confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus list_icps, get_primary_icp, or get_audience_size. The only hint is the schema note that the ID comes 'from list_icps', so usage must be inferred entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_inbox_placement_runGet inbox placement runARead-onlyInspect
Returns one inbox placement run with its complete report: the headline inbox / spam / promotions / missing split, the breakdown per mailbox provider and per sending mailbox, and the content spam rules the copy triggered. 'Promotions' rolls up every Gmail category tab. 'Missing' means the probe was never found in any folder, which usually indicates a silent block: a worse problem than spam, with a different fix. spamScore is in tenths of a point (47 means 4.7) and -1 means not scored.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric inbox placement run ID (from list_inbox_placement_runs) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=false, so the description does not need to cover safety. It adds valuable domain-specific behavioral context: what 'Promotions' and 'Missing' mean, the implication of a silent block, and the spamScore unit convention. This goes beyond structured fields significantly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what is returned and its components. Every sentence adds distinct value: first sentence lists report sections, subsequent sentences define key terms and units. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values, and it does so thoroughly for the main report sections. However, it could be more complete by mentioning pagination or limits, though for a single-run retrieval those are likely not applicable. The tool is complex and the description provides substantial context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single 'id' parameter fully. The description does not add parameter syntax or format details, which is acceptable given the high coverage. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Returns), resource (one inbox placement run), and the full report content. Clearly distinguishes from siblings like list_inbox_placement_runs by specifying singular run retrieval, and no sibling tool retrieves a single run with this detailed report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage by mentioning 'from list_inbox_placement_runs' but doesn't explicitly state when to use this vs list_inbox_placement_run_results or list_inbox_placement_tests. The context of retrieving a single run is clear from the description, but no direct alternatives or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_inbox_placement_stats_by_dateGet inbox placement stats by dateARead-onlyInspect
Rolls completed inbox placement runs up by UTC calendar day over a window, for a trend. Defaults to the last 30 days. Days on which nothing ran are ABSENT from the series rather than zero: never report a missing day as 0% placement, it means nobody ran a test that day.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | End of the window, RFC3339. Defaults to now | |
| from | No | Start of the window, RFC3339, e.g. 2026-08-01T00:00:00Z. Defaults to 30 days ago |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly and closed-world, but the description adds a genuinely non-obvious behavioral trait: the series is sparse, and days with no runs are absent rather than zero. The explicit warning never to read a missing day as 0% placement is exactly the kind of semantics annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core operation and its default, then the crucial sparse-series caveat. No filler; every sentence carries load.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by describing the shape of the series and its absence behavior. An agent can both call it correctly and interpret the result without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (from/to) are already documented with RFC3339 format and defaults. The description restates the 30-day default and adds the UTC-day bucketing concept, but adds little else beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (rolls up), resource (completed inbox placement runs), aggregation dimension (UTC calendar day), and purpose (trend). This cleanly separates it from the raw-run siblings like list_inbox_placement_runs and get_inbox_placement_run, which return individual runs rather than an aggregated series.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context (use it for a trend) and states the window default of the last 30 days. It does not name an alternative tool or give explicit when-not-to-use conditions, but the trend framing makes the intended use unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_leadGet leadBRead-onlyInspect
Retrieves a lead by ID, including contact details and company information.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric lead ID (returned when the lead was created) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the useful non-annotated detail that the response includes contact and company data, which matters because there is no output schema, but it says nothing about error behavior for missing IDs or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tightly written sentence with the resource and return contents front-loaded and zero filler. It is appropriately sized for a single-parameter read, though it is arguably too terse to carry any routing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-by-ID tool with complete schema coverage and readOnly annotations, the description supplies the one missing piece – what the call returns – even without an output schema. Nothing essential to invoking it correctly is absent, though sibling routing would make it airtight.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is fully documented in the schema, including the note that the ID is returned at creation. The description only echoes 'by ID' and adds no format, range, or sourcing detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Retrieves a lead by ID') and even summarizes the payload (contact details, company information), so an agent immediately knows what it fetches. It does not, however, explicitly distinguish itself from siblings like list_leads or get_lead_conversation, leaving that differentiation to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the many sibling read tools (list_leads, get_lead_conversation, get_icp). The 'by ID' phrasing weakly implies single-record lookup, but no exclusions or alternatives are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lead_conversationGet lead conversationARead-onlyInspect
Returns a lead's full email thread in chronological order: outbound emails (sent, scheduled and unsent drafts) plus inbound replies, each with direction and status. AI reply drafts awaiting review are marked isDraft; edit one with update_reply_draft and send it with send_reply_draft (or a human sends it from the app). Use it to read the back-and-forth with one lead before deciding what to do next.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric lead ID (returned when the lead was created) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds useful behavioral context beyond that: chronological ordering, inclusion of unsent/scheduled drafts, and the isDraft marker flagging AI drafts awaiting review, plus who can send them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the return contents and efficiently packs direction, status, and draft behavior into three sentences. The closing sentence about editing/sending drafts is adjacent routing information rather than core to what this tool returns, a minor bit of scope bleed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values and does so thoroughly (thread ordering, direction, status, isDraft marker). An agent has everything needed to call and interpret this read-only tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is fully documented in the schema as the numeric lead ID returned at creation. The description adds no syntax or format detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Returns a lead's full email thread') and details the exact composition of the result: outbound sent/scheduled/unsent drafts plus inbound replies with direction and status. This clearly separates it from siblings like get_lead (lead record) and list_replies (cross-lead reply list).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use condition ('read the back-and-forth with one lead before deciding what to do next') and routes follow-up actions to update_reply_draft and send_reply_draft. It lacks explicit when-not-to-use guidance or a direct comparison against get_lead or list_replies, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lead_finder_filtersLead Finder: filter values and limitsARead-onlyInspect
Returns the values Lead Finder's list filters accept (seniorities, jobFunctions, companySizes, revenue, industries, countries, headquartersCountries, regions, continents), the limits that apply to this workspace (page sizes, deepest page, people per add, first-N count, imports running at once, and the daily browsing allowance with what is left today), and creditsPerProspect, what adding one person to a campaign costs. Free, and makes no call to the data provider. Call it before search_lead_finder so filter values match exactly.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description adds meaningful context: the call is free, makes no call to the data provider, and exposes workspace-specific quotas and remaining daily allowance. This is valuable behavioral/useful context that annotations alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well organized: the filter list is parenthesized and the limits are grouped in a single clause. It is long, but every item names a distinct output category and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-input metadata endpoint, the description covers all returned content: accepted filter values, applicable limits, and credit cost per prospect. It also states free/no-data-provider-call behavior, so an agent has what it needs to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and an empty input schema, so the description carries no parameter-documentation burden. Baseline 4 is appropriate; no parameter semantics are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('Returns') and specific resource (Lead Finder filter values and workspace limits), enumerating all filter categories and limit types. It is immediately distinguishable from sibling tools like search_lead_finder and get_lead_finder_search because it targets filter metadata, not search results or searches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs the agent to call it before search_lead_finder so filter values match exactly. It doesn't enumerate when not to use related Lead Finder getters, but the precondition is unambiguous and clearly signals the intended workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lead_finder_importLead Finder: add progressARead-onlyInspect
Returns one Lead Finder add's progress: status (running, completed, failed or canceled), how many people were added and skipped and why (skipReasons), what email verification said, and credits spent. added, creditsSpent and status can still change until finishedAt is set, so poll it after add_lead_finder_prospects until it is. A completed add's endReason is all_selected, requested_reached, audience_exhausted or fetch_limit, and a canceled one's is canceled. A failed add can still have added and charged people (see added and creditsSpent); lastError says why it stopped: insufficient_credits, contact_cap_reached, results_expired, budget_exhausted, provider_blocked, invalid_filters, campaign_not_found or internal. A new filters add to finish a partial one starts from the top again and counts the people already added toward its fetch limit (five people per one requested, at least 1,000), so after a large add, pick the missing people by ref from search_lead_finder results instead.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The import ID (importId from add_lead_finder_prospects or list_lead_finder_imports) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark the tool read-only; the description adds non-obvious behavior: fields can change until finishedAt, a failed add may still have added and charged people, and endReason/lastError values are enumerated. There is no contradiction with readOnlyHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five dense sentences, front-loaded with the return contract before caveats and alternatives. The status and error enums are inline rather than padded, and the final strategic note about search_lead_finder earns its place as usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return semantics and does so thoroughly: statuses, completion reasons, failure reasons, credits, and the polling caveat are all specified. Minor omissions like exact skipReasons values are not needed to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single id parameter is fully documented in the schema ('importId from add_lead_finder_prospects or list_lead_finder_imports'), so the description does not need to repeat it. It adds no syntax or formatting details beyond the schema, so the baseline for full schema coverage applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names the exact resource ('one Lead Finder add's progress') and lists the specific fields returned: status, added, skipped, skipReasons, email verification, and creditsSpent. This clearly distinguishes it from siblings like get_lead_finder_search and list_lead_finder_imports.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs the agent to poll this tool after add_lead_finder_prospects until finishedAt is set, and provides a conditional alternative: after a large add, use search_lead_finder to pick missing people. This is concrete when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lead_finder_searchLead Finder: read a searchARead-onlyInspect
Reads a search started with search_lead_finder, by its searchId. While it is still running, or while its exact total is still being counted (totalStatus pending), the call waits up to about six seconds, so calling again straight away is fine. A finished page can be re-read for 10 minutes after it was fetched (less if it came from cache), and re-reading it in that window keeps its refs valid for another 60 minutes. After that the call answers status failed with error expired, though refs already shown stay usable until 60 minutes after they were last shown; running search_lead_finder again shows them afresh and uses browsing rows. After 15 minutes the search is not found.
| Name | Required | Description | Default |
|---|---|---|---|
| searchId | Yes | The searchId returned by search_lead_finder |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the readOnlyHint annotation: the ~6-second wait, page re-read window, cache nuance, ref validity extension, failed/expired status, and the 15-minute disappearance. No annotation contradiction exists; the description enriches what the agent needs to know about side effects and lifecycle.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with its core action, and nearly every clause contributes needed behavioral detail. It is somewhat long and could be organized into bullets, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description covers the whole call-and-refresh cycle: when to retry, what windows apply, what error to expect, and what to do after expiry. Nothing required to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents searchId at 100% coverage, so the description only adds provenance by tying it to search_lead_finder. This is mildly useful but does not meaningfully expand parameter understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Reads a search started with search_lead_finder, by its searchId.' This clearly distinguishes it from the sibling search_lead_finder and other get_lead_finder_* tools, and the title reinforces the read action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit lifecycle guidance: calling again while the search is still running is fine, re-reading within 10 minutes extends ref validity, and after 15 minutes the search is not found. It also points to running search_lead_finder again as the refresh path, which effectively covers both when to use this tool and how to recover after expiry.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_outcomes_reportGet outcomes reportARead-onlyInspect
Puts money against results for a time window: credit spend (grouped by reason, valued at the $0.033 list price) plus done-for-you order costs on one side, and sent emails, replies, positive replies and booked meetings on the other, with cost per reply, per positive reply and per meeting. Subscription fees are NOT included. campaignId narrows the outcome side only — spend stays workspace-level — so per-campaign cost figures are partial attribution, not a true campaign cost. Use it to answer 'what did we spend and what did we get'.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Window start, RFC3339 or YYYY-MM-DD (inclusive) | |
| until | No | Window end, RFC3339 or YYYY-MM-DD (inclusive; a bare date means midnight UTC at the start of that day) | |
| campaignId | No | Narrow the outcome side to one campaign (spend stays workspace-level) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses non-obvious calculation behavior: spend is valued at a fixed list price, subscription fees are excluded, and campaignId narrows only the outcome side while spend remains workspace-level. This is exactly the kind of behavioral context an agent needs and could not infer from the schema alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences carry all the essential information with no filler. Metric groups, exclusions, the campaignId caveat, and the intended use case are each given their due, and the most important scoping context is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only reporting tool without an output schema, this description covers the metric families, pricing basis, exclusions, and attribution semantics. It could go slightly further by stating default time-window behavior and return shape, but the provided information is sufficient for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already defines all three parameters. The description adds useful context about campaignId's partial-attribution caveat, but it mostly reinforces what the schema already states about spend remaining workspace-level rather than introducing substantial new parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific operation—mapping money spent to results for a time window—and enumerates the exact metric groups included. The closing 'Use it to answer what did we spend and what did we get' makes the tool's purpose unmistakable and clearly separates it from campaign-stat or credit-balance tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear when-to-use signal and includes important exclusions: subscription fees are not included, and campaignId produces partial attribution rather than a true campaign cost. However, it does not explicitly name sibling alternatives or say when NOT to use this tool, so it stops just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_primary_icpGet primary ICPARead-onlyInspect
Returns the workspace's active (primary) Ideal Customer Profile — the profile the workspace currently targets; at most one is primary. Errors with 404 when none is set.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuine behavioral context beyond that: the 'at most one primary' invariant and the explicit 404 error when none is set, which is not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the return value, with the invariant and error behavior appended. No filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description identifies the returned entity but does not describe the ICP's shape or fields. For a no-param read tool whose safety profile is already in annotations, this is nearly complete, with only minimal detail on the returned object missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document and the baseline of 4 applies. The description appropriately spends no words on parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Returns) and resource (the workspace's active/primary ICP), and the parenthetical clarifies the entity's semantics. It is cleanly distinguishable from siblings like get_icp and list_icps by the 'primary/active' qualifier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when the tool applies (you want the single active profile, not an arbitrary one) and states the boundary condition that at most one exists. It stops short of naming the alternatives (get_icp, list_icps, set_primary_icp) that an agent should choose between.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sender_emailGet sender emailARead-onlyInspect
Retrieves a connected sending email account by ID, including its health score. healthScore is the account's health score: 0-100, higher is better, the share of its warm-up emails over the last 7 full days that landed in the inbox rather than spam. It moves daily as warm-up emails land in the inbox (up) or in spam (down), and is null while warm-up is off or before 20 warm-up emails were checked.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric sender email ID (from list_sender_emails) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real behavioral context beyond that: the healthScore range (0-100), its definition (share of warm-up emails landing in inbox over the last 7 full days), that it moves daily, and that it returns null when warm-up is off or before 20 emails are checked. It does not describe error behavior for invalid IDs, but the added semantics are substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first clause, then spends the remainder on the one piece of non-obvious semantics (healthScore). Every sentence earns its place, though the healthScore explanation is dense and runs long for a single-parameter lookup.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by explaining the most important returned field, healthScore, including its null conditions. Annotations cover the read-only profile and there is only one fully documented parameter. Only error handling and any other returned fields remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is already documented as 'The numeric sender email ID (from list_sender_emails)' with a minimum bound. The description only restates 'by ID', adding nothing the schema doesn't already carry. Baseline 3 applies when the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with scope: 'Retrieves a connected sending email account by ID, including its health score.' The 'by ID' qualifier implicitly separates it from list_sender_emails, though it never names that sibling. An agent can identify what it does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, no named alternatives. The description never mentions that list_sender_emails is the bulk counterpart, nor how this differs from get_sender_email_warmup or get_sender_reputation, which are the obvious near-neighbors for sender-account data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sender_email_dnsGet sender email DNS vitalsARead-onlyInspect
Checks the sender domain's SPF, DKIM, DMARC and MX records and reports each as OK, WARNING, MISSING or ERROR with human-readable issues (the same checks behind the app's DNS vitals view), plus the mailbox's warm-up status: whether it is enrolled, today's ramp target, the daily cap and its healthScore (0-100, higher is better, the share of its warm-up emails over the last 7 full days that landed in the inbox rather than spam). Use it to diagnose deliverability. DNS lookups run live at request time, so the call can take a few seconds but never returns stale records.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric sender email ID (from list_sender_emails) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds meaningful behavior beyond them: live DNS lookups at request time, a few-second latency, guarantee of no stale records, the OK/WARNING/MISSING/ERROR status vocabulary, and how healthScore is computed (7-day inbox share). This is rich context an agent cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is checked before the usage hint and latency note. Dense but every clause carries information; the healthScore parenthetical is slightly long but justified since no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully compensates by explaining the returned record statuses, human-readable issues, and warm-up fields (enrollment, ramp target, daily cap, healthScore). An agent has everything needed to call and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single id parameter is documented with its source (list_sender_emails). The description adds no parameter-level detail, so baseline 3 is appropriate when the schema fully carries that burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (checks) and specific resources (SPF, DKIM, DMARC, MX records plus warm-up status) for the sender domain. It clearly distinguishes itself from siblings like get_sender_email_warmup and get_sender_reputation by enumerating exactly which signals it returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use case ('Use it to diagnose deliverability'), which orients the agent. However, it never names when NOT to use it or routes to alternatives such as get_deliverability_insights or get_sender_email_warmup, so selection among similar siblings still requires inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sender_email_warmupGet sender email warm-up settingsARead-onlyInspect
Reads a mailbox's warm-up settings: whether warm-up is on, the ramp (startLimit warm-up emails on the first day, increaseBy more each sending day, up to capLimit a day), weekdays-only, timezone, today's ramp target (currentPerDay) and the account's healthScore (0-100, higher is better, the share of its warm-up emails over the last 7 full days that landed in the inbox rather than spam; null while warm-up is off or before 20 were checked). Warm-up volume is separate from the campaign sending limits changed with update_sender_email. When configured is false the mailbox has never warmed, and the values shown are the defaults that switching warm-up on would use.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric sender email ID (from list_sender_emails) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only readOnlyHint/openWorldHint from annotations, the description carries the burden and delivers: it explains the healthScore formula and scale, the null conditions (off, or fewer than 20 checked), what currentPerDay represents, and that when configured is false the values are defaults. This is substantial disclosure well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and resource, and every clause carries real information (ramp semantics, healthScore definition, defaults). It is a dense run-on sentence but not padded; slightly heavy for a single-parameter read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description fully compensates by describing the return shape and the meaning/edge cases of each field. An agent can interpret the response correctly without any additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single 'id' parameter already documented (from list_sender_emails). The description adds no further meaning to the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Reads) and resource (a mailbox's warm-up settings) and enumerates the exact fields returned (ramp, weekly limits, timezone, currentPerDay, healthScore). It even distinguishes itself from the mutation siblings by noting warm-up volume is separate from campaign sending limits changed with update_sender_email.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the agent can infer this is the read counterpart to update_sender_email_warmup, and the note about update_sender_email clarifies a boundary. But there is no explicit when-to-use guidance versus the other get_sender_email*, get_sender_reputation, or get_deliverability_insights siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sender_reputationGet sender reputation and mailbox healthARead-onlyInspect
Returns the latest health check for every connected mailbox: SPF, DKIM, DMARC and MX status (OK, WARNING, MISSING or ERROR), blacklist listings with their delisting links, measured inbox placement over recent runs and a combined 0-100 health score. blacklistsChecked is reported next to blacklistsListed so the count is never read as a total. A placementScore or healthScore of -1 means not measured, not zero.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and closed-world safety. The description goes further with two high-value semantics the agent cannot get elsewhere: blacklistsChecked vs blacklistsListed to avoid misreading counts as totals, and -1 meaning not measured rather than zero. That's genuinely useful, though it omits freshness/rate-limit details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the verb and the exact fields returned, then the two semantic caveats. Efficient, though the long field enumeration makes it a bit dense; a bulleted structure would scan better.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden and largely does it: it enumerates the status fields and explains the two ambiguous value conventions. It could still say what a mailbox-scoped result looks like and what 'recent runs' means in terms of time window.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description correctly adds no parameter content, and instead spends its words on output semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (returns) and resource (latest health check for every connected mailbox) with concrete scope: SPF, DKIM, DMARC, MX, blacklists, placement, health score. Distinguishable from siblings like get_sender_email_dns and get_inbox_placement_run, which cover subsets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage: it's the aggregate health-check tool, so an agent would pick it when needing a full mailbox reputation snapshot. But it never says when to prefer this over get_sender_email_dns or get_inbox_placement_run, nor whether it must be called before certain mutations. Context is implied, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workspace_detailsGet workspace detailsARead-onlyInspect
Returns the workspace behind the API key: campaign counts broken down by status and the list of members. Useful as a first call to confirm which workspace a key belongs to.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds the concrete return shape (counts by status, members), but says nothing about auth requirements, rate limits, or whether the member list is paginated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the return contents come first and the usage hint second, which is the right front-loading for a zero-argument tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does list the payload contents. It stops short of noting whether the member list can be large or how counts are grouped, which a caller might want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the schema provides nothing to interpret and the baseline is 4. The description correctly implies the tool is parameterless by scoping it to 'the workspace behind the API key'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (returns) and resource (the workspace behind the API key), and enumerates the payload: campaign counts by status and member list. Implicitly distinguished from the sibling list_workspaces because it takes no scoping arguments and operates on the caller's own key.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the tool as 'useful as a first call to confirm which workspace a key belongs to', giving a clear use context. It does not name list_workspaces as the alternative when enumerating multiple workspaces, so the routing guidance is incomplete but substantive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_instantlyImport from InstantlyAInspect
Imports your Instantly.ai campaigns and leads into Emailchaser using your Instantly API key. Imported campaigns are created as DRAFTS — nothing sends until you review and launch them. Runs in the background and returns a job id.
| Name | Required | Description | Default |
|---|---|---|---|
| instantlyApiKey | Yes | Your Instantly.ai API key (Instantly → Settings → API keys) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (openWorldHint, non-idempotent, non-destructive), and the description adds materially new behavior: imported campaigns land as DRAFTS, nothing sends until manually launched, and the call is asynchronous returning a job id. It does not call out duplicate-import risk on re-invocation, which the non-idempotent hint implies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then the safety-relevant draft behavior, then the async return. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter import tool with no output schema, the description covers purpose, prerequisite credential, side effects (drafts only), and execution model (background job with id). Everything an agent needs to call and interpret the result is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter (instantlyApiKey) already documents where to find the key, matching the description's mention. The description adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Imports) and resources (Instantly.ai campaigns and leads) with the destination system (Emailchaser) named. This clearly distinguishes it from read-oriented siblings like list_instantly_accounts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to invoke it (bringing Instantly data into Emailchaser, requiring an API key) and sets expectations about the draft-review workflow. It does not name alternatives or say when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kill_autopilot_runAutopilot: kill a runADestructiveIdempotentInspect
Engages the kill switch: the run halts wherever it is and can NEVER be resumed, and any campaign the run launched is paused so nothing more sends. Use pause_autopilot_run instead when you might want to continue later. The optional reason is recorded in the run's audit trail.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric autopilot run ID (from list_autopilot_runs) | |
| reason | No | Why the run is being killed, for the audit trail |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=true, but the description adds crucial context beyond them: the action is irreversible ('can NEVER be resumed') and it has a cascading side effect (launched campaigns are paused so nothing more sends), plus that the reason lands in the audit trail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the irreversible halt and the campaign side effect, then the alternative. Efficient, though wording is slightly colloquial rather than purely informational.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive tool with no output schema, the description covers irreversibility, side effects, and the routing alternative – everything an agent needs to call it safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – both 'id' and 'reason' are documented in the schema. The description repeats the audit-trail purpose of 'reason' but adds no syntax, format, or ID-sourcing detail beyond what the schema and list_autopilot_runs reference already provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (killing an autopilot run) with vivid scope: 'the run halts wherever it is and can NEVER be resumed.' It is clearly distinguishable from the sibling pause/resume/start autopilot tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('Use pause_autopilot_run instead when you might want to continue later'), giving a concrete condition that selects between the two tools. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
launch_campaignLaunch campaignAInspect
Launches a new campaign or resumes a paused one. The campaign configuration is validated before launching; emails start sending per the campaign schedule.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare openWorldHint=false and destructiveHint=false, so the safety profile is partially covered. The description adds meaningful behavioral context: validation occurs before launching and emails start sending per the campaign schedule. However, it does not disclose what happens on validation failure, required permissions, or side effects like sending emails to leads.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action, no filler. Efficient and clear, though it could be tightened by combining the validation and sending behavior into fewer words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation tool with no output schema and basic annotations, the description covers the core action and validation behavior. It does not explain return values (none expected), but that is acceptable. Some gaps remain around failure handling and idempotency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description provides no parameter syntax or format beyond the schema. The description does mention the campaign configuration, but there is no additional parameter detail. Baseline 3 is appropriate when the schema already documents the single 'id' parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Launches a new campaign or resumes a paused one') with a clear dual behavior. It is distinguishable from pause_campaign, resume_autopilot_run, and copilot_launch, but does not explicitly name or differentiate against neighboring sibling tools like pause_campaign or copilot_launch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to launch or resume), but does not state preconditions or alternatives. It does not explain when to use launch_campaign vs resume behavior vs paused-campaign handling, or how it relates to pause_campaign.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_autopilot_runsAutopilot: list runsARead-onlyInspect
Lists the workspace's autopilot runs, newest first, with status (pending, building_icp, writing_sequence, sourcing_prospects, awaiting_approval, running, paused, completed, failed) and linked ICP/campaign IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| limit | No | Page size (default 50, max 200) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safe-read profile is covered. The description adds the ordering guarantee ('newest first') and the status taxonomy, which is useful behavioral context, but says nothing about pagination totals, default page size behavior, or empty-result handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the core action, and the parenthetical status list is dense but informative. The inline enumeration of nine statuses is a little heavy but each item carries meaning for interpreting results.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so well, naming ordering, status field, and linked ICP/campaign IDs. It omits pagination metadata (total counts, has-more) that an agent listing paginated results would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters (page, limit) have 100% schema description coverage, including defaults and max. The description adds no parameter-level detail, which is acceptable at this coverage level but earns only the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the workspace's autopilot runs'), adds ordering ('newest first') and enumerates the returned status lifecycle and linked IDs. It is clearly the collection-level counterpart to get_autopilot_run among the listed siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the shape of the tool (browse all runs in the workspace), but there is no explicit guidance on when to prefer this over get_autopilot_run for a single run, nor on how paging interacts with filtering. No alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_blocklist_entriesList blocklist entriesARead-onlyInspect
Returns the suppression entries (blocked domains and email addresses) that apply to this workspace, newest first. Blocked entries are never emailed by any campaign. By default that spans both lists: this workspace's own entries and the account-wide ones inherited from the main workspace. Each entry reports which list it came from.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (default 100, max 1,000) | |
| scope | No | Which list to return: "workspace" for this workspace's own entries, "global" for the account-wide list, "all" for both (default) | |
| offset | No | Number of entries to skip (default 0) | |
| search | No | Case-insensitive substring match on the domain or address |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuine domain context beyond that: blocked entries are never emailed by any campaign, the default spans workspace plus account-wide lists, and each entry reports its source list. It omits pagination/return-shape detail, but that is minor given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the return semantics, then the default scope, then the per-entry provenance field. Every sentence earns its place with no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A read-only list tool with no output schema and fully documented parameters; the description conveys the domain meaning of the data and default scope. Minor gap: it does not hint at result ordering/paging interaction with limit/offset, though the schema covers those parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents limit, scope, offset, and search, including defaults and enum meanings. The description restates the default scope behavior but adds no syntax or format detail beyond the schema, matching the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('returns the suppression entries... blocked domains and email addresses') plus scope and ordering ('newest first'). An agent can distinguish this list tool from the singular get_blocklist_entry sibling without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the description explains the default scope ('spans both lists') but never says when to call this instead of get_blocklist_entry or get_blocklist_entry's filtered variants. No explicit when-not guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_campaignsList campaignsARead-onlyInspect
Lists cold email campaigns in the workspace with their status and stats (sent, replies, bounces). Supports pagination and filtering by status or folder.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 (20 campaigns per page) | |
| limit | No | Page size (default 20, max 200) | |
| status | No | Filter by campaign status | |
| folderId | No | Filter by folder ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real value by saying what comes back (status plus sent/replies/bounce stats) and that pagination/filtering are supported, which matters because there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the primary action and followed by the payload/supporting capabilities. No filler or restated title boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, describing the returned fields (status, sent, replies, bounces) is exactly the missing piece, and it is provided. Minor gaps remain around default sort order and what an empty workspace returns, but the definition is broadly sufficient for a read-only list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the four parameters (page, limit, status, folderId) is already documented in the schema. The description only echoes the filtering and pagination capability without adding format or default details, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists cold email campaigns in the workspace') and summarizes the returned payload. It distinguishes itself implicitly from get_campaign/get_campaign_stats via the 'Lists' scope, though it never names a sibling explicitly to draw the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions the supported capabilities (pagination, filtering by status or folder), which implies when to use it, but gives no explicit when-to-use vs when-not, no guidance on choosing list_campaigns over get_campaign or get_campaign_stats, and no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_credit_transactionsList credit transactionsARead-onlyInspect
Lists the workspace's credit ledger, newest first: every grant, spend, reservation and refund with the balance after each entry. Filter by movement kind, reason or a time window. Use it to see where credits went, or to check for a stripe_topup entry after an unconfirmed purchase_credits call. An entry with free set to true was written while the workspace's credits were free: its amount is what the action would have cost, and no credits moved.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by movement type | |
| page | No | Page number, starting at 1 | |
| limit | No | Page size (default 50, max 200) | |
| since | No | Only entries at or after this time (RFC3339 or YYYY-MM-DD) | |
| until | No | Only entries at or before this time (RFC3339 or YYYY-MM-DD; a bare date means midnight UTC at the start of that day) | |
| reason | No | Filter by what the credits were spent on or granted for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral context beyond the readOnlyHint annotation: entries are ordered 'newest first', each entry includes the balance after it, and the semantics of the 'free' flag are explained (amount reflects hypothetical cost, no credits moved). It also notes that filters exist. This fully discloses the tool's behavior without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences, each adding value: the core purpose, filtering capability, a specific use case, and a nuanced output detail (free flag). It is front-loaded with the primary action and avoids fluff. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with six optional parameters and no output schema, the description covers purpose, ordering, filtering, and a special field semantics. It doesn't explicitly mention pagination (though page/limit are in the schema) or the return format, but these are standard for list tools. The description is comprehensive enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% coverage for all six parameters, so the description doesn't need to explain them. It does mention filtering by 'movement kind, reason or a time window', which maps to parameters, but adds no detail beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Lists', the resource 'credit ledger', and specifies the contents (grants, spends, reservations, refunds with balance). It also differentiates from get_credit_balance by emphasizing the ledger history. This is precise and distinguishes it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit use cases: 'see where credits went' and 'check for a stripe_topup entry after an unconfirmed purchase_credits call'. While it doesn't explicitly name alternatives, it implies the tool is for detailed history rather than a balance snapshot, and it mentions filtering to narrow down. This provides clear context but lacks an explicit 'instead of X' comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dfy_ordersList done-for-you ordersBRead-onlyInspect
Lists the workspace's done-for-you orders, newest first, each with status, cost breakdown and the ordered domains and mailboxes. Internal billing and provider records are never exposed.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| limit | No | Page size (default 50, max 200) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a useful data-scope guarantee ('Internal billing and provider records are never exposed'), which is real value beyond the annotations, but says nothing about pagination behavior or result volume for a paged list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no waste; scope, ordering, and payload are front-loaded and the data-exclusion caveat follows naturally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with no output schema and fully documented parameters, the description covers what is returned and what is deliberately withheld. Only the pagination/traversal behavior is left unaddressed, which is a minor gap given the schema documents page and limit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — page and limit are both documented with defaults and bounds in the schema. The description adds no syntax, default, or constraint detail beyond it, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists the workspace's done-for-you orders') plus ordering and returned contents (status, cost breakdown, domains, mailboxes). The plural scope implicitly separates it from get_dfy_order, but no sibling or alternative is named explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use context, no prerequisites, and no routing to alternatives such as get_dfy_order for a single order or search_dfy_domains for domain lookup. The agent must infer usage purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_email_verification_jobsList email verification jobsARead-onlyInspect
Lists the workspace's standalone email verification jobs, newest first, with live counts and what each has billed so far (meterEvents and spentUsd).
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| size | No | Page size (default 25, max 200) | |
| search | No | Filter by job name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds valuable context about what the tool returns: live counts and billing data (meterEvents and spentUsd), which goes beyond annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, efficient sentence that front-loads the action and resource, then provides key details. No filler words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing operation with all optional parameters documented in the schema, the description provides sufficient detail about output content (live counts, billing fields) and ordering. No missing information that an agent would need to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents page, size, and search. The description adds no additional parameter context, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (lists), the resource (workspace's standalone email verification jobs), and adds distinguishing details like newest-first ordering, live counts, and billing information. This differentiates it from singular get_email_verification_job and results-oriented list_email_verification_results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by noting 'standalone' jobs, but does not explicitly state when to use this over alternatives like get_email_verification_job or list_email_verification_results. No exclusions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_email_verification_resultsList an email verification job's resultsARead-onlyInspect
Returns one row per address with its verdict: valid, catchall_validated (the domain accepts everything, so delivery is likely but not proven), invalid, or unknown (no verdict, and not billed). Each row also carries what each provider said and how many verification credits it cost. Filter with result, where 'deliverable' means valid plus catch-all, which is what a campaign will actually send to.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| size | No | Page size (default 25, max 200) | |
| jobId | Yes | The job ID | |
| result | No | Filter by verdict |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description doesn't need to repeat that. It adds valuable behavioral context: explains the meaning of each verdict, notes that 'unknown' is not billed, and mentions that rows include provider statements and credit costs. This goes beyond the structured data, providing insight into the response content. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff. The first sentence states the core output and defines verdicts; the second explains the filter. All information is pertinent, and the most important details are front-loaded. Efficient and structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema, the description covers what the response contains (one row per address, verdict, provider statements, credit cost) and how to filter. Pagination is standard and self-explanatory via schema. It also clarifies the 'deliverable' filter's meaning, which is crucial for correct usage. No significant gaps for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are documented. The description adds semantic depth for the 'result' parameter by explaining what each enum value means and defining 'deliverable' as a combination. It also hints at the output rows carrying provider details and credit costs, which indirectly informs the 'page' and 'size' usage. The added meaning elevates it above the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Returns one row per address') and clearly identifies the resource (email verification results). It enumerates the verdict values and explains the 'deliverable' filter, distinguishing it from sibling tools like list_email_verification_jobs (lists jobs) and get_email_verification_job (single job). The purpose is unambiguous and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to use the 'result' filter, clarifying that 'deliverable' means valid plus catch-all, which is what campaigns send to. This guides the agent on when to use this filter. However, it does not explicitly state when to choose this tool over alternatives (e.g., get_email_verification_job for aggregate stats), though the sibling names imply it. Slight gap in exclusions, but the core usage is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_icpsList ICPsARead-onlyInspect
Lists the workspace's stored Ideal Customer Profiles (ICPs), newest first, with their targeting criteria, whether each is AI-generated or human-edited, and the estimated audience size.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| limit | No | Page size (default 50, max 200) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds useful behavioral context: sorting order ('newest first') and the content of returned records (targeting criteria, AI vs human-edited, audience size). It doesn't mention pagination limits, but the schema covers those.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-structured sentence that front-loads the core action and efficiently lists the key returned attributes without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description adequately conveys what the tool returns (ICPs with targeting criteria, generation type, audience size) and implies pagination via parameters. It could mention default/max limits, but those are in the schema, so this is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both page and limit fully documented in the input schema. The description adds no parameter meaning beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (workspace's stored ICPs), and clarifies the scope and sorting ('newest first') plus the returned fields. Clearly distinguishable from siblings like get_icp, create_icp, delete_icp, and update_icp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it's a list/read operation for multiple ICPs, but doesn't explicitly state when to use this versus get_icp (single) or get_primary_icp (default). No exclusions or alternatives named, leaving usage context partially inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_inbox_placement_run_resultsList inbox placement run probesARead-onlyInspect
Returns one row per probe email of a run: which of your mailboxes sent it, which seed mailbox received it, where it landed and how many seconds it took to arrive. deliverySeconds is worth watching on its own: greylisting and throttling show up there before they show up in a placement number.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric inbox placement run ID (from list_inbox_placement_runs) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Read-only annotations mean the safety bar is already met, and the description adds genuinely useful behavioral context: row granularity, the columns returned, and the note that deliverySeconds surfaces greylisting/throttling before placement numbers move. That is real operator context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the row semantics, then a focused tip about deliverySeconds. No filler, though the second sentence is somewhat advisory rather than structural.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param read tool with no output schema, the description covers granularity and key fields adequately, and the annotations cover the safety profile. It stops short of pagination or return-value shape details, but nothing essential for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single id parameter is well documented in the schema itself, including the source list_inbox_placement_runs. The description adds no syntax or format detail beyond that, so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Returns one row per probe email'), specific resource (probe results of a run), and explicit row granularity. An agent can tell it apart from get_inbox_placement_run and list_inbox_placement_runs without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The id description routes back to list_inbox_placement_runs, implying this is the drill-down step. However, no explicit when-to-use versus get_inbox_placement_run or get_inbox_placement_stats_by_date, and no stated prerequisite that a run must have completed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_inbox_placement_runsList inbox placement runsARead-onlyInspect
Lists inbox placement runs, newest first, optionally only those of one test. Each run counts how many probe emails landed in the inbox, in spam, in a Gmail category tab (promotions) or nowhere (missing, which usually means a silent block). Counters are only final once status is completed. Any rate of -1 means NOT MEASURED yet and must never be reported as 0%.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many runs to return (default 50, max 200) | |
| testId | No | Only runs of this test (from list_inbox_placement_tests) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds substantial behavioral context the annotations don't cover: counters are only final once status is completed, 'missing' usually means a silent block, and any -1 must never be reported as 0%. This directly prevents a serious downstream reporting error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all dense and front-loaded: identity and ordering first, then what the counters mean, then the critical -1 caveat. No filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry return-value semantics, and it does well by explaining the inbox/spam/promotions/missing counters and the -1 sentinel. It leaves the rest of the run object's fields (ids, dates, test association) and pagination unspecified, a modest remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both limit and testId are fully documented, including the default/max for limit and that testId comes from list_inbox_placement_tests. The description's 'optionally only those of one test' largely restates the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists inbox placement runs') plus meaningful scope and ordering ('newest first, optionally only those of one test'). This cleanly distinguishes it from the singular get_inbox_placement_run and from list_inbox_placement_run_results, which returns per-probe results rather than runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the optional testId filter is appropriate ('optionally only those of one test') but never states when to use this tool versus list_inbox_placement_tests or list_inbox_placement_run_results. Usage is inferable from the resource description rather than explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_inbox_placement_testsList inbox placement testsARead-onlyInspect
Lists every inbox placement test defined in the workspace, newest first. A test is the definition: the email content, which mailboxes send it and, for a recurring test, how often it runs. Each execution of a test is a run (list_inbox_placement_runs). Inbox placement is a paid add-on: a workspace without it gets an error, not an empty list.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds meaningful non-schema behavior: newest-first ordering and the paid-add-on error case (not an empty list), which prevents the agent from misreading an error as 'no tests'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with purpose, then the definition, then the routing and error caveat. Every sentence carries distinct information with no repetition of the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a no-param list tool: purpose, ordering, the definition/run boundary, the sibling for runs, and the add-on failure mode are all present. No output schema exists, but the description implies the shape via 'newest first' and the test definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description correctly documents there are no filtering or pagination inputs, matching the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Lists') and resource ('inbox placement tests'), and explicitly defines what a test is (content, mailboxes, recurrence), distinguishing the definition from its executions. This clearly separates it from siblings like list_inbox_placement_runs and get_inbox_placement_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to list_inbox_placement_runs for executions and warns that a no-add-on workspace gets an error, not an empty list. It doesn't state when to prefer this over per-test getters, but the definition/run distinction covers the main selection decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_instantly_accountsList Instantly sending accountsARead-onlyInspect
Fetches the sending accounts of an Instantly.ai workspace using your Instantly API key, together with a pre-filled CSV for Emailchaser's bulk IMAP/SMTP connection flow. A read-only preview before an import: nothing is written to the workspace. Mailbox passwords cannot be exported from any provider, so each account is either reconnected via Google/Microsoft OAuth in the app or bulk-uploaded with app passwords using the CSV. The backend treats all non-GET calls as writes, so this needs a read & write API key.
| Name | Required | Description | Default |
|---|---|---|---|
| instantlyApiKey | Yes | Your Instantly.ai API key (Instantly → Settings → API keys) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite readOnlyHint=true, the description adds rich behavioral context the annotation doesn't: it confirms nothing is written to the workspace, explains that mailbox passwords cannot be exported and thus accounts must be reconnected via OAuth or bulk-uploaded with app passwords, and discloses the important quirk that the backend treats all non-GET calls as writes requiring a read & write API key. This is exactly the kind of non-obvious behavioral detail that goes beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then adds necessary operational context. It's a bit dense and packs multiple ideas (read-only preview, password limitation, OAuth flow, API key requirement) into a few sentences, but every sentence earns its place and there's no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with annotations already covering safety and no output schema, the description is quite complete. It explains the return artifact (a CSV) and the downstream flow, and covers the API key requirement. It could theoretically mention what the accounts list contains, but the CSV context largely covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only one parameter and 100% schema description coverage, the schema already documents the instantlyApiKey fully (including where to find it in Instantly's settings). The description adds no additional parameter syntax or meaning. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fetches sending accounts of an Instantly.ai workspace, plus a CSV for Emailchaser's IMAP/SMTP import flow. This distinguishes it from siblings like list_sender_emails (which appears to be the native sender email list) and import_instantly. It's clear what it does, though the dual nature (accounts + CSV) and the relationship to the native list_sender_emails sibling isn't fully clarified.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description frames the tool as a read-only preview before an import, which implies when to use it. However, it does not explicitly say when NOT to use it or compare it to alternatives like list_sender_emails or import_instantly. The context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_lead_finder_importsLead Finder: list addsARead-onlyInspect
Lists the workspace's Lead Finder adds (people added to campaigns from the contact database), newest first, each with its progress: requested, added, skipped and why, what email verification said, and credits spent.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many to return (default 10, max 50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already provide the read-only safety profile (readOnlyHint=true) and workspace scoping (openWorldHint=false). The description adds behavioral context beyond that: newest-first ordering and the specific progress data included (requested, added, skipped, email verification, credits spent), which helps the agent set expectations without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-structured sentence that front-loads the core action and resource, then efficiently lists the return fields. No filler or redundancy; every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional parameter and no output schema, the description is quite complete: it conveys scope, ordering, and the fields returned. It doesn't describe pagination or total counts, but that is a minor gap given the tool's simplicity and the read-only annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% because the only parameter (limit) is fully documented with type, bounds, and default. The description does not mention the limit parameter or add any semantics beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Lists') and a clear resource ('workspace's Lead Finder adds'), then disambiguates what that means ('people added to campaigns from the contact database'). It also states ordering ('newest first') and enumerates the returned data fields, making it unmistakable from siblings like list_leads or get_lead_finder_import.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scoping ('workspace's') and the detailed field list imply when this tool is useful, but there is no explicit guidance on when to use it versus alternatives like get_lead_finder_import (singular) or search_lead_finder. No exclusions or alternative conditions are stated, so the guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_leadsList leadsARead-onlyInspect
Lists the leads in the workspace, newest first, 20 per page. Filter by exact email to look up a single lead, or by campaign to list only that campaign's leads. This is how you find the ID of a lead you added earlier.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| No | Filter by exact email address, which returns at most one lead | ||
| limit | No | Page size (default 20, max 200) | |
| campaignId | No | Filter to the leads in one campaign |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds genuinely useful behavior beyond that: newest-first ordering, 20-per-page pagination, and that an exact-email filter collapses the result to at most one lead.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what the tool does before the optional filters and the use case. No filler, and each sentence adds something.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden but only partially: it conveys ordering and page size and hints at the ID field, without describing the lead record shape. Adequate for a simple list tool, but not fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters. The description restates the email and campaign filters but adds no format or syntax detail beyond what the schema provides, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Lists the leads in the workspace") plus scope details (newest first, 20 per page). The purpose is unambiguous, but it never names or distinguishes itself from the obvious sibling get_lead, which also targets a single lead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use: filter by exact email to look up a single lead, by campaign for one campaign's leads, and to recover an ID. It stops short of naming an alternative tool or stating when not to use it, so it lacks explicit exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_repliesList repliesARead-onlyInspect
Lists inbound emails (replies received from prospects) in the workspace, newest first, 20 per page. Filter by AI response category, campaign, lead or a time floor. responseCategory is null while AI categorization is still pending. This tool only lists: to answer a reply, review its AI draft (list_reply_drafts), edit it if needed (update_reply_draft) and send it with send_reply_draft.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| limit | No | Page size (default 20, max 200) | |
| since | No | Only replies received at or after this time (RFC3339 or YYYY-MM-DD) | |
| leadId | No | Filter by lead ID | |
| category | No | Filter by AI response category | |
| campaignId | No | Filter by campaign ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description complements this by specifying behavioral details: 'newest first, 20 per page' and the null-state behavior of responseCategory during AI categorization. It also clarifies the tool's scope ('This tool only lists'), reinforcing the read-only nature. These details add value beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, with the core purpose and behavior stated in the first sentence, followed by filtering options, the null-state caveat, and the workflow routing. Every sentence contributes information; there is no fluff or redundancy. It is appropriately front-loaded and efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with six optional parameters, no output schema, and read-only annotations, the description covers all essential aspects: what is returned (inbound emails), ordering, pagination default, available filters, a special null state, and the intended follow-up workflow. Nothing critical is missing for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does mention the filtering dimensions ('AI response category, campaign, lead or a time floor') which maps to category, campaignId, leadId, and since parameters, but it does not add new meanings beyond what the schema already documents. The page and limit parameters are self-explanatory from the schema, so no extra explanation is needed. The description adds minimal semantic value over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('lists') and resource ('inbound emails (replies received from prospects)') with clear scoping ('in the workspace'). It differentiates itself from the workflow tools list_reply_drafts, update_reply_draft, and send_reply_draft by explicitly stating 'This tool only lists' and referencing the follow-up actions. This makes it unambiguous which tool to pick for listing replies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly frames the intended workflow: to answer a reply, first list replies, then review the AI draft, edit, and send. It names the specific sibling tools for each subsequent step, effectively instructing when to use this tool vs. the alternatives. The statement 'This tool only lists' serves as a clear exclusion, preventing misuse for drafting or sending.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_reply_draftsList AI reply draftsARead-onlyInspect
Lists AI-suggested reply drafts awaiting review, newest first, 20 per page. Each draft answers the inbound reply referenced by inReplyToEmailId. Nothing here has been sent: a draft stays a draft until a human sends it from the app or send_reply_draft is called. Edit one first with update_reply_draft if the wording needs work.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| limit | No | Page size (default 20, max 200) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false), but the description adds genuinely useful state semantics: drafts are unsent and remain drafts until a human acts, plus ordering and page-size behavior. This is real context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, purpose front-loaded, and each clause earns its place by clarifying the draft lifecycle or the follow-up tools. No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description could say more about what a draft object contains (fields returned), but it does convey the key semantics: what a draft is, its ordering, pagination, and the inReplyToEmailId relationship. Adequate for correct invocation, slightly light on return content.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents page and limit. The description still adds the sort order ('newest first') and reinforces the default page size ('20 per page'), which the schema only partially conveys. It adds modest but non-redundant meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Lists') and resource ('AI-suggested reply drafts awaiting review'), and frames the draft lifecycle so an agent distinguishes it from list_replies (inbound messages) and the send/update siblings. The scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the review context and names the downstream actions (send_reply_draft to send, update_reply_draft to edit), effectively telling the agent where this fits in the workflow. It does not explicitly state when to prefer this over list_replies, so it falls short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sender_emailsList sender emailsARead-onlyInspect
Lists the sending email accounts in the workspace, 20 per page, with their connection status, daily limits and health score. Filter to one campaign's senders, to connected or disconnected mailboxes only, or to one workspace on the account. healthScore is the account's health score: 0-100, higher is better, the share of its warm-up emails over the last 7 full days that landed in the inbox rather than spam. It moves daily as warm-up emails land in the inbox (up) or in spam (down), and is null while warm-up is off or before 20 warm-up emails were checked. Use it to pick the accounts to add to a campaign.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number, starting at 1 | |
| limit | No | Page size (default 20, max 200) | |
| spaceId | No | Only the sender emails of this workspace (from list_workspaces) | |
| campaignId | No | Only the sender emails attached to this campaign | |
| isConnected | No | true for mailboxes that are connected and able to send, false for disconnected ones |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description goes well beyond them by disclosing pagination (20 per page), the exact semantics of healthScore (0-100, share of warm-up emails landing in inbox over the last 7 full days), how it moves daily, and the null conditions (warm-up off or fewer than 20 emails checked). This is rich behavioral context an agent could not infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and filters before the longer healthScore explanation, and each sentence carries information. The healthScore block is lengthy but justified given there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains the key derived return field (healthScore) and its nullability, plus page size defaults. Other returned fields (daily limits, connection status) are named but not elaborated, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema, and the description largely restates the same filter intents (campaign, connected/disconnected, workspace). It adds only marginal value, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Lists the sending email accounts in the workspace') and immediately enumerates the returned fields (connection status, daily limits, health score). An agent can clearly separate this bulk-listing tool from the singular get_sender_email sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit filtering scenarios ('Filter to one campaign's senders, to connected or disconnected mailboxes only, or to one workspace on the account') and a concrete selection rationale ('Use it to pick the accounts to add to a campaign'). It lacks an explicit when-not-to-use or a named alternative such as get_sender_email for a single account.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_webhooksList webhooksARead-onlyInspect
Lists the webhook endpoints registered for the workspace, with their event type and status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered structurally. The description adds modest value by disclosing the returned fields (event type and status), but says nothing about ordering, result size, or empty-state behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every clause (scope, returned fields) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, naming the returned fields (event type, status) is enough for an agent to understand the result. Only minor gaps remain, such as whether the list is paginated or ordered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document and the baseline of 4 applies. Nothing in the description contradicts or under-specifies the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Lists) and resource (webhook endpoints) scoped to the workspace, and names what the results include (event type and status). It is clearly distinguishable from the write siblings create_webhook/update_webhook/delete_webhook, though it does not explicitly contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The listing purpose implies when to use it, but there is no explicit guidance about when this is preferable to alternatives or any prerequisites (e.g. workspace context). Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspacesList workspacesARead-onlyInspect
Lists the workspaces owned by the account behind the API key: the main workspace and its sub-workspaces, each with id, name, icon and creation date. isCurrent marks the workspace this key is bound to. Every other tool acts on the current workspace only; to work in another workspace, connect with that workspace's own key (create_workspace_api_key mints one).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds non-obvious behavioral context that annotations cannot: any other tool operates solely on the key-bound workspace, and cross-workspace work requires a different credential. It does not discuss pagination or list-size limits, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what is listed and what comes back, followed by the constraint and the alternative. No filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description names every returned field (id, name, icon, creation date, isCurrent), so the agent knows what it will get. For a zero-parameter read tool with annotations covering safety, nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description does not need to explain inputs, and it usefully documents the shape of the returned items instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Lists the workspaces owned by the account behind the API key') and specifies the exact scope (main workspace plus sub-workspaces). It even enumerates the returned fields and the meaning of isCurrent, so an agent can distinguish it from get_workspace_details and create_workspace without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the cross-cutting constraint that every other tool acts on the current workspace only, and names the alternative path (connect with that workspace's own key, minted by create_workspace_api_key). This is when-to-use and when-you-need-something-else guidance in one place.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_lead_meetingMark meeting booked on leadAIdempotentInspect
Records that a meeting was actually booked with this lead: sets meetingBookedAt to now, sets the lead's category to meeting_booked and fires the LeadCategoryUpdate webhook. Emailchaser never infers meetings from reply text or calendars, so this explicit mark is the only way a meeting is counted (in get_campaign_stats totals.meetings and get_outcomes_report). Idempotent: repeating the call keeps the original timestamp and fires no second webhook.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric lead ID (returned when the lead was created) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply idempotentHint/destructiveHint/openWorldHint booleans, while the description discloses the concrete side effects: timestamp assignment, category mutation, and an outbound LeadCategoryUpdate webhook. It also details the idempotency semantics that annotations merely flag — repeated calls keep the original timestamp and fire no second webhook.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core action and effects before the routing rationale and idempotency caveat. No filler or repetition of the tool title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema, the description covers effects, idempotency, webhook behavior, and where the result surfaces in reporting. An agent has everything needed to call it correctly and predict its consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'id' parameter is already documented in the schema as the numeric lead ID. The description only refers to 'this lead' generically, adding no format, constraint, or lookup guidance beyond the structured field, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Records that a meeting was actually booked with this lead') and enumerates the concrete effects (sets meetingBookedAt, sets category to meeting_booked, fires LeadCategoryUpdate webhook). It is clearly separable from siblings like update_lead or update_lead_category, which could otherwise be confused with this state transition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear condition for use and a strong rationale: Emailchaser never infers meetings from reply text or calendars, so this explicit mark is the only way a meeting counts. It also names the downstream consumers (get_campaign_stats totals.meetings, get_outcomes_report). It stops short of explicitly naming the sibling alternative unmark_lead_meeting or stating when not to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_lead_to_campaignMove lead to another campaignAInspect
Moves a lead from one campaign to another: the lead's unsent emails in the source campaign are deleted, and if the target campaign is running, emails for the lead are enqueued there. Use it to re-route a lead into a better-fitting campaign.
| Name | Required | Description | Default |
|---|---|---|---|
| leadId | Yes | The numeric lead ID (returned when the lead was created) | |
| sourceCampaignId | Yes | The campaign the lead is currently in | |
| targetCampaignId | Yes | The campaign to move the lead into |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses destructive behavior (unsent emails in the source campaign are deleted) and conditional behavior (emails enqueued only if the target campaign is running). This meaningfully exceeds the annotations, which only declare destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and its consequences. Efficient, though the second sentence is somewhat redundant with the first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutation tool with no output schema; the description covers prerequisites (source vs. target campaigns), side effects (deletion, enqueueing), and the conditional running-state requirement. An agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and all three parameters are documented in the schema; the description adds the semantic distinction that sourceCampaignId is the current campaign and targetCampaignId the destination, but does not add syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (moves a lead between campaigns) and describes the exact side effects. No sibling tool performs this re-routing action, so it is clearly distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage scenario ('re-route a lead into a better-fitting campaign'). Does not name alternatives or state when not to use it, but the purpose is narrow enough that this is nearly unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_autopilot_runAutopilot: pause a runAInspect
Suspends an autopilot run that is still in flight; resume later with resume_autopilot_run. Pausing never skips the approval gate. A completed or failed run cannot be paused.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric autopilot run ID (from list_autopilot_runs) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover openWorldHint=false and destructiveHint=false. The description adds meaningful behavioral context: pausing is reversible via resume_autopilot_run, it will not bypass the approval gate, and it cannot apply to completed or failed runs. It stops short of permissions, idempotency, or immediate-effect details, so 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, front-loaded with the action and scope, then the follow-up alternative, then the constraint. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter pause operation with no output schema, the description covers purpose, reversibility, approval-gate behavior, and invalid-state exclusion. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'id' parameter is already documented as the numeric autopilot run ID from list_autopilot_runs. The description adds no additional syntax, formatting, or source detail beyond what the schema provides, so the baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Suspends') and resource ('autopilot run'), plus the condition ('still in flight'). It distinguishes this from resume_autopilot_run and from a terminal-state run, so an agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the follow-up alternative ('resume later with resume_autopilot_run') and states the excluding condition ('A completed or failed run cannot be paused'). This gives clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_campaignPause campaignAIdempotentInspect
Pauses a running campaign and cancels all currently scheduled emails.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover idempotency and non-destructiveness; the description usefully discloses the side effect of cancelling scheduled emails, which annotations don't convey. It doesn't mention reversibility via a resume tool or scheduling data loss.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and the primary consequence. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple single-param mutation, but with no output schema and no annotation for reversibility, the description could note whether a paused campaign can be resumed and via which tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is fully documented in the schema. The description adds no parameter detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb + resource ('Pauses a running campaign') plus a meaningful effect ('cancels all currently scheduled emails') that clearly distinguishes it from siblings like delete_campaign and pause_autopilot_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description doesn't state when to use this over delete_campaign or resume flows, but the 'running campaign' scope implies the precondition. Usage is implied rather than explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_autopilotAutopilot: size a budget planAInspect
Turns a monthly budget in dollars into a concrete outbound setup: how many domains and mailboxes it buys, monthly sending volume, prospects per month, and costs (one-off setup, recurring, first month). creditsUsd, recurringUsd and firstMonthUsd value prospect credits at the list price, 1 credit ($0.033 at list) per revealed prospect, and a note says when that puts the monthly cost above the budget. Pure computation — nothing is created or charged. expectedMeetingsPerMonth is a planning estimate on pessimistic assumptions, not a promise. Use it to answer 'what does $X/month get me' before start_autopilot_run.
| Name | Required | Description | Default |
|---|---|---|---|
| budgetUsd | Yes | Total monthly budget in US dollars, inclusive of the platform subscription | |
| platformUsd | No | Subscription cost to reserve before sizing infrastructure. Defaults to 0 | |
| maxMailboxes | No | Cap on mailboxes regardless of budget. 0 means no cap |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare non-destructive and closed-world, but the description goes well beyond: it discloses that nothing is created or charged, that credits are valued at list price ($0.033) per revealed prospect, that a note fires when monthly cost exceeds budget, and that expectedMeetingsPerMonth is a pessimistic estimate rather than a promise. This is exactly the behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but it is front-loaded with the core purpose and every sentence earns its place: computation scope, pricing model, side-effect guarantee, estimate caveat, and sibling routing. Slightly dense, but complexity justifies the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return semantics — and it does so thoroughly, naming output fields (creditsUsd, recurringUsd, firstMonthUsd), what they value, the above-budget note behavior, and the caveat on expectedMeetingsPerMonth. For a pure-computation tool with fully documented inputs, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents budgetUsd, platformUsd, and maxMailboxes including defaults and constraints. The description references 'monthly budget in dollars' which maps to budgetUsd but adds no per-parameter meaning beyond the schema; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource pairing ('Turns a monthly budget in dollars into a concrete outbound setup') with a precise enumeration of what it computes. It explicitly differentiates itself from the sibling start_autopilot_run as a pre-flight planning step, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence gives an explicit usage directive — 'Use it to answer "what does $X/month get me" before start_autopilot_run' — naming both the trigger question and the sibling it precedes. The 'Pure computation' clause reinforces when to choose this over execution tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
purchase_creditsBuy credits (charges real money)AInspect
SPENDS REAL MONEY. Immediately charges the workspace's saved default card (the one behind the active subscription), off-session, with no confirmation step beyond this call — every successful call is a new charge. Buys 1,000 to 10,000 prospect credits at the tiered list price: $33/1k · $20/1k from 5k. The server prices the charge, so treat this as the list price and not a quote. Errors carry a machine-readable code: billing_required and payment_failed mean no money moved (fix billing or the card, then retry); credits_free means this workspace's credits are free, so nothing was charged and there is nothing to buy; temporarily_unavailable means the purchase stopped before any charge, so it is safe to retry shortly; purchase_incomplete means the card WAS charged but crediting failed — do NOT retry, support is already notified; purchase_unconfirmed means the outcome is unknown or the purchase completed without a balance to report — check list_credit_transactions for a stripe_topup entry before retrying. After a timeout or a dropped connection, check the same way before calling this again.
| Name | Required | Description | Default |
|---|---|---|---|
| credits | Yes | How many prospect credits to buy (1,000 to 10,000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the sparse annotations, disclosing that the charge happens off-session with no confirmation step, that the server sets the price, and that every successful call is a separate charge. It enumerates machine-readable error codes with precise money-movement semantics (e.g., 'purchase_incomplete' means the card WAS charged but crediting failed). This is exemplary behavioral disclosure for a high-stakes financial tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but every sentence earns its place for a money-spending operation: the critical warning is front-loaded, the purchase range and pricing are compact, and the error-code details are dense with actionable information. No filler is present, and the structure leads with the most safety-critical fact before moving to operational details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description tells the agent what outcomes to expect, how to interpret unusual cases, and how to verify completion via list_credit_transactions. It covers success paths implicitly, failure modes explicitly, and retry protocol thoroughly. For a mutating tool with real-world money movement, this is complete enough to call correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'credits' is fully described in the schema with the same range (1,000–10,000) and the same meaning ('prospect credits'). The tool description adds pricing-tier context and the note that the server prices the charge, which is useful but not essential to understanding the parameter itself. With 100% schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action: buying 1,000–10,000 prospect credits, while clearly flagging that this charges real money. It is strongly differentiated from siblings like get_credit_balance and list_credit_transactions by emphasizing that every successful call creates a new charge. The opening warning 'SPENDS REAL MONEY' leaves no ambiguity about the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use and when-not-to-use guidance, especially around retry behavior: do NOT retry on 'purchase_incomplete', safe to retry on 'temporarily_unavailable', and 'credits_free' means there is nothing to buy. It also names list_credit_transactions as the verification alternative for timeouts, dropped connections, or 'purchase_unconfirmed' outcomes, giving the agent a concrete routing path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_icp_audience_sizeRefresh ICP audience sizeAInspect
Re-counts how many prospects match an ICP's targeting criteria against the live data provider and caches the result on the profile. Sizing is free — it never spends credits; only revealing contact details is metered.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric ICP ID (from list_icps) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover openWorldHint and destructiveHint, so the description adds real value: it discloses that the result is cached on the profile and, crucially, that sizing is free / never spends credits while contact reveals are metered. It stops short of explaining cache staleness or the response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and immediately followed by the cost model, which is the most decision-relevant fact. No wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter refresh tool with no output schema, the description covers the action, the caching side effect, and the billing behavior. The main missing element is whether the call can be slow or rate-limited given it hits a live provider.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single id parameter is fully documented with its source (list_icps), so the schema does the heavy lifting. The description adds no extra meaning about the ID beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: re-counts prospects matching an ICP's targeting criteria against the live provider and caches the result. It is clearly distinguishable from listing/getting tools, though it never explicitly names the sibling get_audience_size it complements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly the tool is for refreshing a stale cached count, but it gives no explicit when-to-use vs get_audience_size or other sizing paths, and no prerequisites beyond requiring an ICP ID. Usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_blocklist_entriesRemove blocklist entriesADestructiveIdempotentInspect
Unblocks up to 1,000 domains or email addresses in one call, so campaigns can email them again. Pass values (domains or addresses, matched case-insensitively) and/or ids (entry IDs from list_blocklist_entries); a value that is not blocked is counted as skipped rather than failing the call. Scope defaults to "workspace" and never touches the account-wide list unless asked; removing an account-wide entry only works from the main workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | No | Blocklist entry IDs to remove (from list_blocklist_entries). Provide values, ids or both | |
| scope | No | Which list to remove from: "workspace" (default), "global", or "all" | |
| values | No | Email addresses and/or domains to unblock, e.g. ["competitor.com", "jane@acme.com"] (1 to 1,000 per request) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but the description adds non-obvious partial-success semantics: a value that is not blocked is counted as skipped rather than failing the call. It also discloses that scope defaults to workspace and that removing an account-wide entry only works from the main workspace, which is a constraint the annotations do not carry. Return shape and any rate limits remain undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and capacity, then the input contract, then the scoping constraint. No filler, and the most decision-relevant facts (bulk limit, skipped-not-failed, workspace default) come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description adequately implies the outcome (skipped counts, scope removed from). For a destructive bulk tool with annotation coverage for safety, this is nearly complete; only the exact response fields and any throughput limits are unaddressed, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by stating that values are matched case-insensitively, that values and ids can be combined, and by explaining the skipped-on-miss behavior for unmatched values, which the schema does not say.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (unblock domains/email addresses) with a concrete capacity bound (up to 1,000 per call) and an outcome (campaigns can email them again). It is clearly distinguishable from the singular remove_blocklist_entry and from list_blocklist_entries, which it references for sourcing IDs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when to pass values, ids, or both, and where IDs come from. It also clarifies the default scope behavior and the account-wide constraint. It stops short of explicitly telling the agent when to prefer the singular sibling remove_blocklist_entry over this bulk tool, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_blocklist_entryRemove one blocklist entryADestructiveIdempotentInspect
Removes one suppression entry by ID, so campaigns can email that domain or address again. This cannot be undone: the entry and the record of when it was added are gone. A sub-workspace key cannot remove an account-wide entry it merely inherits. To unblock many values at once use remove_blocklist_entries.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric blocklist entry ID (from list_blocklist_entries) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the description's added value is the irreversibility detail ('the entry and the record of when it was added are gone') and the scope/permission constraint. It does not contradict the annotation, but doesn't add significantly beyond what annotations plus the destruction note cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences front-loaded with the action, the irreversibility warning, and the sibling routing. Nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param destructive tool with no output schema, the description covers effect, irreversibility, permission limits, and the bulk alternative – everything an agent needs to call it correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the id parameter is fully documented, so the schema carries parameter meaning. The description reinforces 'by ID' and the effect, but adds no new parameter syntax. Slightly above baseline 3 for reinforcing the ID-based retrieval source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Removes one suppression entry by ID') with the effect stated ('so campaigns can email that domain or address again'). Clearly distinguishes itself from the singular/mass sibling remove_blocklist_entries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes to the alternative: 'To unblock many values at once use remove_blocklist_entries.' It also names a precondition for when this tool cannot be used (sub-workspace key inheriting an account-wide entry).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replace_campaign_sequenceReplace campaign sequenceAIdempotentInspect
Replaces a campaign's WHOLE sequence with the steps provided. This is not a partial edit: steps you leave out are removed, so send the full sequence every time. Sending the same payload twice leaves the same result, so a retry cannot duplicate steps. Editing a running campaign is allowed and changes future sends only; already-sent emails are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) | |
| steps | Yes | The full sequence, first email first. The first step must have delayDays 0 | |
| signature | No | Optional signature appended to every step |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already give idempotentHint=true and destructiveHint=false, but the description adds real substance: omitted steps are removed (whole-sequence overwrite semantics), the idempotency rationale ("same payload twice leaves the same result, so a retry cannot duplicate steps"), and the effect on a live campaign (future sends change, already-sent emails are untouched). This is exactly the behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, each carrying distinct information, with the destructive whole-replacement semantics front-loaded. No filler, though the idempotency and running-campaign points could be slightly compressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the critical decision points: it is a full overwrite, retries are safe, and live-campaign edits affect only future sends. Combined with 100%-covered schema and safety annotations, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents id, steps, signature, delayDays, variants, etc. in detail. The description's "send the full sequence" reiterates what the schema already states ("The full sequence, first email first"), so it adds little parameter-level meaning beyond the structured data. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (replaces) and resource (a campaign's WHOLE sequence) with emphasis on the whole-vs-partial distinction. An agent can immediately tell this apart from update_campaign (field edits) and get_campaign_sequence (read). No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"This is not a partial edit: steps you leave out are removed, so send the full sequence every time" gives clear usage context and implicitly steers away from incremental-edit assumptions. It stops short of naming the sibling alternative (e.g. update_campaign) or stating preconditions like the campaign needing to exist, so it is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_autopilot_runAutopilot: resume a runAInspect
Resumes a paused autopilot run. A run paused before approval goes back to awaiting approval — resuming can never skip the gate. A failed run resumes from the stage it failed at once the cause (e.g. insufficient credits) is fixed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric autopilot run ID (from list_autopilot_runs) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds meaningful behavior beyond the annotations: a run paused before approval returns to awaiting approval and resuming can never skip the gate, and a failed run resumes from its failure stage after the cause is fixed. Annotations only say openWorldHint=false and destructiveHint=false; the description supplies the state-machine semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and then the two state cases. Every sentence adds a distinct fact with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter state-transition tool with no output schema, the description covers what an agent needs to select and call it correctly. Return behavior and idempotency are unstated, but nothing required for invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single `id` parameter is fully documented as the numeric run ID from list_autopilot_runs. The description adds no further parameter detail, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Resumes a paused autopilot run"), which implicitly distinguishes it from start_autopilot_run, pause_autopilot_run, and kill_autopilot_run. It does not explicitly name those alternatives, so the differentiation is inferable rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real when-to-use context by describing the two resumable states (paused-before-approval and failed) and what resuming does in each. No explicit exclusions or named alternative tools, but the state conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_dfy_domainsSearch available sending domainsARead-onlyInspect
Checks availability and pricing of sending domains derived from a brand name (e.g. 'acme' yields acme-mail.com and similar). Only .com and .org are supported. Free to call — use it to pick domains and see exact prices before create_dfy_order.
| Name | Required | Description | Default |
|---|---|---|---|
| tlds | No | Comma-separated TLDs to check (default "com,org"; only com and org are supported) | |
| limit | No | Maximum suggestions to return | |
| query | Yes | Brand name to derive domain suggestions from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the description adds real value beyond them: the cost model ('Free to call'), the supported TLD restriction (.com/.org only), and the derivation behavior from a brand name. It does not describe pagination or suggestion ordering, but for a read-only search that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero filler, leading with what it does, then limits, then the call-to-action routing to create_dfy_order. Every sentence carries a distinct piece of information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still conveys the return content (availability and pricing) and the key constraint (.com/.org), which is enough for an agent to call and interpret it. It could say more about result shape or suggestion count, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents query, tlds, and limit in detail. The description only reinforces the brand-name input and the com/org constraint, adding no syntax or format detail beyond the schema, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (checks availability and pricing) plus resource (sending domains derived from a brand name) and gives a concrete derivation example ('acme' -> acme-mail.com). An agent can distinguish it from create_dfy_order and other domain/order tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('use it to pick domains and see exact prices before create_dfy_order') and names the sibling that follows it, plus notes it is free to call so there is no cost risk in exploring. Nothing about invocation timing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_lead_finderLead Finder: search the contact database (free)ARead-onlyInspect
Searches Emailchaser's B2B contact database (Lead Finder) with explicit filters and returns one page of matching people, masked: first name, last initial, title, seniority, job function, company name, industry, size and revenue band, and location. Email addresses, domains, LinkedIn URLs and phone numbers are never shown; contact details are revealed only when people are added to a campaign with add_lead_finder_prospects. Spends no credits, but every result row uses the account's daily browsing allowance (rowsLeftToday; 2,000 rows a day by default, 250 on a trial). An account may start 20 searches a minute. Searches that find nobody also draw on a bucket of 120 that refills one every 30 seconds; while it is empty every new search is refused for a moment (rate_limited, reason empty_searches, with retryAfterSeconds). The same page asked for again within 10 minutes costs nothing. total is the exact audience size only when totalIsExact is true; otherwise it is a lower bound (often 50,000). If status is running or totalStatus is pending, call get_lead_finder_search with the searchId: total, hasMore and maxPage are recomputed then. status failed means the search did not run, and error says why (rate_limited, timeout, provider_blocked, budget_exhausted, invalid_filters, expired or internal): it is not an empty audience, so try again later. Each result's ref is what add_lead_finder_prospects takes, valid for 60 minutes after its page was last shown. inWorkspace means Lead Finder added that person to this workspace and their lead is still there, so adding them again is skipped for free; someone who is a lead from another source shows false, and adding them is skipped for free too. A read-only key may use this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Results page, starting at 1 (default 1). The deepest page is 100 by default, and a search's maxPage says how deep that audience goes | |
| filters | Yes | Who to look for. Every field is optional but at least one include filter is required. Free-text fields take up to 50 values of up to 100 characters each (get_lead_finder_filters has the limits that apply); a value with commas is split into one value per comma, and blank values are ignored. A filter the API cannot use is refused with the field, value and reason named. | |
| pageSize | No | People per page: 25 (default) or 50 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description discloses detailed runtime behavior: no credit spend, daily browsing allowance, per-minute search limits, empty-search rate limiting with retryAfterSeconds, 10-minute page caching, total exactness semantics, status and error handling, ref expiry, and inWorkspace behavior. This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries a distinct behavioral fact with no filler or repetition. It is front-loaded with the core purpose before diving into limits and edge cases. A more structured layout would improve scannability, but the density is justified for this tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description covers the critical output fields (total, totalIsExact, status, error, ref, inWorkspace) and all relevant operational constraints (rate limits, daily allowance, cache, ref validity). An agent has enough context to call the tool correctly and handle the main follow-up paths.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents every parameter thoroughly, including filter constraints, defaults, and allowed values. The description adds no parameter-specific detail beyond the schema, so the baseline of 3 applies. It does not harm, but the schema is doing the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact action (search Emailchaser's B2B contact database), the resource (Lead Finder), and the output (one page of masked matching people). It also differentiates itself from add_lead_finder_prospects and get_lead_finder_search by explaining how contact details are revealed and how status is checked.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool and routes the agent to add_lead_finder_prospects for revealing contacts and to get_lead_finder_search when a search is still running. It does not explicitly contrast with get_lead_finder_filters or get_audience_size, but the intended use is clear enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_reply_draftSend AI reply draftAInspect
Sends an AI reply draft to the prospect through the conversation's sender mailbox; it is scheduled for delivery right away and cannot be recalled. The draft's CURRENT subject and body are what goes out, so read it first (list_reply_drafts or get_lead_conversation) and fix it with update_reply_draft if needed. Refused with 409 for anything that is not a sendable AI reply draft (already sent or scheduled emails, inbound replies, sequence templates) and with 422 when the conversation has no connected sender mailbox or no resolvable recipient.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric reply draft ID (from list_reply_drafts) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint/idempotentHint=false/destructiveHint=false; the description adds far more — delivery is immediate and irrevocable, the CURRENT subject/body (not a snapshot) is what is sent, and precise 409/422 failure semantics are enumerated. This is exactly the extra behavioral context annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and its irreversibility, then prerequisites and error codes. Three dense sentences, each carrying load, though the error-code enumeration makes it slightly heavy for a single-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema, it covers the action, the irreversibility, the required pre-read/draft-fix workflow, the sibling tools to use, and both refusal modes — nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'id' parameter already documents itself as the numeric reply draft ID. The description only reinforces the provenance ('from list_reply_drafts') without adding format or validation detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+mechanism: sends an AI reply draft through the conversation's sender mailbox, scheduled immediately. It is clearly distinguishable from siblings like update_reply_draft or list_reply_drafts, which it explicitly names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit pre-flight workflow ('read it first via list_reply_drafts or get_lead_conversation and fix it with update_reply_draft if needed') and names the exact refusal conditions (409 for non-sendable drafts, 422 for missing sender mailbox or recipient). When-to-use and when-it-will-fail are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_billing_profileSet workspace billing profileAIdempotentInspect
Creates or replaces the workspace's registrant and postal contact details. Do this before create_dfy_order: domain registration files these details with the registrar, so every field except addressLineTwo is required. Idempotent: sending the same body twice leaves the same state.
| Name | Required | Description | Default |
|---|---|---|---|
| city | Yes | City | |
| phone | Yes | Phone number without the country code | |
| state | Yes | State, province or region | |
| company | Yes | Company name | |
| country | Yes | ISO 3166-1 alpha-2 country code, e.g. US | |
| phoneCc | Yes | Telephone country calling code without the plus, e.g. 1 | |
| lastName | Yes | Registrant last name | |
| firstName | Yes | Registrant first name | |
| postalCode | Yes | Postal or ZIP code | |
| addressLineOne | Yes | Street address | |
| addressLineTwo | No | Suite, floor or unit (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=false, but the description adds genuine context: the operation replaces existing registrant/postal data and the body must be complete because registrar filing consumes it. It does not cover auth/permission requirements, which is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the operation and followed by the ordering prerequisite and idempotency guarantee. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a pure setter with no output schema and full schema coverage, the description supplies the sequencing rule, completeness requirement, and idempotency behavior an agent needs. Only permission/auth context and the safe way to read back the stored profile are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every one of the 11 fields, including formats like ISO 3166-1 alpha-2 and the phone country-code convention, is already documented. The description's only parameter note (everything except addressLineTwo is required) merely restates the schema's required array.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (creates/replaces the workspace's registrant and postal contact details) and distinguishes itself from read-side get_billing_profile by being the write counterpart. An agent knows exactly what state this mutates without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to call this before create_dfy_order and explains why (domain registration files these details with the registrar). That is a clear ordering prerequisite, though it does not name a when-not condition or an alternative (e.g., get_billing_profile for reading existing values).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_primary_icpSet primary ICPAIdempotentInspect
Promotes an Ideal Customer Profile to the workspace's active one, demoting any existing primary. At most one profile is primary per workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric ICP ID (from list_icps) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover openWorld, idempotency, and destructiveness. The description adds valuable behavior beyond them: demotion of any existing primary and the invariant that at most one profile is primary per workspace. It does not cover permissions or edge cases, but the side effect is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded and free of filler. The primary action comes first, followed by the key invariant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter mutation with full schema coverage and annotations, the description is nearly complete. It communicates purpose, side effect, and invariant, though it lacks usage context or error behavior details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter is fully documented in the schema. The description adds no parameter-level meaning, which matches the baseline of 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Promotes an Ideal Customer Profile to the workspace's active one.' The demotion side effect further clarifies scope and distinguishes it from read-only siblings like get_primary_icp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but gives no explicit when-to-use guidance, no prerequisites, and no alternatives. It does not say when to use set_primary_icp instead of update_icp or get_primary_icp.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
source_prospectsSource prospects into a campaign (spends credits)AInspect
SPENDS CREDITS. Queues a background job that searches Emailchaser's contact database with an Ideal Customer Profile's targeting criteria, reveals the matching people and adds them to the campaign as leads. Uses the workspace's primary ICP when icpId is omitted, and adds 50 prospects unless count says otherwise (max 500 per call). Every stored prospect costs the reveal price in credits (1 at the time of writing; the response reports creditsPerProspect and estimatedCredits, the most this batch can cost). Duplicates, blocklisted domains and contacts without an email address are filtered out before any credit is spent. Asynchronous: poll list_leads with campaignId to watch the prospects arrive. Calling again for the same campaign and profile pages deeper into the audience instead of re-revealing (and re-paying for) the same people. Check get_credit_balance and get_audience_size first.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | How many prospects to add in this call (default 50, max 500) | |
| icpId | No | The ICP whose targeting criteria drive the search (from list_icps). Defaults to the primary ICP | |
| campaignId | Yes | The campaign the sourced prospects are added to (from list_campaigns) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (openWorldHint, non-idempotent, non-destructive), the description discloses credit costs, the asynchronous background-job nature, duplicate/blocklist/email filtering before credits are spent, and response fields like creditsPerProspect and estimatedCredits. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but purposeful, opening with 'SPENDS CREDITS' and then covering the core function, cost, filtering, async behavior, and usage guidance. It is longer than minimal but every sentence contributes important context; slightly tighter organization would earn a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, a credit-costing side effect, and no output schema, the description covers required inputs, defaults, maximums, cost reporting, async completion via list_leads, and preflight checks. It does not detail error handling or full response shape, but provides enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema coverage is 100%, so the schema already documents campaignId, icpId, and count, including defaults and maximums. The description restates these defaults and adds related workflow hints (list_icps, list_campaigns, cost implications), but it does not add substantially deeper parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool sources prospects from Emailchaser's contact database using an ICP's targeting criteria and adds them to a campaign as leads. The verb 'source' plus the campaign resource is specific, but it does not explicitly name or differentiate from sibling tools like add_leads, which would push it to a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for use: check get_credit_balance and get_audience_size first, poll list_leads for async results, and on subsequent calls page deeper into the audience rather than re-revealing the same people. It does not explicitly state when to choose add_leads or another alternative, so it stops short of full exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_autopilot_runAutopilot: start a runAInspect
Starts an autopilot run for a company website: the AI builds an ICP, writes a sequence and checks that prospects match, then STOPS at an approval gate. Nothing is revealed, charged or sent and no sending accounts are bought until approve_autopilot_run is called; approval then reveals the first batch at 1 credit ($0.033 at list) per prospect, up to targetProspects and at most 2,000. The call is refused when the wallet cannot cover that batch. When budgetUsd is given, the sized budget plan is snapshotted on the run and returned in this response, the only response that carries it, with credits valued at list, and targetProspects defaults to the plan's monthly prospects. Without budgetUsd the run can never buy sending accounts, so the workspace needs one connected before approval or the run fails. Requires an active subscription (any plan).
| Name | Required | Description | Default |
|---|---|---|---|
| website | Yes | The company website to seed the run from, e.g. https://acme.com | |
| budgetUsd | No | Monthly budget in US dollars; sizes a plan that is snapshotted on the run | |
| replyMode | No | How inbound replies are handled: off, draft (default: AI drafts a reply for a human to send), approve, or auto. auto SENDS AI replies to interested prospects without review, so only pass it when the user has asked for that explicitly | |
| maxMailboxes | No | Cap on the budget plan's mailboxes. 0 means no cap | |
| targetProspects | No | How many prospects approval reveals first, at 1 credit ($0.033 at list) each (at most 2,000 in one batch). Defaults to the budget plan's monthly prospects when budgetUsd is given, else 50 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only openWorldHint and destructiveHint false in annotations, the description carries the behavioral burden and does so thoroughly: no reveal/charge/send until approval, credit pricing and caps, wallet-refusal behavior, budget plan snapshot and response exclusivity, and the no-sending-account implication when budgetUsd is absent. This is the type of side-effect disclosure that annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and approval gate are front-loaded, and the subsequent sentences each add a distinct operational constraint, cost, default, or prerequisite. It is dense rather than concise, and the long clauses could be easier to scan, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema and only generic annotations, the description covers the invocation flow, side effects, cost model, defaults, failure modes, and response-specific caveats. An agent has enough operational context to call the tool and to decide whether it should call approve_autopilot_run next.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is complete, so the baseline is 3, but the description adds cross-parameter semantics beyond the schema: targetProspects defaults derive from budgetUsd's plan, approval reveals cost 1 credit per prospect up to a 2,000 cap, and wallet coverage is assessed against that batch. These relationships materially affect how an agent should set parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Starts an autopilot run for a company website') and summarizes the run's lifecycle: build ICP, write sequence, check prospects, then stop at an approval gate. This makes it unmistakable how start_autopilot_run relates to approve_autopilot_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly places the tool in the approval-gate flow and names approve_autopilot_run as the next step that reveals and charges for the first batch. It also provides hard usage conditions: refusal when wallet cannot cover the batch and the requirement to have a sending account connected when budgetUsd is omitted. It does not explicitly contrast with plan_autopilot, but the start/approve lifecycle is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unmark_lead_meetingUnmark booked meeting on leadADestructiveIdempotentInspect
Removes the booked-meeting mark from a lead, e.g. when it was recorded by mistake or the meeting was canceled: clears meetingBookedAt and, when the category is still meeting_booked, reverts it to interested. The original booked time is lost. Idempotent: unmarking a lead with no meeting is a no-op.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric lead ID (returned when the lead was created) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=true, but the description goes well beyond them by disclosing the concrete destructive consequence ('the original booked time is lost') and the revert target of the category, plus confirming the no-op idempotency behavior. This is exactly the extra context the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then side effects, then the idempotency edge case. Every clause carries load — no filler and nothing repeated from the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param mutation tool with no output schema, the description covers the semantic effect, the destructive loss of data, and the idempotent no-op case. An agent has everything needed to decide and call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100% ('The numeric lead ID (returned when the lead was created)'), so the schema already carries the meaning. The description adds no additional parameter detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (removes the booked-meeting mark from a lead) and enumerates the exact state changes (clears meetingBookedAt, reverts category from meeting_booked to interested). This is clearly distinguishable from mark_lead_meeting and update_lead_category without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use context: 'when it was recorded by mistake or the meeting was canceled'. It implies the counterpart tool mark_lead_meeting but never names it explicitly, nor states when NOT to use this instead of updating category directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_blocklist_entryUpdate blocklist entryAIdempotentInspect
Changes one suppression entry's blocked value, its scope, or both; provide at least one. A value containing @ is stored as an email address, otherwise as a domain, so an entry can be converted between the two. Moving an entry to or from the account-wide list ("global") requires the main workspace's key and moves the entry onto the main workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric blocklist entry ID (from list_blocklist_entries) | |
| scope | No | "workspace" blocks for this workspace only; "global" blocks for every workspace on the account | |
| value | No | The new domain or email address to block |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover idempotency and non-destructiveness, and the description adds context those annotations cannot express: the @-based email/domain interpretation rule, the fact that a value can be converted between forms, and that a global-scope change both requires the main workspace key and relocates the entry onto the main workspace. That side effect and auth requirement are exactly what an agent needs before mutating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses, front-loaded with the core action before the interpretation rule and the global-scope caveat. Dense but every sentence carries information; the parenthetical enumeration of "global" is slightly redundant with the schema enum.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter mutation tool with no output schema, the description covers the operation, the required-at-least-one rule, value interpretation, and the global-scope side effect. It does not address failure behavior (e.g. an unknown id), a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning the schema lacks: the at-least-one-of constraint between value and scope, the @ rule governing how value is interpreted, and the behavioral consequence of scope="global". The enum values themselves are already documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource+scope: it changes one blocklist entry's value, scope, or both. A reader can immediately separate it from add_blocklist_entries and remove_blocklist_entry, which create and delete rather than mutate in place.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the operative constraint ("provide at least one" of value/scope) and the prerequisite for the global case (main workspace's key). It does not explicitly route the agent toward alternatives such as remove_blocklist_entry when the intent is deletion, so guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_campaignUpdate campaignAIdempotentInspect
Updates campaign properties such as name, emoji, timezone, sending limits, the minimum email account health score and deliverability settings. Only the provided fields are changed. A new minimum health score re-plans a running campaign at once.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) | |
| name | No | New campaign name | |
| emoji | No | Campaign emoji | |
| timezone | No | IANA timezone, e.g. America/New_York | |
| dailyLimit | No | Cap on the total emails (initial + follow-ups) the campaign may schedule per calendar day in its timezone. Omit for no campaign-level cap; per-mailbox limits still apply | |
| isEnabledLlm | No | Enable AI (LLM) features for this campaign | |
| minimumHealthScore | No | Minimum email account health score for this campaign, 1-100 (see healthScore in list_sender_emails). An account below it, or with no score yet, sends nothing in this campaign, first emails and follow-ups alike, until its score is back at or above it; its conversations wait for it and never move to another account. 0 means no minimum; omit it to leave the setting as it is | |
| allowNonBusinessEmails | No | Allow sending to free mailbox providers (gmail.com, etc.) | |
| isEnabledEmailVerifier | No | Verify lead emails before sending | |
| ignoreOutOfOfficeReplies | No | Do not stop follow-ups on out-of-office replies | |
| maximumTimeBetweenEmails | No | Maximum gap between two sends, in minutes | |
| minimumTimeBetweenEmails | No | Minimum gap between two sends, in minutes | |
| isEnabledCatchallValidated | No | Send to catch-all validated addresses | |
| isEnabledStopFollowUpsOnReply | No | Stop follow-ups to a lead once they reply | |
| isEnabledIgnoreHardBouncedLeads | No | Skip leads that previously hard-bounced | |
| isEnabledSkipLeadIfAlreadyExists | No | Skip leads that already exist in the workspace | |
| maximumSendingLimitPerSenderEmail | No | Daily sending limit per sender email account | |
| isEnabledStopFollowUpsForSameCompany | No | Company reply stop: once a lead replies (out-of-office and other automatic replies don't count), stop emailing the other leads at the same company in this campaign. Subdomains count as the same company; personal addresses like @gmail.com never do. On by default | |
| isEnabledStopFollowUpsAcrossCampaigns | No | Stop follow-ups to a lead in this campaign once they reply in any campaign | |
| isEnabledIgnoreLeadsWhoAlreadyResponded | No | Skip leads who already responded in another campaign | |
| maximumSendingLimitPerSenderEmailVariation | No | Random daily variation applied to the sending limit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so safety is partly covered. The description adds two genuinely useful behavioral facts beyond that: only supplied fields mutate (partial-update semantics), and setting a new minimum health score immediately re-plans a running campaign. That side-effect disclosure is the kind of thing an agent cannot infer from structured data alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, and the core mutation semantics are front-loaded before the field list and side-effect note. Slightly better routing information could have been added without cost, but nothing here is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 21-parameter mutation tool with no output schema, the description covers the essential patch semantics and the one notable side effect, and annotations carry the safety profile. It stops short of explaining what the tool returns or how it relates to the adjacent update_campaign_schedule tool, leaving a real gap in routing completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 21 parameters in detail (including the minimumHealthScore semantics that the description also gestures at). The description's field enumeration adds little beyond what the schema says; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Updates campaign properties') and enumerates the categories of fields affected (name, emoji, timezone, sending limits, health score, deliverability settings). It is clearly distinguishable from create_campaign, delete_campaign, and get_campaign, but it does not distinguish itself from the sibling update_campaign_schedule, whose territory (timezone, sending limits) it appears to overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The patch semantics are made explicit ('Only the provided fields are changed'), which is useful context for when to use this over a full-replace operation, and the re-planning note hints at when the tool has side effects. However, there is no guidance on when to prefer update_campaign vs update_campaign_schedule, nor any stated prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_campaign_scheduleUpdate campaign scheduleBIdempotentInspect
Updates a campaign's sending schedule: timezone, sending days, daily time window and frequency.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric campaign ID (from list_campaigns) | |
| sendAt | No | Exact send time for single-lead scheduled campaigns (ISO 8601) | |
| timezone | Yes | IANA timezone the schedule runs in, e.g. Europe/London | |
| endSchedule | No | Daily sending window end time, e.g. 17:00 | |
| daysSchedule | No | Days of the week to send on, e.g. ["monday", "tuesday", "wednesday"] | |
| everySchedule | No | Sending frequency interval | |
| startSchedule | No | Daily sending window start time, e.g. 09:00 | |
| maximumTimeBetweenEmails | No | Maximum gap between two sends, in minutes | |
| minimumTimeBetweenEmails | No | Minimum gap between two sends, in minutes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds a useful scoping signal that only schedule fields are touched, but says nothing about partial-update semantics (do omitted fields reset?) or whether it takes effect on an actively sending campaign.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence of roughly twenty words, front-loaded with the verb and resource, with no filler or redundancy. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nine-parameter mutation tool with no output schema and no annotations covering update semantics, the description covers the 'what fields' question but leaves the crucial merge/override behavior and prerequisite state unanswered. Adequate but with a real gap for a write tool of this size.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (timezone, daysSchedule, startSchedule, endSchedule, everySchedule, sendAt, min/max gap) is already documented with format hints like IANA names and ISO 8601. The description only restates the field groups and adds no format or constraint detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Updates) plus the exact resource (a campaign's sending schedule) and enumerates the affected facets: timezone, days, time window, frequency. That is enough to separate it from generic update_campaign, though it never names the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites (Does the campaign need to be paused? Does it apply to running campaigns?), and no routing to update_campaign for non-schedule fields. The agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_icpUpdate ICPAIdempotentInspect
Applies a partial edit to an Ideal Customer Profile: omitted fields are left unchanged, and sending an empty list clears that list. Any edit marks the profile as human-authored.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric ICP ID (from list_icps) | |
| name | No | New profile name | |
| titles | No | Job titles to target, e.g. ["CEO", "Head of Sales"] | |
| summary | No | New one-to-three-sentence summary | |
| keywords | No | Free-text keywords, e.g. ["b2b saas"] | |
| locations | No | Locations, e.g. ["United States"] | |
| industries | No | Industries, e.g. ["Software"] | |
| seniorities | No | Seniority levels, e.g. ["owner", "director"] | |
| companySizes | No | Company headcount ranges, e.g. ["11-50", "51-200"] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give idempotentHint=true, destructiveHint=false, openWorldHint=false. The description adds meaningful behavioral context beyond that: partial-update semantics ('omitted fields are left unchanged'), the destructive edge case that empty lists clear that list, and the side effect of marking the profile as human-authored. Good value-add; doesn't cover auth or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single tight sentence, front-loaded with the verb and resource, zero filler. Every clause earns its place (partial semantics, clear-list edge case, authorship side effect).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Nine params with 100% schema coverage and no output schema. The description covers the essential non-obvious semantics (partial update, empty-list clearing) that an agent must know to avoid unintended wipes or missed updates. Missing explicit guidance on when versus a full replace, but otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with per-field examples, so the schema already carries parameter meaning. The description adds the crucial update semantic (omitted vs empty-list) that the schema itself doesn't state, but no per-field guidance. Baseline 3 appropriate given the rich schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Applies a partial edit to an Ideal Customer Profile') and immediately defines the edit semantics (partial, empty-list-clears). Distinguishes it from create_icp/delete_icp/get_icp in the sibling list by the 'partial edit' framing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear on what it does (partial edit, omitted fields unchanged, empty list clears, marks as human-authored) but does not name alternatives or state when to prefer this over a full replace or when not to use it. Context is implied rather than explicitly routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_leadUpdate leadBIdempotentInspect
Updates a lead's contact details. Only the provided fields are changed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric lead ID (returned when the lead was created) | |
| No | Lead email address | ||
| phone | No | Phone number | |
| title | No | Job title | |
| company | No | Company name | |
| website | No | Company website URL | |
| lastName | No | Lead last name | |
| No | LinkedIn profile URL | ||
| firstName | No | Lead first name | |
| middleName | No | Lead middle name | |
| customVariables | No | Any attribute the named fields do not cover, e.g. { pageVisited: '/pricing' }. Each key becomes a merge tag usable in email copy as {key}. Keys match case-insensitively with spaces treated as underscores, so 'Page Visited' and 'page_visited' are the same tag. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact beyond that: 'Only the provided fields are changed,' clarifying partial-update semantics. It says nothing about permissions, error behavior, or whether clearing a field is possible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the partial-update constraint front-loaded as the second sentence. Nothing padding-like or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutation tool with a nested object and no output schema, the description is thin: it never states the required 'id' is the target lead, nor which field categories are editable. The exhaustive schema and the annotation set fill most of the gap, leaving the definition minimally viable rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including a rich explanation of customVariables and its merge-tag semantics, so the schema carries parameter meaning. The description adds no per-parameter detail beyond the generic 'contact details' framing, which is the baseline expectation at full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Updates a lead's contact details.' An agent can distinguish it from get_lead or delete_lead easily. However, it does not differentiate from closely-related siblings such as update_lead_category, move_lead_to_campaign, or mark_lead_meeting, which all mutate a lead in different ways.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus update_lead_category or move_lead_to_campaign, and no stated prerequisites or preconditions. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_lead_categoryUpdate lead categoryAIdempotentInspect
Sets or corrects a lead's category (tag), typically to fix an AI misclassification of a reply. Setting meeting_booked behaves exactly like mark_lead_meeting. Changing the category away from meeting_booked does NOT clear the booked-meeting mark; use unmark_lead_meeting for that. A real change fires the LeadCategoryUpdate webhook; setting the value the lead already has is a no-op.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric lead ID (returned when the lead was created) | |
| category | Yes | The category to set |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare idempotentHint/destructiveHint/openWorldHint; the description adds real behavior beyond them: a genuine change fires the LeadCategoryUpdate webhook, and re-setting the existing value is a no-op. It also discloses the important side effect that changing away from meeting_booked does NOT clear the booked-meeting mark, preventing a plausible misuse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then each following sentence carries a distinct, non-redundant fact (alternate tool equivalence, non-obvious side effect, webhook/no-op behavior). No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and no destructive permission flag, the description supplies everything an agent needs: side effects (webhook), idempotency semantics beyond the hint, and cross-tool consequences for the meeting_booked value. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning the schema lacks: the enum value meeting_booked is documented as behaving exactly like mark_lead_meeting, which clarifies semantics of a specific allowed value. It does not further explain the id parameter, but that is already covered by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb+resource ('Sets or corrects a lead's category (tag)') and immediately names the sibling tools it overlaps with (mark_lead_meeting, unmark_lead_meeting). An agent can distinguish it from update_lead, get_lead, and mark_lead_meeting without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit primary use case ('typically to fix an AI misclassification of a reply') and routes to alternatives with conditions: use mark_lead_meeting for the same effect as meeting_booked, and use unmark_lead_meeting to clear the booked-meeting mark. When-to-use and when-not-to-use are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_reply_draftEdit AI reply draftAIdempotentInspect
Edits the subject and/or body of an AI reply draft before it goes out. Provide at least one field; provided fields must be non-empty. This never changes the draft's status or sends anything: the draft stays a draft until a human sends it from the app or send_reply_draft is called. Emails that are not AI reply drafts (sent or scheduled emails, inbound replies, sequence templates) are refused.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric reply draft ID (from list_reply_drafts) | |
| body | No | New body | |
| subject | No | New subject line |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuine context beyond that: the draft's status is never mutated and unsupported email types (sent/scheduled emails, inbound replies, sequence templates) are rejected. It stops short of describing auth requirements or the response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler, and the core action plus its at-least-one-field constraint lead. The negative guarantees follow and each sentence earns its place by preventing a distinct misuse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter edit tool with full schema coverage and non-destructive annotations, the description covers the action, constraints, and side-effect profile adequately. With no output schema present, a brief note on the return value would close the only remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the individual fields and their minLength are already documented, which sets a baseline of 3. The description adds a real constraint not encoded in the schema: at least one of subject/body must be supplied, plus a non-empty requirement for whichever is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (edits) and resource (subject and/or body of an AI reply draft) with a clear scope boundary ('before it goes out'). It implicitly separates itself from send_reply_draft by explicitly noting it never sends, so an agent can route correctly without opening sibling schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions ('provide at least one field; provided fields must be non-empty'), states what the tool will never do, and names the alternative path (a human sending from the app, or send_reply_draft). It also enumerates what is refused, which is exactly the when-not guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_sender_emailUpdate sender emailAIdempotentInspect
Updates a connected sending email account's display name, signature or daily sending limits. Only the provided fields are changed. Warm-up settings are separate: use update_sender_email_warmup. The returned account includes its healthScore (see get_sender_email).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric sender email ID (from list_sender_emails) | |
| givenName | No | Sender first name | |
| signature | No | Email signature appended to sends from this mailbox | |
| familyName | No | Sender last name | |
| currentDailyLimit | No | Today's effective daily limit | |
| toggleGradualBuildUp | No | Ramp the daily limit up gradually (recommended for new mailboxes) | |
| maximumSendingsLimitPerDay | No | Maximum emails this mailbox may send per day | |
| minimumSendingsLimitPerDay | No | Minimum emails per day during gradual build-up |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so safety is covered. The description adds genuinely new behavior: partial-update semantics ('only the provided fields are changed') and that the returned account carries a healthScore, pointing to get_sender_email for interpretation. Auth requirements and limit-conflict behavior remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the partial-update rule, then sibling routing and return note. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an idempotent mutation with full schema coverage and no output schema, the description covers scope, partial-update behavior, sibling routing, and hints at the return payload via healthScore. Missing only edge details like validation failures or permission requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, making 3 the baseline, and the description goes slightly beyond by grouping parameters into display name, signature, and daily-limit categories and clarifying the partial-update contract. It adds real meaning but does not explain the interaction between currentDailyLimit and the build-up limits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (updates) and resource (a connected sending email account) plus the exact fields touched (display name, signature, daily sending limits). It also distinguishes itself from the similarly-named update_sender_email_warmup, so an agent can route without opening both schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes warm-up changes to update_sender_email_warmup and notes that only provided fields are modified, giving clear context for use. It stops short of naming the full when/when-not set (e.g., prerequisites like connected status or how it differs from connect_sender_email), so it is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_sender_email_warmupUpdate sender email warm-upAIdempotentInspect
Switches warm-up on or off and changes the warm-up ramp for up to 100 mailboxes in one call, applying the same settings to each. Only the settings provided change. Warm-up sends startLimit emails on the first day and adds increaseBy each sending day until it reaches capLimit a day; this is separate from campaign sending limits. A mailbox that has never warmed needs enabled: true to take any other setting, and is then enrolled with the defaults (2 a day, 2 more each day, up to 10 a day, weekdays only, UTC) plus the settings given. Raising capLimit on a mailbox that is already warming lifts today's target without restarting its ramp. Each mailbox succeeds or fails on its own: the result lists the updated mailboxes' settings, each with its healthScore (see get_sender_email_warmup), under updated, and any failures with the reason under failed.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | Sender email IDs to update (from list_sender_emails) | |
| enabled | No | true switches warm-up on, false switches it off | |
| capLimit | No | Most warm-up emails per day; the ramp stops here | |
| timezone | No | IANA timezone for the warm-up send window, e.g. America/New_York | |
| increaseBy | No | Warm-up emails added each sending day until the cap | |
| startLimit | No | Warm-up emails on the first day | |
| weekdaysOnly | No | Send warm-up emails Monday to Friday only | |
| enableReplies | No | Let warm-up recipients reply, adding reply signals |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (idempotent, non-destructive, closed-world), and the description adds substantial behavior beyond them: ramp mechanics (startLimit on day one, +increaseBy per sending day to capLimit), enrollment defaults for new mailboxes, the fact that raising capLimit lifts today's target without restarting the ramp, and per-mailbox independent success/failure semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the opening sentence states what the tool does before the mechanics. Every sentence carries operational information, though the description is long enough that some ramp detail could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 8 parameters, the description carries the load well: it describes the response shape (updated entries with settings and healthScore, failures with reason) and covers the ramp/enrollment edge cases an agent must know to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes further by explaining how startLimit, increaseBy and capLimit interact in the ramp, and by stating that other settings are inert until enabled: true. This adds meaning the per-field schema text alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (switch warm-up on/off, change ramp) and scopes it to up to 100 mailboxes in one call. It is clearly distinguishable from sibling update_sender_email, which handles general sender settings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong context: partial-update semantics ('only the settings provided change'), the prerequisite that a never-warmed mailbox needs enabled: true, and the pointer to get_sender_email_warmup for healthScore. It stops short of explicitly contrasting when to use this versus update_sender_email, so it is not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_webhookUpdate webhookAIdempotentInspect
Updates a registered webhook's name, URL, event type or enabled state. Only the provided fields are changed. Disabling a webhook stops deliveries without deleting it.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The numeric webhook ID (from list_webhooks) | |
| url | No | New HTTPS URL to receive events | |
| name | No | New display name | |
| type | No | New event type | |
| isEnabled | No | Enable or disable deliveries without deleting the webhook |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so safety is covered structurally. The description adds genuinely useful behavior beyond them: it is a partial update ('Only the provided fields are changed') and disabling has a side effect on delivery without removing the resource. It does not cover permissions or delivery-timing implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the mutation scope, with zero filler. Each sentence carries distinct information (what changes vs. the disable-without-delete behavior).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter update tool with full schema coverage and no output schema or contradictions, the description supplies the extra semantics an agent needs: partial update and non-destructive disable. Only a note on required 'id' provenance or auth expectations is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (id, url, name, type enum, isEnabled) is already documented in the schema. The description restates which fields are updatable but adds no format, constraint or default detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Updates') and resource ('a registered webhook'), and even enumerates the mutable fields (name, URL, event type, enabled state). That is enough for an agent to distinguish it from create_webhook, delete_webhook and list_webhooks, though it never explicitly references a sibling to reinforce the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operational context: only supplied fields change (partial/PATCH semantics), and disabling stops deliveries without deletion — effectively steering the agent toward isEnabled=false instead of delete_webhook. It stops short of naming that alternative or stating prerequisites, so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
copilot_launch3 fields changed- changed
Input schema / properties / salesNavSearchUrl / descriptionPrevious value: -"A LinkedIn Sales Navigator people-search URL to source leads from"New value: +"Optional LinkedIn Sales Navigator people-search URL to source leads from. Leave it out to create the draft with its emails only. Required when salesNavData is given." - removed
Input schema / properties / salesNavSearchUrl / minLengthRemoved value: -1 - changed
Input schema / requiredPrevious value: -[ - "name", - "salesNavSearchUrl", - "sequence" -]New value: +[ + "name", + "sequence" +]
2 tool updates
- Changed
create_campaign1 field changed- added
Input schema / properties / minimumHealthScoreAdded value: +{ + "description": "Minimum email account health score for this campaign, 1-100 (see healthScore in list_sender_emails). An account below it, or with no score yet, sends nothing in this campaign, first emails and follow-ups alike, until its score is back at or above it; its conversations wait for it and never move to another account. 0 means no minimum; omit it to leave the setting as it is", + "maximum": 100, + "minimum": 0, + "type": "integer" +}
- Changed
update_campaign1 field changed- added
Input schema / properties / minimumHealthScoreAdded value: +{ + "description": "Minimum email account health score for this campaign, 1-100 (see healthScore in list_sender_emails). An account below it, or with no score yet, sends nothing in this campaign, first emails and follow-ups alike, until its score is back at or above it; its conversations wait for it and never move to another account. 0 means no minimum; omit it to leave the setting as it is", + "maximum": 100, + "minimum": 0, + "type": "integer" +}
2 tool updates
- Changed
create_campaign1 field changed- changed
Input schema / properties / isEnabledStopFollowUpsForSameCompany / descriptionPrevious value: -"Stop follow-ups to a company once someone there replies"New value: +"Company reply stop: once a lead replies (out-of-office and other automatic replies don't count), stop emailing the other leads at the same company in this campaign. Subdomains count as the same company; personal addresses like @gmail.com never do. On by default"
- Changed
update_campaign1 field changed- changed
Input schema / properties / isEnabledStopFollowUpsForSameCompany / descriptionPrevious value: -"Stop follow-ups to a company once someone there replies"New value: +"Company reply stop: once a lead replies (out-of-office and other automatic replies don't count), stop emailing the other leads at the same company in this campaign. Subdomains count as the same company; personal addresses like @gmail.com never do. On by default"
1 tool update
- Changed
list_replies1 field changed- changed
Input schema / properties / category / enumPrevious value: -[ - "interested", - "not_interested", - "wrong_person", - "bounced", - "out_of_office", - "delivery_incomplete", - "dmarc_report", - "mixmax", - "warmup_email", - "unsubscribe" -]New value: +[ + "interested", + "not_interested", + "wrong_person", + "bounced", + "out_of_office", + "delivery_incomplete", + "dmarc_report", + "mixmax", + "warmup_email", + "unsubscribe", + "newsletter" +]
1 tool update
- Changed
create_dfy_order2 fields changed- changed
Input schema / properties / domains / descriptionPrevious value: -"Domains to purchase"New value: +"Domains to purchase. Each one needs at least one mailbox in mailboxes with the same domainName." - changed
Input schema / properties / mailboxes / descriptionPrevious value: -"Mailboxes to provision"New value: +"Mailboxes to provision: at least one on every domain in domains. A mailbox can also go on a domain from an earlier order."
6 tool updates
- Added
cancel_email_verification_job - Added
create_email_verification_job - Added
get_email_verification_job - Added
get_email_verification_rates - Added
list_email_verification_jobs - Added
list_email_verification_results
8 tool updates
- Added
add_lead_finder_prospects - Added
cancel_lead_finder_import - Changed
create_workspace_api_key1 field changed- changed
Input schema / properties / readOnly / descriptionPrevious value: -"Mint a read-only key (list/get tools plus get_audience_size, nothing else). Defaults to false, a read & write key"New value: +"Mint a read-only key. It can use the read (list and get) tools plus confirm_connection, get_audience_size and search_lead_finder, so it can read, size audiences and browse Lead Finder (which still uses the account's daily browsing allowance), but it cannot spend credits or change anything else. Defaults to false, a read & write key"
- Added
get_lead_finder_filters - Added
get_lead_finder_import - Added
get_lead_finder_search - Added
list_lead_finder_imports - Added
search_lead_finder
1 tool update
- Changed
start_autopilot_run1 field changed- changed
Input schema / properties / targetProspects / descriptionPrevious value: -"How many prospects approval reveals first, at 5 credits ($0.40 at list) each (at most 2,000 in one batch). Defaults to the budget plan's monthly prospects when budgetUsd is given, else 50"New value: +"How many prospects approval reveals first, at 1 credit ($0.033 at list) each (at most 2,000 in one batch). Defaults to the budget plan's monthly prospects when budgetUsd is given, else 50"
1 tool update
- Changed
start_autopilot_run1 field changed- changed
Input schema / properties / targetProspects / descriptionPrevious value: -"How many prospects to reveal, at 5 credits ($0.40 at list) each, before asking for approval (at most 2,000 in one batch). Defaults to the budget plan's monthly prospects when budgetUsd is given, else 50"New value: +"How many prospects approval reveals first, at 5 credits ($0.40 at list) each (at most 2,000 in one batch). Defaults to the budget plan's monthly prospects when budgetUsd is given, else 50"
1 tool update
- Changed
start_autopilot_run1 field changed- changed
Input schema / properties / targetProspects / descriptionPrevious value: -"How many prospects to source before asking for approval. Defaults to the budget plan's capacity when budgetUsd is given, else the runner default"New value: +"How many prospects to reveal, at 5 credits ($0.40 at list) each, before asking for approval (at most 2,000 in one batch). Defaults to the budget plan's monthly prospects when budgetUsd is given, else 50"
1 tool update
- Changed
purchase_credits2 fields changed- changed
Input schema / properties / credits / descriptionPrevious value: -"How many prospect credits to buy (1,000 to 100,000)"New value: +"How many prospect credits to buy (1,000 to 10,000)" - changed
Input schema / properties / credits / maximumPrevious value: -100000New value: +10000
88 tool updates
- First observed
add_blocklist_entries - First observed
add_leads - First observed
approve_autopilot_run - First observed
attach_campaign_sender_emails - First observed
confirm_connection - First observed
connect_sender_email - First observed
copilot_launch - First observed
copilot_plan - First observed
create_campaign - First observed
create_dfy_order - First observed
create_icp - First observed
create_webhook - First observed
create_workspace - First observed
create_workspace_api_key - First observed
delete_campaign - First observed
delete_icp - First observed
delete_lead - First observed
delete_webhook - First observed
detach_campaign_sender_email - First observed
get_audience_size - First observed
get_autopilot_run - First observed
get_billing_profile - First observed
get_blocklist_entry - First observed
get_campaign - First observed
get_campaign_sequence - First observed
get_campaign_stats - First observed
get_credit_balance - First observed
get_deliverability_insights - First observed
get_dfy_order - First observed
get_icp - First observed
get_inbox_placement_run - First observed
get_inbox_placement_stats_by_date - First observed
get_lead - First observed
get_lead_conversation - First observed
get_outcomes_report - First observed
get_primary_icp - First observed
get_sender_email - First observed
get_sender_email_dns - First observed
get_sender_email_warmup - First observed
get_sender_reputation - First observed
get_workspace_details - First observed
import_instantly - First observed
kill_autopilot_run - First observed
launch_campaign - First observed
list_autopilot_runs - First observed
list_blocklist_entries - First observed
list_campaigns - First observed
list_credit_transactions - First observed
list_dfy_orders - First observed
list_icps - First observed
list_inbox_placement_run_results - First observed
list_inbox_placement_runs - First observed
list_inbox_placement_tests - First observed
list_instantly_accounts - First observed
list_leads - First observed
list_replies - First observed
list_reply_drafts - First observed
list_sender_emails - First observed
list_webhooks - First observed
list_workspaces - First observed
mark_lead_meeting - First observed
move_lead_to_campaign - First observed
pause_autopilot_run - First observed
pause_campaign - First observed
plan_autopilot - First observed
purchase_credits - First observed
refresh_icp_audience_size - First observed
remove_blocklist_entries - First observed
remove_blocklist_entry - First observed
replace_campaign_sequence - First observed
resume_autopilot_run - First observed
search_dfy_domains - First observed
send_reply_draft - First observed
set_billing_profile - First observed
set_primary_icp - First observed
source_prospects - First observed
start_autopilot_run - First observed
unmark_lead_meeting - First observed
update_blocklist_entry - First observed
update_campaign - First observed
update_campaign_schedule - First observed
update_icp - First observed
update_lead - First observed
update_lead_category - First observed
update_reply_draft - First observed
update_sender_email - First observed
update_sender_email_warmup - First observed
update_webhook
Related MCP Connectors
Cold email API for AI agents: connect mailboxes, warm up, run campaigns and handle replies.
Run B2B outreach from your AI agent: 250+ tools for campaigns, leads, LinkedIn and email workflows.
Email, WhatsApp and Telegram for AI agents: send, campaigns, automations, contacts, agent inboxes.
Cold email infrastructure — campaigns, prospects, mailbox health and replies via Claude.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to programmatically manage email outreach campaigns, leads, domains, senders, and webhook events, as well as send emails, through the Model Context Protocol over HTTP.-
- AlicenseNot gradedqualityDmaintenanceGenerate hyper-personalized cold email sequences per lead via AI. List campaigns, generate emails for individual leads or batches, and export results as CSV or JSON.2MIT
- AlicenseAqualityCmaintenanceProvides AI agents with tools for structuring and auditing cold outreach messages, building follow-up sequences, handling replies, and detecting red flags.748 npmMIT

PumpGTM MCP serverofficial
AlicenseNot gradedqualityBmaintenanceEnables AI agents to find intent-signaling buyers and run LinkedIn, email, and X outreach from connected accounts while respecting platform limits. Every reply is handed back to a human for approval or decision.1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.