Skip to main content
Glama

agent-cold-email

Server Details

Coldrig — cold-email infra run by your agent: 28 MCP tools, live sending, free sandbox, from $99/mo.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Uptime
100.0% over 51 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Server Listing
agent-cold-email

TDQS

A4.2/5.0

Scored across 30 tools

Disambiguation4/5

Most tools target distinct resources or actions, and descriptions explicitly redirect agents to alternatives (e.g. metrics vs campaign_results vs list_campaigns vs account). Overlap remains among account/activity/metrics/infrastructure_status and inbox/thread/list_messages, but the descriptions make the boundaries workable.

Naming Consistency4/5

All names use lowercase snake_case with common prefixes like get_, list_, configure_, launch_, setup_, and update_. A few standalone nouns/verbs (account, inbox, thread, metrics, mark, pause, reply) deviate from strict verb_noun, but the convention is still mostly predictable.

Tool Count3/5

At 30 tools, this is heavy compared with the 3-15 sweet spot, though the platform's scope covers campaigns, leads, infrastructure, billing, BYO domains, webhooks, dashboards, and support. Most tools appear to earn their place, but several peripheral admin/config tools make the set borderline over-scoped.

Completeness4/5

The surface covers the core cold-email lifecycle well: infrastructure setup, campaigns, leads, messages, replies, suppression, metrics, activity, billing, BYO domains, dashboards, and webhooks. Gaps exist, notably no campaign resume/update/delete and no un-suppress, but they are either intentional irreversible actions or can be worked around.

Available Tools

30 tools
accountA
Read-only
Inspect

Account overview: brand, plan, status, vertical ('crowdfunding' or null — set at signup), billingState, activationState, resource counts, usageCents, quota, deliverability (loop state: paused/throttled mailboxes, burning domains, auto-replacements, recentActions[]), and teardown (reclaim summary once canceled, else null). Billing is per-provisioned-mailbox on the live provisioned count: $49 platform + per mailbox, graduated: 1-5 at $10, 6-20 at $8, 21+ at $7; 5-mailbox minimum ($99 at 5); the billed quantity tracks the real provisioned count (deprovision lowers it). activationState is the HONEST send state — trust it over 'sent' counts: 'active' = real sending live; 'pending_provisioning' = paid but infrastructure still being armed, sends shown are sandbox previews that DON'T leave; 'capacity_pending' = provisioning held at a spend limit; 'screening_hold' = account under review; 'sandbox' = demo/free. Use metrics for counts, infrastructure_status for per-mailbox health.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint, and the description carries the real behavioral weight: the exact billing formula (graduated $10/$8/$7 with a $99 five-mailbox minimum, per provisioned mailbox, deprovision lowers billed quantity), what each activationState value actually means, and the critical warning that 'pending_provisioning' sends are sandbox previews that 'DON'T leave'. That is unusually useful disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the field inventory and tight throughout; each clause adds a distinct fact. It is a single very dense paragraph, and the billing and state-semantics sections might read better as separate blocks, but there is little genuine filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only, no-output-schema tool, the description supplies everything an agent needs: the full field set, null semantics ('vertical' null, teardown null unless canceled), billing math, and the state machine that determines whether sending is real.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to document beyond the schema (baseline 4). It correctly spends its words on return-field semantics instead of inventing parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (account) and enumerates the exact fields returned (brand, plan, status, vertical, billingState, activationState, counts, usage, quota, deliverability, teardown). The retrieval verb is only implied by 'overview', but sibling differentiation is explicit ('Use metrics for counts, infrastructure_status for per-mailbox health'), so an agent can place it correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes to two siblings with the condition that selects them ('Use metrics for counts, infrastructure_status for per-mailbox health') and tells the agent which field to trust ('trust activationState over sent counts'). There is no explicit 'when not to use this' statement, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ack_messageAInspect

Acknowledge a message by id (from list_messages or infrastructure_status's messages[]) — sets it read so it stops surfacing as unacked. THIS IS THE ONLY THING THAT SETS readAt: reading or listing messages never marks them, so a message stays unacked until you call this explicitly. Idempotent: acking an already-acked id returns success with no second effect, never an error. 404 if the id doesn't exist for this tenant.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageIdYesThe message id, from list_messages or infrastructure_status's messages[].

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation only provides destructiveHint=false, so the description carries the burden of explaining the tool's side effect. It does so well: it sets readAt, stops the message surfacing as unacked, is idempotent, and returns 404 for invalid IDs. No contradiction exists, though it stops short of discussing any broader state or rate-limit implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is slightly longer than strictly necessary, but every sentence adds actionable information: purpose, exclusivity vs reading/listing, idempotence, and error behavior. The core how-to is front-loaded and the extra details support reliable calling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter ack operation with no output schema, this description is complete. It tells the agent what the effect is, confirms repeated calls are harmless, gives the error condition, and notes the ID provenance. No essential behavioral information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description fully covers the only parameter, messageId, including where the ID originates. The description repeats that same source rather than adding new semantic detail, but because schema coverage is 100%, the parameter is already well documented and no further compensation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Acknowledge a message by id'), the resource ('message'), and the effect ('sets it read so it stops surfacing as unacked'). It also distinguishes this tool from message-reading/list operations by emphasizing that it is the only thing that sets readAt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent when to use this tool: after getting message IDs from list_messages or infrastructure_status, when intentionally marking a message as read. It also clarifies when not to expect side effects via idempotence, and what happens with a nonexistent ID. This gives clear usage context without assuming alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

activityA
Read-only
Inspect

Unified activity feed: campaign events (sent/reply/bounce/...) merged with deliverability loop actions (pause/throttle/replace-domain). Cursor-paginated → { items[], nextCursor }; each item { id, kind:'event'|'deliverability', label, ts, target, detail }. Filters: kind, limit (default 50, max 200). Use inbox for replies only.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
limitNo
cursorNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true. Description adds behavioral context: cursor-pagination, merged feed, response structure { items[], nextCursor }, and item fields. No contradictions. Adds value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise single paragraph front-loaded with purpose, then details on pagination, item structure, and filters. Every sentence earns its place. Slightly dense but efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully specifies the return structure, pagination, filters, and default/max values. Covers all needed context for a read-only feed tool with simple parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%. Description mentions filters (kind, limit) and defaults/max (50, 200) already in schema. Adds meaning for kind enum values ('event' vs 'deliverability') and cursor usage. Does not fully compensate for missing schema descriptions but adds moderate value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a 'Unified activity feed' merging campaign events and deliverability actions, and distinguishes itself from sibling 'inbox' tool by noting 'Use inbox for replies only.' The specificity of verb and resource is high.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to use this tool vs inbox for replies. Provides filter guidance (kind, limit, cursor). Does not mention other sibling tools like get_dashboard, but the context is sufficient for typical usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

campaign_resultsA
Read-only
Inspect

Outcome counts for ONE campaign. Input: campaignId (from launch_campaign). Returns { campaignId, sent, reply, bounce, complaint, unsubscribe, failed, soft_bounce } — bounce = HARD only, soft_bounce separate, opens not tracked. 404 if unknown. Use metrics for account-wide totals, list_campaigns for every campaign at once.

ParametersJSON Schema
NameRequiredDescriptionDefault
campaignIdYesThe campaign id returned by launch_campaign.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds significant detail beyond readOnlyHint annotation: clarifies bounce is HARD only, opens not tracked, and 404 response. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three efficient sentences covering purpose, input/output format, and explicit usage guidance. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Fully describes a simple read operation: input, output fields, error case, and relationship to sibling tools. No output schema needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with description for campaignId. Description restates that input comes from launch_campaign but adds no new meaning beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states 'Outcome counts for ONE campaign' - specific verb and resource. Clearly distinguishes from sibling tools by naming 'metrics' and 'list_campaigns' as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use this tool (single campaign results) and alternatives (metrics for account-wide, list_campaigns for all). Also specifies input source and 404 handling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

configure_byo_domainA
Destructive
Inspect

Register or advance a BYO domain/mailbox intake. action = register (needs domain + domainRelationship: fresh_standalone|subdomain_of_primary|is_primary — runs pre-flight scan + abuse gate + reputation ladder, returns starting byoStatus) | poll_dns (needs id — re-checks DNS delegation, advances pending_dns -> active, or -> abandoned after 7 idle days) | acknowledge_consent (needs id + acknowledged:true — required before a primary domain can proceed past pending_consent; documents informed consent, does not remove exposure) | request_managed_mailboxes (needs id + count — platform-provisioned mailboxes on an ALREADY-ACTIVE domain; every response carries a billing projection { provisionedAfter, projectedMonthlyCents, formula }; quoteOnly:true previews without provisioning) | connect_mailbox (needs id + email + transport — declares an EXISTING OAuth/SMTP+IMAP connection, bypassing provisioning; transport is smtp | gmail_api | ms_graph, each with its own credential fields).

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRequired for poll_dns/acknowledge_consent/request_managed_mailboxes/connect_mailbox — the domainId from register.
countNoRequired for request_managed_mailboxes — how many platform-provisioned mailboxes to attach.
emailNoRequired for connect_mailbox — the existing mailbox address.
actionYes
domainNoRequired for register.
quoteOnlyNoOptional for request_managed_mailboxes — true previews the new mailbox count + projected monthly price WITHOUT provisioning (SPEC §18 quote-before-add).
transportNoRequired for connect_mailbox — { kind: 'smtp', host, port, secure, user, pass } | { kind: 'gmail_api', clientId, clientSecret, refreshToken } | { kind: 'ms_graph', mode: 'delegated'|'app_only', tenantId, clientId, clientSecret, refreshToken? }.
personaSlugNoOptional for request_managed_mailboxes — defaults to a slug of the domain.
acknowledgedNoRequired (must be true) for acknowledge_consent — SPEC.md §20.4's separate, unbundled risk acknowledgment.
domainRelationshipNoRequired for register: fresh_standalone | subdomain_of_primary | is_primary.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only indicate destructiveHint=true and a title. The description meaningfully adds context: pre-flight scan + abuse gate + reputation ladder on register, DNS re-checks and 7-day idle timeout on poll_dns, consent documentation that doesn't remove exposure, billing projection details, and bypassing provisioning. This goes well beyond the single annotation, though it could state rate limits or idempotency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence with pipe-delimited action blocks. It's front-loaded with the tool's overall purpose and then enumerates each action's needs, but it's quite long and slightly convoluted with parentheses and semicolons.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a complex multi-action tool with 10 params, 90% schema coverage, no output schema, and only a destructive hint annotation, the description covers a lot of ground: action semantics, prerequisites, state transitions, billing, and consent nuances. It's largely complete, though it could explicitly note that register returns a starting byoStatus and the overall domain lifecycle progression.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 90% and quite verbose, so baseline is 3, but the description adds per-action conditions (e.g., which params are needed for each action, transport credential field variation) that tie parameters to actions in a way the schema's individual param descriptions do not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific multi-action intake tool for BYO domains/mailboxes and enumerates all five actions with their semantics. It's clear what the tool does, though it doesn't explicitly contrast with the sibling get_byo_domains which handles the read side.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Effectively outlines prerequisites: register needs domain+relationship, poll_dns needs id, acknowledge_consent needs id+acknowledged, request_managed_mailboxes needs an ALREADY-ACTIVE domain, connect_mailbox bypasses provisioning. It provides clear when-to-use per action but doesn't state when NOT to use it or point to alternatives like get_byo_domains for read-only status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

configure_dashboardA
Destructive
Inspect

Write a saved dashboard view. action = create (needs name+layout) | update (needs id+rev+layout; optional name renames) | promote (id → default) | delete (id). update is rev-CAS: a stale rev returns { currentRev, currentLayout } to rebase and retry. Optional note. Read the current rev+layout via get_dashboard first.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRequired for update/promote/delete.
revNoRequired for update — the rev this edit is based on; stale vs. the view's CURRENT rev returns a structured conflict with currentRev/currentLayout to rebase onto.
nameNoRequired for create. Optional for update — pass it to rename the view; omit to leave the name unchanged.
noteNoOptional human-readable note recorded alongside this edit (edited_by_note).
actionYes
layoutNoRequired for create/update.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the 'destructiveHint' annotation, the description discloses the compare-and-swap behavior for update: a stale rev returns '{ currentRev, currentLayout }' to rebase and retry. Does not mention rate limits or auth, but the conflict resolution is a key behavioral trait.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Efficiently written in one paragraph with no filler. Front-loaded with the main purpose. Could be improved with bullet points for the actions, but is very compact and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description explains the conflict response for stale revisions, which is critical. No description of success return values, but the tool's complexity and the CAS behavior are well-covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Adds action-specific required parameters (e.g., create needs name+layout, update needs id+rev+layout) and explains the rev parameter's CAS semantics. The schema covers 83% of parameters, but the description reinforces and provides context beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Begins with 'Write a saved dashboard view,' clearly identifying the verb and resource. Lists four specific actions (create, update, promote, delete), and distinguishes from sibling 'get_dashboard' which is read-only.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to 'Read the current rev+layout via get_dashboard first' before update. Explains when each action is appropriate (create needs name+layout, update needs id+rev+layout, etc.) and details the CAS conflict handling for update.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

configure_webhookA
Destructive
Inspect

Manage an outbound webhook subscription. action = create (needs url + eventTypes: reply|bounce|soft_bounce|complaint|unsubscribe; optional secret/active) | update (needs id + one changed field; active:true re-enables an auto-disabled one, active:false pauses; secret rotates) | delete (needs id). create/rotate return the HMAC signing secret ONCE. URLs must be https to a public host (private/metadata IPs rejected). Deliveries are signed X-Coldrig-Signature: sha256=HMAC-SHA256(secret, raw body).

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRequired for update/delete.
urlNoRequired for create. HTTPS endpoint; private/link-local/metadata IPs are rejected.
noteNoIgnored placeholder for symmetry; webhooks record no provenance note.
actionYes
activeNoOptional. On update, active:true re-enables an auto-disabled subscription; active:false pauses delivery.
secretNoOptional signing secret (>=16 chars). Omit on create to have one generated; pass on update to rotate.
eventTypesNoRequired for create: which events to push (reply | bounce | soft_bounce | complaint | unsubscribe).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the destructiveHint annotation by disclosing that create/rotate return the HMAC secret only once, that URLs must be public HTTPS with private/metadata IPs rejected, that deliveries are signed with HMAC-SHA256, and that active:true can re-enable an auto-disabled subscription. These behavioral details are valuable and not available in the annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries unique high-value information. The front-loaded purpose phrase is followed by compact action-specific syntax, and the security details are appended without unnecessary prose. It is long enough to be complete but not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-action mutation tool with no output schema, the description covers the core behaviors: required inputs per action, URL validation, secret rotation and one-time return, and the delivery signature. The only notable gap is that the success/return behavior for update and delete beyond the 'one changed field' is not described, but this is minor given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 86%, so the schema already documents most parameters. The description adds useful semantic relationships: id is required only for update/delete, url + eventTypes are the create prerequisites, update takes exactly one changed field, and secret on update rotates the signing secret. This additions help the agent assemble the correct action-specific parameter sets.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description announces 'Manage an outbound webhook subscription' and then enumerates the three concrete actions: create, update, and delete. This clearly distinguishes the tool from siblings like get_webhooks, which exists for reading subscriptions rather than mutating them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance for when to use each action: create needs URL + eventTypes, update needs id + one changed field, delete needs id. It also clarifies edge cases like re-enabling via active:true, but it does not explicitly mention 'use get_webhooks to view existing webhooks' as an alternative, so it misses an opportunity for a stronger routing cue.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contact_operatorAInspect

Reach a human operator for anything list_messages/infrastructure_status cannot answer (a stuck vendor issue, a billing question, an account-level ask). Inputs: body (1-2000 chars), urgency (normal|needs_human, default normal). Files a ticket, notifies the operator; returns { ticketId, note, deduplicated }. Works in every account state a token authenticates in (dunning-suspended, canceling, canceled) — admin-terminated is the one exception, rejected at auth. The reply arrives as a message on THIS account (poll list_messages/infrastructure_status — no reply-fetch call). Sending the IDENTICAL body+urgency again within an hour returns the SAME ticketId, no second ticket/alert (deduplicated: true). TEXT MATCH, NOT AN INTENT MATCH: a genuinely new message with identical wording collapses the same way, silently — vary the wording (or raise urgency, always a new ticket). needs_human bypasses the ~10-min ops-email throttle. Rate-limited to 5/hour/tenant — a 429 names retryAfter (seconds).

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
urgencyNonormal

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only carry destructiveHint=false, so the description bears the load — and it does: ticket creation, operator notification, account-state eligibility (dunning-suspended/canceling/canceled, admin-terminated rejected at auth), reply delivery mechanism (poll list_messages, no reply-fetch call), 5/hour/tenant rate limit with 429 retryAfter, and throttle bypass semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose then packs dense, distinct operational facts; almost every clause earns its place. It is long and reads as one wall of prose rather than clearly segmented, but little is expendable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description supplies the return shape ({ ticketId, note, deduplicated }) and explains the dedup field's meaning. Combined with the state/rate-limit/throttle coverage, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only gives bare min/max/enum, so the description must compensate — and it does by giving the 1-2000 char bound, the urgency values with their default, and crucially the semantic consequence of urgency (needs_human bypasses the ops-email throttle) and of body identity (text-match dedup).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Reach a human operator') and explicitly scopes it against the two siblings (list_messages/infrastructure_status) it is the fallback for. The agent can distinguish it from every other tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers (stuck vendor issue, billing, account-level ask) and names the alternatives that cannot answer them. Also states the one exception (admin-terminated) where it will not work, which is a when-not condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_byo_domainsA
Read-only
Inspect

List your BYO (bring-your-own) domains, or (with id) one domain's full intake detail. No id → [{ domainId, domain, isPrimary, dnsMode, byoStatus, breakerTier, reputationBranch, mailboxCount }]. With id → adds the pre-flight scan result, abuse-gate verdict, and consent-acknowledgment status. byoStatus progresses pending_kyc|pending_consent|pending_dns → active (or rejected/abandoned). Use configure_byo_domain to register a new one or advance it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoOmit to list every BYO domain; pass an id for that domain's full intake detail (scan result, abuse verdict, consent status).

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true; description adds rich behavioral details: lists output fields per mode, explains status progression, and notes that configure_byo_domain is for mutation. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five sentences, front-loaded with main action, then specifics. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 1-parameter tool with no output schema, the description fully covers behavior, output structure, status progression, and sibling tool linkage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage for 'id' parameter; description goes beyond schema by explaining the effect of omitting vs providing id, and enumerates additional returned fields when id is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it lists BYO domains or gets full intake detail for a specific domain. Distinguishes between two modes and differentiates from sibling tool 'configure_byo_domain'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to omit 'id' vs pass it, and directs users to 'configure_byo_domain' for registration or advancement. No ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dashboardA
Read-only
Inspect

Read saved dashboard views. No id → list all: [{ id, name, isDefault, rev, editedBy }]. With id → that view's full layout + rev (pass this rev as the CAS base to configure_dashboard update). Views are both agent- and human-editable; write them with configure_dashboard.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoOmit to list every saved view (summary); pass a view id for its full layout + rev.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so description's reading behavior is consistent. It adds value by explaining the rev role and that views are both agent- and human-editable, going beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is three concise sentences: core purpose, mode behavior, and additional context. No unnecessary words, front-loaded with key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one optional parameter and no output schema, the description fully explains return structures for both modes and connects to sibling tool, providing complete context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema already describes id parameter well, but description adds meaning by specifying the exact fields returned in each mode and how rev connects to configure_dashboard, enhancing schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Read saved dashboard views' and distinguishes between two modes: listing all with no id, and retrieving full layout with id. It contrasts with sibling tool configure_dashboard, making purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use each mode ('No id → list all', 'With id → that view's full layout + rev') and advises to write views with configure_dashboard, providing clear alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_webhooksA
Read-only
Inspect

List your outbound webhook subscriptions, or (with id) one subscription plus its recent delivery + attempt log. No id → [{ id, url, eventTypes, active, status, disabledReason, consecutiveFailures }]. With id → { subscription, recentDeliveries[], recentAttempts[] }. Secrets are never returned on reads — they are shown once at create/rotate.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoOmit to list every subscription; pass an id for that subscription plus its recent delivery + attempt log.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals that secrets are never returned on reads, a behavioral trait beyond the 'readOnlyHint' annotation. This is consistent with the annotation and adds important context for the agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet comprehensive, covering both use cases and a critical note about secrets without unnecessary verbosity. It is well-structured with front-loaded main action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description details return shapes for both cases, compensating for the lack of an output schema. It lacks information on pagination or limits, but the simplicity of the tool makes this acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, and the parameter description already explains the id behavior. The tool description does not add additional meaning beyond the schema, meeting the baseline expectation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists webhook subscriptions, with distinct behaviors for omitting or providing an id. It specifies the return format for each case, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the id parameter versus omitting it, providing clear usage patterns. It does not explicitly mention when not to use this tool or alternatives among siblings, but the guidance is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inboxA
Read-only
Inspect

Unified reply inbox across mailboxes. Cursor-paginated → { threads[], nextCursor }; each row: threadId, campaignName, leadEmail, subject, mailboxEmail, label, lastEventType, markStatus. Filters: mailbox, campaign, label, read, includeNonreply (bounces/OOO, default true), archived (exclude|include|only). Use thread for one thread's history.

ParametersJSON Schema
NameRequiredDescriptionDefault
readNo
labelNo
limitNo
cursorNo
mailboxNo
archivedNoexclude
campaignNo
includeNonreplyNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true. The description adds value by detailing pagination behavior and the structure of return data (threads[], nextCursor, fields). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences plus a filter list and an alternative tip. Every sentence adds value, and the most important information (purpose, pagination, key filters) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description compensates by specifying return structure and fields. It covers all major filters and provides a usage pointer to a sibling. Missing details like error handling or ordering, but still adequate for a read-only paginated list tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description explains the meaning of 6 out of 8 parameters (mailbox, campaign, label, read, includeNonreply, archived) including the enum options for archived and default for includeNonreply. Two parameters (limit, cursor) are only implicitly mentioned via pagination context, which is acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns a unified reply inbox across mailboxes, specifies cursor-paginated output with fields, and distinguishes from sibling 'thread' by telling to use that for one thread's history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It lists available filters and explicitly mentions an alternative tool ('Use thread for one thread's history'), providing clear context on when to use this tool. However, it does not explicitly state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

infrastructure_statusA
Read-only
Inspect

Warmup + provisioning progress per mailbox. Returns { domains, mailboxes, sendReady, mailboxHealth[], messages[], nextSteps }. New mailboxes ramp-limit server-side: 5 sends/day week 1, 40/day week 4 — dailyCap/sentToday show real capped sending from day 1, never zero. Per-mailbox sendReady is a FULLY-RAMPED flag, NOT a send gate — below-ramp mailboxes still send, capped at their own dailyCap; read dailyCap/sentToday for capacity. Top-level sendReady is the AND across ALL mailboxes (true once every one is ramped) — for one mailbox's capacity use its own flag. Each mailbox carries delivStatus (healthy/throttled/paused), complaint/bounce/softBounce rates; vendorReputationScore/vendorPlacementRate are VENDOR-REPORTED, null when unmeasured. Vendor-pool warmup is FEED-INVISIBLE by design. messages[] previews 5 unacked notices (severities: list_messages); reading never marks read. nextSteps names the next action (see setup_infrastructure). Use account/metrics for rollups.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true is complemented by rich context the annotation can't carry: sendReady is a fully-ramped flag and NOT a send gate, top-level sendReady is an AND across all mailboxes, vendorReputationScore/vendorPlacementRate are vendor-reported and null when unmeasured, vendor-pool warmup is feed-invisible, and reading never marks messages read. This preempts real misinterpretation traps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and fairly long, but front-loaded with purpose and each sentence disambiguates a specific semantic trap (dual sendReady flags, null vendor fields, read-only messages). It could be broken into clearer groupings, but little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full burden of describing returns and does so by naming the return shape plus the semantics of ambiguous flags and severity sources. Nothing an agent needs to call or interpret this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document; the baseline of 4 applies. The description instead explains the meaning of returned fields, which is more valuable here than parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and scope: warmup + provisioning progress per mailbox, and enumerates the returned shape. It explicitly differentiates from siblings by naming account/metrics for rollups and setup_infrastructure for next actions, so an agent can tell it apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent to alternatives ('Use account/metrics for rollups', 'see setup_infrastructure', severities via list_messages) and clarifies that this tool is for per-mailbox ramp capacity. It never states explicit when-not-to-use conditions, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

label_threadAInspect

Set or clear a triage LABEL on an inbox thread — the same chip the dashboard shows. Inputs: threadId, label (string; pass label:null to clear). Distinct from mark (read/unread/archived state): a label is a free-form category, not a read flag. Filterable via inbox's label param.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNo
threadIdYesThe thread id, e.g. from inbox() or campaign events.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Describes the mutation behavior (set or clear, with null to clear) and is consistent with destructiveHint=false. Adds value beyond annotations by explaining how to clear a label.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose, no wasted words. Efficiently conveys all necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (2 params, no output schema), the description provides enough context: it links to dashboard chips, explains clearing, and notes filterability. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%; the description adds meaning for the 'label' parameter by explaining how to clear it. The 'threadId' parameter is already described in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'set or clear' and the resource 'triage LABEL on an inbox thread', and distinguishes it from the sibling tool 'mark' by explaining that a label is a free-form category, not a read flag.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit differentiation from 'mark' and mentions filterability via inbox's label param. Does not list all alternative tools but gives clear context for when to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

launch_campaignA
Destructive
Inspect

Create and activate a campaign on a lead list. You supply name, offer, leads[], sequence[] (per step: subject, body, delayDays — subject is REQUIRED on step 1; a LATER step may omit it to keep step 1's subject line [recommended], which resolves at launch to "Re: " + step 1's subject (unchanged if that already starts with "Re: "), or set one explicitly to change only that step's subject line — every later step is still sent In-Reply-To step 1, though some mail clients display a changed subject as a separate conversation), and optionally timezone, sendWindow, stopOnReply, listSource (one sentence on where the list came from — REQUIRED on a crowdfunding account, where a missing or "unknown" source is refused 400) — the platform does not write copy. Each step's step number must be unique within the sequence; a duplicate step number is refused (400) naming it. Subjects and bodies may use {{firstName}} and {{company}} (filled from each lead); a launch is refused (400) if any subject or body uses any other {{token}}, a single-brace form of a known token such as {firstName} or {company} (a typo for the double-brace form — it would otherwise reach the recipient unsubstituted), or a run of 3 or more consecutive { or } anywhere (e.g. {{{firstName}}} or a lopsided run like {{{firstName}} — it would otherwise reach the recipient with stray braces attached). The send window is evaluated in timezone (IANA, e.g. America/New_York; default UTC) — set it to your recipients' zone. sendWindow is { startHour, endHour, days? }: integer hours 0-23, endHour INCLUSIVE (the last hour a send may start in), days 0=Sunday … 6=Saturday. Each omitted field (or all of them, by omitting sendWindow) gets the platform's RECOMMENDED default {"startHour":8,"endHour":16,"days":[1,2,3,4,5]} — Monday-Friday business hours — except on a sandbox account, where omitted fields are open every hour of every day until it upgrades and they take that default; the decision is yours, e.g. include 0 and 6 in days to send on weekends. Setting exactly ONE of startHour/endHour combines it with the recommended default for the OTHER (08/16, including on a sandbox account, since that value re-resolves to the same default on upgrade) — if that combination would wrap the window past midnight (start > end), the launch is refused; set both explicitly, including for a deliberate overnight window such as {startHour:22, endHour:6}. Step 1 goes out from the least-loaded mailbox; each later step goes out delayDays after the previous step ACTUALLY sent, from the SAME mailbox, with In-Reply-To/References set to step 1's Message-ID — it waits while that mailbox is at its daily cap, and the lead's remaining steps are cancelled (a 'failed' event says why) once that mailbox is paused or released, since neither ever lifts on its own, or when a paid account's thread went out from a connected BYO mailbox, which this build never sends from; a step whose previous step never went out is skipped. Suppressed leads are skipped. Returns { campaignId, sendWindow, timezone, nextSteps } — sendWindow and timezone are what actually applied. Campaigns send real mail, so a launch identical to one this account made in the last 60 seconds is REFUSED with 409 { code:'duplicate_campaign', existingCampaignId } rather than contacting the same prospects twice — check that campaign instead of relaunching. Resend the same idempotencyKey to retry a call whose response you lost: that replays the original result instead of being refused. Campaigns that differ in any field other than listSource (which is not part of that check), and deliberate relaunches after those 60 seconds, are never blocked.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
endByNoOptional fixed end date, as an ISO 8601 instant with a zone. From then on nothing more in this campaign sends: every step still waiting is skipped and recorded as a 'failed' event with reason campaign_ended. Paid accounts only: a trial account that sends one is refused 400. Must be later than now and later than every step's notBefore. Omit to let the sequence run to its end.
leadsYes
offerYes
sequenceYes
timezoneNoIANA time zone the send window is evaluated in — set it to your recipients' zone. Default UTC.UTC
listSourceNoWhere this lead list came from, in a sentence (e.g. "our own past-donor CRM export"). REQUIRED on a crowdfunding account: a launch with no source, or "unknown", is refused (400) — cold lists are allowed only with a stated source. Optional elsewhere. Max 500 characters.
sendWindowNoInteger hours 0-23 (endHour inclusive: the last hour a send may start in) and weekdays, evaluated in `timezone`. Each field you omit (or all of them, by omitting sendWindow) gets the recommended {"startHour":8,"endHour":16,"days":[1,2,3,4,5]} — except on a sandbox account, where omitted fields are open every hour of every day until it upgrades and they take the recommended default. If you set exactly ONE of startHour/endHour, the other still defaults to the recommended value (08/16) — if that combination would wrap past midnight (start > end), the launch is refused; set both explicitly, including for a deliberate overnight window.
stopOnReplyNo
idempotencyKeyNoOptional idempotency key: resend the SAME key when retrying this call so a dropped-response retry is not applied twice (no duplicate campaign/provision/send).

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only destructiveHint=true available as annotation, the description carries the full behavioral burden and does so richly: real mail is sent, 409 duplicate semantics with the 60s window, idempotency replay behavior, mailbox rotation and cap/pause effects on later steps, suppression skipping, and explicit 400 refusal reasons. This is far beyond what the annotation covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded, but the body is a single sprawling block that runs to roughly 2,000 characters and repeats material already in the schema (the step-1 subject rule appears in both the description and the schema property). Much of the content is valuable but oversized and hard to parse for an agent deciding whether to call it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-param, deeply nested mutation tool with no output schema, the description covers the mutation semantics, default resolution, failure modes, and even the return shape ({ campaignId, sendWindow, timezone, nextSteps }). Nothing an agent needs to invoke this correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% across 10 params, and the description compensates with rules the schema omits entirely: unique step numbers, token whitelist ({{firstName}}/{{company}}), single-brace and 3+ brace refusal, and the startHour/endHour combination default. Some blocks (subject, timezone, sendWindow, listSource) largely restate the schema descriptions rather than extend them, preventing a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ("Create and activate a campaign on a lead list") and enumerates the required inputs, so the agent knows exactly what this does. It does not explicitly name or distinguish itself from the adjacent siblings plan_campaign or list_campaigns, keeping it at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-not conditions (identical launch within 60s is refused, sandbox vs paid differences, trial accounts refused for endBy/notBefore) and tells the agent to check the existing campaign instead of relaunching, which implies the listing/results siblings. It stops short of explicitly naming the alternative tool for planning or inspecting campaigns.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_campaignsA
Read-only
Inspect

List every campaign at once: [{ campaignId, name, status, counts{sent,reply,bounce,complaint,unsubscribe,failed,soft_bounce} }], newest first — no per-campaign lookup needed. Use campaign_results for one campaign's counts, metrics for account-wide totals.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description adds behavioral context (returns list with specific fields, newest first, bulk operation). Could mention limitations like pagination, but not necessary given no parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with purpose, zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description fully covers what the tool returns, ordering, and provides sibling guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, and schema coverage is 100% (trivially). Baseline is 4; description does not add parameter info but provides output structure, which is helpful but not required for this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'List every campaign at once' and provides the output structure, clearly distinguishing it from siblings like campaign_results and metrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use (when you need all campaigns) and when not to, offering alternatives: 'Use campaign_results for one campaign's counts, metrics for account-wide totals.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_leadsA
Read-only
Inspect

List/export leads with their contact-level disposition, cursor-paginated. Returns { leads[], nextCursor }; each row: leadId, email, firstName, company, campaignId, campaignName, globalStatus, interestStatus, notes, tags, suppressed, lastEventType, lastEventTs, createdAt. Filters: campaign, interestStatus, suppressed, replied. This IS the export surface — paginate to dump the full book of business as JSON (no separate CSV endpoint). Use update_lead to write disposition, suppress_lead to opt an address out.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cursorNo
repliedNo
campaignNo
suppressedNo
interestStatusNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond the readOnlyHint annotation by specifying cursor-pagination, return structure (leads[], nextCursor), and available filters. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with purpose, each providing essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given six parameters and no output schema, the description covers purpose, response structure, pagination, and sibling alternatives. It lacks detailed parameter explanations but is otherwise comprehensive for a list tool. Minor gap: no explicit mention of default limit value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description partially compensates by listing four filter parameters (campaign, interestStatus, suppressed, replied). However, it does not explain the pagination parameters (limit, cursor) beyond mentioning cursor-pagination. Enums are only in schema, not described. Adequate but not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'List/export leads' with specific verb and resource, and mentions cursor-pagination, making the purpose immediately obvious. It also differentiates from siblings by referencing update_lead and suppress_lead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool: 'This IS the export surface — paginate to dump the full book of business as JSON (no separate CSV endpoint).' It also points to alternatives: 'Use update_lead to write disposition, suppress_lead to opt an address out.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_messagesA
Read-only
Inspect

List this tenant's system + operator messages (a retryable setup step, a credential going live, an operator notice), cursor-paginated. Unacked messages sort first (newest within that group), then acked (also newest first). Returns { messages[], nextCursor }; each: id, kind, severity ('info' = resolves on its own | 'action_required' = act and it progresses | 'operator_pending' = platform stopped, only an operator can restart it, retrying the SAME call then works | 'terminal' = platform stopped, only a human can move it, do NOT retry), body, actionHint, source (system|operator), createdAt, readAt. readAt is set ONLY by ack_message — listing never marks messages read, so a null readAt does not mean never seen, only not yet acknowledged. Use ack_message to stop one resurfacing. infrastructure_status inlines the newest 5 unacked messages; this is the full paginated surface.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cursorNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, but the description adds substantial behavioral context beyond that: unacked-first sort order, cursor pagination, the ack/readAt distinction ('listing never marks messages read'), and the retry semantics encoded in severity levels. This is exactly the kind of disclosure that structured fields cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but densely packed and front-loaded: purpose first, then pagination and sort order, then the return shape, then the readAt caveat and sibling routing. Every clause (severity table, readAt semantics) earns its place, though the prose is heavy enough to risk skimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the full burden and does so: it specifies the response envelope { messages[], nextCursor } and every field (id, kind, severity, body, actionHint, source, createdAt, readAt). Nothing an agent needs to interpret results or call correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for two parameters, so the description must compensate. It explains cursor pagination and the returned nextCursor, which adds real meaning to the cursor parameter, but gives no guidance on limit (default 50, max 200). Partial compensation justifies a mid score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List this tenant's system + operator messages') with explicit scope and the three message kinds enumerated. It also distinguishes itself from infrastructure_status as 'the full paginated surface' versus the inline newest 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool and condition: use ack_message to stop a message resurfacing, and infrastructure_status is only the inline 5 while this is the full paginated list. The severity semantics supply clear when-to-act guidance (do NOT retry on 'terminal', retrying the same call works on 'operator_pending').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

markAInspect

Set a thread's READ-STATE for inbox triage. Inputs: threadId, status = 'read' | 'unread' | 'archived' (archived hides it from the default inbox; refetch with inbox archived='include'/'only'). Returns { marked: true }. 404 if unknown. This is the read/archive flag ONLY — use label_thread for a triage label chip, reply to respond.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYes
threadIdYesThe thread id, e.g. from inbox() or campaign events.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds value beyond annotations by detailing the return value ({marked: true}), error handling (404 for unknown thread), and the effect of 'archived' status (hides from default inbox, refetching behavior). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (three sentences plus a front-loaded purpose). Every sentence adds necessary information with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the two parameters, no output schema, and available sibling tools, the description provides complete context: input details, output, errors, and usage boundaries.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While schema coverage is 50%, the description adds meaning by listing the valid status values and explaining the archival behavior. It also provides context for threadId ('e.g. from inbox() or campaign events'). This compensates for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Set' and the resource 'a thread's READ-STATE' for inbox triage. It also explicitly distinguishes from sibling tools 'label_thread' and 'reply'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (for read/unread/archived flags) and provides alternatives for other actions (label_thread for labels, reply for responses). However, it does not explicitly state when not to use it beyond those comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

metricsA
Read-only
Inspect

Account-wide outcome totals across ALL campaigns: { sent, reply, bounce, complaint, unsubscribe, failed, soft_bounce } — same shape as campaign_results but summed tenant-wide (bounce = hard only, opens not tracked). Use campaign_results for one campaign, list_campaigns per-campaign, or account for billing/quota.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true. The description adds context on the data shape (fields), scope (tenant-wide), and exclusions (bounce = hard only, opens not tracked), without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with a clear list and alternative references. No superfluous information; well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description explains return fields, scope, and tracking exclusions. Sufficient for understanding tool behavior, though could hint at data types.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist (schema has zero properties, 100% coverage). Baseline of 4 for zero-parameter tools is appropriate; description adds no parameter info but clarifies return semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies 'Account-wide outcome totals across ALL campaigns' with a clear list of fields, and explicitly distinguishes from sibling tools (campaign_results, list_campaigns, account).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when to use this tool vs alternatives: 'Use campaign_results for one campaign, list_campaigns per-campaign, or account for billing/quota.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pauseA
Destructive
Inspect

Pause ONE campaign: its status → 'paused', so the tick schedules no further steps (already-sent mail is unaffected; there is no resume tool). Input: campaignId. Returns { paused: true }. 404 if not found. Add leadEmails to stop ONLY those leads in this campaign instead (e.g. a donor who already gave): their remaining steps are cancelled, the campaign keeps running for everyone else, and nobody is suppressed — returns { paused: false }. Matching ignores capitalisation; if any address is not a lead of this campaign, nothing is stopped and the call is refused (404) naming them. Use pause_all to pause every active campaign at once.

ParametersJSON Schema
NameRequiredDescriptionDefault
campaignIdYesThe campaign id returned by launch_campaign.
leadEmailsNoOptional. Stop ONLY these leads in this campaign (any capitalisation): their remaining steps are cancelled, the campaign keeps running for everyone else, and nobody is added to the suppression list. Every address must be a lead of this campaign, or nothing is stopped and the call is refused naming the ones that are not. Omit to pause the whole campaign.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=true; the description goes well beyond by disclosing that already-sent mail is unaffected, that no resume tool exists, the two distinct return shapes ({ paused: true } vs { paused: false }), 404 on missing campaign, case-insensitive matching, the all-or-nothing refusal when an address isn't a lead, and that nobody is added to the suppression list. This is exactly the behavioral detail annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the primary action and its effect, then the optional leadEmails path, then the sibling routing. Dense and largely waste-free, though the leadEmails parenthetical runs long enough that it duplicates the schema description's content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-mode tool with no output schema, the description covers both invocation modes, both return payloads, error conditions (404 naming offending addresses), and irreversibility (no resume). Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3; the description adds genuine meaning by explaining the semantic effect of leadEmails (only those leads stop, campaign continues for everyone else, no suppression) and the omit-to-pause-everything rule, which the schema phrasing only partly captures. It stops short of describing the email-matching normalization rules in full.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Pause') and resource ('ONE campaign') with the exact resulting state ('status → paused'), and explicitly names the sibling it must not be confused with ('Use pause_all to pause every active campaign at once'). An agent can distinguish it from pause_all and suppress_lead without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use routing: pause the whole campaign vs. supply leadEmails to stop only certain leads (e.g. 'a donor who already gave'), plus the pause_all alternative for the multi-campaign case. The condition selecting each mode is spelled out, not left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pause_allA
Destructive
Inspect

Pause EVERY active campaign for the tenant at once (each active status → 'paused'; the tick then schedules no further sends). No inputs. Returns { pausedAll: true }. Use pause to pause a single campaign by id.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructive intent, and description adds specific behavioral details (status change, stops sends, return value). No mention of reversibility or prerequisites, but sufficient for the simple operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise with two sentences plus a note. Front-loaded with purpose and effect.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All necessary context provided: purpose, effect, return value, no inputs, and relationship to sibling tool. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters, schema fully covers it. Description explicitly states 'No inputs', adding confirmation beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it pauses every active campaign for the tenant, with specific effect on status and scheduling. Distinguishes from sibling tool 'pause' for single campaign.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to use 'pause' for single campaign, providing clear when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_campaignA
Read-only
Inspect

Plan a dated outreach campaign (for example a fundraising drive) and get its price. Read-only: nothing is booked, charged or changed. Inputs: contacts, campaignStart (first campaign day, YYYY-MM-DD, UTC), campaignDays, bookingDate (the day you would book, normally today), and optionally domainsAlreadyHeld (default 0; infrastructure_status shows domains) and listIsOwnOptIn (your own opted-in donors; default false). Returns { mailboxes, domainsToBuy, bookBy, advanceStartOn, safeReach, priceCents, warnings[] }. safeReach is how many contacts the warm-up ramp can safely reach if you book on bookingDate; the price is for that many. warnings: over_60_mailboxes (above the 60-address self-serve ceiling), late_booking (safeReach is below contacts), own_list_cheaper_on_esp (an opted-in list may cost less on an email-service provider). A bookingDate after the last campaign day is refused 400. Campaign checkout and booking are not available yet: to book, contact support (contact_operator).

ParametersJSON Schema
NameRequiredDescriptionDefault
contactsYes
bookingDateYes
campaignDaysYes
campaignStartYes
listIsOwnOptInNo
domainsAlreadyHeldNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, but the description goes well beyond that: it enumerates the returned fields, names and explains each warning code, discloses the 60-address self-serve ceiling, explains the price/safeReach relationship, and states the 400 refusal for a bookingDate after the last campaign day. This is unusually rich behavioral disclosure for a read-only tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and safety, then inputs, then returns, then the routing note. Dense but every sentence carries information. It is on the long side and slightly packed, keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 0% schema description coverage, this description still fully explains inputs, return shape, and error behavior. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the full burden, and it does: it defines every parameter in prose (contacts, campaignStart as first campaign day in UTC YYYY-MM-DD, campaignDays, bookingDate as the day you would book and normally today, domainsAlreadyHeld with default and a pointer to infrastructure_status, listIsOwnOptIn with default and meaning). This compensates fully for the empty schema annotations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Plan') and resource ('dated outreach campaign') with a concrete example (a fundraising drive), plus the immediate outcome ('get its price'). It also distinguishes itself from the sibling launch_campaign by declaring that checkout and booking are not available. An agent can identify this as a read-only estimation/quotation tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use ('Read-only: nothing is booked, charged or changed') and an explicit alternative for the adjacent job ('to book, contact support (contact_operator)'). It also guides input choice ('bookingDate ... normally today'). It stops short of a full when-not-to-use-this-vs-sibling clause, so it lands just under a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_mailboxesA
Destructive
Inspect

Downgrade: release your N NEWEST live mailboxes, lowering the billed quantity. Inputs: count, acknowledged (must be true — release is immediate but irreversible; no mid-cycle credit, lower price starts next renewal, minimum 5 mailboxes/$99). Returns { releasedCount, failedCount, unreleased, billing, deduplicated }. count is RELATIVE (that many MORE, not a fleet target). ALWAYS pass an idempotencyKey: the FIRST call under a key fixes WHICH mailboxes to release; a later call with the SAME key can only finish that set. releasedCount can be less than count — failedCount names how many the vendor refused (still live, billed; unreleased lists them). Resend with the same key until failedCount is 0. deduplicated: true = NO new work — an earlier recorded outcome under that key was replayed. A genuinely NEW downgrade needs a NEW key. 409 = a release already running; re-check infrastructure_status before retrying. To ADD mailboxes use setup_infrastructure or configure_byo_domain.

ParametersJSON Schema
NameRequiredDescriptionDefault
countYes
acknowledgedYes
idempotencyKeyNoOptional idempotency key: resend the SAME key when retrying this call so a dropped-response retry is not applied twice (no duplicate campaign/provision/send).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=true; the description adds substantial context the agent could not otherwise know: release is immediate and irreversible, no mid-cycle credit, lower price starts at next renewal, a 5-mailbox/$99 minimum, and the tricky idempotency semantics where the first call under a key fixes which mailboxes are released.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and scope, and nearly every sentence carries non-obvious operational detail. It is dense and long for a three-parameter tool, but the length is justified by idempotency and partial-failure semantics rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description enumerates the return shape (releasedCount, failedCount, unreleased, billing, deduplicated) and explains how to interpret the partial-failure and dedup cases, leaving nothing essential unstated for a destructive billing operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (count and acknowledged are undocumented in the schema), but the description compensates fully: count is RELATIVE not a fleet target, acknowledged must be true, and an idempotencyKey should always be passed with the replay/dedup semantics explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('release your N NEWEST live mailboxes, lowering the billed quantity') and immediately frames it as a downgrade, which cleanly separates it from the add-side siblings setup_infrastructure and configure_byo_domain named at the end.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (downgrade billing), when-not (to ADD mailboxes use setup_infrastructure or configure_byo_domain), plus retry protocol (resend with the same key until failedCount is 0) and error handling (409 = release already running, re-check infrastructure_status).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replyA
Destructive
Inspect

Send a reply on an existing thread, from the mailbox that sent it. Inputs: threadId, body. Returns { messageId, deduplicated }. A reply is real send volume, governed like campaign sends: counts against the mailbox daily cap, and is REFUSED (never silently dropped) when the recipient is suppressed, the mailbox is deliverability-paused, or the cap is used up — a refusal returns { error, code: send_blocked, reason: suppressed|mailbox_paused|daily_cap_reached, retryable }; retryable (cap) clears next daily rollover, non-retryable does not — stop retrying, do not loop replies to manufacture volume. Idempotent: identical retries collapse to one send — pass a stable idempotencyKey (else a body hash is used, matched only 10 minutes) so a dropped-response retry cannot double-send. deduplicated: true = NO new email sent; messageId is from an earlier matching send. Pass a key, or vary the body, if a repeat must genuinely go out. 404 if no sending mailbox is on record for the thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
threadIdYesThe thread id, e.g. from inbox() or campaign events.
idempotencyKeyNoOptional idempotency key: resend the SAME key when retrying this call so a dropped-response retry is not applied twice (no duplicate campaign/provision/send).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply destructiveHint=true; the description carries far more: cap accounting, refusal error shape { error, code: send_blocked, reason, retryable }, retryability rules, idempotency collapse with a 10-minute body-hash fallback window, deduplicated semantics, and the 404 no-mailbox case. Nothing contradicts the destructive hint and it richly extends it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense paragraph, but front-loaded with purpose then inline 'Inputs:'/'Returns:' structure, and every sentence carries operational value (refusal codes, idempotency). Slightly heavy for a 3-param tool, but no filler to cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly spells out return values ({ messageId, deduplicated }, the error envelope), plus the 404 case. For a side-effecting send tool, an agent has everything needed to call, retry, and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description adds real meaning the schema lacks: the idempotencyKey's purpose, the body-hash fallback matched only 10 minutes, and how to force a genuine resend ('pass a key, or vary the body'). threadId/body are named but not elaborated beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Send a reply on an existing thread, from the mailbox that sent it'), which is unambiguous and distinct from siblings like launch_campaign (new sends) or inbox/thread (reads). An agent can select it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-and-when-not guidance: it is 'governed like campaign sends,' REFUSED rather than silently dropped on suppression/pause/cap, retryable only for cap and cleared at daily rollover, and explicitly warns 'stop retrying, do not loop replies to manufacture volume.' This is unusually actionable routing/retry guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

setup_infrastructureAInspect

Provision sending infra: buy lookalike domains, create mailboxes, start warmup. Inputs: brand, primaryDomain, domains + inboxesEach (or distribution), persona, physicalAddress, senderIdentity. Billing is per provisioned mailbox; quoteOnly:true previews cost first. domains/inboxesEach are the infra you want to HAVE, not an amount to add — a repeat call never buys twice; each ordinal 0..domains-1 fills to its own count. Mailbox addresses are DETERMINISTIC from persona+ordinal+slot — keep persona unchanged on a retry. registerDomains:true is required consent before any NEW domain purchase; omitting it on a buy call is refused 400 registrar_optin_missing (self-correct by resending true, never an operator escalation) — also needs registrant, unless already on file. NO background retry exists: a provisioning result (pending|capacity_pending) needs YOU to retry to progress; capacity_pending needs contact_operator instead. Returns { jobId, billing, provisioning?, nextSteps }.

ParametersJSON Schema
NameRequiredDescriptionDefault
brandYes
domainsYes
personaYes
quoteOnlyNo
registrantNo
inboxesEachNo
distributionNo
primaryDomainYes
idempotencyKeyNoOptional idempotency key: resend the SAME key when retrying this call so a dropped-response retry is not applied twice (no duplicate campaign/provision/send).
senderIdentityYes
physicalAddressYes
registerDomainsNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=false, so the description carries the behavioral burden and delivers: per-mailbox billing, absence of background retry, deterministic mailbox addresses tied to persona, idempotency semantics, and the exact failure/refusal mode. This is unusually rich disclosure well beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then flows into inputs, billing, and failure semantics. It is dense and long, but given the 12-param complexity nearly every clause earns its place. Slightly information-dense at the cost of scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description specifies the return shape { jobId, billing, provisioning?, nextSteps } and the full retry/error workflow. For a high-complexity provisioning tool this is complete enough for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is very low (8%), so the description must compensate, and it does for the ambiguous params: domains/inboxesEach are 'the infra you want to HAVE, not an amount to add', their ordinal fill behavior, distribution as an alternative, and persona determinism on retry. Minor params (senderIdentity, physicalAddress, brand) are left to the schema but are largely self-descriptive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Provision sending infra') and enumerates the concrete sub-actions (buy lookalike domains, create mailboxes, start warmup). This clearly distinguishes it from siblings like configure_byo_domain, remove_mailboxes, and infrastructure_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when/when-not guidance: quoteOnly:true previews cost first, registerDomains:true is required consent before new purchases, and a pending provisioning result needs the caller to retry while capacity_pending routes to contact_operator instead. It even covers the 400 registrar_optin_missing self-correction path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_feedbackAInspect

Send private feedback to the Coldrig founders about how the product is working for you. Inputs: score (integer 0-10, optional) and/or text (1-2000 chars, optional); at least one is required. It is stored privately and emailed to the founders, never published, and it opens no support ticket and changes nothing on your account (use contact_operator for a problem that needs an answer). Each call is a separate submission. Returns { feedbackId, received }.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
scoreNo

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply destructiveHint=false, so the description carries the real burden and does: it discloses the storage/visibility model (private, emailed to founders, never published), the side-effect profile (no ticket, no account changes), and the append semantics ('Each call is a separate submission'). It even states the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core action, then constraints, then behavioral caveats. Every clause carries information an agent needs; nothing is restated from the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description names the return fields ({ feedbackId, received }), covers the at-least-one-input rule, and resolves the main ambiguity (feedback vs. support). Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description compensates fully: it gives the type and range for score (integer 0-10, optional), the length bounds for text (1-2000 chars, optional), and — critically — the cross-field constraint that at least one is required, which the schema itself does not enforce.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Send private feedback to the Coldrig founders') plus the channel it lives in (private, emailed, never published). It also distinguishes itself from the nearby sibling contact_operator, so an agent can pick between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the alternative case: 'use contact_operator for a problem that needs an answer,' which names both the alternative and the condition that selects it. It also clarifies that this opens no support ticket and changes nothing on the account.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suppress_leadA
Destructive
Inspect

Permanently suppress an email address tenant-wide (every current and future campaign) — the manual/free-text 'stop emailing me' path for opt-outs the strict typed-unsubscribe matcher misses. Inputs: email, reason (fixed 'manual' — the only value this tool honestly claims; bounce/complaint/unsubscribe are recorded automatically elsewhere), note (accepted, not persisted). Cancels every pending send + marks every campaign-lead row 'suppressed'. Last-write-wins: re-suppressing a bounce/complaint/unsubscribe row relabels its reason to 'manual'. There is no un-suppress tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
emailYes
reasonNomanual

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the `destructiveHint` annotation by detailing side effects: cancels pending sends, marks rows suppressed, last-write-wins relabeling, and the fact that notes are accepted but not persisted. This provides complete transparency beyond what annotations offer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise but packs in all critical details. The structure is front-loaded with the core purpose, then details. Minor improvement: could break into bullet-style for even quicker scanning, but overall every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description fully covers input semantics, side effects, behavioral traits (last-write-wins), and caveats (no un-suppress). It leaves no ambiguity about what the tool does or its consequences for a 3-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description fully compensates: it explains that `reason` is fixed to 'manual' (const), `note` is accepted but not persisted, and `email` is required. This adds significant meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies a clear verb ('suppress') and resource ('email address tenant-wide'), and distinguishes its purpose as the manual opt-out path for cases the typed-unsubscribe matcher misses. It names the exact inputs and effects, making the tool's role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use ('manual opt-out path') and implicitly when not to use (bounce/complaint/unsubscribe are handled automatically elsewhere). It also warns there is no un-suppress tool and explains the last-write-wins behavior, guiding appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

threadA
Read-only
Inspect

Full message history for ONE thread. Input: threadId (from inbox). Returns { threadId, campaignId, leadId, leadEmail, mailboxEmail (null before first send), messages[] }, each message { type (sent/reply/bounce/...), ts, messageId, metadata }, oldest first. 404 if unknown. Use inbox to LIST threads; reply to respond; mark/label_thread to triage.

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdYesThe thread id, e.g. from inbox() or campaign events.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already set readOnlyHint=true. Description adds specifics: return shape with nullable mailboxEmail, oldest-first order, and error behavior. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single paragraph covering purpose, input, output, error, and sibling hints. No redundancy; efficient but could be slightly more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one parameter, no output schema, and high schema coverage, the description provides complete context: purpose, input, output shape, error, and usage alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage with clear description for threadId. Description mentions it but adds no new semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Explicitly states 'Full message history for ONE thread' with specific verb and resource. Distinguishes from siblings by naming inbox, reply, and label_thread.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear when-to-use: 'Use inbox to LIST threads; reply to respond; mark/label_thread to triage.' Also notes 404 error for unknown threads.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_leadAInspect

Record what you learned about a contact (their reply, your triage) as a durable, contact-level disposition — keyed by email, visible across every campaign that lists them. Inputs: email, interestStatus (none|interested|meeting_booked|not_now|not_interested|bad_fit|out_of_office|wrong_person — a server-enforced enum; 'do not contact' is NOT a member, use suppress_lead instead), notes, tags (free-form). A PARTIAL patch — only the fields you pass are changed; at least one of interestStatus/notes/tags is required. Filterable via list_leads.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
emailYes
notesNo
interestStatusNo

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses key behavioral traits beyond annotations: it states it's a partial patch (only passed fields changed), requires at least one of interestStatus/notes/tags, and mentions server-enforced enum validation. Annotations only provide destructiveHint=false, so the description adds useful context. However, it does not mention authorization needs or rate limits, which are minor gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear lead sentence followed by a parameter list. It is relatively long but each sentence adds value. Minor redundancy: 'Inputs: ' could be integrated. Overall, it's concise for the amount of detail provided.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 parameters, no output schema, and low schema coverage, the description thoroughly covers inputs, behavior (partial patch), and relationships to sibling tools (suppress_lead, list_leads). It also notes the contact-level scope and cross-campaign visibility. No missing aspects for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description fully carries the burden. It explains each parameter's role: email as key, interestStatus with explicit enum values and the note that it's server-enforced, notes and tags with constraints (maxLength, maxItems). It also clarifies that at least one optional param is required, adding meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: recording contact-level disposition keyed by email. It uses specific verbs ('record what you learned') and resource ('contact-level disposition'), and distinguishes itself from sibling tool 'suppress_lead' by explicitly noting that 'do not contact' is not a valid interestStatus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool versus alternatives: it notes that 'do not contact' should be handled via suppress_lead. It also mentions that the tool is filterable via list_leads, offering context for integration. No when-not-to-use scenarios are omitted.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updates
    • Changedlaunch_campaign3 fields changed
      • addedInput schema / properties / endBy
        Added value: +{
        +  "description": "Optional fixed end date, as an ISO 8601 instant with a zone. From then on nothing more in this campaign sends: every step still waiting is skipped and recorded as a 'failed' event with reason campaign_ended. Paid accounts only: a trial account that sends one is refused 400. Must be later than now and later than every step's notBefore. Omit to let the sequence run to its end.",
        +  "format": "date-time",
        +  "pattern": "^(?:(?:\\d\\d[2468][048]|\\d\\d[13579][26]|\\d\\d0[48]|[02468][048]00|[13579][26]00)-02-29|\\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\\d|30)|(?:02)-(?:0[1-9]|1\\d|2[0-8])))T(?:(?:[01]\\d|2[0-3]):[0-5]\\d(?::[0-5]\\d(?:\\.\\d+)?)?(?:Z|([+-](?:[01]\\d|2[0-3]):[0-5]\\d)))$",
        +  "type": "string"
        +}
      • addedInput schema / properties / listSource
        Added value: +{
        +  "description": "Where this lead list came from, in a sentence (e.g. \"our own past-donor CRM export\"). REQUIRED on a crowdfunding account: a launch with no source, or \"unknown\", is refused (400) — cold lists are allowed only with a stated source. Optional elsewhere. Max 500 characters.",
        +  "maxLength": 500,
        +  "type": "string"
        +}
      • addedInput schema / properties / sequence / items / properties / notBefore
        Added value: +{
        +  "description": "Optional fixed date this step is held until, as an ISO 8601 instant with a zone (e.g. \"2026-11-03T09:00:00-05:00\" or \"...Z\"). The step goes out at the later of notBefore and delayDays after the previous step actually sent, still inside sendWindow. Paid accounts only: a trial account that sends one is refused 400, so use delayDays while on trial. Omit to send on delayDays alone.",
        +  "format": "date-time",
        +  "pattern": "^(?:(?:\\d\\d[2468][048]|\\d\\d[13579][26]|\\d\\d0[48]|[02468][048]00|[13579][26]00)-02-29|\\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\\d|30)|(?:02)-(?:0[1-9]|1\\d|2[0-8])))T(?:(?:[01]\\d|2[0-3]):[0-5]\\d(?::[0-5]\\d(?:\\.\\d+)?)?(?:Z|([+-](?:[01]\\d|2[0-3]):[0-5]\\d)))$",
        +  "type": "string"
        +}
    • Changedpause1 field changed
      • addedInput schema / properties / leadEmails
        Added value: +{
        +  "description": "Optional. Stop ONLY these leads in this campaign (any capitalisation): their remaining steps are cancelled, the campaign keeps running for everyone else, and nobody is added to the suppression list. Every address must be a lead of this campaign, or nothing is stopped and the call is refused naming the ones that are not. Omit to pause the whole campaign.",
        +  "items": {
        +    "format": "email",
        +    "pattern": "^(?!\\.)(?!.*\\.\\.)([A-Za-z0-9_'+\\-\\.]*)[A-Za-z0-9_+-]@([A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$",
        +    "type": "string"
        +  },
        +  "maxItems": 5000,
        +  "minItems": 1,
        +  "type": "array"
        +}
    • Addedplan_campaign
    • Addedsubmit_feedback
  2. 1 tool update
    • Changedlaunch_campaign7 fields changed
      • removedInput schema / properties / sendWindow / default
        Removed value: -{
        -  "endHour": 23,
        -  "startHour": 0
        -}
      • addedInput schema / properties / sendWindow / description
        Added value: +"Integer hours 0-23 (endHour inclusive: the last hour a send may start in) and weekdays, evaluated in `timezone`. Each field you omit (or all of them, by omitting sendWindow) gets the recommended {\"startHour\":8,\"endHour\":16,\"days\":[1,2,3,4,5]} — except on a sandbox account, where omitted fields are open every hour of every day until it upgrades and they take the recommended default. If you set exactly ONE of startHour/endHour, the other still defaults to the recommended value (08/16) — if that combination would wrap past midnight (start > end), the launch is refused; set both explicitly, including for a deliberate overnight window."
      • addedInput schema / properties / sendWindow / properties / days
        Added value: +{
        +  "description": "Weekdays sends may go out on, in `timezone`: 0 = Sunday … 6 = Saturday. Omit for [1,2,3,4,5] (every day while the account is a sandbox); include 0 and 6 to send on weekends.",
        +  "items": {
        +    "maximum": 6,
        +    "minimum": 0,
        +    "type": "integer"
        +  },
        +  "maxItems": 7,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • removedInput schema / properties / sendWindow / required
        Removed value: -[
        -  "startHour",
        -  "endHour"
        -]
      • addedInput schema / properties / sequence / items / properties / subject / description
        Added value: +"Required on step 1. A LATER step (step > 1) may omit it to keep step 1's subject line (recommended) — it resolves at launch to \"Re: \" + step 1's subject (unchanged if that already starts with \"Re: \"). Set one explicitly on a later step to change only its subject line: every later step is still sent In-Reply-To step 1, though some mail clients display a changed subject as a separate conversation."
      • changedInput schema / properties / sequence / items / required
        Previous value: -[
        -  "step",
        -  "subject",
        -  "body",
        -  "delayDays"
        -]New value: +[
        +  "step",
        +  "body",
        +  "delayDays"
        +]
      • addedInput schema / properties / timezone / description
        Added value: +"IANA time zone the send window is evaluated in — set it to your recipients' zone. Default UTC."
  3. 1 tool update
    • Changedconfigure_webhook1 field changed
      • changedInput schema / properties / eventTypes / description
        Previous value: -"Required for create: which events to push (reply | bounce | soft_bounce | complaint)."New value: +"Required for create: which events to push (reply | bounce | soft_bounce | complaint | unsubscribe)."
  4. 1 tool update
    • Changedsetup_infrastructure2 fields changed
      • addedInput schema / properties / distribution
        Added value: +{
        +  "items": {
        +    "maximum": 10,
        +    "minimum": 1,
        +    "type": "integer"
        +  },
        +  "maxItems": 20,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "brand",
        -  "primaryDomain",
        -  "domains",
        -  "inboxesEach",
        -  "persona",
        -  "physicalAddress",
        -  "senderIdentity"
        -]New value: +[
        +  "brand",
        +  "primaryDomain",
        +  "domains",
        +  "persona",
        +  "physicalAddress",
        +  "senderIdentity"
        +]
  5. 1 tool update
    • Addedcontact_operator
  6. 3 tool updates
    • Addedack_message
    • Addedlist_messages
    • Changedremove_mailboxes1 field changed
      • addedInput schema / properties / idempotencyKey
        Added value: +{
        +  "description": "Optional idempotency key: resend the SAME key when retrying this call so a dropped-response retry is not applied twice (no duplicate campaign/provision/send).",
        +  "maxLength": 200,
        +  "minLength": 1,
        +  "type": "string"
        +}
  7. 1 tool update
    • Changedsetup_infrastructure1 field changed
      • removedInput schema / properties / registerDomains / default
        Removed value: -false
  8. 1 tool update
    • Changedsetup_infrastructure2 fields changed
      • addedInput schema / properties / registerDomains
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
      • addedInput schema / properties / registrant
        Added value: +{
        +  "properties": {
        +    "addressLine1": {
        +      "maxLength": 500,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "city": {
        +      "maxLength": 200,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "country": {
        +      "maxLength": 100,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "email": {
        +      "format": "email",
        +      "pattern": "^(?!\\.)(?!.*\\.\\.)([A-Za-z0-9_'+\\-\\.]*)[A-Za-z0-9_+-]@([A-Za-z0-9][A-Za-z0-9\\-]*\\.)+[A-Za-z]{2,}$",
        +      "type": "string"
        +    },
        +    "firstName": {
        +      "maxLength": 200,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "lastName": {
        +      "maxLength": 200,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "organization": {
        +      "maxLength": 200,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "phone": {
        +      "maxLength": 50,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "postalCode": {
        +      "maxLength": 20,
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    "state": {
        +      "maxLength": 200,
        +      "minLength": 1,
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "firstName",
        +    "lastName",
        +    "email",
        +    "phone",
        +    "addressLine1",
        +    "city",
        +    "state",
        +    "country",
        +    "postalCode"
        +  ],
        +  "type": "object"
        +}
  9. 3 tool updates
    • Changedconfigure_byo_domain1 field changed
      • addedInput schema / properties / quoteOnly
        Added value: +{
        +  "description": "Optional for request_managed_mailboxes — true previews the new mailbox count + projected monthly price WITHOUT provisioning (SPEC §18 quote-before-add).",
        +  "type": "boolean"
        +}
    • Addedremove_mailboxes
    • Changedsetup_infrastructure1 field changed
      • addedInput schema / properties / quoteOnly
        Added value: +{
        +  "default": false,
        +  "type": "boolean"
        +}
  10. 4 tool updates
    • Changedconfigure_webhook2 fields changed
      • changedInput schema / properties / eventTypes / items / enum
        Previous value: -[
        -  "reply",
        -  "bounce",
        -  "soft_bounce",
        -  "complaint"
        -]New value: +[
        +  "reply",
        +  "bounce",
        +  "soft_bounce",
        +  "complaint",
        +  "unsubscribe"
        +]
      • changedInput schema / properties / eventTypes / maxItems
        Previous value: -4New value: +5
    • Addedlist_leads
    • Addedsuppress_lead
    • Addedupdate_lead

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Email-deliverability tools for AI agents — 12 MCP tools across email verification, DNSBL across 50 zones, SPF/DKIM/DMARC analysis, spam-trap scoring, domain intelligence, and email finder. Free tier with no credit card.
    12
    38 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for managing cold-email infrastructure: buy domains via Porkbun, import them into CheapInboxes, provision Google/Microsoft mailboxes, and sync credentials to platforms like Instantly and Smartlead.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Argorant MCP Server — give your AI agent direct access to 614M verified B2B contacts. Query by industry, role, geography, and 100+ filters, then export emails verified by a live SMTP probe at request time (catch-alls flagged, invalids free). OAuth-secured. Works with Claude, ChatGPT, Cursor, and any MCP client. Endpoint: https://mcp.argorant.com/mcp
    1
    -
  • A
    license
    A
    quality
    A
    maintenance
    MCP server for SentVia that provides email infrastructure for AI agents, enabling them to create inboxes, send, reply, forward, search messages, manage drafts, domains, webhooks, and allow/block rules through 21 tools.
    21
    237 npm
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources