AgentPhone MCP Server
OfficialAllows sending iMessage messages with support for media, threaded replies, send effects, and group chats.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AgentPhone MCP ServerBuy me a phone number in the 415 area code"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AgentPhone MCP Server
Give AI agents real phone numbers, SMS, and voice calls via the Model Context Protocol.
AgentPhone lets your AI agent buy phone numbers, send/receive SMS, and place voice calls — all through natural language in Cursor, Claude Desktop, or any MCP-compatible client.
Agents are the core concept — each agent gets its own phone numbers, voice personality, system prompt, and webhook. Think of an agent as a virtual team member with its own phone line. You can create agents for different purposes (support, sales, scheduling) and configure how they sound and behave on calls.
Quick Start
1. Get your API key
Sign up at agentphone.ai and create an API key from Settings.
2. Connect via MCP
Option A: Remote server (recommended)
Point your MCP client at the hosted endpoint — no install needed:
{
"mcpServers": {
"agentphone": {
"type": "streamable-http",
"url": "https://mcp.agentphone.ai/mcp",
"headers": {
"Authorization": "Bearer your_api_key_here"
}
}
}
}Works with any MCP client that supports Streamable HTTP transport (Switchboard, remote agent platforms, etc.).
Option B: Local server (stdio)
Runs locally via npx — works with Cursor, Claude Desktop, Windsurf, and Claude Code:
Cursor: Settings > MCP or ~/.cursor/mcp.json
Claude Desktop: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
{
"mcpServers": {
"agentphone": {
"command": "npx",
"args": ["-y", "agentphone-mcp"],
"env": {
"AGENTPHONE_API_KEY": "your_api_key_here"
}
}
}
}Option C: Self-hosted HTTP server
Run your own HTTP MCP endpoint:
AGENTPHONE_API_KEY=your_api_key npx agentphone-mcp --http --port 3000Then connect to http://localhost:3000/mcp.
Related MCP server: agentline-mcp
What Can It Do?
Once configured, just ask your AI agent things like:
"Buy me a phone number in the 415 area code"
"Create a support agent that greets callers and helps with billing"
"Call +14155551234 and have a conversation about scheduling a dentist appointment"
"Text +14155551234 saying 'Your appointment is confirmed for 3pm tomorrow'"
"Show me my recent calls and transcripts"
"List the available voices and switch my agent to a different one"
"Set up a webhook so I get notified when someone calls or texts my number"
"Show me this month's usage breakdown"
Transport & hosting
The server is built on the mcp-use server framework,
which owns the HTTP layer: Streamable HTTP, the SSE stream, session management,
and the OAuth discovery endpoints. npm start (or the Docker image) runs it as
an HTTP server on PORT (default 3000), reachable at /mcp.
Hosted:
https://mcp.agentphone.ai/mcpSelf-hosted:
PORT=3000 npm start→http://localhost:3000/mcp
Authentication
OAuth (recommended for end users): the framework proxies an Authorization Code + PKCE flow to the AgentPhone authorization server, so the client opens a browser to sign in at agentphone.ai — no key to paste. Enable it by setting
MCP_OAUTH_CLIENT_ID(a client pre-registered with the AgentPhone AS).MCP_OAUTH_CLIENT_SECRETis optional: set it to run the gateway as a confidential client, or leave it unset to run as a public client (the gateway must then be registered withtoken_endpoint_auth_method=none). Public mode advertisesnoneto downstream clients, which strict OAuth clients require.API key (scripts / single-tenant): set
AGENTPHONE_API_KEY. Used as the fallback credential when no OAuth token is present.
The per-request access token is forwarded to the AgentPhone REST API, so the server stores no credentials.
Environment
Var | Purpose |
| HTTP port (default 3000) |
| Fallback API key when OAuth is off |
| Enable OAuth; a gateway client pre-registered with the AgentPhone AS |
| Optional. Set = confidential gateway; unset = public client (gateway registered with |
| API base (default |
| Override authorize/consent URL (default |
Highlights
Phone numbers — buy and manage numbers in any US/CA area code
SMS — send and receive text messages, view conversation threads
Voice calls — place outbound calls with built-in AI conversation (no webhook needed) or bring your own webhook
Inbound handling — set up webhooks to receive and respond to inbound calls and texts in real time
Agents — create agents with custom voices, system prompts, call transfer, and voicemail
Usage & billing — monitor number usage, message/call/webhook volume, and daily/monthly breakdowns
All Tools (28)
Account
Tool | Description |
| Get a full snapshot of your account — agents, numbers, webhook, and usage |
| Get usage stats — number usage, billed SMS segments, and message/call/webhook volume. Use |
Phone Numbers
Tool | Description |
| List all phone numbers in your account |
| Purchase a new phone number with optional |
SMS
Tool | Description |
| Send SMS or iMessage. Supports media, threaded replies ( |
| Get messages for a specific number |
| List SMS conversations. Pass |
| Get a conversation with full message history |
| Set metadata on a conversation |
Contacts
Tool | Description |
| List saved contacts (address book). Filter with a |
| Create, update, or delete a contact (set |
Voice Calls
Tool | Description |
| List calls. Filter by |
| Get call details and transcript |
| Place an outbound call (webhook-driven) |
| Place a call with built-in AI conversation — no webhook needed |
Agents
Tool | Description |
| List all agents with their numbers and voice config |
| Create an agent with voice, system prompt, call transfer, voicemail, and voice tuning (speed, interruption sensitivity, backchannel, language, and more) |
| Update an agent's configuration |
| Delete an agent (numbers are kept but unassigned) |
| Get agent details including phone numbers and voice config |
| Assign a phone number to an agent |
| Remove a phone number from an agent |
| List available voices for agents |
Webhooks
All webhook tools accept an optional agent_id — pass it to manage an agent-specific webhook, omit it for the project-level default. Agent webhooks take priority over project-level.
Tool | Description |
| Get webhook configuration |
| Set a webhook URL for inbound messages and call events |
| Remove a webhook |
| Send a test event to verify your webhook works |
| View delivery history for debugging |
Environment Variables
Variable | Required | Description |
| stdio: yes, HTTP: no | Your AgentPhone API key (HTTP mode can use Authorization header instead) |
| No | Override the API base URL (defaults to |
| No | Port for HTTP mode (defaults to |
Development
git clone https://github.com/AgentPhone-AI/agentphone-mcp.git
cd agentphone-mcp
npm install
npm run dev # Run with tsx (hot reload)
npm run build # Compile TypeScript
npm start # Run compiled JS (stdio)How It Works
This MCP server connects your AI assistant to the AgentPhone API. Your assistant talks to the MCP server, which calls the AgentPhone API, which talks to the phone network.
Your AI Assistant <--> agentphone-mcp <--> AgentPhone API <--> Phone NetworkOutbound: your assistant places calls and sends texts through AgentPhone's API.
Inbound: when someone calls or texts your number, AgentPhone sends a webhook event to your server — you can then respond programmatically or let your agent's built-in AI handle it.
License
MIT
Available Tools
28 toolsaccount_overviewARead-onlyIdempotentInspect
Get a complete snapshot of your AgentPhone account: agents, phone numbers, webhook status, and usage limits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds real value beyond that by disclosing what the snapshot contains (agents, numbers, webhook status, usage limits), which is important since no output schema exists to reveal the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the verb and scope, with the returned contents listed after the colon. No filler and no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only aggregation tool with no output schema, the description gives enough to call it correctly and to know roughly what comes back. It stops short of noting relationships to the granular sibling tools, which would have made routing unambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to clarify about arguments, and the schema is trivially complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Get') plus resource ('a complete snapshot of your AgentPhone account') with an explicit enumeration of the contents (agents, phone numbers, webhook status, usage limits). This aggregate framing implicitly separates it from the many per-resource siblings (list_agents, list_numbers, get_usage), though it never names an alternative outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the entry-point/orientation call rather than a per-resource query, but there is no explicit when-to-use, no when-not-to-use, and no pointer to the sibling tools that give finer-grained data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_numberBIdempotentInspect
Attach a phone number to an agent so the agent handles calls/SMS on that number. Takes an agent ID and a phone number ID.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | The agent ID | |
| number_id | Yes | The phone number ID to attach |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare idempotentHint, so the description carries most of the burden. It usefully explains the observable consequence of attaching (the agent starts handling calls/SMS on that number), but omits permissions, conflict/replacement behavior, and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose and effect front-loaded with no filler. The second sentence is largely redundant with the required-parameter list in the schema, a minor waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with full schema coverage and no output schema, the core purpose is covered. However, nothing is said about where to obtain a valid number_id, whether the number must be purchased first, or what happens if the number is already attached.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented in the schema; the description merely restates 'takes an agent ID and a phone number ID' without adding format, source-of-ID, or constraint details. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (attach) and resource (phone number to an agent) plus the concrete effect: the agent then handles calls/SMS on that number. The contrast with the sibling detach_number is implied by the verb, but the description never names alternatives or differentiates explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to attach vs. when not to, no prerequisites (e.g. the number must already exist in the account, whether an existing attachment is replaced), and no mention of the sibling detach_number or list_numbers for discovering number IDs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
buy_numberAInspect
Purchase a new phone number. This is a paid purchase that charges the account's balance. area_code requests a specific region (e.g. '415' for San Francisco); agent_id attaches it on purchase. Requires confirm=true to proceed.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Must be true to complete the purchase. When false, no number is bought. | |
| country | No | 2-letter ISO country code (e.g. US, CA, GB) | US |
| agent_id | No | Agent ID to attach this number to immediately | |
| area_code | No | 3-digit area code for a specific region (e.g. '415', '212', '310') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply openWorldHint=true, so the description must carry the rest, and it does: it discloses that this is a paid purchase that charges the account balance and that confirm must be true to complete it. It stops short of stating reversibility or what happens to a number once purchased, but the financial-impact disclosure is the key behavior an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action and cost, then the two notable parameters, then the confirm gate. Slightly repetitive with the schema (which already documents confirm and area_code), but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 100% schema coverage, minimal annotations, and no output schema, the description covers the essentials: what it does, that it costs money, the confirm gate, and the two behaviorally interesting parameters. It omits what a successful result returns and whether country defaults matter, which are minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds interpretive value: area_code maps to a region with a concrete example ('415' for San Francisco) and agent_id is described as attaching the number on purchase. That clarifies the effect of agent_id beyond the schema's terse wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Purchase a new phone number'), and the word 'new' implicitly separates it from attach_number, which attaches an existing number. It never names that sibling explicitly, so it falls short of the 5-tier bar for sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one concrete prerequisite ('Requires confirm=true to proceed') and notes the cost, which is genuine usage guidance. However, it never says when to prefer this over attach_number or how it relates to list_numbers for selecting a number, leaving the choice to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_agentAInspect
Create a new agent. An agent owns phone numbers and handles calls/SMS.
A newly created agent has no phone number until one is attached. Set voice_mode to 'hosted' with a system_prompt for autonomous AI voice calls, or 'webhook' (default) to forward call transcripts to your webhook URL.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name for the agent (e.g. 'Customer Support', 'Sales Bot') | |
| voice | No | Voice ID for the agent. Defaults to 'Skylar - Friendly Guide'. | |
| language | No | BCP-47 locale for speech recognition and synthesis, e.g. 'en-US', 'es-ES', 'ja-JP'. | |
| stt_mode | No | Speech-to-text mode: 'fast' (default) lowest latency, 'accurate' for exact names/numbers (~200ms slower). | |
| model_tier | No | Model quality/speed tier for hosted-mode agents. 'turbo' = fastest/cheapest, 'balanced' (default) = general use, 'max' = highest quality. | |
| voice_mode | No | 'webhook' (default) forwards transcripts to your webhook. 'hosted' uses built-in AI with system_prompt. | |
| description | No | Description of what this agent does | |
| voice_speed | No | Speech speed multiplier. 1.0 = normal, 0.5 = half speed, 2.0 = double. | |
| ambient_sound | No | Optional background ambience for call audio. Not a substitute for the AI self-identification, which is always disclosed. | |
| begin_message | No | What the AI says when a call connects. Only used in 'hosted' mode. | |
| system_prompt | No | Required when voice_mode is 'hosted'. The AI's personality and instructions for voice calls. | |
| denoising_mode | No | Audio denoising. The aggressive mode helps callers in cars/cafes (small surcharge). | |
| max_silence_ms | No | Hang up after this many ms of caller silence. Default 600000 (10 min). | |
| transfer_number | No | Phone number to transfer calls to (E.164 format). Enables call transfer during conversations. | |
| enable_messaging | No | Whether a hosted agent can send and read texts during a call. Default true. | |
| voicemail_message | No | Voicemail greeting text. When set, unanswered calls hear this message and can leave a voicemail. | |
| enable_backchannel | No | Whether the agent interjects 'uh-huh'/'mhmm' while the caller speaks. Default true. | |
| interruption_sensitivity | No | How easily callers can interrupt (barge in). 0 = never, 1 = at first sound. Default 0.8. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries full disclosure burden. It usefully reveals that a newly created agent has no phone number until attached and clarifies the voice_mode branching, but says nothing about permissions, side effects, or what the creation returns — significant gaps for an 18-parameter mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then the phone-number caveat, then mode selection. Zero waste and well ordered for scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers the key branching decision (hosted vs webhook) and the post-creation prerequisite (attach a number). It could mention the response shape or default behavior, but it is adequate for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 18 parameters thoroughly. The description reinforces the voice_mode/system_prompt relationship but adds little beyond what the schema states. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a new agent') and adds real domain context — an agent owns phone numbers and handles calls/SMS. This distinguishes it from list_agents/update_agent/get_agent, though it does not explicitly name a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides actionable mode-selection guidance: 'hosted' with a system_prompt for autonomous AI calls, or 'webhook' (default) to forward transcripts. It also implies the follow-up step (attach_number) by noting a new agent has no phone number. No explicit when-not or sibling routing, but clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_agentADestructiveInspect
Delete an agent permanently. Phone numbers attached to it are kept but unassigned. This cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | The agent ID to delete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=true; the description goes further by disclosing the side effect on attached phone numbers (kept but unassigned) and that the action is irreversible. This is exactly the kind of relational consequence an agent needs before committing a destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, with the core action front-loaded and the irreversibility warning last for emphasis. Every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with no output schema and destructiveHint already set, the description supplies the missing pieces: permanence and the effect on related phone numbers. Nothing needed to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the schema already explains agent_id. The description adds no format or sourcing detail beyond it, making 3 the appropriate baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete an agent') plus scope ('permanently'), which cleanly separates it from the update_agent, get_agent, and delete_webhook siblings. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The permanence warning implies this should be used only for true removal, but there is no explicit when-to-use guidance or mention of alternatives such as deactivating via update_agent. Usage is inferable from the name rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_webhookADestructiveInspect
Remove a webhook. Pass agent_id to remove an agent's webhook (falls back to project default). Omit agent_id to remove the project-level webhook.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID to delete that agent's webhook. Omit for project-level. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, which sets the safety floor. The description adds the non-obvious fallback/scope semantics (agent-level delete falls back to project default; omitting targets project-level) that annotations cannot express. Does not state irreversibility explicitly, but the destructive hint plus 'remove' wording covers the essential traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the verb and resource, then the branching guidance. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a one-parameter destructive tool: what it does, the scope branching, and the fallback are all present. An output schema is absent but a delete's return value is conventionally minimal; no critical gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents agent_id, so baseline would be 3. The description adds the fallback behavior ('falls back to project default') that the schema field description does not convey, which is meaningful semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Remove a webhook') and immediately clarifies the two scopes (agent-level vs project-level). Distinguishable from siblings like get_webhook, set_webhook, and test_webhook without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains precisely when to pass agent_id versus omit it, which is the sole decision point for this tool. No explicit exclusions or named alternatives, but the branching logic is fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detach_numberBIdempotentInspect
Detach a phone number from an agent. The number is kept in your account but becomes unassigned.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | The agent ID that currently owns the number | |
| number_id | Yes | The phone number ID to detach |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation idempotentHint=true already tells the agent that repeated detaches are safe. The description adds that the number is kept but becomes unassigned, which is useful (non-destructive in terms of data retention), but doesn't cover permissions, side effects on the agent, or what happens if the number is already detached.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and result. No wasted words, though it could be slightly more structured for usage guidelines.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple mutation tool with full schema coverage and an idempotency annotation, the description covers the essential outcome. However, it lacks guidelines on when to use it and doesn't mention any side effects on the agent, leaving some gaps for an agent's decision-making.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already fully documents both parameters. The description adds no additional parameter information, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (detach) and resource (phone number), and clarifies the result. Clearly distinguishes from the sibling attach_number.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It doesn't mention prerequisites or when you might detach a number (e.g., before reassigning). The sibling attach_number is not referenced.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agentARead-onlyIdempotentInspect
Get details for a specific agent including its phone numbers, voice configuration, and instructions.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | The agent ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds useful context about what the response contains (phone numbers, voice config, instructions) but says nothing about auth requirements, error behavior for unknown IDs, or payload size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence that names the resource and the payload; no boilerplate or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only getter with annotations covering safety and a fully described single-param schema, the description is nearly sufficient; it helpfully enumerates the returned fields even though no output schema exists. Only the absence of any error/not-found behavior keeps this from being complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single agent_id parameter is already documented in the schema as 'The agent ID'. The description adds no format, prefix, or sourcing guidance beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (a specific agent) and enumerates the content returned: phone numbers, voice configuration, instructions. It clearly is not a list operation, but it never explicitly differentiates itself from list_agents or get_call siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'a specific agent' implicitly signals that this is the lookup to use when you already have an agent_id, versus list_agents for enumeration, but no when/when-not guidance or alternative is stated explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_callARead-onlyIdempotentInspect
Get details and transcript for a specific call, identified by its call ID. Pass wait=true to block until an in-progress call finishes before returning.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | When true, long-polls until the call completes before returning. Useful for checking back on a call you initiated earlier. | |
| call_id | Yes | The call ID | |
| timeout | No | Max seconds to wait when wait=true. Defaults to 120. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint and idempotentHint, so safety is covered; the description adds the non-obvious long-poll blocking behavior of wait=true, which the agent needs for scheduling. It stops short of describing pagination or how a still-in-progress call is represented on timeout.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the core retrieval purpose comes first and the wait modifier second, correctly front-loading the primary semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return burden and does name the payload ('details and transcript'). That is sufficient for correct invocation, though it could say more about what happens when the call is not yet complete or when timeout is reached.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so call_id, wait, and timeout are already documented in the schema, establishing a baseline of 3. The description restates wait=true's effect but adds no syntax or interaction detail (e.g., how timeout relates to wait) beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get details and transcript for a specific call') and pins the identifier to call_id. It implicitly separates itself from list_calls by scoping to a single call, though it doesn't name the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives actionable context for the wait=true mode ('block until an in-progress call finishes'), which implies when to use it, but it never states when not to use the tool or points to alternatives like list_calls or get_conversation for adjacent lookups.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_conversationBRead-onlyIdempotentInspect
Get a specific SMS conversation with full message history, identified by its conversation ID.
| Name | Required | Description | Default |
|---|---|---|---|
| message_limit | No | Max messages to include | |
| conversation_id | Yes | The conversation ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, covering the safety profile. The description adds that it returns full message history, but it does not clarify that message_limit caps the returned messages or describe pagination/rate behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-formed sentence with the key action and identifier front-loaded. There is no redundant or wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with full schema coverage and annotations, the description is nearly complete, stating the return scope. A minor gap is that 'full message history' is bounded by the message_limit parameter, which the description does not mention.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters including message_limit. The description mentions the conversation ID but adds no additional syntax or format details beyond the schema, making a baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Get a specific SMS conversation with full message history.' This clearly communicates the tool's core action. It does not explicitly distinguish itself from siblings like list_conversations or get_messages, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a conversation ID is available, but it provides no explicit when-to-use guidance, no when-not-to-use conditions, and no alternatives such as list_conversations or get_messages. This leaves routing to the agent's inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_messagesARead-onlyIdempotentInspect
Get SMS messages for a specific phone number, identified by its number ID. Returns individual messages rather than grouped conversation threads.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max messages to return | |
| number_id | Yes | The ID of the phone number |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds the useful behavioral detail that results are individual messages, not threads, but says nothing about pagination, ordering, or time-window behavior despite the limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero waste; the scope and the distinguishing return-granularity note are front-loaded and immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool this is essentially complete: it tells the agent what it fetches and how the output is shaped. Only minor gaps remain, such as ordering or pagination behavior, which the absent output schema would otherwise need to cover.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description restates that the number is identified by its ID but adds no syntax, format, or usage nuance beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) plus resource (SMS messages) scoped to a phone number identified by number ID. It also distinguishes itself from the conversation-thread siblings by explicitly saying it returns individual messages rather than grouped threads.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'rather than grouped conversation threads' implicitly routes the agent to the conversation tools when threads are wanted, giving clear context for selection. It does not name the alternative tool explicitly, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageARead-onlyIdempotentInspect
Get account usage statistics. By default returns a summary with number usage plus message, call, and webhook volume. Use breakdown='daily' or 'monthly' for time-series data.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Number of days to look back (only used with breakdown='daily') | |
| months | No | Number of months to look back (only used with breakdown='monthly') | |
| breakdown | No | 'summary' for number usage and totals, 'daily' for per-day breakdown, 'monthly' for per-month breakdown | summary |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the read-only contract is covered by structured data. The description adds genuinely useful context by stating exactly what the default response contains and how the breakdown modes change the shape of the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler: the subject and default behavior come first, then the mode-switching instruction. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by naming the contents of the default summary and the alternate breakdown outputs. Combined with full schema coverage of parameters, an agent has enough to call the tool correctly, though a note on which sibling to prefer would complete the picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – all three parameters, including the enum values and the days/months applicability notes, are documented in the schema. The description only restates the breakdown choices and adds no syntax, format, or constraint detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get account usage statistics') and enumerates what the return contains (number usage plus message, call, and webhook volume). It does not distinguish itself from the sibling account_overview, which an agent could plausibly confuse with a usage tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the default behavior and tells the agent to use breakdown='daily' or 'monthly' for time-series data, which is implied guidance on selecting a mode. There is no explicit when-to-use-this-vs-alternative statement (e.g., versus account_overview), so usage remains implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_webhookARead-onlyIdempotentInspect
Get the webhook configuration. Pass agent_id to get an agent-specific webhook, or omit for the project-level default.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID to get that agent's webhook. Omit for project-level webhook. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered and the bar is lower. The description's only added behavior is the omit-defaults-to-project-level rule, which is really parameter semantics; it says nothing about the return shape or behavior when no webhook is configured.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the parameter branching immediately after. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless-required read tool with full schema coverage and annotations covering safety, the description is nearly complete. The main gap is that no output schema exists, yet the description doesn't hint at what the returned configuration contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional parameter, so the schema carries the burden and the baseline is 3. The description's phrasing about agent_id closely mirrors the schema's own text, adding little new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (webhook configuration), and clarifies the two retrieval scopes: agent-specific vs project-level default. It does not explicitly contrast with siblings like set_webhook or list_webhook_deliveries, so it falls short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear situational guidance for the branch: pass agent_id for an agent-specific webhook, omit it for the project-level default. It stops short of naming alternatives or exclusions (e.g., when to use set_webhook or list_webhook_deliveries instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agentsARead-onlyIdempotentInspect
List all agents with their phone numbers and voice configuration. An agent is required before you can make calls — it owns phone numbers and handles voice/SMS.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds domain context that agents own phone numbers and handle voice/SMS, but says nothing about pagination or result ordering despite the limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and returned data, with the domain note following. No wasted words, though the second sentence drifts slightly into general platform context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with full annotation coverage and no output schema, the description adequately conveys what is returned and why it matters. Minor gaps around pagination and ordering remain, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional parameter fully documented in the schema ('Max results to return'). The description adds no parameter detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all agents') and even names the returned fields (phone numbers, voice configuration), which is more than the bare name. It distinguishes itself reasonably from get_agent by scope ('all'), but never explicitly contrasts with the sibling list/get tools, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful context ('an agent is required before you can make calls') that implies why you'd list agents, but gives no explicit when-to-use, when-not-to-use, or alternative routing versus get_agent, create_agent, or list_numbers. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_callsARead-onlyIdempotentInspect
List recent calls. Scope by agent_id or number_id, or use status/direction/search to filter globally.
When agent_id or number_id is passed, status/direction/search filters are not applied. Returns call summaries including each call's ID, direction, participants, and status.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return | |
| offset | No | Number of results to skip (for pagination) | |
| search | No | Search by phone number or keyword | |
| status | No | Filter by status: ringing, in-progress, completed, failed, busy, no-answer | |
| agent_id | No | Filter to calls for a specific agent | |
| direction | No | Filter by direction: 'inbound' or 'outbound' | |
| number_id | No | Filter to calls for a specific phone number |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint, idempotentHint), so the description adds value with the non-obvious silent-ignore behavior: status/direction/search are dropped when agent_id or number_id is supplied. It also discloses the returned summary fields, which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the primary action, then the scoping rule, then return contents. Every sentence carries load and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param, all-optional listing tool with no output schema, the description covers the tricky parameter interaction and the shape of the results. Pagination behavior (limit/offset) is left entirely to the schema, but that is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline would be 3, but the description contributes the mutual-exclusivity rule between scoping parameters and filter parameters that the schema does not express. That is genuine added meaning beyond per-field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recent calls') and enumerates the scoping/filtering dimensions available. It is clearly distinguishable from the sibling get_call (single call) by its listing nature, though it does not name siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing between two modes: scope by agent_id/number_id, or filter globally with status/direction/search. It further states the when-not case (filters are ignored when scoping params are present), which is exactly the guidance an agent needs. No named alternative functions, but the selection rule is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_contactsARead-onlyIdempotentInspect
List saved contacts (your address book). Optionally filter with a search term that matches name or phone number.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return | |
| offset | No | Number of results to skip (for pagination) | |
| search | No | Filter by name or phone number |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered without description help. The description adds only that contacts are an address book and that search spans name/phone; it says nothing about pagination behavior, result ordering, or result caps beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler, and the core purpose is front-loaded before the optional filter detail. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, fully-annotated, zero-required-parameter list tool with 100% schema coverage and no output schema, the description covers what an agent needs to call it. The only gap is the absent relationship to the manage_contact sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (limit, offset, search) are already documented with defaults and bounds. The description's note that search "matches name or phone number" restates the schema's own wording rather than adding new semantics, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List saved contacts") and disambiguates the resource with "(your address book)." It does not, however, distinguish itself from the sibling manage_contact, so an agent must infer the read-vs-write split on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the optional filter and what it matches, which implies how to use the tool, but it gives no explicit when-to-use guidance or exclusions relative to manage_contact or any search alternative. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_conversationsARead-onlyIdempotentInspect
List SMS conversations. Optionally filter by agent_id to see conversations for a specific agent.
Each conversation is a thread between your number and an external contact, with its ID, participant, message count, and last-message preview.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return | |
| offset | No | Number of results to skip (for pagination) | |
| agent_id | No | Filter to conversations for a specific agent |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds value by enumerating what each returned conversation contains (ID, participant, message count, last-message preview), which matters because there is no output schema. It stops short of describing pagination or ordering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, purpose front-loaded before the filter and the return-shape detail. No filler, though the trailing clause about conversation contents could be trimmed without losing selection-critical meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by describing the returned conversation fields, and all three params are covered by the schema. It is largely sufficient, missing only pagination/ordering behavior for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so limit, offset, and agent_id are already documented in the schema. The description restates the agent_id filter without adding format, default, or range guidance beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List SMS conversations') that cleanly separates it from the singular get_conversation and mutating update_conversation siblings. It does not explicitly name an alternative tool, but the operation type is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes the optional agent_id filter and its effect, which implies a usage context, but it never states when to prefer this tool over get_conversation or how agent-scoped listing differs from unfiltered listing. Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_numbersARead-onlyIdempotentInspect
List all phone numbers in your account, each with its ID, status, country, and assigned agent.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return | |
| offset | No | Number of results to skip (for pagination) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds the shape of the result (ID, status, country, assigned agent), which is genuinely useful since no output schema exists, but it says nothing about pagination behavior, result caps, or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action and the returned fields with zero filler; nothing is repeated from structured data unnecessarily.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates returned fields and, with only two well-documented optional parameters and strong read-only annotations, an agent has enough to call it correctly. Only the pagination/result-limit behavior is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (limit, offset) are documented in the schema, so the baseline is 3. The description adds no syntax, default, or interaction detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (phone numbers in your account) and enumerates the returned fields, so it is clearly distinct from buy_number or attach_number. It does not explicitly name siblings, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name and the 'in your account' scoping, but there is no explicit when-to-use statement, no mention of prerequisites, and no routing to related tools such as attach_number or buy_number when the desired number is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_voicesARead-onlyIdempotentInspect
List available voices for agents. Each voice has a voice_id used in an agent's voice field.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds genuinely useful behavioral context by explaining the shape of the result (each voice carries a voice_id) and how that value is consumed downstream in an agent's voice field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the purpose is front-loaded before the return-value detail. Every sentence earns its place by adding either scope or consumption guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only lookup with no output schema, the description covers what is returned and how the key field is used. It stops short of saying whether the list is filtered or scoped to the account, but nothing essential to a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. Nothing in the description is needed to compensate for schema gaps, and the schema coverage is fully reported as 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List available voices for agents'), so an agent immediately knows what the tool returns. It lacks explicit differentiation from siblings like list_agents, but the resource is distinct enough that confusion is unlikely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: by noting the voice_id feeds an agent's voice field, it hints this is the lookup step before create_agent/update_agent. There is no explicit when-to-use or when-not-to-use guidance, and no named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_webhook_deliveriesARead-onlyIdempotentInspect
View recent webhook delivery history. Shows which events were delivered, HTTP status codes, and timing.
Pass agent_id to see deliveries for that agent's webhook. Omit for project-level.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | Only show deliveries from the last N hours | |
| limit | No | Max results to return | |
| agent_id | No | Agent ID to see deliveries for that agent's webhook. Omit for project-level. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds the return content (event, status code, timing), which matters because there is no output schema, but it says nothing about ordering, pagination behavior, retention window, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences plus a one-line parameter note, with the core purpose front-loaded before the scoping detail. Efficient overall, though the agent_id sentence duplicates schema text rather than adding anything new.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with full schema coverage, covered safety annotations, and no output schema, the description supplies the essentials: what is listed, what each record shows, and how to scope the query. Missing pagination and retention details are minor omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema; the agent_id prose merely restates the schema text almost verbatim. The hours and limit parameters receive no additional prose meaning beyond their schema descriptions, matching the baseline 3 for a fully covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (view webhook delivery history) and enumerates what the listing contains — delivered events, HTTP status codes, timing. It is clearly distinguishable from config-oriented siblings like get_webhook or set_webhook, though it never names them explicitly to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives real scoping guidance — pass agent_id for an agent's webhook, omit for project-level — which is the main usage decision for this tool. It offers no when-to-use versus alternatives such as test_webhook or get_webhook, and no framing around when one would inspect delivery history at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_callAInspect
Initiate an outbound, webhook-driven phone call where your backend handles the conversation logic.
The agent must have a phone number attached and a webhook configured. Every call automatically opens with an automated-assistant disclosure (added server-side and not removable).
| Name | Required | Description | Default |
|---|---|---|---|
| voice | No | Voice ID override for this call | |
| agent_id | Yes | The agent ID (the agent must have a phone number attached) | |
| to_number | Yes | Recipient phone number in E.164 format (e.g. +14155551234) | |
| from_number_id | No | Specific phone number ID to call from (if agent has multiple numbers) | |
| initial_greeting | No | What the agent says when the call connects |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply openWorldHint=true, so the description carries the behavioral load and does it well: it discloses that the conversation is driven by the caller's webhook, and that every call opens with an automated-assistant disclosure added server-side and not removable. The latter is non-obvious, non-configurable behavior an agent must know before calling. It stops short of describing failure modes or call outcome.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler. The core action leads, and the prerequisite/disclosure constraints follow in a scannable second paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a call-initiation tool with no output schema and fully documented parameters, the description covers action, prerequisites, and a mandatory behavioral quirk. The remaining gap is that it never points to make_conversation_call as the alternative, which matters given how similar the two siblings are.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented, making 3 the baseline. The description adds one meaningful nuance: the opening disclosure is server-side and cannot be removed, which implicitly constrains what initial_greeting can control, but it does not explain voice override or from_number_id selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Initiate an outbound ... phone call') and adds the defining mechanism: webhook-driven, with the caller's backend handling conversation logic. That mechanism implicitly separates it from make_conversation_call, but the sibling is never named, so the differentiation requires inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a real precondition ('The agent must have a phone number attached and a webhook configured'), which tells the agent when the call can succeed. However, it never states when to pick this over make_conversation_call or what happens if the webhook is missing, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_conversation_callAInspect
Place a phone call where the AI has an autonomous conversation about a given topic — scheduling, surveys, follow-ups, etc. No webhook setup needed.
The agent must have a phone number attached. Every call automatically opens with an automated-assistant disclosure (added server-side; cannot be disabled via topic or initial_greeting). By default this blocks until the call finishes and returns the full transcript. Set wait=false for fire-and-forget.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | When true (default), blocks until the call ends and returns the full transcript. | |
| topic | Yes | The conversation goal for the assistant to pursue. Embedded as the call objective within locked instructions that always enforce the automated-assistant disclosure (it can't be overridden). Be specific about what to discuss and any goals. | |
| voice | No | Voice ID override for this call | |
| agent_id | Yes | The agent ID (must have a phone number attached) | |
| to_number | Yes | Recipient phone number in E.164 format (e.g. +14155551234) | |
| from_number_id | No | Specific phone number ID to call from (if agent has multiple numbers) | |
| initial_greeting | No | What the AI says when the call connects. If not set, the AI will generate one from the topic. | |
| max_wait_seconds | No | Maximum seconds to wait for the call to complete. Defaults to 300 (5 minutes). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only carry openWorldHint, so the description does the heavy lifting and does it well: it discloses a server-side, non-disableable automated-assistant disclosure, the default blocking behavior with transcript return, and the wait=false escape hatch. It omits cost, concurrency, and failure/retry behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, each carrying distinct information, with the core behavior and the blocking default front-loaded. Nothing is redundant with the schema or wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers prerequisites, side effects, blocking semantics, and return behavior (full transcript) for an 8-param tool with no output schema. Minor gaps remain around transcript format and any cost/latency expectations, but an agent has what it needs to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all 8 parameters are documented in the schema, so the baseline is 3. The description reinforces wait semantics and the locked nature of topic but adds no syntax or format detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Place a phone call') plus the distinguishing capability ('AI has an autonomous conversation about a given topic'), which separates it from the plain make_call sibling without naming it. The scope is clear enough that an agent can pick it over make_call from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives prerequisites ('The agent must have a phone number attached') and notes 'No webhook setup needed', which is useful context. However, it never explicitly says when to use this versus make_call or send_message, so the routing guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_contactADestructiveInspect
Create, update, or delete a saved contact. Set action to choose the operation.
create: requires phone_number and name
update: requires contact_id; only the fields you pass are changed
delete: requires contact_id (permanent)
contact_id refers to a saved contact.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Contact name. Required for 'create'. | |
| No | Contact email address | ||
| notes | No | Freeform notes about the contact | |
| action | Yes | The operation to perform | |
| contact_id | No | Contact ID. Required for 'update' and 'delete'. | |
| phone_number | No | Phone number in E.164 format (e.g. +14155551234). Required for 'create'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations only supply destructiveHint=true; the description adds real behavioral context on top of that, notably that delete is permanent and that update is a partial/PATCH-style change rather than a full replace. It still omits auth/permission requirements and what happens on a bad contact_id, keeping it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences: a front-loaded purpose line followed by a compact per-action bullet list. Every line carries actionable information (required fields, partial-update semantics, permanence) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with destructiveHint and no output schema, the description covers operation selection, per-action required fields, partial-update behavior, and delete permanence. Missing only error behavior (unknown contact_id) and permission context, which are minor for this scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter's per-action requirement (e.g. 'Required for create.') is already stated in the schema. The description restates these conditionals in one consolidated block, which aids readability but adds essentially no semantics beyond the structured fields, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set and resource ('Create, update, or delete a saved contact') and immediately enumerates the three operations, so the agent knows exactly what the tool does. It does not explicitly contrast itself with the read-side sibling list_contacts, but the mutation-only scope is unambiguous from the verbs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides per-action guidance: create requires phone_number and name, update requires contact_id and only changes passed fields, delete requires contact_id. This tells the agent which operation to select and what each needs. It does not name alternatives (e.g. use list_contacts to look up a contact_id first), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_messageAInspect
Send an SMS or iMessage from one of your agent's phone numbers, or react to a message.
The agent must have at least one phone number attached. If the agent has multiple numbers, number_id or from_number selects which one to send from. iMessage extras (silently ignored on SMS): reply_to_message_id threads the reply under an earlier message, and send_style adds an expressive screen/bubble effect. To react to a message instead of sending one, set reaction and react_to_message_id (iMessage only); to_number and body are not needed in that case.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | The message text to send (may be empty when sending media only) | |
| agent_id | No | The agent ID to send from (the agent must have a phone number attached). Optional if you pass from_number or number_id instead. | |
| reaction | No | iMessage only. Tapback or single emoji applied to react_to_message_id: one of love, like, dislike, laugh, emphasize, question, or a single emoji (e.g. '🔥'). When set, this reacts instead of sending a message. | |
| media_url | No | URL of a single image/media file to attach | |
| number_id | No | Specific phone number ID to send from (if agent has multiple numbers) | |
| to_number | No | Recipient: a phone number in E.164 format (e.g. +14155551234), a US short code, or a group ID (grp_...) to post into an iMessage group chat. Required when sending a message (not needed for a reaction). | |
| media_urls | No | Multiple media URLs to attach (delivered as an image carousel on iMessage) | |
| send_style | No | iMessage only. Expressive screen/bubble effect to send with the message | |
| from_number | No | Exact number to send from in E.164 format (alternative to number_id) | |
| react_to_message_id | No | iMessage only. ID of a message to react to. Set together with `reaction`. | |
| reply_to_message_id | No | iMessage only. ID of an earlier message to reply to inline (threaded reply) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare openWorldHint=true, so the description carries most of the burden and delivers real behavioral context: iMessage-only fields are 'silently ignored on SMS', reply_to_message_id threads a reply, and reaction mode changes which parameters are required. It omits permissions, delivery/failure behavior, and rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the send/extras/react structure; each sentence carries distinct information. Dense but readable, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with zero required params and no output schema, the description covers the primary send path, the number-selection path, and the reaction path, which is what an agent needs to invoke it. Remaining gaps (single vs. carousel media, error surfaces) are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds cross-parameter conditional semantics the schema does not state on its own: number_id/from_number select the sending number, and to_number/body are explicitly unnecessary when reacting. It does not clarify media_url vs media_urls interaction.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb and resource ('Send an SMS or iMessage') plus an explicit second mode ('or react to a message'), scoped to 'one of your agent's phone numbers'. An agent can immediately distinguish this from siblings like make_call or get_messages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the prerequisite ('agent must have at least one phone number attached') and the selection rule when multiple numbers exist, plus when to switch to reaction mode instead of sending. It does not name sibling alternatives (e.g. make_call for voice), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_webhookAIdempotentInspect
Set a webhook URL to receive inbound messages and call events.
Pass agent_id to set a webhook for a specific agent (overrides project default). Omit agent_id to set the project-level webhook for all agents. The webhook secret is returned — use it to verify signatures.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The publicly accessible webhook URL (must be HTTPS in production) | |
| timeout | No | Webhook response timeout in seconds | |
| agent_id | No | Agent ID to set webhook for that agent only. Omit for project-level. | |
| context_limit | No | Number of recent messages to include as conversation context (0-50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only idempotentHint in annotations, the description adds meaningful context: that a webhook secret is returned and must be used to verify signatures. That is a behavioral detail not captured by annotations. It still omits permission/auth requirements and whether an existing webhook is overwritten, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the scoping rule, then the return-value note. No filler, though the scoping sentence is somewhat verbose for the point it makes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema mutation tool, the description covers the key operational facts an agent needs: what it sets, the scope switch, and that a secret is returned. It leaves minor gaps (auth requirements, overwrite semantics) that are not critical to invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented. The description restates the agent_id scope behavior ('overrides project default') and adds nothing new for url, timeout, or context_limit. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set a webhook URL') and immediately clarifies its purpose ('to receive inbound messages and call events'). It does not explicitly name the sibling webhook tools (delete_webhook, get_webhook, test_webhook), but the agent/scope distinction makes the operation's identity clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the branching condition: pass agent_id for a per-agent webhook, omit it for the project default. It does not mention when to prefer alternatives such as delete_webhook or test_webhook, but the core usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
test_webhookAIdempotentInspect
Send a test event to verify a webhook is working. Returns the HTTP status code and response time.
Pass agent_id to test that agent's webhook. Omit to test the project-level webhook.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID to test that agent's webhook. Omit for project-level. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare openWorldHint=true and idempotentHint=true, so the external-effect and repeat-safety profile comes for free. The description adds value beyond that by disclosing the return payload (HTTP status code and response time), which is the key thing an agent needs to interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: what it does, what it returns, and how to scope the target. The purpose is front-loaded and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one optional parameter and no output schema, the description covers the essential contract, including the return values the schema cannot express. It does not mention failure modes (e.g., what a non-2xx status implies) or any prerequisite like a previously configured webhook.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single agent_id parameter is already documented in the schema. The description restates the same agent-level vs project-level semantics, adding emphasis but no new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('send a test event to verify a webhook'), and the scope distinction (agent-level vs project-level) sets it apart from get_webhook, set_webhook, delete_webhook, and list_webhook_deliveries without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Makes the usage condition explicit: use it to verify a webhook is working, and the agent_id-omit rule tells the caller exactly which webhook gets tested. It stops short of naming a sibling alternative or stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_agentAIdempotentInspect
Update an agent's configuration — name, description, voice settings, instructions, greeting, call transfer, or voicemail. Only provided fields are updated. Switching voice_mode to 'hosted' requires a system_prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name for the agent | |
| voice | No | Voice ID for the agent. | |
| agent_id | Yes | The agent ID to update | |
| language | No | BCP-47 locale for speech recognition and synthesis, e.g. 'en-US', 'es-ES', 'ja-JP'. | |
| stt_mode | No | Speech-to-text mode: 'fast' lowest latency, 'accurate' for exact names/numbers (~200ms slower). | |
| model_tier | No | Model quality/speed tier for hosted-mode agents. 'turbo' = fastest/cheapest, 'balanced' = general use, 'max' = highest quality. | |
| voice_mode | No | 'webhook' forwards transcripts to your webhook. 'hosted' uses built-in AI with system_prompt. | |
| description | No | New description | |
| voice_speed | No | Speech speed multiplier. 1.0 = normal, 0.5 = half speed, 2.0 = double. | |
| ambient_sound | No | Optional background ambience for call audio. Not a substitute for the AI self-identification, which is always disclosed. | |
| begin_message | No | What the AI says when a call connects (hosted mode only). | |
| system_prompt | No | The AI's personality and instructions. Required when voice_mode is 'hosted'. | |
| denoising_mode | No | Audio denoising. The aggressive mode helps callers in cars/cafes (small surcharge). | |
| max_silence_ms | No | Hang up after this many ms of caller silence. Default 600000 (10 min). | |
| transfer_number | No | Phone number to transfer calls to (E.164 format), or empty string to remove. | |
| enable_messaging | No | Whether a hosted agent can send and read texts during a call. Default true. | |
| voicemail_message | No | Voicemail greeting text, or empty string to disable voicemail. | |
| enable_backchannel | No | Whether the agent interjects 'uh-huh'/'mhmm' while the caller speaks. Default true. | |
| interruption_sensitivity | No | How easily callers can interrupt (barge in). 0 = never, 1 = at first sound. Default 0.8. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare idempotentHint, so the description carries the remaining burden. The partial-update disclosure is genuinely useful behavior beyond the annotation. However, the voice_mode/system_prompt dependency largely restates the schema's own 'Required when voice_mode is hosted' note, and nothing is said about permissions or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the verb+resource, and the second sentence carries the highest-value constraint. Nothing is padded or repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a rich 19-parameter input schema, the description supplies the two things the schema can't: patch semantics and a cross-field precondition. Complete enough for correct invocation, though it omits any note on what the response returns after an update.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents ranges, enums, defaults, and formats for all 19 fields. The description names field groups but adds no syntax or format detail beyond the schema, so it sits at the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (agent configuration), then enumerates the configurable surface: name, description, voice, instructions, greeting, transfer, voicemail. An agent can immediately tell this is the mutation counterpart to create_agent/get_agent. It stops short of explicitly naming a sibling, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Only provided fields are updated' gives the caller clear patch-semantics guidance and the co-requirement for hosted mode is a concrete precondition. There is no explicit when-to-use-this-vs-sibling routing (e.g., when to prefer update_agent over create_agent), which keeps it out of the 5 band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_conversationAIdempotentInspect
Set metadata on a conversation. Use this to store custom state, tags, or context that persists between messages. Pass null to clear metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| metadata | Yes | JSON metadata object to store on the conversation, or null to clear | |
| conversation_id | Yes | The conversation ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply idempotentHint=true, and the description adds real value on top: metadata persists between messages and null clears it. However, for an update tool it omits whether a new metadata object replaces or merges with existing keys, which is the key behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero waste, and the core purpose is front-loaded before the usage note and the clearing rule. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, full schema coverage, and an idempotency annotation, the description covers purpose, usage, and clearing behavior adequately. The only gap is replace-vs-merge semantics, which is minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters documented in the schema, so the baseline is 3. The description's note about passing null to clear metadata largely restates the schema's own 'or null to clear' description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Set') and resource ('metadata on a conversation'), and usefully narrows the broad name 'update_conversation' to a metadata-scoped operation. It does not explicitly contrast with siblings like get_conversation or list_conversations, but the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this to store custom state, tags, or context that persists between messages' gives clear context for when the tool applies. There are no overlapping siblings to exclude and no explicit when-not guidance, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
28 tool updates
v0.7.0- First observed
account_overview - First observed
attach_number - First observed
buy_number - First observed
create_agent - First observed
delete_agent - First observed
delete_webhook - First observed
detach_number - First observed
get_agent - First observed
get_call - First observed
get_conversation - First observed
get_messages - First observed
get_usage - First observed
get_webhook - First observed
list_agents - First observed
list_calls - First observed
list_contacts - First observed
list_conversations - First observed
list_numbers - First observed
list_voices - First observed
list_webhook_deliveries - First observed
make_call - First observed
make_conversation_call - First observed
manage_contact - First observed
send_message - First observed
set_webhook - First observed
test_webhook - First observed
update_agent - First observed
update_conversation
TDQS
Scored across 28 tools
Most tools have clearly distinct resource+action targets, and the two call tools (make_call vs make_conversation_call) are well-differentiated by their descriptions. Minor overlap exists in the messaging domain where get_messages (raw per-number messages) could be confused with get_conversation (threaded history), but descriptions clarify the boundary.
The set overwhelmingly follows a consistent verb_noun pattern (list_agents, create_agent, get_webhook, send_message). The main deviations are account_overview (noun-first) and manage_contact (a bundled multi-operation CRUD tool) versus the otherwise granular per-verb style.
28 tools is on the heavy side of the recommended range, though the domain (agents, numbers, calls, SMS, conversations, contacts, webhooks, usage) is genuinely broad. Almost every tool earns its place with little redundancy, but the count pushes into borderline-heavy territory.
Coverage is strong: full CRUD for agents and contacts, complete webhook lifecycle (get/set/delete/test/deliveries), and call/message/conversation operations. The notable gap is the phone-number lifecycle — buy_number exists but there is no release/delete or update number operation.
Maintenance
Related MCP Connectors
Give AI agents a phone number. Voice calls, SMS, and phone number management for MCP clients.
Give AI agents a phone layer for consent-based calls, transcripts, summaries, and outcomes.
Build, test, deploy and run AI phone agents: agents, numbers, calls, tests, knowledge, campaigns.
Give AI agents secure access to RevDesk calling, SMS, phone numbers, caller IDs, and usage.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to make real-world phone calls with AI voice technology and provides tools to track call status, transcripts, and summaries. It supports automated communication with both live numbers and simulated businesses for testing and demonstration purposes.-
- AlicenseAqualityCmaintenanceGives AI agents phone numbers, email, SMS, and voice calls as MCP tools, enabling them to provision numbers, capture 2FA codes, send messages, and make calls.15MIT
- AlicenseAqualityDmaintenanceGives AI agents real phone numbers to receive SMS and extract verification codes through tool calls.631 npmMIT
- AlicenseAqualityDmaintenanceEnables AI assistants to control PSTN phone calls via TelePath, including dialing, hangup, and phone number management.115 npmMIT