Skip to main content
Glama
AgentPhone-AI

AgentPhone MCP Server

Official

AgentPhone MCP Server

Give AI agents real phone numbers, SMS, and voice calls via the Model Context Protocol.

AgentPhone lets your AI agent buy phone numbers, send/receive SMS, and place voice calls — all through natural language in Cursor, Claude Desktop, or any MCP-compatible client.

Agents are the core concept — each agent gets its own phone numbers, voice personality, system prompt, and webhook. Think of an agent as a virtual team member with its own phone line. You can create agents for different purposes (support, sales, scheduling) and configure how they sound and behave on calls.

Quick Start

1. Get your API key

Sign up at agentphone.ai and create an API key from Settings.

2. Connect via MCP

Option A: Remote server (recommended)

Point your MCP client at the hosted endpoint — no install needed:

{
  "mcpServers": {
    "agentphone": {
      "type": "streamable-http",
      "url": "https://mcp.agentphone.ai/mcp",
      "headers": {
        "Authorization": "Bearer your_api_key_here"
      }
    }
  }
}

Works with any MCP client that supports Streamable HTTP transport (Switchboard, remote agent platforms, etc.).

Option B: Local server (stdio)

Runs locally via npx — works with Cursor, Claude Desktop, Windsurf, and Claude Code:

Cursor: Settings > MCP or ~/.cursor/mcp.json Claude Desktop: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)

{
  "mcpServers": {
    "agentphone": {
      "command": "npx",
      "args": ["-y", "agentphone-mcp"],
      "env": {
        "AGENTPHONE_API_KEY": "your_api_key_here"
      }
    }
  }
}

Option C: Self-hosted HTTP server

Run your own HTTP MCP endpoint:

AGENTPHONE_API_KEY=your_api_key npx agentphone-mcp --http --port 3000

Then connect to http://localhost:3000/mcp.

Related MCP server: agentline-mcp

What Can It Do?

Once configured, just ask your AI agent things like:

  • "Buy me a phone number in the 415 area code"

  • "Create a support agent that greets callers and helps with billing"

  • "Call +14155551234 and have a conversation about scheduling a dentist appointment"

  • "Text +14155551234 saying 'Your appointment is confirmed for 3pm tomorrow'"

  • "Show me my recent calls and transcripts"

  • "List the available voices and switch my agent to a different one"

  • "Set up a webhook so I get notified when someone calls or texts my number"

  • "Show me this month's usage breakdown"

Transport & hosting

The server is built on the mcp-use server framework, which owns the HTTP layer: Streamable HTTP, the SSE stream, session management, and the OAuth discovery endpoints. npm start (or the Docker image) runs it as an HTTP server on PORT (default 3000), reachable at /mcp.

  • Hosted: https://mcp.agentphone.ai/mcp

  • Self-hosted: PORT=3000 npm start → http://localhost:3000/mcp

Authentication

  1. OAuth (recommended for end users): the framework proxies an Authorization Code + PKCE flow to the AgentPhone authorization server, so the client opens a browser to sign in at agentphone.ai — no key to paste. Enable it by setting MCP_OAUTH_CLIENT_ID (a client pre-registered with the AgentPhone AS). MCP_OAUTH_CLIENT_SECRET is optional: set it to run the gateway as a confidential client, or leave it unset to run as a public client (the gateway must then be registered with token_endpoint_auth_method=none). Public mode advertises none to downstream clients, which strict OAuth clients require.

  2. API key (scripts / single-tenant): set AGENTPHONE_API_KEY. Used as the fallback credential when no OAuth token is present.

The per-request access token is forwarded to the AgentPhone REST API, so the server stores no credentials.

Environment

Var

Purpose

PORT

HTTP port (default 3000)

AGENTPHONE_API_KEY

Fallback API key when OAuth is off

MCP_OAUTH_CLIENT_ID

Enable OAuth; a gateway client pre-registered with the AgentPhone AS

MCP_OAUTH_CLIENT_SECRET

Optional. Set = confidential gateway; unset = public client (gateway registered with token_endpoint_auth_method=none, advertises none downstream)

AGENTPHONE_BASE_URL

API base (default https://api.agentphone.ai)

AGENTPHONE_OAUTH_AUTHORIZE

Override authorize/consent URL (default https://agentphone.ai/oauth/authorize)

Highlights

  • Phone numbers — buy and manage numbers in any US/CA area code

  • SMS — send and receive text messages, view conversation threads

  • Voice calls — place outbound calls with built-in AI conversation (no webhook needed) or bring your own webhook

  • Inbound handling — set up webhooks to receive and respond to inbound calls and texts in real time

  • Agents — create agents with custom voices, system prompts, call transfer, and voicemail

  • Usage & billing — monitor number usage, message/call/webhook volume, and daily/monthly breakdowns

All Tools (28)

Account

Tool

Description

account_overview

Get a full snapshot of your account — agents, numbers, webhook, and usage

get_usage

Get usage stats — number usage, billed SMS segments, and message/call/webhook volume. Use breakdown for daily or monthly time-series.

Phone Numbers

Tool

Description

list_numbers

List all phone numbers in your account

buy_number

Purchase a new phone number with optional area_code and agent_id

SMS

Tool

Description

send_message

Send SMS or iMessage. Supports media, threaded replies (reply_to_message_id), iMessage send effects (send_style), and group chats

get_messages

Get messages for a specific number

list_conversations

List SMS conversations. Pass agent_id to filter by agent.

get_conversation

Get a conversation with full message history

update_conversation

Set metadata on a conversation

Contacts

Tool

Description

list_contacts

List saved contacts (address book). Filter with a search term.

manage_contact

Create, update, or delete a contact (set action)

Voice Calls

Tool

Description

list_calls

List calls. Filter by agent_id, number_id, status, direction, or keyword.

get_call

Get call details and transcript

make_call

Place an outbound call (webhook-driven)

make_conversation_call

Place a call with built-in AI conversation — no webhook needed

Agents

Tool

Description

list_agents

List all agents with their numbers and voice config

create_agent

Create an agent with voice, system prompt, call transfer, voicemail, and voice tuning (speed, interruption sensitivity, backchannel, language, and more)

update_agent

Update an agent's configuration

delete_agent

Delete an agent (numbers are kept but unassigned)

get_agent

Get agent details including phone numbers and voice config

attach_number

Assign a phone number to an agent

detach_number

Remove a phone number from an agent

list_voices

List available voices for agents

Webhooks

All webhook tools accept an optional agent_id — pass it to manage an agent-specific webhook, omit it for the project-level default. Agent webhooks take priority over project-level.

Tool

Description

get_webhook

Get webhook configuration

set_webhook

Set a webhook URL for inbound messages and call events

delete_webhook

Remove a webhook

test_webhook

Send a test event to verify your webhook works

list_webhook_deliveries

View delivery history for debugging

Environment Variables

Variable

Required

Description

AGENTPHONE_API_KEY

stdio: yes, HTTP: no

Your AgentPhone API key (HTTP mode can use Authorization header instead)

AGENTPHONE_BASE_URL

No

Override the API base URL (defaults to https://api.agentphone.ai)

PORT

No

Port for HTTP mode (defaults to 3000, overridden by --port)

Development

git clone https://github.com/AgentPhone-AI/agentphone-mcp.git
cd agentphone-mcp
npm install
npm run dev     # Run with tsx (hot reload)
npm run build   # Compile TypeScript
npm start       # Run compiled JS (stdio)

How It Works

This MCP server connects your AI assistant to the AgentPhone API. Your assistant talks to the MCP server, which calls the AgentPhone API, which talks to the phone network.

Your AI Assistant  <-->  agentphone-mcp  <-->  AgentPhone API  <-->  Phone Network

Outbound: your assistant places calls and sends texts through AgentPhone's API.

Inbound: when someone calls or texts your number, AgentPhone sends a webhook event to your server — you can then respond programmatically or let your agent's built-in AI handle it.

License

MIT

Available Tools

28 tools
account_overviewA
Read-onlyIdempotent
Inspect

Get a complete snapshot of your AgentPhone account: agents, phone numbers, webhook status, and usage limits.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds real value beyond that by disclosing what the snapshot contains (agents, numbers, webhook status, usage limits), which is important since no output schema exists to reveal the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the verb and scope, with the returned contents listed after the colon. No filler and no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only aggregation tool with no output schema, the description gives enough to call it correctly and to know roughly what comes back. It stops short of noting relationships to the granular sibling tools, which would have made routing unambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to clarify about arguments, and the schema is trivially complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Get') plus resource ('a complete snapshot of your AgentPhone account') with an explicit enumeration of the contents (agents, phone numbers, webhook status, usage limits). This aggregate framing implicitly separates it from the many per-resource siblings (list_agents, list_numbers, get_usage), though it never names an alternative outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the entry-point/orientation call rather than a per-resource query, but there is no explicit when-to-use, no when-not-to-use, and no pointer to the sibling tools that give finer-grained data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attach_numberB
Idempotent
Inspect

Attach a phone number to an agent so the agent handles calls/SMS on that number. Takes an agent ID and a phone number ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent ID
number_idYesThe phone number ID to attach

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare idempotentHint, so the description carries most of the burden. It usefully explains the observable consequence of attaching (the agent starts handling calls/SMS on that number), but omits permissions, conflict/replacement behavior, and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose and effect front-loaded with no filler. The second sentence is largely redundant with the required-parameter list in the schema, a minor waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with full schema coverage and no output schema, the core purpose is covered. However, nothing is said about where to obtain a valid number_id, whether the number must be purchased first, or what happens if the number is already attached.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema; the description merely restates 'takes an agent ID and a phone number ID' without adding format, source-of-ID, or constraint details. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (attach) and resource (phone number to an agent) plus the concrete effect: the agent then handles calls/SMS on that number. The contrast with the sibling detach_number is implied by the verb, but the description never names alternatives or differentiates explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to attach vs. when not to, no prerequisites (e.g. the number must already exist in the account, whether an existing attachment is replaced), and no mention of the sibling detach_number or list_numbers for discovering number IDs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

buy_numberAInspect

Purchase a new phone number. This is a paid purchase that charges the account's balance. area_code requests a specific region (e.g. '415' for San Francisco); agent_id attaches it on purchase. Requires confirm=true to proceed.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNoMust be true to complete the purchase. When false, no number is bought.
countryNo2-letter ISO country code (e.g. US, CA, GB)US
agent_idNoAgent ID to attach this number to immediately
area_codeNo3-digit area code for a specific region (e.g. '415', '212', '310')

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply openWorldHint=true, so the description must carry the rest, and it does: it discloses that this is a paid purchase that charges the account balance and that confirm must be true to complete it. It stops short of stating reversibility or what happens to a number once purchased, but the financial-impact disclosure is the key behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and cost, then the two notable parameters, then the confirm gate. Slightly repetitive with the schema (which already documents confirm and area_code), but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 100% schema coverage, minimal annotations, and no output schema, the description covers the essentials: what it does, that it costs money, the confirm gate, and the two behaviorally interesting parameters. It omits what a successful result returns and whether country defaults matter, which are minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds interpretive value: area_code maps to a region with a concrete example ('415' for San Francisco) and agent_id is described as attaching the number on purchase. That clarifies the effect of agent_id beyond the schema's terse wording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Purchase a new phone number'), and the word 'new' implicitly separates it from attach_number, which attaches an existing number. It never names that sibling explicitly, so it falls short of the 5-tier bar for sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one concrete prerequisite ('Requires confirm=true to proceed') and notes the cost, which is genuine usage guidance. However, it never says when to prefer this over attach_number or how it relates to list_numbers for selecting a number, leaving the choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agentAInspect

Create a new agent. An agent owns phone numbers and handles calls/SMS.

A newly created agent has no phone number until one is attached. Set voice_mode to 'hosted' with a system_prompt for autonomous AI voice calls, or 'webhook' (default) to forward call transcripts to your webhook URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName for the agent (e.g. 'Customer Support', 'Sales Bot')
voiceNoVoice ID for the agent. Defaults to 'Skylar - Friendly Guide'.
languageNoBCP-47 locale for speech recognition and synthesis, e.g. 'en-US', 'es-ES', 'ja-JP'.
stt_modeNoSpeech-to-text mode: 'fast' (default) lowest latency, 'accurate' for exact names/numbers (~200ms slower).
model_tierNoModel quality/speed tier for hosted-mode agents. 'turbo' = fastest/cheapest, 'balanced' (default) = general use, 'max' = highest quality.
voice_modeNo'webhook' (default) forwards transcripts to your webhook. 'hosted' uses built-in AI with system_prompt.
descriptionNoDescription of what this agent does
voice_speedNoSpeech speed multiplier. 1.0 = normal, 0.5 = half speed, 2.0 = double.
ambient_soundNoOptional background ambience for call audio. Not a substitute for the AI self-identification, which is always disclosed.
begin_messageNoWhat the AI says when a call connects. Only used in 'hosted' mode.
system_promptNoRequired when voice_mode is 'hosted'. The AI's personality and instructions for voice calls.
denoising_modeNoAudio denoising. The aggressive mode helps callers in cars/cafes (small surcharge).
max_silence_msNoHang up after this many ms of caller silence. Default 600000 (10 min).
transfer_numberNoPhone number to transfer calls to (E.164 format). Enables call transfer during conversations.
enable_messagingNoWhether a hosted agent can send and read texts during a call. Default true.
voicemail_messageNoVoicemail greeting text. When set, unanswered calls hear this message and can leave a voicemail.
enable_backchannelNoWhether the agent interjects 'uh-huh'/'mhmm' while the caller speaks. Default true.
interruption_sensitivityNoHow easily callers can interrupt (barge in). 0 = never, 1 = at first sound. Default 0.8.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries full disclosure burden. It usefully reveals that a newly created agent has no phone number until attached and clarifies the voice_mode branching, but says nothing about permissions, side effects, or what the creation returns — significant gaps for an 18-parameter mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then the phone-number caveat, then mode selection. Zero waste and well ordered for scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers the key branching decision (hosted vs webhook) and the post-creation prerequisite (attach a number). It could mention the response shape or default behavior, but it is adequate for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 18 parameters thoroughly. The description reinforces the voice_mode/system_prompt relationship but adds little beyond what the schema states. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a new agent') and adds real domain context — an agent owns phone numbers and handles calls/SMS. This distinguishes it from list_agents/update_agent/get_agent, though it does not explicitly name a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides actionable mode-selection guidance: 'hosted' with a system_prompt for autonomous AI calls, or 'webhook' (default) to forward transcripts. It also implies the follow-up step (attach_number) by noting a new agent has no phone number. No explicit when-not or sibling routing, but clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_agentA
Destructive
Inspect

Delete an agent permanently. Phone numbers attached to it are kept but unassigned. This cannot be undone.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent ID to delete

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=true; the description goes further by disclosing the side effect on attached phone numbers (kept but unassigned) and that the action is irreversible. This is exactly the kind of relational consequence an agent needs before committing a destructive call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, no filler, with the core action front-loaded and the irreversibility warning last for emphasis. Every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with no output schema and destructiveHint already set, the description supplies the missing pieces: permanence and the effect on related phone numbers. Nothing needed to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage, so the schema already explains agent_id. The description adds no format or sourcing detail beyond it, making 3 the appropriate baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete an agent') plus scope ('permanently'), which cleanly separates it from the update_agent, get_agent, and delete_webhook siblings. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The permanence warning implies this should be used only for true removal, but there is no explicit when-to-use guidance or mention of alternatives such as deactivating via update_agent. Usage is inferable from the name rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_webhookA
Destructive
Inspect

Remove a webhook. Pass agent_id to remove an agent's webhook (falls back to project default). Omit agent_id to remove the project-level webhook.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoAgent ID to delete that agent's webhook. Omit for project-level.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, which sets the safety floor. The description adds the non-obvious fallback/scope semantics (agent-level delete falls back to project default; omitting targets project-level) that annotations cannot express. Does not state irreversibility explicitly, but the destructive hint plus 'remove' wording covers the essential traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the verb and resource, then the branching guidance. Zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a one-parameter destructive tool: what it does, the scope branching, and the fallback are all present. An output schema is absent but a delete's return value is conventionally minimal; no critical gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents agent_id, so baseline would be 3. The description adds the fallback behavior ('falls back to project default') that the schema field description does not convey, which is meaningful semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Remove a webhook') and immediately clarifies the two scopes (agent-level vs project-level). Distinguishable from siblings like get_webhook, set_webhook, and test_webhook without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains precisely when to pass agent_id versus omit it, which is the sole decision point for this tool. No explicit exclusions or named alternatives, but the branching logic is fully spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detach_numberB
Idempotent
Inspect

Detach a phone number from an agent. The number is kept in your account but becomes unassigned.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent ID that currently owns the number
number_idYesThe phone number ID to detach

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation idempotentHint=true already tells the agent that repeated detaches are safe. The description adds that the number is kept but becomes unassigned, which is useful (non-destructive in terms of data retention), but doesn't cover permissions, side effects on the agent, or what happens if the number is already detached.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and result. No wasted words, though it could be slightly more structured for usage guidelines.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple mutation tool with full schema coverage and an idempotency annotation, the description covers the essential outcome. However, it lacks guidelines on when to use it and doesn't mention any side effects on the agent, leaving some gaps for an agent's decision-making.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already fully documents both parameters. The description adds no additional parameter information, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (detach) and resource (phone number), and clarifies the result. Clearly distinguishes from the sibling attach_number.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. It doesn't mention prerequisites or when you might detach a number (e.g., before reassigning). The sibling attach_number is not referenced.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agentA
Read-onlyIdempotent
Inspect

Get details for a specific agent including its phone numbers, voice configuration, and instructions.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent ID

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds useful context about what the response contains (phone numbers, voice config, instructions) but says nothing about auth requirements, error behavior for unknown IDs, or payload size.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence that names the resource and the payload; no boilerplate or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only getter with annotations covering safety and a fully described single-param schema, the description is nearly sufficient; it helpfully enumerates the returned fields even though no output schema exists. Only the absence of any error/not-found behavior keeps this from being complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single agent_id parameter is already documented in the schema as 'The agent ID'. The description adds no format, prefix, or sourcing guidance beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (a specific agent) and enumerates the content returned: phone numbers, voice configuration, instructions. It clearly is not a list operation, but it never explicitly differentiates itself from list_agents or get_call siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'a specific agent' implicitly signals that this is the lookup to use when you already have an agent_id, versus list_agents for enumeration, but no when/when-not guidance or alternative is stated explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_callA
Read-onlyIdempotent
Inspect

Get details and transcript for a specific call, identified by its call ID. Pass wait=true to block until an in-progress call finishes before returning.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoWhen true, long-polls until the call completes before returning. Useful for checking back on a call you initiated earlier.
call_idYesThe call ID
timeoutNoMax seconds to wait when wait=true. Defaults to 120.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and idempotentHint, so safety is covered; the description adds the non-obvious long-poll blocking behavior of wait=true, which the agent needs for scheduling. It stops short of describing pagination or how a still-in-progress call is represented on timeout.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the core retrieval purpose comes first and the wait modifier second, correctly front-loading the primary semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return burden and does name the payload ('details and transcript'). That is sufficient for correct invocation, though it could say more about what happens when the call is not yet complete or when timeout is reached.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so call_id, wait, and timeout are already documented in the schema, establishing a baseline of 3. The description restates wait=true's effect but adds no syntax or interaction detail (e.g., how timeout relates to wait) beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get details and transcript for a specific call') and pins the identifier to call_id. It implicitly separates itself from list_calls by scoping to a single call, though it doesn't name the sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives actionable context for the wait=true mode ('block until an in-progress call finishes'), which implies when to use it, but it never states when not to use the tool or points to alternatives like list_calls or get_conversation for adjacent lookups.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conversationB
Read-onlyIdempotent
Inspect

Get a specific SMS conversation with full message history, identified by its conversation ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
message_limitNoMax messages to include
conversation_idYesThe conversation ID

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, covering the safety profile. The description adds that it returns full message history, but it does not clarify that message_limit caps the returned messages or describe pagination/rate behavior. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-formed sentence with the key action and identifier front-loaded. There is no redundant or wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with full schema coverage and annotations, the description is nearly complete, stating the return scope. A minor gap is that 'full message history' is bounded by the message_limit parameter, which the description does not mention.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters including message_limit. The description mentions the conversation ID but adds no additional syntax or format details beyond the schema, making a baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'Get a specific SMS conversation with full message history.' This clearly communicates the tool's core action. It does not explicitly distinguish itself from siblings like list_conversations or get_messages, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a conversation ID is available, but it provides no explicit when-to-use guidance, no when-not-to-use conditions, and no alternatives such as list_conversations or get_messages. This leaves routing to the agent's inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_messagesA
Read-onlyIdempotent
Inspect

Get SMS messages for a specific phone number, identified by its number ID. Returns individual messages rather than grouped conversation threads.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax messages to return
number_idYesThe ID of the phone number

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds the useful behavioral detail that results are individual messages, not threads, but says nothing about pagination, ordering, or time-window behavior despite the limit parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with zero waste; the scope and the distinguishing return-granularity note are front-loaded and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool this is essentially complete: it tells the agent what it fetches and how the output is shaped. Only minor gaps remain, such as ordering or pagination behavior, which the absent output schema would otherwise need to cover.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented. The description restates that the number is identified by its ID but adds no syntax, format, or usage nuance beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) plus resource (SMS messages) scoped to a phone number identified by number ID. It also distinguishes itself from the conversation-thread siblings by explicitly saying it returns individual messages rather than grouped threads.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'rather than grouped conversation threads' implicitly routes the agent to the conversation tools when threads are wanted, giving clear context for selection. It does not name the alternative tool explicitly, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_usageA
Read-onlyIdempotent
Inspect

Get account usage statistics. By default returns a summary with number usage plus message, call, and webhook volume. Use breakdown='daily' or 'monthly' for time-series data.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoNumber of days to look back (only used with breakdown='daily')
monthsNoNumber of months to look back (only used with breakdown='monthly')
breakdownNo'summary' for number usage and totals, 'daily' for per-day breakdown, 'monthly' for per-month breakdownsummary

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the read-only contract is covered by structured data. The description adds genuinely useful context by stating exactly what the default response contains and how the breakdown modes change the shape of the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler: the subject and default behavior come first, then the mode-switching instruction. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by naming the contents of the default summary and the alternate breakdown outputs. Combined with full schema coverage of parameters, an agent has enough to call the tool correctly, though a note on which sibling to prefer would complete the picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – all three parameters, including the enum values and the days/months applicability notes, are documented in the schema. The description only restates the breakdown choices and adds no syntax, format, or constraint detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get account usage statistics') and enumerates what the return contains (number usage plus message, call, and webhook volume). It does not distinguish itself from the sibling account_overview, which an agent could plausibly confuse with a usage tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the default behavior and tells the agent to use breakdown='daily' or 'monthly' for time-series data, which is implied guidance on selecting a mode. There is no explicit when-to-use-this-vs-alternative statement (e.g., versus account_overview), so usage remains implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_webhookA
Read-onlyIdempotent
Inspect

Get the webhook configuration. Pass agent_id to get an agent-specific webhook, or omit for the project-level default.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoAgent ID to get that agent's webhook. Omit for project-level webhook.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered and the bar is lower. The description's only added behavior is the omit-defaults-to-project-level rule, which is really parameter semantics; it says nothing about the return shape or behavior when no webhook is configured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action front-loaded and the parameter branching immediately after. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless-required read tool with full schema coverage and annotations covering safety, the description is nearly complete. The main gap is that no output schema exists, yet the description doesn't hint at what the returned configuration contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single optional parameter, so the schema carries the burden and the baseline is 3. The description's phrasing about agent_id closely mirrors the schema's own text, adding little new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (webhook configuration), and clarifies the two retrieval scopes: agent-specific vs project-level default. It does not explicitly contrast with siblings like set_webhook or list_webhook_deliveries, so it falls short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear situational guidance for the branch: pass agent_id for an agent-specific webhook, omit it for the project-level default. It stops short of naming alternatives or exclusions (e.g., when to use set_webhook or list_webhook_deliveries instead).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agentsA
Read-onlyIdempotent
Inspect

List all agents with their phone numbers and voice configuration. An agent is required before you can make calls — it owns phone numbers and handles voice/SMS.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds domain context that agents own phone numbers and handle voice/SMS, but says nothing about pagination or result ordering despite the limit parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and returned data, with the domain note following. No wasted words, though the second sentence drifts slightly into general platform context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with full annotation coverage and no output schema, the description adequately conveys what is returned and why it matters. Minor gaps around pagination and ordering remain, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single optional parameter fully documented in the schema ('Max results to return'). The description adds no parameter detail, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all agents') and even names the returned fields (phone numbers, voice configuration), which is more than the bare name. It distinguishes itself reasonably from get_agent by scope ('all'), but never explicitly contrasts with the sibling list/get tools, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides useful context ('an agent is required before you can make calls') that implies why you'd list agents, but gives no explicit when-to-use, when-not-to-use, or alternative routing versus get_agent, create_agent, or list_numbers. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_callsA
Read-onlyIdempotent
Inspect

List recent calls. Scope by agent_id or number_id, or use status/direction/search to filter globally.

When agent_id or number_id is passed, status/direction/search filters are not applied. Returns call summaries including each call's ID, direction, participants, and status.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return
offsetNoNumber of results to skip (for pagination)
searchNoSearch by phone number or keyword
statusNoFilter by status: ringing, in-progress, completed, failed, busy, no-answer
agent_idNoFilter to calls for a specific agent
directionNoFilter by direction: 'inbound' or 'outbound'
number_idNoFilter to calls for a specific phone number

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint, idempotentHint), so the description adds value with the non-obvious silent-ignore behavior: status/direction/search are dropped when agent_id or number_id is supplied. It also discloses the returned summary fields, which matters since there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary action, then the scoping rule, then return contents. Every sentence carries load and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param, all-optional listing tool with no output schema, the description covers the tricky parameter interaction and the shape of the results. Pagination behavior (limit/offset) is left entirely to the schema, but that is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline would be 3, but the description contributes the mutual-exclusivity rule between scoping parameters and filter parameters that the schema does not express. That is genuine added meaning beyond per-field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List recent calls') and enumerates the scoping/filtering dimensions available. It is clearly distinguishable from the sibling get_call (single call) by its listing nature, though it does not name siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing between two modes: scope by agent_id/number_id, or filter globally with status/direction/search. It further states the when-not case (filters are ignored when scoping params are present), which is exactly the guidance an agent needs. No named alternative functions, but the selection rule is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_contactsA
Read-onlyIdempotent
Inspect

List saved contacts (your address book). Optionally filter with a search term that matches name or phone number.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return
offsetNoNumber of results to skip (for pagination)
searchNoFilter by name or phone number

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered without description help. The description adds only that contacts are an address book and that search spans name/phone; it says nothing about pagination behavior, result ordering, or result caps beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler, and the core purpose is front-loaded before the optional filter detail. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, fully-annotated, zero-required-parameter list tool with 100% schema coverage and no output schema, the description covers what an agent needs to call it. The only gap is the absent relationship to the manage_contact sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (limit, offset, search) are already documented with defaults and bounds. The description's note that search "matches name or phone number" restates the schema's own wording rather than adding new semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List saved contacts") and disambiguates the resource with "(your address book)." It does not, however, distinguish itself from the sibling manage_contact, so an agent must infer the read-vs-write split on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the optional filter and what it matches, which implies how to use the tool, but it gives no explicit when-to-use guidance or exclusions relative to manage_contact or any search alternative. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_conversationsA
Read-onlyIdempotent
Inspect

List SMS conversations. Optionally filter by agent_id to see conversations for a specific agent.

Each conversation is a thread between your number and an external contact, with its ID, participant, message count, and last-message preview.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return
offsetNoNumber of results to skip (for pagination)
agent_idNoFilter to conversations for a specific agent

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds value by enumerating what each returned conversation contains (ID, participant, message count, last-message preview), which matters because there is no output schema. It stops short of describing pagination or ordering behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, purpose front-loaded before the filter and the return-shape detail. No filler, though the trailing clause about conversation contents could be trimmed without losing selection-critical meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully compensates by describing the returned conversation fields, and all three params are covered by the schema. It is largely sufficient, missing only pagination/ordering behavior for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so limit, offset, and agent_id are already documented in the schema. The description restates the agent_id filter without adding format, default, or range guidance beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List SMS conversations') that cleanly separates it from the singular get_conversation and mutating update_conversation siblings. It does not explicitly name an alternative tool, but the operation type is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes the optional agent_id filter and its effect, which implies a usage context, but it never states when to prefer this tool over get_conversation or how agent-scoped listing differs from unfiltered listing. Usage is implied rather than guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_numbersA
Read-onlyIdempotent
Inspect

List all phone numbers in your account, each with its ID, status, country, and assigned agent.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results to return
offsetNoNumber of results to skip (for pagination)

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds the shape of the result (ID, status, country, assigned agent), which is genuinely useful since no output schema exists, but it says nothing about pagination behavior, result caps, or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action and the returned fields with zero filler; nothing is repeated from structured data unnecessarily.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates returned fields and, with only two well-documented optional parameters and strong read-only annotations, an agent has enough to call it correctly. Only the pagination/result-limit behavior is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (limit, offset) are documented in the schema, so the baseline is 3. The description adds no syntax, default, or interaction detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (phone numbers in your account) and enumerates the returned fields, so it is clearly distinct from buy_number or attach_number. It does not explicitly name siblings, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and the 'in your account' scoping, but there is no explicit when-to-use statement, no mention of prerequisites, and no routing to related tools such as attach_number or buy_number when the desired number is absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_voicesA
Read-onlyIdempotent
Inspect

List available voices for agents. Each voice has a voice_id used in an agent's voice field.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds genuinely useful behavioral context by explaining the shape of the result (each voice carries a voice_id) and how that value is consumed downstream in an agent's voice field.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the purpose is front-loaded before the return-value detail. Every sentence earns its place by adding either scope or consumption guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only lookup with no output schema, the description covers what is returned and how the key field is used. It stops short of saying whether the list is filtered or scoped to the account, but nothing essential to a correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. Nothing in the description is needed to compensate for schema gaps, and the schema coverage is fully reported as 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List available voices for agents'), so an agent immediately knows what the tool returns. It lacks explicit differentiation from siblings like list_agents, but the resource is distinct enough that confusion is unlikely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: by noting the voice_id feeds an agent's voice field, it hints this is the lookup step before create_agent/update_agent. There is no explicit when-to-use or when-not-to-use guidance, and no named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_webhook_deliveriesA
Read-onlyIdempotent
Inspect

View recent webhook delivery history. Shows which events were delivered, HTTP status codes, and timing.

Pass agent_id to see deliveries for that agent's webhook. Omit for project-level.

ParametersJSON Schema
NameRequiredDescriptionDefault
hoursNoOnly show deliveries from the last N hours
limitNoMax results to return
agent_idNoAgent ID to see deliveries for that agent's webhook. Omit for project-level.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds the return content (event, status code, timing), which matters because there is no output schema, but it says nothing about ordering, pagination behavior, retention window, or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences plus a one-line parameter note, with the core purpose front-loaded before the scoping detail. Efficient overall, though the agent_id sentence duplicates schema text rather than adding anything new.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool with full schema coverage, covered safety annotations, and no output schema, the description supplies the essentials: what is listed, what each record shows, and how to scope the query. Missing pagination and retention details are minor omissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema; the agent_id prose merely restates the schema text almost verbatim. The hours and limit parameters receive no additional prose meaning beyond their schema descriptions, matching the baseline 3 for a fully covered schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (view webhook delivery history) and enumerates what the listing contains — delivered events, HTTP status codes, timing. It is clearly distinguishable from config-oriented siblings like get_webhook or set_webhook, though it never names them explicitly to sharpen the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives real scoping guidance — pass agent_id for an agent's webhook, omit for project-level — which is the main usage decision for this tool. It offers no when-to-use versus alternatives such as test_webhook or get_webhook, and no framing around when one would inspect delivery history at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

make_callAInspect

Initiate an outbound, webhook-driven phone call where your backend handles the conversation logic.

The agent must have a phone number attached and a webhook configured. Every call automatically opens with an automated-assistant disclosure (added server-side and not removable).

ParametersJSON Schema
NameRequiredDescriptionDefault
voiceNoVoice ID override for this call
agent_idYesThe agent ID (the agent must have a phone number attached)
to_numberYesRecipient phone number in E.164 format (e.g. +14155551234)
from_number_idNoSpecific phone number ID to call from (if agent has multiple numbers)
initial_greetingNoWhat the agent says when the call connects

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply openWorldHint=true, so the description carries the behavioral load and does it well: it discloses that the conversation is driven by the caller's webhook, and that every call opens with an automated-assistant disclosure added server-side and not removable. The latter is non-obvious, non-configurable behavior an agent must know before calling. It stops short of describing failure modes or call outcome.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, zero filler. The core action leads, and the prerequisite/disclosure constraints follow in a scannable second paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a call-initiation tool with no output schema and fully documented parameters, the description covers action, prerequisites, and a mandatory behavioral quirk. The remaining gap is that it never points to make_conversation_call as the alternative, which matters given how similar the two siblings are.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented, making 3 the baseline. The description adds one meaningful nuance: the opening disclosure is server-side and cannot be removed, which implicitly constrains what initial_greeting can control, but it does not explain voice override or from_number_id selection.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Initiate an outbound ... phone call') and adds the defining mechanism: webhook-driven, with the caller's backend handling conversation logic. That mechanism implicitly separates it from make_conversation_call, but the sibling is never named, so the differentiation requires inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a real precondition ('The agent must have a phone number attached and a webhook configured'), which tells the agent when the call can succeed. However, it never states when to pick this over make_conversation_call or what happens if the webhook is missing, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

make_conversation_callAInspect

Place a phone call where the AI has an autonomous conversation about a given topic — scheduling, surveys, follow-ups, etc. No webhook setup needed.

The agent must have a phone number attached. Every call automatically opens with an automated-assistant disclosure (added server-side; cannot be disabled via topic or initial_greeting). By default this blocks until the call finishes and returns the full transcript. Set wait=false for fire-and-forget.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoWhen true (default), blocks until the call ends and returns the full transcript.
topicYesThe conversation goal for the assistant to pursue. Embedded as the call objective within locked instructions that always enforce the automated-assistant disclosure (it can't be overridden). Be specific about what to discuss and any goals.
voiceNoVoice ID override for this call
agent_idYesThe agent ID (must have a phone number attached)
to_numberYesRecipient phone number in E.164 format (e.g. +14155551234)
from_number_idNoSpecific phone number ID to call from (if agent has multiple numbers)
initial_greetingNoWhat the AI says when the call connects. If not set, the AI will generate one from the topic.
max_wait_secondsNoMaximum seconds to wait for the call to complete. Defaults to 300 (5 minutes).

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only carry openWorldHint, so the description does the heavy lifting and does it well: it discloses a server-side, non-disableable automated-assistant disclosure, the default blocking behavior with transcript return, and the wait=false escape hatch. It omits cost, concurrency, and failure/retry behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, each carrying distinct information, with the core behavior and the blocking default front-loaded. Nothing is redundant with the schema or wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers prerequisites, side effects, blocking semantics, and return behavior (full transcript) for an 8-param tool with no output schema. Minor gaps remain around transcript format and any cost/latency expectations, but an agent has what it needs to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all 8 parameters are documented in the schema, so the baseline is 3. The description reinforces wait semantics and the locked nature of topic but adds no syntax or format detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Place a phone call') plus the distinguishing capability ('AI has an autonomous conversation about a given topic'), which separates it from the plain make_call sibling without naming it. The scope is clear enough that an agent can pick it over make_call from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives prerequisites ('The agent must have a phone number attached') and notes 'No webhook setup needed', which is useful context. However, it never explicitly says when to use this versus make_call or send_message, so the routing guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_contactA
Destructive
Inspect

Create, update, or delete a saved contact. Set action to choose the operation.

  • create: requires phone_number and name

  • update: requires contact_id; only the fields you pass are changed

  • delete: requires contact_id (permanent)

contact_id refers to a saved contact.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoContact name. Required for 'create'.
emailNoContact email address
notesNoFreeform notes about the contact
actionYesThe operation to perform
contact_idNoContact ID. Required for 'update' and 'delete'.
phone_numberNoPhone number in E.164 format (e.g. +14155551234). Required for 'create'.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations only supply destructiveHint=true; the description adds real behavioral context on top of that, notably that delete is permanent and that update is a partial/PATCH-style change rather than a full replace. It still omits auth/permission requirements and what happens on a bad contact_id, keeping it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences: a front-loaded purpose line followed by a compact per-action bullet list. Every line carries actionable information (required fields, partial-update semantics, permanence) with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with destructiveHint and no output schema, the description covers operation selection, per-action required fields, partial-update behavior, and delete permanence. Missing only error behavior (unknown contact_id) and permission context, which are minor for this scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter's per-action requirement (e.g. 'Required for create.') is already stated in the schema. The description restates these conditionals in one consolidated block, which aids readability but adds essentially no semantics beyond the structured fields, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb set and resource ('Create, update, or delete a saved contact') and immediately enumerates the three operations, so the agent knows exactly what the tool does. It does not explicitly contrast itself with the read-side sibling list_contacts, but the mutation-only scope is unambiguous from the verbs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides per-action guidance: create requires phone_number and name, update requires contact_id and only changes passed fields, delete requires contact_id. This tells the agent which operation to select and what each needs. It does not name alternatives (e.g. use list_contacts to look up a contact_id first), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_messageAInspect

Send an SMS or iMessage from one of your agent's phone numbers, or react to a message.

The agent must have at least one phone number attached. If the agent has multiple numbers, number_id or from_number selects which one to send from. iMessage extras (silently ignored on SMS): reply_to_message_id threads the reply under an earlier message, and send_style adds an expressive screen/bubble effect. To react to a message instead of sending one, set reaction and react_to_message_id (iMessage only); to_number and body are not needed in that case.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoThe message text to send (may be empty when sending media only)
agent_idNoThe agent ID to send from (the agent must have a phone number attached). Optional if you pass from_number or number_id instead.
reactionNoiMessage only. Tapback or single emoji applied to react_to_message_id: one of love, like, dislike, laugh, emphasize, question, or a single emoji (e.g. '🔥'). When set, this reacts instead of sending a message.
media_urlNoURL of a single image/media file to attach
number_idNoSpecific phone number ID to send from (if agent has multiple numbers)
to_numberNoRecipient: a phone number in E.164 format (e.g. +14155551234), a US short code, or a group ID (grp_...) to post into an iMessage group chat. Required when sending a message (not needed for a reaction).
media_urlsNoMultiple media URLs to attach (delivered as an image carousel on iMessage)
send_styleNoiMessage only. Expressive screen/bubble effect to send with the message
from_numberNoExact number to send from in E.164 format (alternative to number_id)
react_to_message_idNoiMessage only. ID of a message to react to. Set together with `reaction`.
reply_to_message_idNoiMessage only. ID of an earlier message to reply to inline (threaded reply)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare openWorldHint=true, so the description carries most of the burden and delivers real behavioral context: iMessage-only fields are 'silently ignored on SMS', reply_to_message_id threads a reply, and reaction mode changes which parameters are required. It omits permissions, delivery/failure behavior, and rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the send/extras/react structure; each sentence carries distinct information. Dense but readable, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with zero required params and no output schema, the description covers the primary send path, the number-selection path, and the reaction path, which is what an agent needs to invoke it. Remaining gaps (single vs. carousel media, error surfaces) are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds cross-parameter conditional semantics the schema does not state on its own: number_id/from_number select the sending number, and to_number/body are explicitly unnecessary when reacting. It does not clarify media_url vs media_urls interaction.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource ('Send an SMS or iMessage') plus an explicit second mode ('or react to a message'), scoped to 'one of your agent's phone numbers'. An agent can immediately distinguish this from siblings like make_call or get_messages.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the prerequisite ('agent must have at least one phone number attached') and the selection rule when multiple numbers exist, plus when to switch to reaction mode instead of sending. It does not name sibling alternatives (e.g. make_call for voice), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_webhookA
Idempotent
Inspect

Set a webhook URL to receive inbound messages and call events.

Pass agent_id to set a webhook for a specific agent (overrides project default). Omit agent_id to set the project-level webhook for all agents. The webhook secret is returned — use it to verify signatures.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe publicly accessible webhook URL (must be HTTPS in production)
timeoutNoWebhook response timeout in seconds
agent_idNoAgent ID to set webhook for that agent only. Omit for project-level.
context_limitNoNumber of recent messages to include as conversation context (0-50)

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only idempotentHint in annotations, the description adds meaningful context: that a webhook secret is returned and must be used to verify signatures. That is a behavioral detail not captured by annotations. It still omits permission/auth requirements and whether an existing webhook is overwritten, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then the scoping rule, then the return-value note. No filler, though the scoping sentence is somewhat verbose for the point it makes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema mutation tool, the description covers the key operational facts an agent needs: what it sets, the scope switch, and that a secret is returned. It leaves minor gaps (auth requirements, overwrite semantics) that are not critical to invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented. The description restates the agent_id scope behavior ('overrides project default') and adds nothing new for url, timeout, or context_limit. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Set a webhook URL') and immediately clarifies its purpose ('to receive inbound messages and call events'). It does not explicitly name the sibling webhook tools (delete_webhook, get_webhook, test_webhook), but the agent/scope distinction makes the operation's identity clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the branching condition: pass agent_id for a per-agent webhook, omit it for the project default. It does not mention when to prefer alternatives such as delete_webhook or test_webhook, but the core usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test_webhookA
Idempotent
Inspect

Send a test event to verify a webhook is working. Returns the HTTP status code and response time.

Pass agent_id to test that agent's webhook. Omit to test the project-level webhook.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoAgent ID to test that agent's webhook. Omit for project-level.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare openWorldHint=true and idempotentHint=true, so the external-effect and repeat-safety profile comes for free. The description adds value beyond that by disclosing the return payload (HTTP status code and response time), which is the key thing an agent needs to interpret results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: what it does, what it returns, and how to scope the target. The purpose is front-loaded and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one optional parameter and no output schema, the description covers the essential contract, including the return values the schema cannot express. It does not mention failure modes (e.g., what a non-2xx status implies) or any prerequisite like a previously configured webhook.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single agent_id parameter is already documented in the schema. The description restates the same agent-level vs project-level semantics, adding emphasis but no new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('send a test event to verify a webhook'), and the scope distinction (agent-level vs project-level) sets it apart from get_webhook, set_webhook, delete_webhook, and list_webhook_deliveries without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Makes the usage condition explicit: use it to verify a webhook is working, and the agent_id-omit rule tells the caller exactly which webhook gets tested. It stops short of naming a sibling alternative or stating when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_agentA
Idempotent
Inspect

Update an agent's configuration — name, description, voice settings, instructions, greeting, call transfer, or voicemail. Only provided fields are updated. Switching voice_mode to 'hosted' requires a system_prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoNew name for the agent
voiceNoVoice ID for the agent.
agent_idYesThe agent ID to update
languageNoBCP-47 locale for speech recognition and synthesis, e.g. 'en-US', 'es-ES', 'ja-JP'.
stt_modeNoSpeech-to-text mode: 'fast' lowest latency, 'accurate' for exact names/numbers (~200ms slower).
model_tierNoModel quality/speed tier for hosted-mode agents. 'turbo' = fastest/cheapest, 'balanced' = general use, 'max' = highest quality.
voice_modeNo'webhook' forwards transcripts to your webhook. 'hosted' uses built-in AI with system_prompt.
descriptionNoNew description
voice_speedNoSpeech speed multiplier. 1.0 = normal, 0.5 = half speed, 2.0 = double.
ambient_soundNoOptional background ambience for call audio. Not a substitute for the AI self-identification, which is always disclosed.
begin_messageNoWhat the AI says when a call connects (hosted mode only).
system_promptNoThe AI's personality and instructions. Required when voice_mode is 'hosted'.
denoising_modeNoAudio denoising. The aggressive mode helps callers in cars/cafes (small surcharge).
max_silence_msNoHang up after this many ms of caller silence. Default 600000 (10 min).
transfer_numberNoPhone number to transfer calls to (E.164 format), or empty string to remove.
enable_messagingNoWhether a hosted agent can send and read texts during a call. Default true.
voicemail_messageNoVoicemail greeting text, or empty string to disable voicemail.
enable_backchannelNoWhether the agent interjects 'uh-huh'/'mhmm' while the caller speaks. Default true.
interruption_sensitivityNoHow easily callers can interrupt (barge in). 0 = never, 1 = at first sound. Default 0.8.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare idempotentHint, so the description carries the remaining burden. The partial-update disclosure is genuinely useful behavior beyond the annotation. However, the voice_mode/system_prompt dependency largely restates the schema's own 'Required when voice_mode is hosted' note, and nothing is said about permissions or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the verb+resource, and the second sentence carries the highest-value constraint. Nothing is padded or repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a rich 19-parameter input schema, the description supplies the two things the schema can't: patch semantics and a cross-field precondition. Complete enough for correct invocation, though it omits any note on what the response returns after an update.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents ranges, enums, defaults, and formats for all 19 fields. The description names field groups but adds no syntax or format detail beyond the schema, so it sits at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (update) and resource (agent configuration), then enumerates the configurable surface: name, description, voice, instructions, greeting, transfer, voicemail. An agent can immediately tell this is the mutation counterpart to create_agent/get_agent. It stops short of explicitly naming a sibling, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Only provided fields are updated' gives the caller clear patch-semantics guidance and the co-requirement for hosted mode is a concrete precondition. There is no explicit when-to-use-this-vs-sibling routing (e.g., when to prefer update_agent over create_agent), which keeps it out of the 5 band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_conversationA
Idempotent
Inspect

Set metadata on a conversation. Use this to store custom state, tags, or context that persists between messages. Pass null to clear metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
metadataYesJSON metadata object to store on the conversation, or null to clear
conversation_idYesThe conversation ID

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply idempotentHint=true, and the description adds real value on top: metadata persists between messages and null clears it. However, for an update tool it omits whether a new metadata object replaces or merges with existing keys, which is the key behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste, and the core purpose is front-loaded before the usage note and the clearing rule. Nothing extraneous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, full schema coverage, and an idempotency annotation, the description covers purpose, usage, and clearing behavior adequately. The only gap is replace-vs-merge semantics, which is minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both parameters documented in the schema, so the baseline is 3. The description's note about passing null to clear metadata largely restates the schema's own 'or null to clear' description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Set') and resource ('metadata on a conversation'), and usefully narrows the broad name 'update_conversation' to a metadata-scoped operation. It does not explicitly contrast with siblings like get_conversation or list_conversations, but the scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this to store custom state, tags, or context that persists between messages' gives clear context for when the tool applies. There are no overlapping siblings to exclude and no explicit when-not guidance, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 28 tool updatesv0.7.0
    • First observedaccount_overview
    • First observedattach_number
    • First observedbuy_number
    • First observedcreate_agent
    • First observeddelete_agent
    • First observeddelete_webhook
    • First observeddetach_number
    • First observedget_agent
    • First observedget_call
    • First observedget_conversation
    • First observedget_messages
    • First observedget_usage
    • First observedget_webhook
    • First observedlist_agents
    • First observedlist_calls
    • First observedlist_contacts
    • First observedlist_conversations
    • First observedlist_numbers
    • First observedlist_voices
    • First observedlist_webhook_deliveries
    • First observedmake_call
    • First observedmake_conversation_call
    • First observedmanage_contact
    • First observedsend_message
    • First observedset_webhook
    • First observedtest_webhook
    • First observedupdate_agent
    • First observedupdate_conversation

TDQS

A3.7/5.0

Scored across 28 tools

Disambiguation4/5

Most tools have clearly distinct resource+action targets, and the two call tools (make_call vs make_conversation_call) are well-differentiated by their descriptions. Minor overlap exists in the messaging domain where get_messages (raw per-number messages) could be confused with get_conversation (threaded history), but descriptions clarify the boundary.

Naming Consistency4/5

The set overwhelmingly follows a consistent verb_noun pattern (list_agents, create_agent, get_webhook, send_message). The main deviations are account_overview (noun-first) and manage_contact (a bundled multi-operation CRUD tool) versus the otherwise granular per-verb style.

Tool Count3/5

28 tools is on the heavy side of the recommended range, though the domain (agents, numbers, calls, SMS, conversations, contacts, webhooks, usage) is genuinely broad. Almost every tool earns its place with little redundancy, but the count pushes into borderline-heavy territory.

Completeness4/5

Coverage is strong: full CRUD for agents and contacts, complete webhook lifecycle (get/set/delete/test/deliveries), and call/message/conversation operations. The notable gap is the phone-number lifecycle — buy_number exists but there is no release/delete or update number operation.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to make real-world phone calls with AI voice technology and provides tools to track call status, transcripts, and summaries. It supports automated communication with both live numbers and simulated businesses for testing and demonstration purposes.
    -
  • A
    license
    A
    quality
    C
    maintenance
    Gives AI agents phone numbers, email, SMS, and voice calls as MCP tools, enabling them to provision numbers, capture 2FA codes, send messages, and make calls.
    15
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to control PSTN phone calls via TelePath, including dialing, hangup, and phone number management.
    1
    15 npm
    MIT