Uhura
Integrates with ElevenLabs Agents to place and manage AI voice phone calls: draft call briefings, rehearse conversations, confirm and dial calls, steer ongoing calls, and retrieve transcripts and costs in ElevenLabs credits.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@UhuraCall 030 23125 000 and ask if the shop is open on Friday."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Uhura
Your AI assistant can now make the phone call.
AI assistants can research, write and plan, but some things still only happen on the phone: the doctor's office without online booking, the travel agency that answers emails slowly, the shop that won't say on its website whether it's open on Friday.
Uhura closes that gap. Ask Claude, or any assistant that speaks MCP, to make the call: it drafts the briefing, you approve it, and an AI voice agent places the call. The agent introduces itself as an AI and works through your questions. If it lacks an answer, it asks you while the person waits on the line. You get the transcript, and what the call cost.
Uhura is a small self-hosted service in front of ElevenLabs Agents. It works from your AI assistant through MCP, or from the command line. One instance can serve one person on a laptop or several people from a server.
Status (2026-10-04): built and verified against a real ElevenLabs account. Real phone calls work through a Twilio number, in German, including a company's phone menu and a waiting queue the agent sits through silently; progress reports and the cost of each call are verified on real calls too. Asking you a question during a spoken call has so far only been verified in text conversations, and no English call has been made yet. See docs/verified.md.
How a call works
Draft. Give the number, the language, the goal and the facts the agent may share. Uhura checks the number and shows the exact briefing. Nothing is dialled.
Rehearse (optional). Talk to the real agent in text, playing the person being called, to see how it handles the briefing.
Confirm. A separate step dials.
Follow and steer. The agent reports progress as it goes (menu, queue, person reached, answers). If it lacks an answer, it asks you and waits on the line for your reply. Instructions you send reach it with its next check-in, at the latest just before it hangs up.
Result. The transcript, duration and cost in ElevenLabs credits are stored; no audio is kept.
$ uhura draft --to "030 23125 000" --principal "Erika Musterfrau" --language de \
--topic "eine Frage zu Ihren Öffnungszeiten" \
--goal "Ask whether the shop is open on Friday."
Message composed, Captain. Awaiting your order to transmit.
call id : 74b8167429cf
number : +493023125000
opening : Guten Tag, hier spricht ein KI-Assistent im Auftrag von Erika Musterfrau. …
after menu: Guten Tag, ich bin ein KI-Assistent … Es geht um eine Frage zu Ihren Öffnungszeiten. …
$ uhura rehearse 74b8167429cf # text conversation, nobody is called
$ uhura confirm 74b8167429cf # dials, follows the call, prompts for answers
Hailing frequencies open.The brief can be in English; the agent speaks the call's language (--language, German
here) and always opens with the fixed disclosure line for it. The topic completes the
sentence "Es geht um …" in the line the agent says to the first person it reaches after a
phone menu or queue, so it is written in the call's language.
Related MCP server: PhoneBooth MCP Server
Responsible use
Uhura places real phone calls to real people, with an AI voice. It is meant for calls you would otherwise make yourself: asking a shop, a doctor's office or a travel agency something on your own behalf. It is not meant for marketing, mass calling, surveys, or anything that hides that an AI is calling or on whose behalf.
The guardrails below are built for that use: every call opens with a fixed disclosure, each call is drafted and confirmed separately, and the limits keep volume low. They are not a licence to call anyone about anything. You are responsible for the calls you place and for the laws that apply to them. docs/legal-notes.md records the reasoning for Germany only; it is not legal advice, and other countries' rules on automated calls, recording and transcription differ.
Guardrails
Enforced by the service, whatever the briefing says:
Drafting never dials, and a draft can be confirmed only once.
Emergency, premium-rate, shared-cost and service numbers are refused.
Only countries in
UHURA_ALLOWED_COUNTRIEScan be called.Per person:
UHURA_DAILY_CALLScalls per 24 hours andUHURA_MONTHLY_MINUTESper 30 days. A call is only dialled if the minutes left cover the longest call the agent will make (20 minutes); calls still running count at that length until their result arrives.Per person:
UHURA_DAILY_REHEARSALSrehearsals per 24 hours, each closed after 10 minutes.The opening line is fixed per language. It says that an AI is calling, for whom, and that the call is transcribed, and asks for agreement. A briefing cannot change it.
People only see their own calls. Calls and transcripts are deleted after
UHURA_RETENTION_DAYS.The service does not start with placeholder or short tokens or tool secret (fewer than 16 characters), since it is usually reachable from the internet.
Asked of the agent through its fixed rules (reliable in tests, but a language model follows them, it is not forced to): information only, no bookings or payments; end the call if the person objects to transcription; ask instead of inventing facts; in phone menus, press keys toward the goal and never agree to a recording; stay silent on hold; give the fixed, short re-introduction (shown in the draft) to the first person reached after a menu or queue; keep turns short and do not repeat the briefing.
Quick start
uv sync
uv run uhura init # creates .env with a fresh token and tool secret
# then add your ElevenLabs key to .env
uv run uhura setup-agent # creates the agent; put the printed id into .env
uv run uhura serve # http://127.0.0.1:8787
uv run uhura check # tells you what is still missingThe full walk-through, including the phone number, the public address the agent's tools need, Docker and hosting for several people, is in docs/setup.md.
Commands
Command | What it does |
| Create |
| Run the service |
| Create or update the agent in your ElevenLabs account |
| Point the agent's tools at the service's public address (default: local ngrok tunnel) |
| Verify the setup; with |
| Prepare a call and show the briefing |
| Try a draft in text against the real agent |
| Dial, follow the call, answer its questions |
| Follow a call that is already running |
| Queue a follow-up instruction |
| Transcript, status and cost / recent calls |
Use from an MCP client
claude mcp add --transport http uhura http://localhost:8787/mcp \
--header "Authorization: Bearer <token>"The service serves the tools itself at /mcp. For clients that can only start a local
program there is a stdio adapter, uhura-mcp; see docs/operations.md.
Tools: draft_call, confirm_call, wait_for_event, answer, send_instruction,
get_call, list_calls, rehearse_call, rehearse_say, end_rehearsal. The server
tells the client to show every draft to the user and to confirm only after approval.
Documentation
docs/setup.md: installation, accounts, phone number, hosting
docs/operations.md: starting a session, making calls, MCP in Claude Code, costs, troubleshooting
docs/architecture.md: how it works, HTTP API, events, stored data
docs/verified.md: what has been tested against the real services and what has not
docs/roadmap.md: possible improvements, not built yet
docs/legal-notes.md: the reasoning behind disclosure and transcript-only (Germany; not legal advice)
Development
uv run pytest # no network needed
uv run ruff check . # lint
uv run ruff format . # formatThe name
Lieutenant Nyota Uhura is the communications officer of the USS Enterprise in Star Trek,
played by Nichelle Nichols. She opens channels, hails other ships and relays what they say
to the captain. Uhura does the same job for you: it places the call, passes the agent's
questions to you and your answers back. Its status lines, from "Hailing frequencies open."
to "Channel closed.", are in her voice and live in src/uhura/phrases.py. The ribbed earpiece
in the banner is a nod to the one she wore on the bridge.
License
MIT.
Available Tools
10 toolsanswerC
Answer a question the voice agent asked mid-call. question_id is the event's seq.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| call_id | Yes | ||
| question_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not say whether the text is spoken to the caller, whether the call is blocked or paused while waiting, whether the answer is terminal for that question, or whether multiple answers are allowed — all critical for a mid-call interaction tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the purpose comes first and the non-obvious parameter mapping second. It is efficient, though the extreme brevity is part of why other dimensions fall short.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a mutation-style mid-call action with no annotations and 0% parameter coverage the description leaves too much unsaid: the effect of 'text', the meaning of call_id, and the interaction with wait_for_event/send_instruction are all absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and all three parameters are required, yet the description only clarifies question_id ('the event's seq'), which is genuinely valuable non-obvious information. call_id and text remain completely undefined, so two of three parameters get no semantics from either schema or description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('answer') and a well-scoped resource ('a question the voice agent asked mid-call'), which is far more informative than the bare name 'answer'. However, it never distinguishes itself from the adjacent sibling send_instruction, which an agent could plausibly pick for the same intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'a question the voice agent asked mid-call' implies the trigger condition (a question event arrives during a live call), but there is no explicit when-to-use statement, no mention of the wait_for_event/get_call workflow that would surface such a question, and no exclusion relative to send_instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_callA
Dial a drafted call. Only use after the user approved this specific draft.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the human-approval gate, but says nothing about irreversibility, external side effects (actually placing a phone call), failure behavior, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the core action front-loaded and the constraint immediately after it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, and the approval precondition is covered. What remains thin is the parameter's meaning and the consequences of dialing (cost, irreversibility), which matter for an outward-facing action with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter call_id has 0% schema description coverage and the description never mentions it or clarifies that it must reference an existing draft. The phrase 'a drafted call' hints at the link, but nothing states the ID's origin or format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Dial) and a specific resource (a drafted call), and the qualifier 'drafted' ties it to the draft_call sibling, so an agent can distinguish it from draft_call, rehearse_call, or get_call without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Only use after the user approved this specific draft' is an explicit, unambiguous gating condition that tells the agent both when to invoke it and when not to. This is exactly the kind of prerequisite an agent needs for a side-effecting action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_callA
Prepare a call without dialling. Returns the call id, the normalised number,
the fixed opening line, the voicemail message and the full briefing the voice agent
will receive. voicemail is what to say if a mailbox answers; a fixed AI
introduction is spoken before it. Without it the agent hangs up on voicemail.
topic is a short phrase completing "Es geht um …" / "It is about …" (e.g. "eine
Reiseanfrage"); the agent says it in its fixed re-introduction to the first person
it reaches after a phone menu or queue. No full stops or question marks.
progress switches the agent's progress reports ('progress' events) on or off for
this call; by default the service decides.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| goal | Yes | ||
| facts | No | ||
| topic | No | ||
| region | No | DE | |
| language | No | de | |
| must_not | No | ||
| progress | No | ||
| principal | Yes | ||
| voicemail | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses the return payload (call id, normalised number, opening line, briefing), the critical consequence that omitting voicemail makes the agent hang up, and the progress-event toggle. It omits auth requirements, idempotency, and whether the draft expires or must be confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, then behavior and field semantics follow in a logical order without padding. It is dense but every sentence conveys operational information rather than restating the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return-value detail is a bonus rather than a necessity. For a 10-parameter mutation-ish tool with zero annotation and zero schema-description coverage, however, half the parameters (principal, facts, region, language, must_not) are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 parameters, so the description must compensate, and it only explains voicemail, topic, and progress in depth. Facts, principal, region, language, and must_not receive no semantic guidance, though 'to' and 'goal' are inferable from the purpose statement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening verb+resource+scope, 'Prepare a call without dialling', is precise and immediately signals this is a pre-dial stage. It implicitly separates the tool from confirm_call's dial action, but no sibling is named explicitly, so the differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'without dialling' framing implies this is the step before confirm_call, and the voicemail warning implies when the field is needed. However, the description never names an alternative or states an explicit when/when-not condition, leaving the workflow position to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
end_rehearsalC
End a rehearsal. The draft stays unchanged and can still be confirmed.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It usefully states that the draft is not modified and remains confirmable, but it omits other relevant traits such as idempotency, permissions, or what happens to the rehearsal session itself.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action, that contain no filler. The key behavior follows immediately, making the description efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and the operation is simple, but the description omits parameter semantics and usage alternatives. With no annotations and zero schema description coverage, those gaps leave the definition only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions the sole required parameter 'call_id'. It adds no meaning beyond the bare schema title, leaving the agent without guidance on what identifier to supply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('End a rehearsal'), which distinguishes it from confirming or drafting. However, it does not explicitly name a sibling alternative or clarify its relationship to rehearse_call/confirm_call, so it falls short of the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a consequence ('draft stays unchanged and can still be confirmed') but gives no explicit when-to-use guidance or named alternatives. An agent must infer that confirm_call is the tool to use for confirming rather than ending.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_callB
Status, duration, transcript and cost of a call. cost (once the call has ended)
has the ElevenLabs credits charged, the billed minutes after the silence discount,
and ElevenLabs' dollar prices for the voice platform and the language model.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one genuinely useful behavioral trait: `cost` is only populated once the call has ended, which tells the agent the field may be empty mid-call. It says nothing about auth requirements, error behavior, or transcript availability timing, so the coverage is partial rather than complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the returned fields, and the second sentence spends its words on the non-obvious cost breakdown. The nested parenthetical is slightly dense but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, explaining return fields is partly redundant, so the real added value is the cost-timing nuance. Missing usage guidance and any parameter context leaves it adequate but not complete for a tool whose semantics are otherwise obvious.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter call_id is never mentioned in the description, so no meaning is added beyond the schema. The parameter is self-evident, which softens the gap, but the description contributes nothing to it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (a call) and enumerates what the caller gets back: status, duration, transcript and cost. It reads as a retrieval tool and is clearly distinct from the plural list_calls and the action siblings (answer, end_rehearsal), though it never states the verb 'get' or explicitly contrasts with list_calls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as list_calls, nor any prerequisite (e.g. needing a known call_id). Usage is only inferable from the name and the returned fields.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_callsC
The user's recent calls, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It states ordering ('newest first') but says nothing about pagination behavior, default limit semantics, recency/scope of 'recent', or return shape – despite an output schema existing, behavior beyond the return shape (e.g., does it paginate? is it scoped to a session?) is undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It's a single short clause with no waste, which is good, but it is an incomplete sentence fragment rather than a front-loaded full purpose statement. Concise but under-specified rather than tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a listing tool with an undocumented limit parameter and no annotations, the description is too thin. It gives ordering but omits scope, pagination defaults, and what the agent should do with results – not enough for reliable invocation despite the output schema covering return shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'limit' parameter (default 20) is undocumented in either the description or the schema. The description does not explain what limit does, its default, or bounds. With one parameter at 0% coverage, the description should compensate but adds nothing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'The user's recent calls, newest first' is a noun phrase that restates a resource but does not state a verb or action. It's not a full purpose statement – the action ('list/retrieve') must be inferred from the tool name. It's not tautological but it's a fragment, not a clear purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus siblings like get_call or the rehearsal/draft tools. The ordinal hint 'newest first' implies a browsing use case but the description doesn't state when an agent should prefer this listing over fetching a specific call. No exclusions or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rehearse_callA
Start a rehearsal of a drafted call: the real agent with the real briefing, in text,
with nobody being called. Read what the agent says with wait_for_event, passing the
returned events_after as after ('agent_said' events), and reply as the person
being called with rehearse_say. Questions and instructions work as in a real call.
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the real agent runs with the real briefing, that the mode is text, and that nobody is actually called — important safety-relevant context. It omits lifecycle details such as how the rehearsal ends (end_rehearsal) or whether state persists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in the first clause, then gives the operational workflow. Dense but each sentence earns its place; the parenthetical about events_after is slightly compressed but functional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the description explains how to consume its events, so return values need not be restated. It covers start-of-rehearsal and the interaction loop, but not how the rehearsal concludes, which an agent must learn from end_rehearsal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and call_id has no schema description, but the description's framing ('a drafted call') makes the parameter's role inferable as the id of the drafted call. With a single trivial parameter, this is adequate but adds no explicit format or sourcing detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a rehearsal of a drafted call') and immediately scopes it as text-only with no real call placed. This clearly separates it from siblings like confirm_call (real call) and rehearse_say (mid-rehearsal reply).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to the follow-up tools: read agent output with wait_for_event (passing returned events_after as after) and reply with rehearse_say. It stops short of stating when to rehearse instead of placing a real call (e.g., before confirm_call), so it is clear context but not full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rehearse_sayC
Say something to the agent in a running rehearsal, as the person being called.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It hints that this is a simulation via 'rehearsal' and specifies the caller's role, but says nothing about side effects, whether the agent responds, permissions, or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no padding, front-loaded on the action. It is efficient, though its brevity borders on under-specification rather than being a structural flaw.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values are covered, and two required params need guidance. With no annotations and 0% parameter coverage, the description leaves too much unstated for a multi-agent rehearsal tool that sends a message.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — the schema gives only bare titles 'Text' and 'Call Id'. The description implies text is the utterance and call_id identifies the rehearsal but adds no format, constraints, or meaning beyond that obvious mapping.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (say) and scopes it to a running rehearsal and a role (the person being called), which distinguishes it from rehearse_call and end_rehearsal. However 'say something to the agent' remains slightly abstract about the actual effect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'In a running rehearsal' implies the prerequisite that a rehearsal must already be active, which is useful context. It does not explicitly name alternatives or state when-not to use it versus sibling tools like answer or rehearse_call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_instructionA
Queue a follow-up instruction for a running call. It reaches the agent with its next tool call, at the latest just before it hangs up, so delivery is not immediate.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| call_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers the key async trait: delivery is not immediate, it lands on the agent's next tool call or at latest before hangup. That is genuinely useful behavioral context. It still omits permission requirements and behavior when the call has already ended.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and immediately followed by the timing caveat. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the delivery semantics are covered. However, with 0% parameter description coverage and no annotations, the definition leaves call_id/text semantics under-specified for a two-required-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both call_id and text, so the schema documents nothing. The description implies 'text' is the instruction and 'call_id' targets a running call, but adds no format, length, or identification detail to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Queue a follow-up instruction for a running call.' An agent can tell it apart from siblings like draft_call, confirm_call, or answer, though it doesn't explicitly name the differentiating sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for a running call' implies the precondition (call must be in progress), but there is no explicit when-to-use vs alternatives guidance and no when-not conditions, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_eventA
Wait up to timeout seconds for new events on a call (questions from the agent,
'progress' reports with stage and note, call ended). Pass the highest seq seen
so far as after. Returns an empty list if nothing happened; call again until status
is 'done' or 'failed'.
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | ||
| call_id | Yes | ||
| timeout | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the important traits: a bounded blocking wait (timeout), an empty-list return rather than an error on no activity, and the loop-termination condition. It omits auth requirements, whether the wait is a true long-poll versus sleep, and concurrency caveats, so it is good but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what is being waited on, then the parameter contract, then the return/loop semantics. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter polling tool with an output schema, the description covers the essential loop contract and return-on-empty behavior; because an output schema exists it need not detail the event payload shape. Only the call_id semantics and any failure behavior are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for two of three parameters: timeout is in seconds and after is the highest seq already seen (with a default implied by the polling loop). call_id is never explained, though its meaning is inferable from the resource context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (wait) plus resource (new events on a call) and an enumeration of the event kinds that can arrive (questions, progress reports with stage/note, call ended). It is unmistakably a long-poll primitive rather than a state read, though it never names a sibling (e.g. get_call) to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operating instructions: pass the highest seq seen so far as after, expect an empty list when nothing happened, and re-call until status is 'done' or 'failed'. That is clear when-to-use guidance for the polling loop, but no alternative tool or condition for preferring get_call over waiting is mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
answer - First observed
confirm_call - First observed
draft_call - First observed
end_rehearsal - First observed
get_call - First observed
list_calls - First observed
rehearse_call - First observed
rehearse_say - First observed
send_instruction - First observed
wait_for_event
TDQS
Scored across 10 tools
Most tools have clearly distinct roles: draft_call prepares, confirm_call dials, rehearse_call/rehearse_say/end_rehearsal form a coherent rehearsal group, and get_call/list_calls are reads. Minor overlap exists between answer (real-call question), send_instruction (real-call follow-up), and rehearse_say (rehearsal reply), but descriptions clarify the boundaries.
Most names follow a predictable verb_noun pattern (draft_call, confirm_call, get_call, list_calls, send_instruction, rehearse_call, rehearse_say, end_rehearsal). The bare verb 'answer' and the noun-less 'wait_for_event' are minor deviations from the dominant convention.
Ten tools is well-scoped for a voice-call agent covering drafting, dialing, live interaction, and rehearsal. Each tool earns its place with no obvious redundancy.
The lifecycle is well covered: draft, confirm, event-waiting, answering, instructing, and inspecting calls, plus a full rehearsal loop. Gaps are minor, such as no way to edit or cancel a draft, but core workflows are complete.
Maintenance
Related MCP Connectors
- DialMCPOAuthcom.dialmcp
Let AI agents place real phone calls from your verified number, with transcripts and recordings.
Give AI agents a phone: outbound AI calls that return a summary, transcript, and extracted fields.
Give your AI assistant a real phone line. Place and end real phone calls with AI voice agents, read call transcripts, run a live two-way interpreter between two people who share no language (31 languages, browser link or phone), create and edit voice agents, and search, buy and bind phone numbers in 21 countries. OAuth 2.1 with PKCE — the model never sees your API key; a read-only scope is available. Pay as you go from $0.10/min, $5 free credit for new accounts.
Place real phone calls and hold spoken conversations on a user's behalf, billed per minute.
Related MCP Servers
- FlicenseNot gradedqualityFmaintenanceEnables AI assistants to make real phone calls on your behalf using VoIP, handling conversations automatically through OpenAI's Real-Time Voice API. Simply tell Claude what you want to accomplish and it will call and manage the entire conversation for you.25-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to make real-world phone calls with AI voice technology and provides tools to track call status, transcripts, and summaries. It supports automated communication with both live numbers and simulated businesses for testing and demonstration purposes.-
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to place real phone calls, navigate IVR trees, and retrieve structured answers with transcripts and recordings.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to initiate and control outbound telephone calls and answer inbound calls through Telnyx and LiveKit, with realtime voice conversations powered by Gemini Live or OpenAI Realtime.MIT