Call Me
Server Details
Your AI rings your iPhone, speaks its question, and gets your spoken answer back as text.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- radres/call-me
- GitHub Stars
- 20
- Server Listing
- call-me
TDQS
Scored across 6 tools
Call vs text is cleanly separated (ringing vs push notification), and setup/set_thread_title are distinct utilities. The only mild overlap is poll_result (tracks an in-flight call) versus wait_for_reply (awaits incoming human messages), but the descriptions clearly delineate each.
Most names follow a snake_case verb_noun or verb_object pattern (poll_result, set_thread_title, wait_for_reply). The bare names call, text, and setup deviate slightly but remain readable and idiomatic for the domain.
Six tools is well-scoped for a phone-bridge server: each action (ring, poll, text, await reply, label thread, onboard) earns its place without redundancy or bloat.
The surface covers the core lifecycle: make a call, poll for its outcome, send one-way texts, receive replies, and title the thread. Minor gaps exist (no explicit hangup/cancel or thread listing), but agents can work around them.
Available Tools
6 toolscallCall the humanADestructiveInspect
Ring the human's iPhone, read question aloud word for word, and return
what they say back. It does not have to be a question: an update or a
summary of the day ahead works too, up to 2000 characters (about two
minutes of speech). Waits up to ~30s; if the result comes back status
"ringing" or "in_call", follow up with poll_result. On a local client, read to from
~/.aiphone/config.json → user_number; this hosted server cannot read
that file itself. It is a Call Me app number, not a regular phone number.
If missing, returns setup_required with the App Store link. Pass a previous
session_token to stay in the same thread on their phone.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| question | Yes | ||
| from_label | No | Someone's AI | |
| session_token | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint=false, destructiveHint=true, openWorldHint=true) by disclosing the ~30s wait ceiling, the status values that require polling, the 2000-character (~2 min speech) input cap, the setup_required failure mode, and the fact that the hosted server cannot read the local config file. This is real operational context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and its return are front-loaded in the first sentence, and every subsequent sentence carries distinct operational detail. It is dense and slightly sprawling with parentheticals, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotation-level detail on returns, the description supplies the success statuses, the polling fallback, and the setup_required path. It stops short of describing timeout/error behavior when the human never answers, which is the one remaining gap for a 4-parameter outbound-call tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load, and it explains `question` (spoken verbatim, up to 2000 chars, not necessarily interrogative), `to` (a Call Me app number sourced from config), and `session_token` (thread continuity). Only `from_label` is left unexplained, so it nearly compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource chain: ring the human's iPhone, read `question` aloud, return their reply. It also names the follow-up sibling (poll_result) and the alternative content types, so an agent can distinguish it from text/wait_for_reply without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: use poll_result when the result status is "ringing" or "in_call"; read `to` from ~/.aiphone/config.json when running locally; pass a prior session_token to stay in the same thread; expect setup_required with an App Store link if unconfigured. When-to-use and what-to-do-next are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
poll_resultGet call resultARead-onlyInspect
Check on a call that was still ringing or in progress. Waits up to ~25s for it to settle; returns the final status and transcript once the call ends, the human declines, or the ring times out ("missed").
| Name | Required | Description | Default |
|---|---|---|---|
| call_id | Yes | ||
| session_token | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, non-destructive safety profile. The description adds important behavioral detail: it waits up to about 25 seconds, returns final status and transcript, and lists terminal conditions including call end, human decline, and ring timeout ('missed').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tightly written sentences with the core action and timeout behavior front-loaded. Every clause adds useful information with no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description helpfully explains what is returned and when. Still, it omits parameter meanings, especially session_token, which is a notable gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only loosely implies what call_id means by mentioning 'a call' and says nothing about session_token, leaving both required parameters largely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Check on a call') plus a clear scope ('that was still ringing or in progress'). It distinguishes this polling tool from siblings like wait_for_reply by focusing on call status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the context for use: checking a call that is still ringing or in progress. However, it does not name alternatives or state when not to use this tool, such as for already-completed calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_thread_titleSet thread titleBDestructiveInspect
Name this session's conversation thread on the human's phone (e.g. the project or task you are working on). Shown as the thread title in the app.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| session_token | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so the agent knows this mutates state. The description adds the useful detail that the title is user-visible in the app, but never explains the destructive aspect (e.g. that it overwrites an existing title) or any auth/session requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the action and scope; the second sentence adds display context. No wasted words, though the parenthetical example could be folded more tightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-param tool with no output schema this is nearly adequate, but the unexplained destructiveHint and the undocumented session_token leave an agent guessing about overwrite behavior and session prerequisites.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for 2 params, so the description carries the burden. It meaningfully clarifies the 'title' parameter (project/task name, displayed as thread title) but says nothing about session_token or its format/source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (name/set) and resource (this session's conversation thread) plus where it surfaces (the human's phone, app thread title). An agent can tell it apart from call/text/wait_for_reply, though it never explicitly contrasts with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'e.g. the project or task you are working on' implies what content to pass, giving soft usage context. There is no explicit when-to-use trigger, no when-not-to-use, and no alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
setupSet up Call MeARead-onlyInspect
Return first-setup instructions and a clickable App Store download link when no number is available.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful behavior context: it fires only when no number is available and it returns instructions plus a clickable download link, telling the agent what it gets back for a tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the return value and closes with the triggering condition. No filler, no restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only setup helper with no output schema, the description tells the agent when it applies and what it returns, which is enough to invoke it correctly. Minor gaps remain around what happens if a number is already configured, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the empty schema already makes that unambiguous. With no parameters to document, the baseline is 4; the description correctly does not invent parameter semantics that don't exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a concrete verb+resource: returns first-setup instructions and an App Store download link. It doesn't name or differentiate itself from siblings like call or text, but the purpose is unmistakable for a tool named 'setup' with no parameters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a condition — 'when no number is available' — which implies the agent should call this before a number is configured. However, it doesn't say what to do instead once a number exists, or name the alternative tools (call/text) that become usable. Guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
textText the humanADestructiveInspect
Send a one-way text to the human's phone (push notification, no ring).
On a local client, read to from ~/.aiphone/config.json → user_number;
this hosted server cannot read that file itself. If missing, returns
setup_required with the App Store link. Pass a previous session_token to
stay in the same thread.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| body | Yes | ||
| from_label | No | Someone's AI | |
| session_token | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, and the description is consistent with a write/external-send operation. It adds real value beyond the annotations: the no-ring delivery semantics, the setup_required failure path with an App Store link, and the threading behavior of session_token. It does not describe rate limits, delivery latency, or whether the send is retractable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with what the tool does, then setup mechanics, then threading. No filler. Slightly dense but every sentence conveys a distinct, load-bearing fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by naming the setup_required error case and its payload. The remaining gap is the undocumented `from_label` parameter and success-response shape, but for a single-action send tool the coverage is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it only partially does: it documents `to` (resolved from config.json → user_number) and `session_token` (continue an existing thread). `body` is left implicit and `from_label` is never mentioned, so one of four parameters has no semantics anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair ('Send a one-way text to the human's phone') and immediately qualifies the channel ('push notification, no ring'). The 'one-way' framing distinguishes it from reply-oriented siblings like wait_for_reply and poll_result without needing their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable operational context: where `to` comes from on a local client, that the hosted server cannot read that file, and that a previous session_token continues the thread. It stops short of explicitly naming an alternative (e.g. 'use call/wait_for_reply when you need an answer'), so exclusions are implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_replyWait for a text replyARead-onlyInspect
Wait for the human to send something back to this session — texts and voicemail transcripts arrive here. Long-polls up to wait_s (max 30s); returns {events, cursor}. Pass the returned cursor next time to only see new events. An empty events list just means nothing yet — poll again if you are still waiting.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | ||
| wait_s | No | ||
| session_token | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare a safe read (readOnlyHint=true, destructiveHint=false); the description adds the traits that actually matter here — that this long-polls and blocks for up to wait_s (max 30s) and returns {events, cursor}. It also pre-empts the confusing case where an empty list is not an error. Error/timeout behavior is not covered, so it is not fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the purpose, then the blocking behavior, then the cursor contract, then the empty-result caveat. No filler and nothing repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly carries the return shape ({events, cursor}) and the empty-list semantics. Combined with the safety annotations, an agent has what it needs to call this correctly. It would be complete with a note on what session_token identifies or what a timeout looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameters, and it does for two of three: it explains cursor as a resume token from the prior response and wait_s as the long-poll bound with a 30s cap. session_token is left unexplained, which for a session-scoped tool is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: waiting for an inbound human message (texts and voicemail transcripts) on this session. It is clear about what arrives and what the call does. It stops short of explicitly contrasting itself with the sibling poll_result, which is the closest analog and would be the obvious confusion point.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies when to call (while waiting for a reply) and instructs to poll again on an empty result, but never says when to prefer this over siblings like poll_result or call. No prerequisites or exclusions are stated, so the routing decision is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
- First observed
call - First observed
poll_result - First observed
set_thread_title - First observed
setup - First observed
text - First observed
wait_for_reply
Related MCP Connectors
Your AI rings your iPhone, speaks its question, and gets your spoken answer back as text.
- call-meOAuthapp.getcallme
Calls your phone when an AI task finishes or is blocked — hear it, say what's next.
Your agent asks you a question on your phone, waits for the answer, then resumes.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to ask questions to users asynchronously via a local macOS app, allowing users to respond by text or voice.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to place real phone calls, engage in spoken conversations across 12 languages, and report the outcomes for tasks like booking appointments, answering questions, negotiating, or confirming orders.MIT
- AlicenseAqualityAmaintenanceCall, text, or push your phone when an agent needs input mid-task — reply by voice instead of babysitting a long-running or blocked terminal.10MIT
- AlicenseNot gradedqualityCmaintenanceEnables an AI agent to read and send real iMessages—including text, voice notes, stickers, links, images, and reactions—through a managed cloud phone line without needing a Mac.4 npmAGPL 3.0
Glama MCP Gateway
Add one secure layer between your agents and this server.