Skip to main content
Glama

Sats4AI - Bitcoin-Powered AI Tools

open_voice_bridge

Open a phone call you drive one turn at a time: we dial, transcribe what the other side says, and speak whatever text you send. You call voice_bridge_say to talk, poll_voice_bridge to read transcripts, end_voice_bridge to hang up. ⚠ NOT real-time conversation. Each turn costs an HTTP round trip plus your own model's thinking plus speech synthesis — measured at 2.5-4 seconds before audio starts, against the ~0.8s a natural back-and-forth needs, and nothing signals you when the other side stops talking, so you are polling and guessing. Good for: leaving a spoken message, navigating an IVR, reading something out, a slow exchange where a pause is fine. For a real conversation with a human, use ai_call — we run the agent at conversational speed and return the transcript. Unused deposit time is refunded automatically when the call ends — including when the callee hangs up. Pass refundAddress (a Lightning address) and the remainder is SENT there with no further action; without it the refund is held as a claimable LNURL-withdraw link returned by end_voice_bridge. Use this when the content of the call must never leave your side, and the pace can tolerate a pause between turns. When NOT to use: not for fully-managed agent-style calls where we handle the brain (use ai_call). Not for one-shot TTS broadcasts or IVR playback (use place_call). Not when live transcript polling adds no value — the per-turn overhead isn't worth it. Privacy: transcripts held in memory only, garbage-collected 30 minutes after the call ends; call audio is never persisted. Pay with Bitcoin Lightning — no telecom account, no signup. Requires create_payment with toolName='voice_bridge_open', phoneNumber, durationMinutes. Deposit is priced per destination off the carrier rate sheet and is BTC-pegged, so call create_payment for the exact figure; premium-rate and satellite ranges are refused.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
codecNoPCMU 8kHz (default, universal) or L16_16000 for HD voice when both endpoints support it
greetingNoSpoken the instant the callee answers (max 500 chars). Synthesized while the phone rings, so there is no pause before the first word — without it the callee hears silence until your first voice_bridge_say lands. Strongly recommended for any call a human answers.
languageNoBCP-47 language tag (default en-US). See /api/l402/voice-bridge/coverage for the matrix.
paymentIdYesValid payment ID from create_payment (toolName=voice_bridge_open)
sttEnabledNoDefault true. Set false for TTS-only broadcast calls.
ttsEnabledNoDefault true. Set false to bring-your-own-audio via voice_bridge_say.
phoneNumberYesDestination phone number in E.164 format (e.g., +14155550100)
refundAddressNoLightning address (e.g. you@wallet.com). Strongly recommended: the unused deposit is sent here automatically when the call ends, so nobody has to claim anything. Without it the refund waits as an LNURL-withdraw link.
durationMinutesNoDeposit for N minutes, 2-30 (default 3). Unused time refunded.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / greeting
      Added value: +{
      +  "description": "Spoken the instant the callee answers (max 500 chars). Synthesized while the phone rings, so there is no pause before the first word — without it the callee hears silence until your first voice_bridge_say lands. Strongly recommended for any call a human answers.",
      +  "type": "string"
      +}
  2. Changed1 schema field changed
    • changedInput schema / properties / refundAddress / description
      Previous value: -"Lightning address for automatic refund of unused time"New value: +"Lightning address (e.g. you@wallet.com). Strongly recommended: the unused deposit is sent here automatically when the call ends, so nobody has to claim anything. Without it the refund waits as an LNURL-withdraw link."
  3. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses major behavioral traits: not real-time, per-turn latency (2.5-4s), no completion signal, refund mechanics (with/without refundAddress), privacy (in-memory, 30-min GC, no audio persistence), deposit pricing, refuses premium/satellite ranges. Annotations only say openWorldHint, so the description carries full burden and excels.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but every sentence earns its place. Front-loads core workflow (how turns work), then warnings, use cases, exceptions, privacy, payment. Though long, it is structured logically and avoids fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, high-stakes tool (payment, privacy, multiple modes), the description covers workflow, prerequisites (create_payment), refund, privacy, and constraints. No critical gaps found; output schema absent but return format is described via emitted signals.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers 100% of params, so baseline is 3. Description adds meaning for key params (greeting recommendation, refundAddress benefits, duration deposit, codec default). Does not explicitly map all params but schema is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific purpose (open a call you drive turn-by-turn) and contrasts with ai_call (real-time agent) and place_call (one-shot). Distinguishes clearly from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (turn-paced, privacy-critical, refund, etc.) and when-not-to-use (real-time, one-shot, polling overhead). Names alternatives (ai_call, place_call) directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.