Skip to main content
Glama

Sats4AI - Bitcoin-Powered AI Tools

epub_to_audiobook

Convert books (EPUB/PDF/TXT) to full audiobooks with automatic chapter detection, multi-voice narration, and optional translation to any language before narration. 3 voice tiers: OmniVoice Global (602+ langs, ~106 chars/sat), Inworld Premium (#1 ranked TTS ELO 1217, ~16 chars/sat), Minimax Studio (voice cloning from reference clip, ~5 chars/sat). Min 500 sats. Async — returns jobId, poll until completed (5-60+ min). Single payment, full outcome — no multi-step orchestration required. Pay with Bitcoin Lightning — no API key or signup needed. Requires create_payment with toolName='epub_to_audiobook'.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
speedNoSpeech speed 0.5-2.0
voiceNoVoice ID. Must belong to the resolved tier's own voice set — Minimax: English_expressive_narrator, Wise_Woman, Deep_Voice_Man; Inworld: Ashley, Abby. OMIT to get a valid default for whichever tier is used. Ashley is Inworld-ONLY and is rejected on the Minimax tier. On the OmniVoice tier, voice is a voice design instead: words, one per group, separated by commas: male|female; child|teenager|young adult|middle-aged|elderly; very low|low|moderate|high|very high pitch; whisper; american|british|australian|canadian|indian|chinese|korean|japanese|portuguese|russian accent. Omit it for OmniVoice's default voice. Any other word is refused, and the payment stays unused.
modelIdNoOptional. 3 voice tiers: OmniVoice Global (602+ langs), Inworld Premium (#1 ranked), Minimax Studio (voice clone). Omit for default.
fileNameYesOriginal filename with extension (e.g., 'mybook.epub', 'document.pdf', 'story.txt'). Required to detect format.
languageNoNarration language (e.g., English, Spanish, French). NOTE: on the default tier this only affects chapter titles / number expansion — the spoken language comes from the chosen voice. For non-English narration pick a voice whose language matches else it narrates in the voice's own (usually English) accent with no error. translateToLanguage is different: it is checked against the voice tier BEFORE payment, and a tier that cannot speak the target is refused with the tier that can.English
paymentIdYesValid payment ID (must be paid)
epubBase64YesBase64-encoded book file (EPUB, PDF, or TXT)
translateToLanguageNoTranslate book to this language before narration. Accepts English names ('Spanish', 'Chinese (Simplified)') or ISO-639 codes / locale tags ('es', 'en-US', 'pt-BR'). Cost added to price.
selectedChapterIndicesNoChapter indices to include (0-based). Omit to auto-select content chapters. NOTE: auto-select drops front/back matter heuristically and can silently exclude a short (<200 char) wanted chapter near the start/end — pass explicit indices if you need a specific set.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedInput schema / properties / voice / description
      Previous value: -"Voice ID. Must belong to the resolved tier's own voice set — Minimax: English_expressive_narrator, Wise_Woman, Deep_Voice_Man; Inworld: Ashley, Abby. OMIT to get a valid default for whichever tier is used. Ashley is Inworld-ONLY and is rejected on the Minimax tier."New value: +"Voice ID. Must belong to the resolved tier's own voice set — Minimax: English_expressive_narrator, Wise_Woman, Deep_Voice_Man; Inworld: Ashley, Abby. OMIT to get a valid default for whichever tier is used. Ashley is Inworld-ONLY and is rejected on the Minimax tier. On the OmniVoice tier, voice is a voice design instead: words, one per group, separated by commas: male|female; child|teenager|young adult|middle-aged|elderly; very low|low|moderate|high|very high pitch; whisper; american|british|australian|canadian|indian|chinese|korean|japanese|portuguese|russian accent. Omit it for OmniVoice's default voice. Any other word is refused, and the payment stays unused."
  2. Changed1 schema field changed
    • changedInput schema / properties / language / description
      Previous value: -"Narration language (e.g., English, Spanish, French). NOTE: on the default tier this only affects chapter titles / number expansion — the spoken language comes from the chosen voice. For non-English narration pick a voice whose language matches (or use translateToLanguage), else it narrates in the voice's own (usually English) accent with no error."New value: +"Narration language (e.g., English, Spanish, French). NOTE: on the default tier this only affects chapter titles / number expansion — the spoken language comes from the chosen voice. For non-English narration pick a voice whose language matches else it narrates in the voice's own (usually English) accent with no error. translateToLanguage is different: it is checked against the voice tier BEFORE payment, and a tier that cannot speak the target is refused with the tier that can."
  3. Changed2 schema fields changed
    • removedInput schema / properties / voice / default
      Removed value: -"Ashley"
    • changedInput schema / properties / voice / description
      Previous value: -"Voice ID (e.g., Ashley, Deep_Voice_Man, Calm_Woman)"New value: +"Voice ID. Must belong to the resolved tier's own voice set — Minimax: English_expressive_narrator, Wise_Woman, Deep_Voice_Man; Inworld: Ashley, Abby. OMIT to get a valid default for whichever tier is used. Ashley is Inworld-ONLY and is rejected on the Minimax tier."
  4. First observed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so richly: async behavior with jobId polling, expected 5-60+ min duration, minimum 500 sats cost, per-tier throughput (chars/sat), single-payment full-outcome model, and the requirement that payment must be completed first. This is exactly the operational context an agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence front-loads the purpose and capabilities, followed by tier, pricing, async, and payment details in a logical order. It is dense and somewhat repetitive with the schema's own voice/tier text, but nearly every sentence carries operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 9-parameter, no-annotation, no-output-schema tool, the description covers the full lifecycle an agent must understand: prerequisite payment, async polling, timing bounds, cost floor, and voice-tier/translation constraints. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema itself is extremely detailed about voice IDs, tier constraints, language vs translateToLanguage, and chapter indices. The description largely restates tier and payment facts rather than adding new parameter-level meaning, so it sits at the baseline for a fully documented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — converting EPUB/PDF/TXT books into full audiobooks — plus the differentiators (chapter detection, multi-voice narration, optional translation). It is clearly distinguishable from siblings like translate_epub and text_to_speech without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the invocation context well: payment is required first via create_payment with toolName='epub_to_audiobook', the call is async and must be polled, and no API key/signup is needed. It does not explicitly name alternative sibling tools or state when-not-to-use this one, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.