Skip to main content
Glama

EasyVoice text to speech

Server Details

Hosted text-to-speech MCP server: 66 voices in 9 languages, long-form narration, podcasts.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL

TDQS

Score is being calculated.

Available Tools

7 tools
clone_voiceClone a VoiceInspect

Enroll a custom cloned voice from an audio sample (async). Requires a Pro or active 7-day pass key. Provide exactly one of audio_base64 or audio_url (https). Sample limits: 10 MB max upload; wav, mp3 or m4a; a 10–30 second clean single-speaker recording (enrollment rejects clips shorter than 10 or longer than 30 seconds; trim longer recordings first). Account limits: maximum 3 cloned voices, 5 enrollments per day. Returns a voice_id with status "enrolling" — check list_voices until its status is "ready", then use it in create_podcast.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesDisplay name for the cloned voice.
consentYesRequired attestation: you confirm you have the recording rights to this audio sample. The server rejects anything but an explicit true.
audio_urlNohttps URL to the audio sample — fetched server-side, same 10 MB bound.
mime_typeNoSample MIME type — wav, mp3 and m4a are accepted (the server enforces the exact allowlist).audio/wav
audio_base64NoBase64-encoded audio sample (10 MB max decoded).

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
voice_idYes
create_podcastCreate Podcast EpisodeInspect

Create a two-host (A/B) podcast episode as an async job. Requires a Pro subscription key. Up to 60 segments; combined text up to 30,000 characters. voice_a and voice_b set the two speaker voices — your READY cloned voice_* ids are allowed. Returns a job id — poll get_job_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNomp3
voice_aYesVoice id for speaker A — catalog voice or one of your READY cloned voice_* ids.
voice_bYesVoice id for speaker B.
segmentsYesAlternating dialogue segments (max 60).

Output Schema

ParametersJSON Schema
NameRequiredDescription
job_idYes
statusYes
generate_long_formGenerate Long-Form NarrationInspect

Generate long-form narration (up to 500,000 characters ≈ 9 hours of audio) as an async job. Requires a Pro subscription key. Kokoro narration voices only — Arabic (ar_*) and cloned (voice_*) voices are rejected by the Long-Form engine. Returns a job id — poll get_job_status; parts render progressively and audio_url appears when completed.

ParametersJSON Schema
NameRequiredDescriptionDefault
toneNoTone preset.neutral
inputYesThe full script, up to 500,000 characters.
pitchNoPitch shift in semitones (-4 to 4).
speedNo
voiceNoKokoro voice id — call list_voices for the catalog.af_aoede
formatNomp3

Output Schema

ParametersJSON Schema
NameRequiredDescription
job_idYes
statusYes
input_charsYes
get_job_statusGet Job Status
Read-only
Inspect

Check the status of a long-form, podcast or tts job owned by this key. Returns status plus audio URL(s) when completed. Audio links are not guaranteed to persist — download promptly.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesJob id returned by an async tool.

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
statusYes
audio_urlNo
get_usageGet Usage
Read-only
Inspect

Check this API key's plan and usage. Free keys see today's used/remaining characters (shared 5,000/day pool, resets at midnight UTC), the 20 requests/minute rate limit and 2-concurrent cap. Pro keys are unlimited (fair use: 10M characters per rolling 30 days, then slowed to 30 requests and 10,000 characters a minute, 10 and 3,000 past 20M, on every surface -- never blocked or billed) at 60 requests/minute.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
planYes
usedNo
limitNo
reset_atNo
remainingNo
rpm_limitYes
unlimitedNo
concurrency_limitNo
fair_use_chars_per_monthNo
list_voicesList Voices
Read-only
Inspect

List all 66 EasyVoice catalog voices (id, name, language, accent, gender, free/pro tier) plus this key's cloned voices. Call this before text_to_speech or create_podcast to pick voice ids.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
planYes
clonedYes
voicesYes
text_to_speechText to SpeechInspect

Convert text to a downloadable MP3 or WAV audio file with one of EasyVoice's 66 neural voices. Returns a hosted audio URL (no inline audio). Per-call limit: 8,000 characters (4,000 for Arabic ar_* voices) — text over the limit is automatically queued as a background job instead of erroring; poll get_job_status with the returned job_id for the result. Free keys can use the 12 free voices; Pro ($9.99/mo) unlocks the full catalog. A Pro account past the fair-use line is paced in characters a minute: a call may wait up to 20 seconds, and one that would wait longer returns a tool error saying when to retry.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text to speak.
speedNo
voiceNoVoice id — call list_voices for the catalog.af_aoede
formatNomp3

Output Schema

ParametersJSON Schema
NameRequiredDescription
voiceYes
formatYes
audio_urlYes
charactersYes
remaining_free_chars_todayNo

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updates
    • First observedclone_voice
    • First observedcreate_podcast
    • First observedgenerate_long_form
    • First observedget_job_status
    • First observedget_usage
    • First observedlist_voices
    • First observedtext_to_speech

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Hosted text-to-speech MCP server for AI agents with 54 neural voices in 9 languages, including Brazilian Portuguese. Pay-per-use API, no GPU or subscriptions needed.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that converts text into lifelike speech using Microsoft Edge's Text-to-Speech service, supporting customizable voice, rate, volume, and pitch.
    4
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Text to speech in 149 languages: MP3 links from any assistant. Free without an account. 2,253 voices; PRO adds HD voices, WAV, dialogue with a voice per speaker, Script mode timing and transcription.
    7
    62 npm
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources