MCP Coqui TTS Server
Related Servers
Alternatives to MCP Coqui TTS Server
No user-submitted related servers found.
Related Servers
- AlicenseAqualityDmaintenanceProvides text-to-speech synthesis and voice cloning using Coqui TTS, enabling natural speech output from text and audio samples.41MIT
- AlicenseAqualityFmaintenanceEnables text-to-speech conversion using ElevenLabs API with voice management, streaming support, and multiple models.51MIT
- FlicenseNot gradedqualityDmaintenanceEnables text-to-speech synthesis and voice cloning through GPT-SoVITS API integration. Supports multiple languages (Chinese, English, Japanese, Korean, Cantonese), dynamic model switching, and reference audio-based voice quality replication.4-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to generate and play high-quality text-to-speech audio using the Kokoro model, with support for multiple voices, adjustable speaking speed, and audio caching.-
- AlicenseNot gradedqualityCmaintenanceEnables text-to-speech generation using the Groq API, supporting multiple audio formats and optional local playback.18 npm1MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to convert text to high-quality speech audio using MeloTTS. Automatically splits long texts into segments, generates WAV files, and merges them using ffmpeg with support for multiple languages and customizable speech parameters.MIT
TDQS
Scored across 4 tools
Most tools are clearly distinct: list_models, clone_voice, and speak/synthesize_long_text target different actions. The only ambiguity is between speak and synthesize_long_text, though the long-text description clarifies the boundary.
All names use snake_case and are readable. The pattern is mostly verb_noun (list_models, synthesize_long_text, clone_voice), with 'speak' as a minor bare-verb deviation.
Four tools is well-scoped for a focused TTS server: model listing, standard synthesis, long-text synthesis, and voice cloning. Each tool earns its place without redundancy.
The core TTS lifecycle is covered: discovering models, synthesizing speech (standard and long), and cloning voices. Minor gaps exist around voice management or output options, but agents can work around them.