A text-to-speech MCP server with 48 voices across 9 languages, supporting emotion spans, SFX tags, and multi-speaker dialogue. Deployable via a single npx command with built-in guardrails and swappable backends.
Enables local text-to-speech synthesis for Claude and Cursor using Supertonic 3, with support for multiple voices, expressions, and languages. No API key or cloud required.
Converts text or transcripts into MP3 audio using Microsoft Edge's free neural voices. Provides text-to-speech and voice listing tools with no API key required.
Enables AI agents to synthesize natural speech using either platform system voices or premium OpenAI TTS, with automatic engine selection and graceful fallback.
Provides text-to-speech synthesis using Microsoft Edge's free TTS engine, supporting multiple voices, languages, and audio output options (base64 or file).
A feature-rich MCP server for Discord, giving AI companions full presence in Discord communities — reading, responding, reacting, searching, and speaking.
Provides voice summaries after each AI request in Cursor, Cline, or any MCP-supported editor, allowing users to hear what was done and stay informed without staring at the screen.
Enables AI assistants to convert text to high-quality speech audio using MeloTTS. Automatically splits long texts into segments, generates WAV files, and merges them using ffmpeg with support for multiple languages and customizable speech parameters.
Local Korean text-to-speech MCP server using a fine-tuned CosyVoice2 model, enabling voice synthesis directly from Claude or Codex with privacy and no API costs.
Enables MCP clients to speak by running local voice models using Chatterbox Turbo TTS or Kokoro TTS, with support for voice cloning, paralinguistic tags, and multiple voices.
Enables AI assistants to synthesize speech with custom or cloned ElevenLabs voices and display an inline audio player with waveform visualization, transcript toggle, and dark mode support.
Exposes a single transcribe tool over streamable HTTP so containerised agents can send a media file name and receive text transcribed locally by MacWhisper on the host Mac's GPU, with token-gated access and no uploads or API keys. Callers place media in a configured directory, and the blocking call returns the finished transcript.
An MCP server that turns a script into a Hyperplexed-style motion-graphics MP4. Plug the URL into Claude Desktop, Cursor, or any MCP-compatible agent and it gains a render_video tool.
Enables cloning a voice from a local audio or video file and generating shareable audio-reactive waveform MP4 videos of spoken text, with customizable palettes, photos, and layouts, all through natural language from MCP-aware agents.
Hosted text-to-speech MCP server for AI agents with 54 neural voices in 9 languages, including Brazilian Portuguese. Pay-per-use API, no GPU or subscriptions needed.