Enables Claude to act as an autonomous radio DJ, generating live-coded music with Strudel, making text-to-speech announcements, and responding to audience requests through a browser UI.
Voice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.
Enables AI agents to present interactive code walkthroughs with voice narration, opening files, highlighting code, and showing inline explanations with synchronized text-to-speech.
Text to speech for MCP clients. Reads numbers, dates and order IDs correctly. 23 languages, six voices, every render watermarked. Free key with 100,000 characters, no card.
A language expression coach MCP server that provides native-sounding translations with cultural context, tone notes, and audio pronunciation for over 13 languages.
Enables local text-to-speech synthesis for Claude and Cursor using Supertonic 3, with support for multiple voices, expressions, and languages. No API key or cloud required.
MCP server for the Supertone TTS API. Generate natural speech, browse and preview the
voice catalog, predict synthesis cost, and create cloned voices — directly from Claude
Desktop, Cursor, or any MCP-compatible client. Supports Korean, English, Japanese, and
20+ other languages, with speed, pitch, and emotion-style control.
Converts text or transcripts into MP3 audio using Microsoft Edge's free neural voices. Provides text-to-speech and voice listing tools with no API key required.
An MCP server that wraps the WellSaid Labs text-to-speech API, enabling Claude to generate lifelike voiceovers with voice selection, prosody control, and caption support directly from text.
An MCP server implementation that integrates with Minimax API to provide AI-powered image generation and text-to-speech functionality in editors like Windsurf and Cursor.
Enables AI agents to perform local audio tasks such as speech synthesis, voice cloning, music and sound effect generation, and audio editing through MCP, with GPU models loaded on demand and released after idle.
Enables text-to-speech synthesis through a streaming, GPU-accelerated gateway, exposing a single tool that returns playable WAV audio with configurable voice and speed.