Voice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.
Image, video, chat and text-to-speech models (GPT Image, Gemini Image, Veo, Kling, Seedance, Claude, GPT, Gemini) behind one API key. generate_image returns a preview the model can see, and review_image has a vision model critique the result and propose a corrected prompt, so an assistant can generate, check and fix images on its own.
Enables AI agents to perform local audio tasks such as speech synthesis, voice cloning, music and sound effect generation, and audio editing through MCP, with GPU models loaded on demand and released after idle.
Provides text-to-speech and Telegram notification capabilities via the Model Context Protocol. Supports multiple TTS providers and Telegram messaging with optional setup web interface.
Enables spoken consulting case interview practice through Claude, providing real-time cases, grading on structure, analytics, judgment, and communication, and targeted practice drills.
Text-to-speech MCP server that enables AI assistants to read text aloud on the user's computer using Windows SAPI, with no API key or cloud service required.
A cross-platform MCP server that enables Claude to speak using Microsoft Edge TTS with support for over 300 voices across 50+ languages. It requires no API keys and allows for customization of speech rate, volume, and pitch.
Enables text-to-speech functionality on macOS using the say command, offering extensive control over speech parameters like voice, rate, volume, and pitch for a customizable auditory experience.
Exposes a single transcribe tool over streamable HTTP so containerised agents can send a media file name and receive text transcribed locally by MacWhisper on the host Mac's GPU, with token-gated access and no uploads or API keys. Callers place media in a configured directory, and the blocking call returns the finished transcript.
Local MCP server for neural text-to-speech using Kokoro ONNX engine on CPU, supporting SSML tags, multiple voice profiles, and zero-GPU operation for low-latency speech synthesis.
Local voice toolkit over MCP: transcribe audio to text in 25 languages, synthesize speech in 9, and list available voices and languages. Runs fully on-device — no API keys, no cloud.
Enables AI assistants to speak with realistic cloned voices via ElevenLabs TTS on Cloudflare Workers, supporting 29 languages and inline audio playback on both desktop and mobile.
VoiceLayer MCP server enables AI coding assistants to speak and hear via local, on-device speech-to-text and text-to-speech, with no cloud dependencies.
Voice interface for Claude Code enabling hands-free, conversational interaction entirely on-device for Apple Silicon Macs. It provides push-to-talk transcription and automatic spoken responses via local AI models.
Text-to-speech MCP server using the Kokoro-82M model accelerated with MLX on Apple Silicon, enabling local Claude and Codex clients to speak text aloud and convert text to audio.
Enables programmatic control of Yandex smart home from LLMs: text-to-speech on Alice speakers, batch reminders, device control, and complete scenario management.