Enables spoken conversations with Claude on a Mac: Claude speaks through speakers, listens to the user's natural replies, and transcribes them locally. No audio leaves the computer, and it includes tools for voice setup and a hands-free voice mode.
Image, video, chat and text-to-speech models (GPT Image, Gemini Image, Veo, Kling, Seedance, Claude, GPT, Gemini) behind one API key. generate_image returns a preview the model can see, and review_image has a vision model critique the result and propose a corrected prompt, so an assistant can generate, check and fix images on its own.
Enables AI agents to perform local audio tasks such as speech synthesis, voice cloning, music and sound effect generation, and audio editing through MCP, with GPU models loaded on demand and released after idle.
Enables spoken consulting case interview practice through Claude, providing real-time cases, grading on structure, analytics, judgment, and communication, and targeted practice drills.
Text-to-speech MCP server that enables AI assistants to read text aloud on the user's computer using Windows SAPI, with no API key or cloud service required.
A cross-platform MCP server that enables Claude to speak using Microsoft Edge TTS with support for over 300 voices across 50+ languages. It requires no API keys and allows for customization of speech rate, volume, and pitch.
Enables text-to-speech functionality on macOS using the say command, offering extensive control over speech parameters like voice, rate, volume, and pitch for a customizable auditory experience.
A Model Context Protocol server that integrates with VOICEVOX engine to provide text-to-speech synthesis and speaker information retrieval, allowing users to generate and play voice audio from text.
Generates images, videos, and speech via Google Nano Banana, Veo, Omni, and Gemini TTS models with pay-as-you-go crypto payments, no subscription required.
Enables Claude to access Google AI Studio's full Gemini API surface, including text/vision, Nano Banana image generation and editing, Veo 3.1 video with synced audio, and Gemini TTS voiceover.
Exposes a single transcribe tool over streamable HTTP so containerised agents can send a media file name and receive text transcribed locally by MacWhisper on the host Mac's GPU, with token-gated access and no uploads or API keys. Callers place media in a configured directory, and the blocking call returns the finished transcript.
Local MCP server for neural text-to-speech using Kokoro ONNX engine on CPU, supporting SSML tags, multiple voice profiles, and zero-GPU operation for low-latency speech synthesis.
Local voice toolkit over MCP: transcribe audio to text in 25 languages, synthesize speech in 9, and list available voices and languages. Runs fully on-device — no API keys, no cloud.
Enables AI assistants to speak with realistic cloned voices via ElevenLabs TTS on Cloudflare Workers, supporting 29 languages and inline audio playback on both desktop and mobile.
VoiceLayer MCP server enables AI coding assistants to speak and hear via local, on-device speech-to-text and text-to-speech, with no cloud dependencies.
Provides comprehensive access to OpenAI's API capabilities including chat completions, image generation, embeddings, text-to-speech, speech-to-text, vision analysis, and content moderation. Enables users to interact with GPT models, DALL-E, Whisper, and other OpenAI services through natural language commands.
Voice interface for Claude Code enabling hands-free, conversational interaction entirely on-device for Apple Silicon Macs. It provides push-to-talk transcription and automatic spoken responses via local AI models.