Enables spoken conversations with Claude on a Mac: Claude speaks through speakers, listens to the user's natural replies, and transcribes them locally. No audio leaves the computer, and it includes tools for voice setup and a hands-free voice mode.
Image, video, chat and text-to-speech models (GPT Image, Gemini Image, Veo, Kling, Seedance, Claude, GPT, Gemini) behind one API key. generate_image returns a preview the model can see, and review_image has a vision model critique the result and propose a corrected prompt, so an assistant can generate, check and fix images on its own.
Enables AI agents to perform local audio tasks such as speech synthesis, voice cloning, music and sound effect generation, and audio editing through MCP, with GPU models loaded on demand and released after idle.
Enables spoken consulting case interview practice through Claude, providing real-time cases, grading on structure, analytics, judgment, and communication, and targeted practice drills.
Text-to-speech MCP server that enables AI assistants to read text aloud on the user's computer using Windows SAPI, with no API key or cloud service required.
Provides access to Voxtral TTS voice and TTS workflows, FAQ, and official links via MCP. Enables AI clients to retrieve supported voices, languages, and TTS configuration without API keys.
Enables AI apps to transcribe local audio and video files on the user's own GPU, producing SRT/VTT subtitles and text without cloud uploads. It also helps find media files, list available devices, and list supported languages and models.
Enables Claude Desktop and Claude Code to synthesize and play speech using VOICEVOX text-to-speech engine. Supports multiple voice characters, session-based voice assignment, and queue management for audio playback.
Enables text-to-speech functionality on macOS using the say command, offering extensive control over speech parameters like voice, rate, volume, and pitch for a customizable auditory experience.
A PowerShell-based MCP server that enables Claude Desktop to convert text to speech using Windows' built-in Speech API, offering features like playback control, speed and volume adjustment.
Enables MCP agents to retrieve Japanese voice presets and synthesize anime-style speech, including reference-conditioned voice cloning, using the CPU of the computer running the local Gradio app.
Enables AI assistants to generate expressive speech using Hume AI's Octave Text-to-Speech, allowing for natural language-based audio synthesis and voice management.
Enables agents to convert text to speech using OpenAI's TTS models with voice selection, delivery instructions, and queue-based audio playback. Supports both blocking and non-blocking modes for flexible audio generation and playback control.
Enables MCP clients to speak text aloud on macOS using the built-in say command, with no external speech API or keys, and to list available voices and control rate or voice selection. Requests from multiple client threads share a single FIFO playback worker, so speech plays one item at a time and can be inspected, stopped, or muted in hold or discard mode.
Exposes a single transcribe tool over streamable HTTP so containerised agents can send a media file name and receive text transcribed locally by MacWhisper on the host Mac's GPU, with token-gated access and no uploads or API keys. Callers place media in a configured directory, and the blocking call returns the finished transcript.