Enables downloading videos from platforms like YouTube and converting them to text using OpenAI Whisper and ffmpeg. It supports multiple output formats including TXT, JSON, SRT, and VTT for transcriptions.
Enables LLM agents to process local videos into timestamped, citable text documents and then query them through tools for listing videos, retrieving transcripts, and fetching specific segments, all fully offline.
Enables converting text to speech audio in 20+ languages, returning base64-encoded MP3 output via Google TTS with x402 micropayment-based pay-per-call access.
Text-to-speech MCP server that enables AI assistants to read text aloud on the user's computer using Windows SAPI, with no API key or cloud service required.
Provides translation and language detection tools to AI agents, processing text, audio, and Google Meet recordings with emotional voice style preservation via Google's Gemini Live API.
An MCP server that enables voice-to-voice AI conversations using ElevenLabs for speech synthesis and recognition, with tools for voice management, text-to-speech, and speech-to-text.
Enables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.
Enables agents to convert text to speech using OpenAI's TTS models with voice selection, delivery instructions, and queue-based audio playback. Supports both blocking and non-blocking modes for flexible audio generation and playback control.
Enables MCP-capable chat clients to synthesize speech from text, convert it with a local RVC voice model, and render the result in an inline audio player.
MCP server that provides a transcribe_audio tool to convert voice messages from channels into text using OpenAI Whisper, enabling Claude Code to process audio attachments.
MCP server that lets Claude place outbound telephone calls via ElevenLabs Agents over a SIP trunk and report the outcome including status, success evaluation, summary, and transcript.
An MCP server that wraps the WellSaid Labs text-to-speech API, enabling Claude to generate lifelike voiceovers with voice selection, prosody control, and caption support directly from text.