Enables downloading videos from platforms like YouTube and converting them to text using OpenAI Whisper and ffmpeg. It supports multiple output formats including TXT, JSON, SRT, and VTT for transcriptions.
Enables users to search movies and actor profiles, discover soundtracks, and generate curated Spotify playlists, with full playback control and playlist management through natural language.
Enables converting text to speech audio in 20+ languages, returning base64-encoded MP3 output via Google TTS with x402 micropayment-based pay-per-call access.
This MCP server enables AI tools to read Dedao Brain (formerly Get Notes) web notes from a free account, exposing AI summaries, original transcripts, and audio attachment links or downloads, with an accompanying CLI.
Enables Claude AI to control Ableton Live and Max for Live, allowing music production tasks like track management, MIDI editing, and pattern generation directly from conversation.
Enables AI agents to perform local video, audio, and file operations inside an isolated workspace, including cutting/concat videos, extracting audio, transcribing, and managing files, with typed responses and background job support.
Text-to-speech MCP server that enables AI assistants to read text aloud on the user's computer using Windows SAPI, with no API key or cloud service required.
MCP server that bridges Ableton Live with AI models, enabling real-time project inspection and control such as track overview, device parameters, and audio analysis.
Enables conversion of YouTube videos to MP3 format through the Youtube To Mp315 API. Supports checking conversion status, retrieving video titles, and asynchronous video-to-audio conversion with customizable quality and time range settings.
Enables users to convert text into high-quality audio by accessing the OpenAI Text-to-Speech API. It supports customizable model selection and voice options for synthesized speech generation via the MCP protocol.
An MCP server that enables voice-to-voice AI conversations using ElevenLabs for speech synthesis and recognition, with tools for voice management, text-to-speech, and speech-to-text.
Discover audio and LLM services, compare seven voices, and check prices for free. Generate speech, transcribe audio and request LLM text with optional capped USDC payments on Base; voice-cloning requirements are also available.
MCP server that exposes recorded family stories as read-only tools, enabling Alexa+ to retrieve and deliver voice stories to children through natural language.
A server that allows Claude to control audio playback on your computer, supporting MP3, WAV, and OGG files with features like play, list, and stop commands.
Media processing server for video trimming, MP3 audio extraction, watermarking, and 9:16 vertical re-framing via Apify Actor (budding_retrograde/clipforge).
MCP server for parsing Douyin links to fetch video info/download links, transcribing audio locally, and organizing transcripts into polished spoken copy.