Audio Processing
Services for manipulating, generating, and working with audio content. Includes audio synthesis, processing, playback control, and format conversion capabilities.
MCP ServersBrowse all →
AlicenseAqualityAmaintenanceVoice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.646MIT- AlicenseAqualityAmaintenanceRun AI workflows hosted on Glif.app via MCP, including ComfyUI-based image generators, meme generators, selfies, chained LLM calls, and more6383201MIT

RunAPI MCP Serverofficial
AlicenseBqualityBmaintenanceConnects MCP-compatible coding tools to RunAPI for AI image, video, music, text-to-speech, and LLM generation using 130+ models from leading providers.959755Apache 2.0- AlicenseAqualityAmaintenanceText to speech for MCP clients. Reads numbers, dates and order IDs correctly. 23 languages, six voices, every render watermarked. Free key with 100,000 characters, no card.4MIT

ElevenLabs MCP Serverofficial
AlicenseAqualityFmaintenanceAn official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.271,534MIT
MMAudio MCPofficial
AlicenseBqualityCmaintenanceEnables AI-powered video-to-audio and text-to-audio generation using MMAudio's API. Create synchronized audio from video content or generate audio from text descriptions with configurable parameters.3103MIT
ZeroTrue MCP Serverofficial
AlicenseAqualityAmaintenanceEnables detection of AI-generated content in text, images, video, and audio via the ZeroTrue API, supporting multiple analysis tools and MCP-compatible clients.616MIT- AlicenseAqualityCmaintenanceGaudio Lab Audio AI — Stem Separation, DME Separation, AI Text Sync7631MIT
- AlicenseAqualityDmaintenanceOfficial AllVoiceLab Model Context Protocol (MCP) server, supporting interaction with powerful text-to-speech and video translation APIs.1259MIT
- AlicenseAqualityAmaintenanceGive your AI agent a voice with x402 pay-per-call speech synthesis, offering 20 voices, 10 personas, 31 languages, and granular controls.46199MIT

@speechweave/mcpofficial
AlicenseAqualityBmaintenanceMCP server for SpeechWeave transcription, enabling AI assistants to transcribe local files and URLs via wait-first or async tools.6122MIT- AlicenseAqualityAmaintenanceOfficial MCP server for Rendobar. Lets AI agents run serverless media processing and upload local files.871841MIT

Sonilo MCPofficial
AlicenseAqualityBmaintenanceAn MCP (Model Context Protocol) server that exposes Sonilo's licensed music and sound-effects API to MCP-compatible clients (Claude Code, Claude Desktop, Codex).954MIT
Augentofficial
AlicenseBqualityCmaintenanceMCP server that turns any audio or video source into structured, searchable intelligence for agents, enabling download, transcription, semantic search, speaker identification, and more.224MIT- AlicenseAqualityBmaintenanceRemove vocals, extract instrumentals, and split any song into up to six stems — directly from Claude Desktop, Cursor, or any MCP client. Supports local audio files, YouTube URLs, and SoundCloud track1124MIT

GlianaAI MCP Serverofficial
AlicenseAqualityBmaintenanceEnables pay-per-call access to 90+ generative AI models and utility tools via any MCP client, with no signup or API key, using wallet-based USDC payments on Base, Tempo, or Solana.7288MIT- MIT

Clipia MCPofficial
AlicenseAqualityAmaintenanceGenerate AI images, video, speech, and music from Claude, ChatGPT, Cursor, and other MCP clients through the Clipia API.15MIT
mocoVoice MCP Serverofficial
AlicenseAqualityBmaintenanceEnables transcription of audio and video files using mocoVoice API, allowing users to start transcription jobs and retrieve results directly from Claude Desktop.63MIT- AlicenseAqualityCmaintenanceEnables searching for songs and retrieving direct MP3 play URLs from gequbao.net. Supports both simple keyword search and enriched result lookup.21MIT
- AlicenseAqualityCmaintenanceEnables video and audio processing through FFmpeg, supporting format conversion, compression, trimming, audio extraction, frame extraction, video merging, and subtitle burning through natural language commands.826MIT
- AlicenseBqualityBmaintenanceConnects Ableton Live to Claude AI through the Model Context Protocol, enabling AI-assisted music production by allowing Claude to directly interact with and control Ableton Live sessions.162,942MIT
- AlicenseAqualityAmaintenanceAn interactive digital audio workstation as an MCP server, enabling music production with a channel rack, piano roll, mixer, effects, automation, microphone recording, and offline WAV rendering.26493MIT
- AlicenseAqualityAmaintenanceEnables AI agents to inspect and control Q-SYS audio/video systems via the QRC protocol over TCP, against a real Core or Q-SYS Designer emulator.181MIT
- AlicenseAqualityBmaintenanceEnables reading and controlling a Universal Audio Apollo interface: channel names, faders, preamp gain, cue sends, monitor control, and more via the UA Mixer Engine's local TCP API.21MIT
- AlicenseAqualityNot gradedmaintenanceEnables interaction with ElevenLabs Text-to-Speech and audio processing APIs. Supports speech generation, voice cloning, audio transcription, and sound effect creation through natural language.24
- AlicenseAqualityBmaintenanceEnables natural language control of Wwise audio middleware, allowing project auditing, batch editing, event creation, and reference checking through conversational interaction.201MIT
- AlicenseAqualityCmaintenanceAI-powered speech tools by Brainiall: pronunciation assessment with phoneme-level feedback, speech-to-text with language detection, and text-to-speech with multiple voices.41MIT
- AlicenseBqualityBmaintenanceAgent-native media processing: video encoding, image manipulation, document conversion, audio transcription, and more via 86+ cloud Robots.773MIT
- AlicenseBqualityDmaintenanceAn MCP server that integrates with fal.ai to provide AI agents with tools for image generation, text processing, audio synthesis, and model management via a unified interface.813MIT
MCP ConnectorsBrowse all →
Generate marketing images, videos and audios for campaigns, product content, and brand assets.
Deterministic music theory for agents: analyze, voice, reharmonize, conduct — computed, not guessed
Generate highly realistic Text to Speech voiceovers.
Generate game assets with AI: sprites, 3D models, animations, sound effects, music, and voices.
Generate Suno AI music (v5.5) from any MCP client. Async; billed only on success.
One tool surface for music, image, video, and audio generation across Suno, Grok Imagine, Seedance, Kling, Hailuo, Wan, VEO, Ideogram, and GPT Image 2. Generate, edit, upscale, reframe, and master through one credit pool. Connect in one click with OAuth, no API key required.
Deepfake detection, media intelligence, and invisible watermarking for audio, image, and video via the Resemble AI API, plus docs tools. Remote MCP server (Streamable HTTP) — also published in the official MCP registry as io.github.resemble-ai/resemble-mcp.
Find & cut horizontal and vertical video clips (Shorts/Reels), transcribe & summarize. Pay per job.
File conversion: PDF, DOCX, STT, TTS, watermarking
Process video, audio, images, and documents with 86+ cloud media processing robots.
Create, co-edit, analyze, publish, and export collaborative step-sequencer sessions through MCP.
One API for 100+ AI video, image, music and speech models.
264 emisoras de radio colombianas en vivo: busca por ciudad, genero o dial y escuchalas.
Youtube Download API: Download audio and video from YouTube. On the JoJ API marketplace.
MCP server for RiverScript, an AI transcription platform - fetches transcripts shared via a link.
Train portable RVC v2 voice models from audio in the cloud and download the .pth, .index, and ZIP.
Verbatim transcription of public video/audio URLs to clean text, SRT, and timestamped records.
AI-manageable audio CDN: upload, transcode, normalize, stream & deliver audio, plus grounded docs.
Remote MCP for AI video, image, music and speech generation.
Video, audio, and image processing for AI agents: convert, transcribe, upscale - 150+ operations.