Audio Processing
Services for manipulating, generating, and working with audio content. Includes audio synthesis, processing, playback control, and format conversion capabilities.
MCP ServersBrowse all →
- AlicenseAqualityCmaintenanceVerify sealed voice recordings (WAV) from any AI assistant: verdict (verified, unsealed, warning, alert) and the seconds that changed after sealing. Verify-only, no secrets.2Apache 2.0

ZeroTrue MCP Serverofficial
AlicenseAqualityCmaintenanceEnables detection of AI-generated content in text, images, video, and audio via the ZeroTrue API, supporting multiple analysis tools and MCP-compatible clients.629 npmMIT- AlicenseBqualityBmaintenanceEnables MCP-compatible agents to queue, run, and track local AI music generation jobs with dry-run defaults, license recording, and pluggable pipeline adapters for YuE or compatible backends.31AGPL 3.0

convert-online-mcpofficial
AlicenseAqualityBmaintenanceConvert files between 400+ image, video, audio, document, spreadsheet, ebook, archive and font formats, with OCR. A thin stdio client for the hosted Convert.Online REST API — it converts nothing itself and no file passes through the process.5189 npmMIT- AlicenseAqualityCmaintenanceRemove vocals, extract instrumentals, and split any song into up to six stems — directly from Claude Desktop, Cursor, or any MCP client. Supports local audio files, YouTube URLs, and SoundCloud track1119 npmMIT

VoiceStudio MCPofficial
AlicenseBqualityCmaintenanceEnables access to a backend's native MCP tools and broader HTTP API for speech generation, voice cloning/design, transcription, dubbing, translation, subtitles, audiobooks, and related batch/long-form workflows over Streamable HTTP or local stdio.21AGPL 3.0
dtelecom-sttofficial
AlicenseAqualityFmaintenanceEnables AI assistants to transcribe audio files using dTelecom's real-time speech-to-text with pay-per-use USDC micropayments, no API keys required.310 npm1MIT- AlicenseAqualityAmaintenanceText to speech for MCP clients. Reads numbers, dates and order IDs correctly. 23 languages, six voices, every render watermarked. Free key with 100,000 characters, no card.461 PyPIMIT
- AlicenseBqualityCmaintenanceProvides 189 tools across 12 domains to control DaVinci Resolve from any MCP-compatible client, including playback, project management, media, timelines, color grading, rendering, Fusion, gallery, and Fairlight.61001MIT
- AlicenseAqualityCmaintenanceGaudio Lab Audio AI — Stem Separation, DME Separation, AI Text Sync747 npm1MIT

Magic Hour MCP Serverofficial
AlicenseAqualityCmaintenanceEnables image, video, and audio generation via Magic Hour's API, including async polling, file uploads, and project downloads.458Apache 2.0- AlicenseAqualityAmaintenanceEnables AI assistants and local LLMs to interact with BirdNET-Go bioacoustic observatories, providing tools to query recent and historical bird detections, monitor station health, and retrieve audio clips.18815 npm4MIT
- Apache 2.0
- AlicenseAqualityAmaintenanceEnables AI-powered music generation and live coding by providing direct control over Strudel.cc through browser automation. Supports pattern creation, audio analysis, and pattern storage for TidalCycles/Strudel music patterns.42738 npm242AGPL 3.0

Augentofficial
AlicenseBqualityCmaintenanceMCP server that turns any audio or video source into structured, searchable intelligence for agents, enabling download, transcription, semantic search, speaker identification, and more.224MIT
noisy-codingofficial
AlicenseAqualityAmaintenanceVoice interface for Claude Code: you talk, the agent listens, codes, and talks back while it works. Live speech-to-text with turn-taking, Grok/xAI voices with per-subagent personas, and a real-time HUD dashboard.412MIT
Clipia MCPofficial
AlicenseAqualityAmaintenanceGenerate AI images, video, speech, and music from Claude, ChatGPT, Cursor, and other MCP clients through the Clipia API.15MIT
CraftStory MCP Serverofficial
AlicenseAqualityAmaintenanceEnables MCP clients to create talking-avatar videos of any length from a photo plus audio, and short AI clips with generated sound or lip-sync, through the CraftStory API. It also covers voice and avatar listing, speech generation, cost previews, job status and result retrieval, and upscaling.1145 npmMIT- MIT

mocoVoice MCP Serverofficial
AlicenseAqualityCmaintenanceEnables transcription of audio and video files using mocoVoice API, allowing users to start transcription jobs and retrieve results directly from Claude Desktop.63MIT
postward-creative-mcpofficial
AlicenseBqualityBmaintenanceEnables AI assistants to generate images, videos, music, and voiceovers with your own provider keys, and edit video locally with ffmpeg — all on your machine with no account, watermark, or telemetry.35112 npmMIT- AlicenseBqualityBmaintenanceEnables MCP clients to run hosted audio tools such as stem splitting, analysis, segmentation, mastering, sample-pack generation, audio-to-MIDI conversion, music generation, synth-preset creation, and text-to-vox singing.10MIT

MMAudio MCPofficial
AlicenseBqualityCmaintenanceEnables AI-powered video-to-audio and text-to-audio generation using MMAudio's API. Create synchronized audio from video content or generate audio from text descriptions with configurable parameters.39 npm4MIT
RunAPI MCP Serverofficial
AlicenseBqualityBmaintenanceConnects MCP-compatible coding tools to RunAPI for AI image, video, music, text-to-speech, and LLM generation using 130+ models from leading providers.9659 npm56Apache 2.0
@paxalabs/mcpofficial
AlicenseAqualityBmaintenanceEnables agents to speak Thai and English audio through local speakers, manage a playback queue, save speech to files, translate text into Thai, and OCR PDFs and images.10122 npm10MIT
ElevenLabs MCP Serverofficial
AlicenseBqualityFmaintenanceAn official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.271,534MIT
famistudio-mcpofficial
AlicenseAqualityBmaintenanceEnables creating, inspecting, validating, and rendering NES FamiStudio .fms projects from plain JSON without a GUI.13350 npmMIT- AlicenseBqualityBmaintenanceMCP server for headless, agent-native control of openDAW, enabling programmatic music production with tracks, instruments, effects, MIDI, automation, and rendering via 543 tools.8100Apache 2.0
- AlicenseAqualityAmaintenanceRun AI workflows hosted on Glif.app via MCP, including ComfyUI-based image generators, meme generators, selfies, chained LLM calls, and more6383 npm212MIT

@speechweave/mcpofficial
AlicenseAqualityBmaintenanceMCP server for SpeechWeave transcription, enabling AI assistants to transcribe local files and URLs via wait-first or async tools.665 npmMIT
MCP ConnectorsBrowse all →
ScanSing reads printed sheet music. Give Claude a PDF or a photo of a score and it gets back MusicXML, the format MuseScore, Sibelius, Finale and Dorico open. It is the recogniser inside the ScanSing app for choir singers, offered as a connector. Once a score is recognised, Claude can answer questions about it without recognising it again: - what the key, time signature, tempo and length are, and each part's range, clefs and lyrics; - one voice on its own, for example just the alto line of a choir piece; - the score or a part moved to another key, with notes respelled for the new key signature; - a MIDI file, one track per part, to listen to or rehearse with; - a short-lived download link to open the result in a browser or notation app. It handles PDFs up to 40 pages and PNG, JPEG, HEIC, TIFF and WebP images, can recognise selected pages of a long PDF, and can split a songbook into one MusicXML per piece. A file can be given as a web link or a small image; in Claude Code, a file on your computer is sent through a one-time upload link. Recognition spends pages from your ScanSing plan, so it is the only tool that is not read-only. Everything else works on the stored result and is free. The first connection with a new email comes with 20 free pages; paid plans start at $9 a month for 300 pages. Scores and results are deleted about an hour after recognition and are never used for training.
Generate game assets with AI: sprites, animations, 3D models with rigging, video, music, SFX, voices
Cut audio from YouTube, TikTok or Instagram links to MP3/WAV, and fetch video transcripts.
Cut audio from YouTube, TikTok or Instagram links to MP3/WAV, and fetch video transcripts.
Create images, short videos, narration, music and multi-shot video projects with your Everygen account. Connect securely with OAuth and approve generation quotes before using credits.
AI video, images, music & SFX: Seedance 2.5, Veo 3.1, Kling 3.0, Nano Banana Pro, 20+ models.
Royalty-free music for apps, games, video and podcasts: generate a track from a prompt, a genre or an image, swap individual instruments, run an endless stream, or search a ready-made catalogue.
Podcast post-production for an AI agent: transcribe a recording, cut fillers and pauses, clean up the voice, add licensed music that ducks under speech, set chapters and export a finished episode.
Create and edit images, videos, and audio through Magic Hour's hosted Streamable HTTP MCP server.
Write original songs with SoundBreak AI artists. Built with artists, not on them. Co-write in this chat, play a preview (not a release), and save to your SoundBreak account. Lyrics you wrote stay yours. After you save, you can submit to artists, SoundBreak Radio, or distribute (Spotify, Apple Music, and more) from SoundBreak. Terms: https://app.soundbreak.ai/terms Privacy: https://app.soundbreak.ai/privacy-policy
Audio for your agent: transcribe, speak, translate, summarise, plus sound effects and music.
KEEP Highlights is an intelligent multimodal engine that detects key moments and surprises in video and audio recordings. Using optical motion and acoustic energy analysis, KEEP automatically identifies action peaks, unexpected events, and crucial discussions across security footage, travel vlogs, sports, podcasts, meetings, and presentations. Condense hours of recordings into actionable summaries and highlight reels.
One tool surface for music, image, video, and audio generation across Suno, Grok Imagine, Seedance, Kling, Hailuo, Wan, VEO, Ideogram, and GPT Image 2. Generate, edit, upscale, reframe, and master through one credit pool. Connect in one click with OAuth, no API key required.
CC0 sound effects API for AI agents — search, preview, and download via MCP.
Media intelligence analysis for audio, video, and images via the Echosaw MCP server.
Generate AI images, video, voiceovers and music from Claude, ChatGPT or Cursor through 50+ models (Veo 3.1, Kling 3, Seedance, Nano Banana, GPT Image, ElevenLabs). Also image editing, upscaling, background removal, face swap, transcription, voice cloning and UGC-style video ads. Sign in with OAuth — no API key to paste. Tools are annotated (read-only vs. credit-spending); failed generations are refunded.
Magic Hour MCP lets AI agents create and edit images, videos, and audio using Magic Hour’s hosted generation tools.
Deepfake detection, media intelligence, and invisible watermarking for audio, image, and video via the Resemble AI API, plus docs tools. Remote MCP server (Streamable HTTP) — also published in the official MCP registry as io.github.resemble-ai/resemble-mcp.
Edit uploaded footage with AI tools: cuts, captions, reframing and previews. Final export in Studio.
Change lyrics in an existing song. Agents upload authorized MP3 audio, provide a 0.1–6 second phrase and replacement words, preview generated singing, and export MP3/WAV. No website signup required: buy credits via x402 with USDC, then use a scoped key. Caller supplies timing. Resumable CLI and Codex/Claude Code setup: https://lyricpatch.com/agents/integrations. Hear a real example: https://lyricpatch.com/agents/demo.