Enables AI-powered video-to-audio and text-to-audio generation using MMAudio's API. Create synchronized audio from video content or generate audio from text descriptions with configurable parameters.
One MCP server for music, image, video, and audio generation across Suno, Grok Imagine, Seedance, Kling, Hailuo, Wan, VEO, Ideogram, and GPT Image 2. Generate, edit, upscale, reframe, and master through one API key and one credit pool.
MCP server for technical SEO audits, powered by the detail.web engine. Run a site audit straight from your AI agent to get a health score, issues across 8 categories, and a GEO sub-score.
Thin MCP server that enables AI agents to generate video, image, and audio content via the Deepy API using a personal API key, with safe generation flow and no business logic.
Transcribes audio files by referencing them in chat, using OpenAI's speech-to-text models locally without uploading audio, and supports speaker diarization.
Remove vocals, extract instrumentals, and split any song into up to six stems — directly from Claude Desktop, Cursor, or any MCP client. Supports local audio files, YouTube URLs, and SoundCloud track
Enables AI agents to turn clinical transcripts or audio recordings into structured SOAP notes, ICD-10/CPT billing codes, patient summaries, and visit summaries via SOAPNoteAPI.
An official Model Context Protocol (MCP) server that enables AI clients to interact with ElevenLabs' Text to Speech and audio processing APIs, allowing for speech generation, voice cloning, audio transcription, and other audio-related tasks.
Enables detection of AI-generated content in text, images, video, and audio via the ZeroTrue API, supporting multiple analysis tools and MCP-compatible clients.
It turns natural-language requests into Microlink calls: take screenshots, generate PDFs, extract metadata or readable text, detect video/audio sources, run Lighthouse audits, scrape custom fields with CSS selectors and more
Enables transcription of audio and video files using mocoVoice API, allowing users to start transcription jobs and retrieve results directly from Claude Desktop.
MCP server that turns any audio or video source into structured, searchable intelligence for agents, enabling download, transcription, semantic search, speaker identification, and more.
Provides 31 tools and 3 resources to access Venice's uncensored AI models for chat, image, video, audio, music, and more, enabling any MCP host to interact with the Venice API.
The Tencent RTC MCP Server provides real-time communication capabilities such as audio and video chat. By integrating this MCP Server into your application or project, you can easily implement secure, high-quality real-time voice and video communication.