"Best browser-based media player controllers (MPC)" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Media intelligence analysis for audio, video, and images via the Echosaw MCP server.
Deepfake detection, media intelligence, and invisible watermarking for audio, image, and video via the Resemble AI API, plus docs tools. Remote MCP server (Streamable HTTP) — also published in the official MCP registry as io.github.resemble-ai/resemble-mcp.
Generate marketing images, videos and audios for campaigns, product content, and brand assets.
Process video, audio, images, and documents with 86+ cloud media processing robots.
Arabic-first AI creative platform for Egyptian and Arab businesses. Generate social media designs, write marketing copy in Egyptian dialect, build content calendars, produce Sora-2 videos, AI photoshoots, music tracks, and business documents — with your brand identity automatically applied. Requires a Grow or Business subscription at vizzy.space.
25+ AI media generation tools — FLUX Pro, Ideogram v3, Recraft v3, Stable Diffusion XL, MiniMax video, and Kokoro TTS. Images, video, and audio from one server. $0.01/call.
15 media & data tools for AI agents: search, transcribe, subtitles, voiceover, translate & more.
125+ browser tools for PDF, Image, Video, Audio, AI, Scanner. Files never leave your device.
MCP to generate media assets like videos, images, music, sound effects, captions and so on
Generate and edit images, video, voice, lip-sync and 3D models from your AI agent.
Generate image, video, audio, 3D and vector media with 100+ AI models. Pay per generation.
Encode uploads to web-ready media: video to HLS, audio to AAC, images to WebP. Async jobs.
Speech, transcription, voice agents, Trace, Recap, dubbing and narration with browser OAuth.
The Listenetic MCP server is a remote, cloud-hosted server that enables AI assistants like ChatGPT and Claude to convert articles, documents, websites, and videos into high-quality AI-generated audio. It provides multi-format support for text and binary files, natural-sounding text-to-audio conversion using AI, and specialized processing for SSML, markup, markdown, and various media formats through three core tools: listentic_supported_mimetypes, listentic_add_content_text, and listentic_add_content_binary.
Generate images, videos, voiceovers, and captions from a chat prompt.
Generate images, video, and audio with Glif's media-generation agent
Transcript-based audio editing: transcribe audio, edit by word ID, export edited audio.