"Improving the Intelligence of Large Language Models" matching MCP connectors:
Matching Connector Tools:
Media intelligence analysis for audio, video, and images via the Echosaw MCP server.
Deterministic music theory for agents: analyze, voice, reharmonize, conduct — computed, not guessed
Generate game assets with AI: sprites, 3D models, animations, sound effects, music, and voices.
Deepfake detection, media intelligence, and invisible watermarking for audio, image, and video via the Resemble AI API, plus docs tools. Remote MCP server (Streamable HTTP) — also published in the official MCP registry as io.github.resemble-ai/resemble-mcp.
Generate and edit images, videos, and audio with 150+ models from 20+ vendors.
Generate AI music via the Lacuna Music API from MCP clients like Claude Desktop & Code.
Music superpowers for your AI agent. Curated licensed music and AI tools that do the rest. Right where you already work.
Image, video, music and text generation across 100+ models through one endpoint.
LibriVox public-domain audiobooks (~17000 titles in dozens of languages)
Search millions of sound and soundboards on 101soundboards.com
Connect the OneStepTranscribe MCP server to your AI assistant and turn audio or video into text without leaving the chat. It is a remote server, so there is nothing to install, no API key, and no account. Just add one URL and ask your assistant to transcribe a file.
Decode an SSTV audio recording and anchor its fingerprint and image hash to the Knox event chain.
Privacy-first audio intelligence: BPM, key, waveform. Audio never stored. Pay per second.
AudioAlpha turns 100+ daily finance and crypto podcasts into structured intelligence — α-sentiment scores, narrative signals, asset mentions, transcripts, and market snapshots with 40+ custom metrics. Built for AI-driven research and trading workflows.
Financial podcast intelligence platform — sentiment, narrative, and asset signals from 100+ podcasts
Hosted MCP server that gives AI agents (Claude, Cursor, Codex, etc.) access to the full Runware API — image generation, video generation, audio generation, 3D, upscaling, background removal, captioning, and more.
Image, video, audio, face-swap, talking avatars and chat across 300+ AI models, one balance.
Search speech in podcasts, government meetings, and your own audio: speakers, entities, timestamps.
Generate images, video, music, voice and 3D through one API. 30 tools, 200+ models.
One key for AI video, voice, music, and image across 40+ frontier models. BYOK zero markup.