"Understanding local computer vision models like YOLO" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Generate AI images, video, voiceovers and music from Claude, ChatGPT or Cursor through 50+ models (Veo 3.1, Kling 3, Seedance, Nano Banana, GPT Image, ElevenLabs). Also image editing, upscaling, background removal, face swap, transcription, voice cloning and UGC-style video ads. Sign in with OAuth — no API key to paste. Tools are annotated (read-only vs. credit-spending); failed generations are refunded.
Generate game assets with AI: sprites, 3D models, animations, sound effects, music, and voices.
One API for 100+ AI video, image, music and speech models.
Generate AI music via the Lacuna Music API from MCP clients like Claude Desktop & Code.
Image, video, music and text generation across 100+ models through one endpoint.
Generate and edit images, videos, and audio with 150+ models from 20+ vendors.
Generate video, images, audio and speech with Vidofy — Veo 3.1, Kling 3.0, Flux 2 and 570+ models.
MCP to generate media assets like videos, images, music, sound effects, captions and so on
Generate and edit images, video, voice, lip-sync and 3D models from your AI agent.
Generate image, video, audio, 3D and vector media with 100+ AI models. Pay per generation.
Turn any LLM multimodal; generate images, voices, videos, 3D models, music, and more.
The Listenetic MCP server is a remote, cloud-hosted server that enables AI assistants like ChatGPT and Claude to convert articles, documents, websites, and videos into high-quality AI-generated audio. It provides multi-format support for text and binary files, natural-sounding text-to-audio conversion using AI, and specialized processing for SSML, markup, markdown, and various media formats through three core tools: listentic_supported_mimetypes, listentic_add_content_text, and listentic_add_content_binary.
Give your AI a real phone: place calls, send SMS, fetch recordings and transcripts. Local or hosted.
Generate reproducible image, video, and audio assets with leading models and your own provider keys.
Generate game-ready 3D models, textures, and audio from natural language, over MCP.
Generate images, video, music, voice and 3D through one API. 30 tools, 200+ models.
One key for AI video, voice, music, and image across 40+ frontier models. BYOK zero markup.
AI content generation with 50+ models: image, video, TTS, voice cloning, and more.