"Understanding local computer vision models like YOLO" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Generate AI images, video, voiceovers and music from Claude, ChatGPT or Cursor through 50+ models (Veo 3.1, Kling 3, Seedance, Nano Banana, GPT Image, ElevenLabs). Also image editing, upscaling, background removal, face swap, transcription, voice cloning and UGC-style video ads. Sign in with OAuth — no API key to paste. Tools are annotated (read-only vs. credit-spending); failed generations are refunded.
320 AI models + 2,720 pay-per-call APIs. x402 USDC on Base or Solana, no API key.
84+ free local-first tools: image, PDF, docs, dev utils. Wasm, zero upload, x402 API.
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
- runapi-mcpOAuth unavailableio.github.runapi-builder
130+ AI models for image, video, music, and audio — 18 model families, one RunAPI account.
MCP to generate media assets like videos, images, music, sound effects, captions and so on
Turn Claude into a professional creative studio: create DNA-locked AI characters that stay identical across every image and video, then generate, edit, upscale, adapt to any size, and turn stills into video — 55 tools across 7 image and 14 video models, plus voiceover, music and talking avatars.
Generate and edit images, video, voice, lip-sync and 3D models from your AI agent.
Turn any LLM multimodal; generate images, voices, videos, 3D models, music, and more.
Image, video, audio, face-swap, talking avatars and chat across 300+ AI models, one balance.
Run 100+ AI models — image, video, audio, 3D — through one API with pay-per-use billing.
The Listenetic MCP server is a remote, cloud-hosted server that enables AI assistants like ChatGPT and Claude to convert articles, documents, websites, and videos into high-quality AI-generated audio. It provides multi-format support for text and binary files, natural-sounding text-to-audio conversion using AI, and specialized processing for SSML, markup, markdown, and various media formats through three core tools: listentic_supported_mimetypes, listentic_add_content_text, and listentic_add_content_binary.
Multimodal video analysis MCP — transcription, vision, and OCR for any video URL.
AI content generation with 50+ models: image, video, TTS, voice cloning, and more.
Analyze images and videos with Gemini to get fast, reliable visual insights. Handle content from U…
Analyze images from multiple angles to extract detailed insights or quick summaries. Describe visu…