media-gen renders Claude-authored content — SVG markup, HTML documents, or a JavaScript canvas draw() function — into PNG/JPEG/WebP images and MP4 video using headless Chromium and ffmpeg, with no API keys required. It also offers an optional bring-your-own-key path for photorealistic text-to-image and text-to-video generation via OpenAI, Gemini, or fal.ai.
MCP server that enables AI assistants to generate images and videos on demand using free-tier providers like Agnes AI, Cloudflare Workers AI, Hugging Face, and Google Gemini, with tools for creating media, polling async jobs, and listing providers.
Enables AI image and video generation using Google Nano Banana and Veo 3.1 via a LiteLLM gateway, providing tools for synchronous image generation and asynchronous video generation with polling, returning public URLs.
Enables image and video generation across GPT-Image, Gemini, Grok, and Jimeng with file-based outputs, multi-reference support, and model capability lookup.
A local MCP server for Claude Desktop that integrates Agnes Image 2.1 Flash and Agnes Video V2.0, enabling natural language generation of images and videos with automatic local saving.
Enables generating images, video, and audio through a single capability-oriented interface, with server-side routing, safety screening, job lifecycle management, concurrency limits, and retention.
MCP server for generating and editing images using OpenAI, and creating videos using OpenAI Sora and Google Veo. Enables fetching media from URLs or disk with smart output placement.
Provides tools for agent-driven video creation, including image generation via Google's GenAI, video generation, and local video stitching with FFmpeg.
Enables generating images from text prompts using various AI models (FLUX, Recraft, GPT Image, etc.) and automatically delivering them through Cloudinary's platform.
Made-to-order data for AI agents via x402 micropayments on Base. Describe a need in plain language, get a custom quote, pay per call. No signup, no API keys. HTTP + MCP transports. 5 tools.
Enables text-only coding models to read images, PDFs, presentations, spreadsheets, and other non-text files through a single analyze_media tool, combining local document extraction, OCR, and optional vision models with clear evidence labeling.