Enables AI agents to remotely operate visual design tools via MCP protocol, with composable architecture for image processing, file operations, and workflow automation.
Provides image recognition capabilities to Codex/Claude Code by routing 'look at screen/screenshot' requests to a multimodal model via an API relay. Enables UI automation agents to analyze local screenshots and return descriptions or structured JSON coordinates.
Enables AI clients like Codex, Claude Code, and Kimi Code to analyze public video URLs by downloading media, uploading it to Gemini, and returning timestamped production breakdowns covering shots, visual design, animation, motion, narration, music, sound effects, and editing, with follow-up Q&A and session management tools.
Enables intelligent document processing by extracting text, classifying document types, and generating structured summaries from PDFs and images using vision LLMs.
A streamlined MCP server for XMP metadata embedding with beautiful formatting and smart filename indicators, enabling metadata embedding, reading, validation, and report generation for lifestyle, product, and orbit schemas.
This server enables interaction with Google's Video Intelligence API for advanced video analysis, auto-generated using AG2's MCP builder to provide a standardized multi-agent interface.
A multi-agent human-computer interaction system that enables natural interaction through integrated visual recognition, speech recognition, and speech synthesis capabilities.
MCP server for MarkItUp's AI image-annotation pipeline. Generate polished marketing-visual variations of any screenshot, regenerate, AI outpaint, and remove backgrounds —
powered by Claude analysis + Gemini rendering.
Enables text-only agents to process images by accepting image files, base64 data, or URLs, sending them to multimodal models, and returning structured text results via MCP.
Enables image generation using Google's Imagen and other AI models through Nexos.ai platform. Supports single and batch image generation with various quality settings and model options.
MCP server that gives AI agents visual intelligence — search Pinterest, analyze images with LLM vision, build a semantic reference library, and retrieve by style or mood.
Generates blog and social media images using Google's Gemini AI with pre-configured platform presets for Ghost, Medium, Instagram, Twitter, LinkedIn, YouTube, and more.
MCP server for ScanToBill, an AI-powered OCR platform that extracts structured JSON from invoices and business documents in Arabic and English. It enables Claude Desktop to extract invoice data from URLs, list past extractions, and retrieve detailed results.
Enables MCP-compatible AI agents to generate background images and videos through the museav platform, and to perform local image post-processing such as background removal, upscaling, watermark removal, and compression, returning file paths instead of base64.
Exposes information about Z-Image AI image generation and editing platform, including styles, pricing, and official links, for MCP-compatible AI clients.