Provides image recognition capabilities to Codex/Claude Code by routing 'look at screen/screenshot' requests to a multimodal model via an API relay. Enables UI automation agents to analyze local screenshots and return descriptions or structured JSON coordinates.
Enables AI clients like Codex, Claude Code, and Kimi Code to analyze public video URLs by downloading media, uploading it to Gemini, and returning timestamped production breakdowns covering shots, visual design, animation, motion, narration, music, sound effects, and editing, with follow-up Q&A and session management tools.
Enables intelligent document processing by extracting text, classifying document types, and generating structured summaries from PDFs and images using vision LLMs.
This server enables interaction with Google's Video Intelligence API for advanced video analysis, auto-generated using AG2's MCP builder to provide a standardized multi-agent interface.
Enables accounts payable teams to extract invoice data from PDFs and images, detect duplicates, normalize vendor names, calculate payment terms, and validate invoice completeness. Supports local extraction for text PDFs and optional vision providers for scanned documents.
Generate and edit images with GPT Image 2.5 for $0.029 each, inside Claude, Cursor or any MCP client. A free you.bot API key comes with 50 credits (about 17 images) and needs no card.
Enables AI agents to drive 360° cameras by connecting, checking status, reading and changing settings, taking photos, recording video, browsing the gallery, and downloading media as cancellable background jobs, with local stitching and export of 360° footage. It is honest about backend limits, returning structured, explained errors for anything a given camera or connection cannot do.
Enables AI image generation, editing, and composition using Google's Gemini image models (Nano Banana Pro and Nano Banana). Supports text-to-image generation, multi-image composition, flexible aspect ratios, high-resolution output up to 4K, and real-time information grounding.
MCP server that gives AI agents visual intelligence — search Pinterest, analyze images with LLM vision, build a semantic reference library, and retrieve by style or mood.
MCP server that captures short screen recordings as numbered, timestamped still frames and composes them into a single contact sheet returned as image content, letting AI agents reason about motion, animations, and UI transitions instead of inspecting video frame by frame. Also supports single screenshots and listing targetable windows.
MCP server that analyzes images with Google's Gemini vision models, allowing agents to describe or ask questions about images without bloating context.
Enables controlling a local Blender instance through natural language or code, allowing arbitrary Python/bpy execution and querying scene or object information via MCP.
Enables AI assistants to recognize and extract information from images via GLM-4V, supporting automatic screenshot recognition and MCP-based local image file reading for non-vision models like DeepSeek.
Enables coding agents to examine a rendered web page by cross-checking what the browser reports about the DOM/CSS against what is actually painted in pixels, pinpointing which element is broken and the evidence why. Reports visual problems without requiring a baseline, and is designed to grow into design-token extraction and accessibility/UX audits built on Playwright.
Exposes information about Z-Image AI image generation and editing platform, including styles, pricing, and official links, for MCP-compatible AI clients.