Enables MCP clients to generate and edit images through Codex's built-in image generation using a ChatGPT login, returning final files with metadata and optional inline previews.
Provides Claude with detailed image inspection capabilities, including metadata extraction, histogram analysis, tonal and color analysis, sharpness detection, and more, supporting both standard and RAW formats.
Reverse-engineers design videos and images into structured frontend implementation specifications using vision LLMs and FFMPEG for frame-level analysis.
Enables intelligent analysis and organization of image collections with smart filename generation, metadata extraction, and automated folder organization. Supports batch processing, color analysis, EXIF data extraction, and multiple naming styles for efficient photo management.
Enables detection of AI-generated content in images, videos, audio, and text via the AI or Not API. Supports media analysis tools for deepfakes, synthetic voices, and AI-written text.
Enables AI assistants to locally process images with tools for cropping, zooming, enhancement, edge detection, segmentation, and text region extraction, all without external API keys. It uses PIL, OpenCV, and scikit-image for robust image analysis.
Enables high-quality 3D-style cartoon image generation using Google Gemini AI alongside secure file system operations like reading, writing, and managing directories. It provides a dual-purpose toolkit for creative asset generation and local file management via the Model Context Protocol.
A server that accepts image URLs and analyzes their content using GPT-4-turbo, enabling Claude AI assistants to understand and describe images through natural language.
A FastAPI-based server that provides a web UI for uploading, processing, and analyzing trading screenshots with automatic trade data extraction and analytics.
Enables image analysis and recognition through multiple LLM vision models (Gemini, GPT-4o, Qwen-VL, Doubao) by accepting image URLs or Base64 data and returning text descriptions or answers to questions about the images.
MCP server that lets AI assistants like Claude or Cursor extract structured data (e.g., invoices) from images and PDFs via Vision API, with tools for analysis, questions, and preset management.
An MCP server that uses Xiaomi MiMo v2.5 multimodal model to provide image recognition capabilities (description, multi-image analysis, OCR, and image info validation) for text-only main models like deepseek-v4-flash, accepting local paths, URLs, file://, and base64 data inputs.