Enables AI agents and MCP-compliant clients to extract visual data from charts and graphs into structured markdown tables using deterministic, zero-dependency Python.
An MCP server that reads a file's embedded C2PA Content Credential to report whether it declares AI generation or a real-world capture, and checks whether that declaration's signature still verifies against the file's current bytes.
MCP server for the OpenGraph.io API -- extract OG metadata, capture screenshots, scrape pages, query sites with AI, and generate branded images with iterative refinement.
MCP server for ScanToBill, an AI-powered OCR platform that extracts structured JSON from invoices and business documents in Arabic and English. It enables Claude Desktop to extract invoice data from URLs, list past extractions, and retrieve detailed results.
Enables comprehensive video file analysis including extracting metadata, stream information, bitrate calculations, and generating technical reports. Supports all FFmpeg-compatible video formats with output in JSON, text, or Markdown formats.
Generates children's storybook illustrations in multiple art styles (3D cartoon, watercolor, pixel art, hand drawn, claymation) with matching stories using Google's Gemini AI.
Facilitates running Python code in a sandbox and generating images using the FLUX model via an MCP server compatible with clients like Goose and the Claude Desktop App.
Enables multimodal AI capabilities through GLM-4.5V API for image processing, visual querying with OCR/QA/detection modes, and file content extraction from various formats including PDFs, documents, and images.
A Model Context Protocol server that generates images using Replicate's FLUX model and stores them in Cloudflare R2, allowing users to create images through simple prompts and retrieve accessible URLs.
Gives AI agents eyes on in-memory image buffers in live gdb/lldb C/C++ debug sessions: list observable symbols at a breakpoint, view renderings, read exact pixel values, and dump lossless .npy copies. Agents can also drive the human's viewer window (pan, zoom, channels, auto-contrast).
Enables natural language interaction with complex computer vision workflows such as auto-labeling, class mapping, and embedding selection through an LLM-agnostic MCP orchestration layer.