MCP server that converts file contents into compact, line-numbered PNG images for vision models to read, reducing token usage by roughly 7x for large files.
Enables OCR-based text, table, handwriting, and formula recognition from images via MCP, supporting local files and URLs for integration with AI assistants.
An MCP server that provides local image recognition on macOS, including OCR, image classification, comprehensive image analysis, and screenshot recognition, all via Apple's Vision framework without any network requests.
Provides image recognition capabilities using Anthropic Claude Vision and OpenAI GPT-4 Vision APIs, supporting multiple image formats and offering optional text extraction via Tesseract OCR.
A Model Context Protocol server that gives AI assistants OCR with first-class accuracy handling and evaluation. It wraps three engines behind one interface and can score and compare them.
Enables document text recognition and extraction from images and PDFs using Claude Vision, including support for scanned documents and structured output, without requiring local OCR engines.
Extracts text from images using Tesseract OCR with support for local files, URLs, and raw image bytes. It provides production-grade OCR capabilities and multi-language support through the Model Context Protocol.
Identifies anime characters and works from illustrations, and traces back to Pixiv originals, artists, or animation screenshots using vision models and SauceNAO/Trace.moe.
Provides OCR services powered by Google's Gemini API to extract text from images via file paths or base64 strings. It enables high-accuracy text recognition and CAPTCHA processing through simple MCP tools.
Enables AI agents to convert images and PDFs, including large scanned books, into Markdown through GLM-OCR, with support for long-image slicing, progress tracking, and asynchronous OCR tasks.
Enables interaction with Adobe Commerce Catalog Services to retrieve product variants, price overrides, category permissions, and environment details via MCP.
Exposes local Umi-OCR v2 capabilities to AI agents via MCP, enabling image text extraction, batch OCR, PDF OCR, and status checks without manually starting the service.