Image Reader MCP
Enables optional local semantic object detection via Ollama, allowing the read_image tool to return open-vocabulary objects with bounding boxes and confidence scores.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Image Reader MCPread this image and extract all text with coordinates"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Merged into anymd (2026-09-25). anymd reads PDFs, Office files, EPUB, web pages, images and video into clean Markdown for AI agents:
npx -y @sylphx/anymd. This repository is archived.
Iris
Image facts with pixel-level proof.
Dimensions, format, and metadata by default. Local Tesseract OCR only when requested.
npm @sylphx/iris · bin iris · MCP io.github.SylphxAI/iris
The problem
A screenshot is not a paragraph. An agent that guesses the text, the size, or the pixels it never measured will cite something it cannot show.
Related MCP server: Vision MCP Server
The difference
The ask | Iris returns |
“Read this screenshot.” | dimensions, format, metadata, and trust warnings. No OCR unless you ask. |
“Read the text.” | OCR lines with boxes, only when |
“What changed between these two captures?” | changed pixels and a changed box, when both images are the same size. |
“Crop this region.” | pixel bounds and a hash. PNG bytes only if |
Iris does not run a generative vision model. A missing Tesseract binary is a gap, not a made-up transcription.
Install
npx -y @sylphx/irisThat starts a stdio MCP server. No API key. OCR needs the tesseract binary on PATH.
Your client | Setup |
Any agent / CLI |
|
Claude Code |
|
Claude Desktop / Cursor / VS Code / Codex |
|
{
"mcpServers": {
"iris": { "command": "npx", "args": ["-y", "@sylphx/iris"] }
}
}A read, then text only if you ask
{ "path": "/absolute/path/to/screenshot.png" }fast, or no profile, returns dimensions, format, metadata, and trust warnings. It does not run OCR.
{
"path": "/absolute/path/to/screenshot.png",
"include_ocr": true
}include_ocr: true, or profile: "quality", runs local Tesseract. Explicit include_ocr: false wins over quality. If Tesseract is missing, status stays ok and gaps includes OCR_UNAVAILABLE.
Tools
Tool | When to call it |
| One local image. Geometry and metadata. OCR only when requested. An optional |
| Format, dimensions, pixel count, and source hash. No OCR and no crop. |
| One citeable region: bounds and hash. Set |
| Changed pixels between two images of equal dimensions. No OCR. |
What it will not pretend
GPS fields are always redacted. There is no switch that puts them back.
The default file cap is 32 MiB (33,554,432 bytes). A cropped region may not exceed 67,108,864 pixels.
compare_imagesrequires equal dimensions. It does not resize either image.OCR lines are
{text, bbox, confidence}. The default language iseng.Oversized or undecodable files fail with an error. They are not returned as a successful read.
Companion MCP tools
Product | Job |
PDF answers with page-level proof | |
Video timelines and timestamp evidence | |
Repository architecture and impact | |
Exact code-chunk retrieval | |
Web research with source excerpts |
Each product is independent. Install only the tools your agent needs.
Documentation
Website | |
Quickstart | |
Compare |
Development
bun install
bun test
bun run check
bun run docs:build
cargo testLicense
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Verified, sourced, real-time intelligence layer for AI agents.
Give agents instant OG image generation, social metadata audits, and rendering guidance.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Give agents eyes on any web page: structured context, and changes explained in plain language.
Related MCP Servers
- FlicenseAqualityNot gradedmaintenanceEnables AI agents to analyze images through vision AI providers (Gemini, OpenAI, Claude), performing tasks like image description, object detection with bounding boxes, region-specific analysis, and precise color extraction without consuming context window with raw pixels.4-
- AlicenseAqualityDmaintenanceEnables AI agents to analyze images, extract text, compare images, and analyze video through any OpenAI-compatible vision model.4145 npm20MIT
- AlicenseAqualityAmaintenanceEnables text-first agents to inspect images via focused questions, returning compact, checked evidence packets. Reduces visual context by ~90% while preserving expected fields.1MIT
- AlicenseNot gradedqualityAmaintenanceEnables text-only AI agents to see images on demand by calling any OpenAI-compatible vision API for OCR, image analysis, structured extraction, image comparison, and GUI screenshot-to-accessibility-tree conversion.1MIT