Skip to main content
Glama
SylphxAI
by SylphxAI

Merged into anymd (2026-09-25). anymd reads PDFs, Office files, EPUB, web pages, images and video into clean Markdown for AI agents: npx -y @sylphx/anymd. This repository is archived.

Iris

Image facts with pixel-level proof.

Dimensions, format, and metadata by default. Local Tesseract OCR only when requested.

npm license

npm @sylphx/iris · bin iris · MCP io.github.SylphxAI/iris


The problem

A screenshot is not a paragraph. An agent that guesses the text, the size, or the pixels it never measured will cite something it cannot show.

Related MCP server: Vision MCP Server

The difference

The ask

Iris returns

“Read this screenshot.”

dimensions, format, metadata, and trust warnings. No OCR unless you ask.

“Read the text.”

OCR lines with boxes, only when include_ocr is true or profile is quality.

“What changed between these two captures?”

changed pixels and a changed box, when both images are the same size.

“Crop this region.”

pixel bounds and a hash. PNG bytes only if include_region_image is true.

Iris does not run a generative vision model. A missing Tesseract binary is a gap, not a made-up transcription.

Install

npx -y @sylphx/iris

That starts a stdio MCP server. No API key. OCR needs the tesseract binary on PATH.

Your client

Setup

Any agent / CLI

npx -y @sylphx/iris

Claude Code

claude mcp add iris -- npx -y @sylphx/iris

Claude Desktop / Cursor / VS Code / Codex

"command": "npx", "args": ["-y", "@sylphx/iris"]

{
  "mcpServers": {
    "iris": { "command": "npx", "args": ["-y", "@sylphx/iris"] }
  }
}

A read, then text only if you ask

{ "path": "/absolute/path/to/screenshot.png" }

fast, or no profile, returns dimensions, format, metadata, and trust warnings. It does not run OCR.

{
  "path": "/absolute/path/to/screenshot.png",
  "include_ocr": true
}

include_ocr: true, or profile: "quality", runs local Tesseract. Explicit include_ocr: false wins over quality. If Tesseract is missing, status stays ok and gaps includes OCR_UNAVAILABLE.

Reference: tools · defaults

Tools

Tool

When to call it

read_image

One local image. Geometry and metadata. OCR only when requested. An optional region crops inside this read.

image_probe

Format, dimensions, pixel count, and source hash. No OCR and no crop.

crop_region

One citeable region: bounds and hash. Set include_region_image for PNG bytes. No OCR.

compare_images

Changed pixels between two images of equal dimensions. No OCR.

What it will not pretend

  • GPS fields are always redacted. There is no switch that puts them back.

  • The default file cap is 32 MiB (33,554,432 bytes). A cropped region may not exceed 67,108,864 pixels.

  • compare_images requires equal dimensions. It does not resize either image.

  • OCR lines are {text, bbox, confidence}. The default language is eng.

  • Oversized or undecodable files fail with an error. They are not returned as a successful read.

Companion MCP tools

Product

Job

Citra

PDF answers with page-level proof

Cue

Video timelines and timestamp evidence

Spine

Repository architecture and impact

Locus

Exact code-chunk retrieval

Lookout

Web research with source excerpts

Each product is independent. Install only the tools your agent needs.

Documentation

Development

bun install
bun test
bun run check
bun run docs:build
cargo test

License

MIT

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    Not graded
    maintenance
    Enables AI agents to analyze images through vision AI providers (Gemini, OpenAI, Claude), performing tasks like image description, object detection with bounding boxes, region-specific analysis, and precise color extraction without consuming context window with raw pixels.
    4
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables text-only AI agents to see images on demand by calling any OpenAI-compatible vision API for OCR, image analysis, structured extraction, image comparison, and GUI screenshot-to-accessibility-tree conversion.
    1
    MIT