image-analysis-mcp
# image-analysis-mcp
FastMCP server for image analysis — OCR, metadata, and EXIF extraction. Part of the [Palimpsest](https://github.com/palimpsest-labs/palimpsest) intelligence toolkit.
## Why
LLMs can't see images natively. This server fills the gap — extract text from screenshots, photos, and diagrams via OCR, pull EXIF camera data and GPS coordinates, and get full image metadata (format, dimensions, DPI, color space). All results are returned as structured JSON for easy downstream processing.
## Architecture
Pluggable OCR backends with graceful degradation:
* **base install** (`pip install image-analysis-mcp`) — metadata / EXIF only (Pillow + exifread, ~5 MB)
* **with OCR** (`pip install image-analysis-mcp[ocr]`) — adds rapidocr-onnxruntime (~250 MB)
* **with Tesseract** (`pip install image-analysis-mcp[tesseract]`) — adds pytesseract (needs system `tesseract-ocr`)
The server tries backends in priority order: rapidocr → tesseract → none.
## Tools
| Tool | Description |
|---|---|
| `ocr_image(image_path)` | OCR text from an image — returns text blocks with confidence scores and bounding boxes |
| `image_metadata(path)` | Full metadata: file info, image dimensions, EXIF tags, GPS coordinates |
| `extract_text(image_path)` | Convenience wrapper — OCR + metadata combined in one response |
## Installation
```bash
git clone https://github.com/palimpsest-labs/image-analysis-mcp
cd image-analysis-mcp
python3 -m venv .venv
source .venv/bin/activate
# Minimal (metadata only)
pip install -e .
# With OCR support
pip install -e ".[ocr]"
# With Tesseract support (requires system tesseract-ocr)
pip install -e ".[tesseract]"
```
## Usage
```python
from image_analysis_mcp.server import extract_text, image_metadata, ocr_image
# Get everything at once
result = extract_text("~/screenshots/page.png")
# Or separate calls
meta = image_metadata("~/screenshots/page.png")
text = ocr_image("~/screenshots/page.png")
```
## Security
Paths must be absolute and resolve to a location under the user's home directory. Path traversal (`..`) and paths starting with `/` are rejected. Symlinks are resolved before checking home-directory containment.
## License
MIT
TDQS
Scored across 3 tools
ocr_image and image_metadata have clearly distinct purposes, but extract_text is a combined wrapper that overlaps with both, creating potential confusion about which tool to use for a given task. The descriptions help clarify that extract_text is a convenience option, but the presence of a full-coverage tool makes the specific tools somewhat redundant.
Tool names are all snake_case, but the patterns vary: ocr_image uses an acronym as a verb, image_metadata is a noun-noun compound with no action verb, and extract_text is a clear verb-noun pair. While still readable, this mix prevents a predictable verb_noun convention across the set.
Three tools is on the lower end of typical server scope, but is appropriate for a focused image analysis tool that covers OCR and metadata extraction. The count is not excessive, and each tool fills a specific need, though the combined wrapper could be seen as unessential.
For the apparent domain of text extraction and metadata retrieval, the set is fairly complete: ocr_image covers text, image_metadata covers metadata, and extract_text provides a combined result. However, other common image analysis operations (e.g., object detection, format conversion) are absent, though they may be out of scope for this server.