Skip to main content
Glama
README.md
# doc2images-mcp

An MCP server that renders documents to images and returns them directly as
**multimodal content** for the model.

- Deterministic "document -> image" conversion only;
- Does **not** convert documents to Markdown and does **not** do LLM-based
  document understanding / OCR;
- PDF rendering is handled by **pypdfium2** (PDFium), a self-contained Python
  wheel with no external binaries. PDFium ships its own encoding data,
  including the CJK CMaps, so Chinese/Japanese/Korean text renders correctly.

> This version supports PDF only. The DOCX / PPTX branch is reserved in the
> code and will be enabled once LibreOffice is wired in.

## Requirements

| Dependency | Notes |
|---|---|
| Python | `>= 3.14` (see `.python-version`) |
| [uv](https://docs.astral.sh/uv/) | Dependency management and run entry point |

No system packages are required: pypdfium2 bundles the PDFium binaries.

## Install

```shell
uv tool install doc2images-mcp
```

## Wiring it into OpenCode

Add the following to the `mcp.servers` section of `opencode.jsonc`:

```jsonc
"doc2images": {
  "type": "local",
  "command": [
    "uvx",
    "doc2images-mcp",
  ],
},
```

To cap how much image data reaches the context, you can also tighten
OpenCode's image settings:

```jsonc
"media": {
  "image": {
    "auto_resize": true,
    "max_width": 2000,
    "max_height": 2000,
    "max_base64_bytes": 5242880,
  },
},
```

## Tool: `pdf_to_images`

Renders PDF pages and returns a list of images (one per page).

| Parameter | Default | Description |
|---|---|---|
| `file_path` | — | Document path. Relative paths resolve against `DOC2IMG_BASE_DIR` |
| `dpi` | `150` | Rasterization resolution, `72`–`300`. Higher is sharper and larger |
| `first_page` | `1` | First page to render (1-based) |
| `last_page` | `0` | Last page to render; `0` means "up to `max_pages`" |
| `max_pages` | `10` | Hard cap on the number of pages (1–50) to protect the context |
| `max_width` | `3000` | Maximum output width; wider pages are scaled down proportionally |

## Environment variables

| Variable | Default | Description |
|---|---|---|
| `DOC2IMG_BASE_DIR` | Current working directory | Base directory used to resolve relative paths |

## Testing

Sample PDFs live in `data/`; the tests render them through the tool:

```powershell
uv run pytest -v
```

Rendered intermediate images are written to `tmp/`, which is ignored via its
own `.gitignore`.

## Roadmap

- **Office support**: after installing LibreOffice, add
  `soffice --headless --convert-to pdf` for `.docx/.pptx/.ppt` in the dispatch
  function, then reuse the same rendering logic.
- **Embedded image extraction**: extract only the images embedded in a
  document instead of whole pages.
- **Alternative backend**: PyMuPDF (`fitz`) is a drop-in alternative if its
  richer text/embedded-image APIs are ever needed; the tool interface stays
  the same.

TDQS

A4.1/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool is exposed, so there is no possibility of selecting the wrong tool. Its purpose—rendering PDF/document pages to images—is unambiguous.

Naming Consistency5/5

The sole tool uses clear snake_case (`pdf_to_images`), and there are no competing conventions to create inconsistency. Though not a strict verb_noun pattern, the name is readable and predictable.

Tool Count3/5

A single tool is thin for any server, but the server's scope is narrowly conversion. It earns its place, but 1 tool is borderline on the count rubric.

Completeness4/5

The tool covers the core PDF-page-to-image workflow, including page-range limiting. However, the server name suggests broader document-to-image support, while the only tool handles PDFs; this is a minor gap.

Maintenance

ActivityMaintained
ResponsivenessNo issues