doc2images-mcp
by BruceKangCN
README.md
# doc2images-mcp
An MCP server that renders documents to images and returns them directly as
**multimodal content** for the model.
- Deterministic "document -> image" conversion only;
- Does **not** convert documents to Markdown and does **not** do LLM-based
document understanding / OCR;
- PDF rendering is handled by **pypdfium2** (PDFium), a self-contained Python
wheel with no external binaries. PDFium ships its own encoding data,
including the CJK CMaps, so Chinese/Japanese/Korean text renders correctly.
> This version supports PDF only. The DOCX / PPTX branch is reserved in the
> code and will be enabled once LibreOffice is wired in.
## Requirements
| Dependency | Notes |
|---|---|
| Python | `>= 3.14` (see `.python-version`) |
| [uv](https://docs.astral.sh/uv/) | Dependency management and run entry point |
No system packages are required: pypdfium2 bundles the PDFium binaries.
## Install
```shell
uv tool install doc2images-mcp
```
## Wiring it into OpenCode
Add the following to the `mcp.servers` section of `opencode.jsonc`:
```jsonc
"doc2images": {
"type": "local",
"command": [
"uvx",
"doc2images-mcp",
],
},
```
To cap how much image data reaches the context, you can also tighten
OpenCode's image settings:
```jsonc
"media": {
"image": {
"auto_resize": true,
"max_width": 2000,
"max_height": 2000,
"max_base64_bytes": 5242880,
},
},
```
## Tool: `pdf_to_images`
Renders PDF pages and returns a list of images (one per page).
| Parameter | Default | Description |
|---|---|---|
| `file_path` | — | Document path. Relative paths resolve against `DOC2IMG_BASE_DIR` |
| `dpi` | `150` | Rasterization resolution, `72`–`300`. Higher is sharper and larger |
| `first_page` | `1` | First page to render (1-based) |
| `last_page` | `0` | Last page to render; `0` means "up to `max_pages`" |
| `max_pages` | `10` | Hard cap on the number of pages (1–50) to protect the context |
| `max_width` | `3000` | Maximum output width; wider pages are scaled down proportionally |
## Environment variables
| Variable | Default | Description |
|---|---|---|
| `DOC2IMG_BASE_DIR` | Current working directory | Base directory used to resolve relative paths |
## Testing
Sample PDFs live in `data/`; the tests render them through the tool:
```powershell
uv run pytest -v
```
Rendered intermediate images are written to `tmp/`, which is ignored via its
own `.gitignore`.
## Roadmap
- **Office support**: after installing LibreOffice, add
`soffice --headless --convert-to pdf` for `.docx/.pptx/.ppt` in the dispatch
function, then reuse the same rendering logic.
- **Embedded image extraction**: extract only the images embedded in a
document instead of whole pages.
- **Alternative backend**: PyMuPDF (`fitz`) is a drop-in alternative if its
richer text/embedded-image APIs are ever needed; the tool interface stays
the same.
TDQS
A4.1/5.0
Scored across 1 tool
Disambiguation5/5
Only one tool is exposed, so there is no possibility of selecting the wrong tool. Its purpose—rendering PDF/document pages to images—is unambiguous.
Naming Consistency5/5
The sole tool uses clear snake_case (`pdf_to_images`), and there are no competing conventions to create inconsistency. Though not a strict verb_noun pattern, the name is readable and predictable.
Tool Count3/5
A single tool is thin for any server, but the server's scope is narrowly conversion. It earns its place, but 1 tool is borderline on the count rubric.
Completeness4/5
The tool covers the core PDF-page-to-image workflow, including page-range limiting. However, the server name suggests broader document-to-image support, while the only tool handles PDFs; this is a minor gap.
Maintenance
ActivityMaintained
ResponsivenessNo issues