Skip to main content
Glama

doc2images-mcp

An MCP server that renders documents to images and returns them directly as multimodal content for the model.

  • Deterministic "document -> image" conversion only;

  • Does not convert documents to Markdown and does not do LLM-based document understanding / OCR;

  • PDF rendering is handled by pypdfium2 (PDFium), a self-contained Python wheel with no external binaries. PDFium ships its own encoding data, including the CJK CMaps, so Chinese/Japanese/Korean text renders correctly.

This version supports PDF only. The DOCX / PPTX branch is reserved in the code and will be enabled once LibreOffice is wired in.

Requirements

Dependency

Notes

Python

>= 3.14 (see .python-version)

uv

Dependency management and run entry point

No system packages are required: pypdfium2 bundles the PDFium binaries.

Related MCP server: MokuPDF

Install

uv tool install doc2images-mcp

Wiring it into OpenCode

Add the following to the mcp.servers section of opencode.jsonc:

"doc2images": {
  "type": "local",
  "command": [
    "uvx",
    "doc2images-mcp",
  ],
},

To cap how much image data reaches the context, you can also tighten OpenCode's image settings:

"media": {
  "image": {
    "auto_resize": true,
    "max_width": 2000,
    "max_height": 2000,
    "max_base64_bytes": 5242880,
  },
},

Tool: pdf_to_images

Renders PDF pages and returns a list of images (one per page).

Parameter

Default

Description

file_path

—

Document path. Relative paths resolve against DOC2IMG_BASE_DIR

dpi

150

Rasterization resolution, 72–300. Higher is sharper and larger

first_page

1

First page to render (1-based)

last_page

0

Last page to render; 0 means "up to max_pages"

max_pages

10

Hard cap on the number of pages (1–50) to protect the context

max_width

3000

Maximum output width; wider pages are scaled down proportionally

Environment variables

Variable

Default

Description

DOC2IMG_BASE_DIR

Current working directory

Base directory used to resolve relative paths

Testing

Sample PDFs live in data/; the tests render them through the tool:

uv run pytest -v

Rendered intermediate images are written to tmp/, which is ignored via its own .gitignore.

Roadmap

  • Office support: after installing LibreOffice, add soffice --headless --convert-to pdf for .docx/.pptx/.ppt in the dispatch function, then reuse the same rendering logic.

  • Embedded image extraction: extract only the images embedded in a document instead of whole pages.

  • Alternative backend: PyMuPDF (fitz) is a drop-in alternative if its richer text/embedded-image APIs are ever needed; the tool interface stays the same.

Available Tools

1 tool
pdf_to_imagesPdf To ImagesA

Render document pages to images for vision models.

Returns one image per page. Use first_page/last_page/max_pages to limit how much of the document is sent to the model.

ParametersJSON Schema
NameRequiredDescriptionDefault
dpiNoRasterisation resolution, 72-300. Higher is sharper but larger.
file_pathYesPath to the document. Relative paths resolve against DOC2IMG_BASE_DIR (defaults to the current working directory).
last_pageNo1-based last page to render; 0 means "until max_pages".
max_pagesNoHard cap on the number of pages returned (1-50).
max_widthNoMaximum output width in pixels; wider pages are scaled down.
first_pageNo1-based first page to render.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the behavioral burden, and it does disclose return cardinality ('one image per page') and that the volume sent to a model can be bounded. It doesn't cover resource cost, failure behavior on unsupported or corrupt files, or default page cap, so it is incomplete but substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the verb and the output, and the second sentence routes straight to the relevant parameters. Nothing redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter tool with no annotations and no output schema, the definition is close to sufficient: format, per-page cardinality, and page-limiting controls are all covered. It stops short of mentioning default output volume or resolution implications, and with no annotations the safety/cost profile is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already names and constrains every parameter (dpi range, page bounds, max_pages cap, max_width). The description only gestures at first_page/last_page/max_pages as limiting controls, which is baseline-level added value at this coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (render document pages to images) and frames the purpose around vision models, which tells an agent exactly when this output format is appropriate. No siblings exist, so differentiation isn't required.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what the images are for ('for vision models') and points at three parameters for limiting volume, which implies usage context. However, it gives no when-not-to-use guidance or scope limits (e.g., at what point to prefer text extraction instead), so this is implied rather than explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedpdf_to_images

TDQS

A4.1/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool is exposed, so there is no possibility of selecting the wrong tool. Its purpose—rendering PDF/document pages to images—is unambiguous.

Naming Consistency5/5

The sole tool uses clear snake_case (`pdf_to_images`), and there are no competing conventions to create inconsistency. Though not a strict verb_noun pattern, the name is readable and predictable.

Tool Count3/5

A single tool is thin for any server, but the server's scope is narrowly conversion. It earns its place, but 1 tool is borderline on the count rubric.

Completeness4/5

The tool covers the core PDF-page-to-image workflow, including page-range limiting. However, the server name suggests broader document-to-image support, while the only tool handles PDFs; this is a minor gap.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers