MinerU MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| MINERU_API_KEY | Yes | Your MinerU API Bearer token | |
| MINERU_BASE_URL | No | API base URL | https://mineru.net/api/v4 |
| MINERU_DEFAULT_MODEL | No | Default model: pipeline or vlm | pipeline |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| mineru_parseB | Parse a document URL. Returns task_id to check status. |
| mineru_statusB | Check task progress. Returns download URL when done. |
| mineru_batchA | Parse multiple URLs in one batch (max 200). Preferred over mineru_upload_batch — faster and more reliable. Use public URLs (arXiv, SSRN, publisher sites) when available. |
| mineru_batch_statusC | Get batch results. Supports pagination for large batches. |
| mineru_upload_batchA | Upload local files for batch parsing. SLOW: uploads can take minutes and may timeout. Prefer mineru_batch with public URLs (arXiv, SSRN, publisher sites) when available — it's faster and more reliable. Only use this for files not available online. |
| mineru_download_resultsA | Download batch results and extract named paper folders. Each folder contains {name}.md, {name}_content.json (structured TOC), and images/. Output includes parsed title — verify it matches the expected paper. |
| mineru_parse_longA | Parse a document LONGER than 200 pages (MinerU's per-file cap) by submitting it as one batch of ≤200-page slices with page_ranges. Give total_pages (from |
| mineru_merge_slicesA | Stitch the slices of a mineru_parse_long batch into one {name}/{name}.md (+ {name}_content.json with page_idx re-based to the whole document, + images/). Slices are ordered by their page range; each is marked with an HTML comment. Waits for nothing — if any slice is still processing, it reports and you re-run later. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 8 tools
Each tool targets a distinct action or resource: status checking, single parse, batch parse, batch status, local upload, result download, slice merging, and long-document parsing. The descriptions explicitly clarify when to use each, especially the overlap between mineru_parse/mineru_parse_long and mineru_batch/mineru_upload_batch, leaving no ambiguity.
All tool names share the 'mineru_' prefix, but the second part mixes verbs (parse, upload, download, merge) with nouns (status, batch, results). This is readable and predictable, though not strictly verb_noun like the calibration example. The minor inconsistency (e.g., 'mineru_batch_status' vs 'mineru_status') does not cause confusion.
With 8 tools, the server covers the core document parsing workflow without excess. Each tool has a clear role: parsing (single/batch/long), status checks, result retrieval, and local file upload. This is well-scoped for the domain.
The tool set covers the full parse lifecycle: submit, check status, download results, and merge slices for long documents. The only notable gap is the lack of a direct single-local-file parse path (requires using upload_batch even for one file), but this is a minor workaround and does not severely impede agents.