anydoc-serve
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@anydoc-serveconvert https://example.com/report.pdf to markdown"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
anydoc-serve
Office documents, PDFs and scans to Markdown over HTTP and MCP, in one Docker container, with local OCR for English, Japanese and Thai.
anydoc-serve wraps firecrawl/anydoc, the fast Rust converter for Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF, in a server you can run anywhere:
docker run -p 8080:8080 ghcr.io/bon5co/anydoc-serve
curl -F file=@report.docx http://localhost:8080/v1/convertWhat it adds over plain anydoc
anydoc | anydoc-serve | |
Office formats, text PDFs to Markdown | yes | yes, through anydoc, output unchanged |
Run as a service | library and CLI | HTTP API, OpenAPI docs, Docker image (anydoc#11) |
Scanned or image-only PDF pages |
| OCR'd locally, on CPU; nothing leaves the container |
Images ( | not supported | OCR'd locally (anydoc#146) |
Mixed PDFs (some pages scanned) | the whole file fails | text pages go through anydoc, only the scanned pages go through OCR (anydoc#157) |
AI agents | agent skill (CLI) | MCP server at |
Sideways or upside-down scans | turned upright before OCR |
OCR is Tesseract 5 reading every configured language in one pass, so a page that mixes Japanese, Thai and English is read correctly line by line. It was picked by measuring it against RapidOCR/PaddleOCR on English, Japanese, Thai and mixed-script pages: see the benchmark.
Related MCP server: Markitdown MCP Server
Quickstart
docker run -d -p 8080:8080 -e API_KEY=change-me ghcr.io/bon5co/anydoc-serveor docker compose up -d with the docker-compose.yml in this repo.
Convert an upload:
curl -H "Authorization: Bearer change-me" -F file=@scan.pdf http://localhost:8080/v1/convert{
"markdown": "# Quarterly report\n\n...",
"metadata": {
"filename": "scan.pdf", "bytes": 460689, "format": "pdf", "page_count": 4,
"ocr_pages": [3, 4], "ocr_engine": "tesseract", "ocr_langs": ["eng", "jpn", "tha"], "duration_ms": 1735.0
}
}ocr_pages lists the 1-indexed pages that went through OCR; every other page came from anydoc.
Just the Markdown:
curl -H "Authorization: Bearer change-me" -H "Accept: text/markdown" -F file=@report.docx http://localhost:8080/v1/convertHave the server fetch a URL:
curl -H "Authorization: Bearer change-me" -H "Content-Type: application/json" \
-d '{"url": "https://example.com/report.pdf"}' http://localhost:8080/v1/convertInteractive docs are at /docs (OpenAPI JSON at /openapi.json).
API
POST /v1/convert
multipart/form-data:file(required),format,ocr,outputapplication/json:{"url": "...", "format": ..., "ocr": ..., "output": ...}
Option | Values | Default | Meaning |
|
| detected | Force the format. Needed for CSV sent without a |
|
|
|
|
|
|
|
|
Errors answer {"error": {"code": ..., "message": ...}}: 400 bad request or unfetchable URL, 401 missing or wrong key, 413 over MAX_UPLOAD_MB or OCR_MAX_PAGES, 415 unknown format, 422 encrypted, malformed, or needs OCR with OCR off, 504 over CONVERT_TIMEOUT_S.
GET /health answers {"status": "ok", "ocr": {...}, "auth": true} and never needs the key.
MCP
The server speaks MCP over streamable HTTP at /mcp, with one tool, convert_document(url | content_base64, filename?, format?, ocr?), which returns the Markdown.
Claude Code:
claude mcp add --transport http anydoc http://localhost:8080/mcp --header "Authorization: Bearer change-me"Clients configured with an mcpServers JSON block that supports remote HTTP servers (Cursor, for example):
{
"mcpServers": {
"anydoc": {
"type": "http",
"url": "http://localhost:8080/mcp",
"headers": { "Authorization": "Bearer change-me" }
}
}
}Configuration
Variable | Default | Meaning |
|
| Port to listen on. Railway and most platforms set it. |
| empty | When set, every endpoint except |
|
| OCR languages, comma-separated: |
|
|
|
|
| Resolution scanned PDF pages are rendered at for OCR. |
|
| A document needing OCR on more pages than this is refused with |
|
| Largest document accepted, uploaded or fetched. |
|
| Wall-clock limit per conversion, OCR included; over it answers |
|
| Conversions that run at once (the rest wait). Each slot holds its own loaded OCR models. |
|
| Limit for downloading a |
|
| Allow URL sources that resolve to private, loopback or link-local addresses. Off by default so the server cannot be used to reach its own network. |
Resource use
Measured on linux/amd64, 2026-09-28:
Image: 376 MB on disk (94 MB compressed download). Tesseract models for English, Japanese, Thai and script detection are 18 MB of that.
Memory: about 305 MiB RSS after a round of conversions (
docker stats); each conversion slot loads about 90 MiB of OCR models on first use.Speed: a born-digital
.docxconverts in about 2 ms; OCR takes about 1.1 s per page (median, 4 cores of an Intel i5-10500T). A 4-page PDF with 2 scanned pages took 1.5 to 1.7 s.
OCR benchmark
Character error rate on synthetic 200 dpi pages: degraded scans, two pages per language plus one page each mixing Japanese+English, Thai+English and all three. Median seconds per page. Full method, clean-render numbers, every engine row and reproduction steps: docs/ocr-benchmark.md.
Engine (scans) | eng | jpn | tha | jpn+eng | tha+eng | jpn+tha+eng | s/page |
anydoc-serve as shipped (Tesseract 5.5 | 0.3% | 12.3% | 2.6% | 3.1% | 2.8% | 4.3% | 1.12 |
Tesseract, told each page's language (oracle) | 0.3% | 11.1% | 2.6% | 3.9% | 2.8% | 4.3% | 0.39 |
Tesseract, one language per page from script detection (earlier build) | 0.3% | 11.1% | 2.6% | 26.2% | 43.2% | 45.2% | 1.00 |
Tesseract, | 0.2% | 16.4% | 81.2% | 4.4% | 26.8% | 21.4% | 0.51 |
RapidOCR 3.9 PP-OCRv6 small (Thai: PP-OCRv5) | 7.6% | 1.1% | 1.7% | 2.6% | 0.4% | 21.0% | 3.44 |
RapidOCR 3.9 PP-OCRv6 medium | 7.2% | 0.9% | 0.5% | 1.7% | 1.6% | 21.4% | 34.31 |
RapidOCR reads pure Japanese far better and two-script pages slightly better. It fails the three-script page, drops lines of English, and is 3 times slower than the server as shipped. Synthetic pages give exact ground truth but are not your documents: treat these as a comparison between engines, not an accuracy promise.
Deploy on Railway
One service, no volume. API_KEY is generated at deploy (read it from the service's Variables tab) and OCR_LANGS defaults to eng,jpn,tha. Measured on Railway: 246 to 263 MiB of RAM idle, 272 MiB after a round of conversions.
Development
uv sync
uv run pytest # unit tests, no Tesseract needed
docker build -t anydoc-serve .
docker run -d -p 8080:8080 -e API_KEY=test-key anydoc-serve
ANYDOC_SERVE_URL=http://127.0.0.1:8080 ANYDOC_SERVE_API_KEY=test-key uv run pytest -m integrationTest fixtures anydoc-text.docx and anydoc-text.pdf are copied from anydoc's own test corpus (MIT); the scanned fixtures are generated by tests/fixtures/make_fixtures.py.
Credits and license
The document conversion is all anydoc by Firecrawl (MIT). OCR is Tesseract (Apache 2.0) through tesserocr (MIT); PDF pages are rendered with pypdfium2. anydoc-serve is not affiliated with Firecrawl.
MIT, see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Document-to-Markdown MCP server — convert PDF, Office and HTML into LLM-ready Markdown.
Document conversion MCP server: PDF to Markdown, image OCR, spreadsheet parsing.
Convert files, URLs, and documents to clean, AI-ready Markdown via MCP.
Document to Markdown MCP server: PDF, Word, PowerPoint, Excel and HTML, with OCR for large files.
Related MCP Servers
- AlicenseAqualityDmaintenanceExposes the MinerU document-to-markdown API as MCP tools for converting PDF, Word, PPT, and images into Markdown. It supports both local and remote file processing with integrated OCR capabilities for multiple languages.210MIT
- FlicenseNot gradedqualityDmaintenanceConverts documents (PDF, DOCX, images, etc.) to Markdown using Microsoft's Markitdown library, with no local setup required. Integrates with AI agents via MCP for seamless document conversion.1-
- FlicenseAqualityDmaintenanceConverts files (PDF, DOCX, PPTX, XLSX, images via OCR) and URLs to Markdown, enabling AI clients to read them via a single MCP tool.1-
- AlicenseAqualityAmaintenanceConvert PDF, DOCX, HTML, and URLs into clean, LLM-ready markdown with tables preserved and boilerplate stripped, through three MCP tools (URL, local file, or raw bytes). Hosted API with no local dependencies; 50 free conversions with a self-serve key, then $0.002 per call.3380 npmMIT