anymd
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@anymd@anymd convert this PDF to markdown"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
anymd
Any document → clean Markdown your LLM can actually read. Including the Chinese ones that break everything else.
A CLI, Python library and MCP server for AI agents — no cloud, no API key.
$ anymd convert report.docx
# Quarterly Report
...Why anymd
Most converters are tuned for clean, English, born-digital PDFs. anymd is tuned for
the documents that actually break: a GBK-encoded CSV exported from a Chinese Windows
desktop, a scanned exam paper, a .docx somebody renamed to .doc.
Three things it handles that general-purpose converters do not:
East-Asian encoding is a first-class case | UTF-8 / UTF-16 BOM / GB18030 / Big5 / HTML |
Maths survives extraction | PDFs embedding the Adobe Symbol font usually extract with every |
Built for agents, not only humans | stdout carries only the product, |
Plus the boring guarantees: truncation is always stated in the output, never silent;
a .docx renamed .doc is still identified by content sniffing; hostile input
(zip/tar bombs, SSRF, oversized files) is capped everywhere.
Where anymd is not the best choice
Saves you time:
Complex merged-cell tables. A timetable PDF with vertically merged cells comes out with columns misaligned —
table_strategymakes no difference, it is a limitation of the upstream PDF table extractor. If tables are your whole problem, reach fordoclingorMinerU.Formula-heavy scans. LaTeX output needs the optional
[math]extra, which pulls PyTorch. Without it, scanned maths degrades to plain OCR.Legacy binary
.doc/.ppt. No pip-installable reader exists — re-save as.docx/.pptxfirst.
Related MCP server: DocMistral MCP Server
Installation
Not on PyPI yet. Packaging and the release workflow are ready (see
docs/RELEASING.md); until the first upload lands, install from git:
pip install "anymd[all] @ git+https://github.com/Ljf857/anymd.git"Once published, that collapses to:
pip install anymd # core: text, code, csv/json/yaml/xml, html, eml, epub, markdown
pip install "anymd[all]" # everything below, one shot24 converters cover Office, PDF, e-mail, e-books, images, archives, data files and
~60 code formats. Heavy dependencies live in extras, and a missing one fails with the
exact pip install "anymd[x]" command rather than a raw ImportError:
extra | adds | formats unlocked |
| pymupdf, pymupdf4llm |
|
| python-docx, openpyxl, xlrd, python-pptx, odfpy, striprtf |
|
| rapidocr[onnxruntime], pillow | text recognition in images + scanned PDF pages |
| pix2text | LaTeX output ( |
| extract-msg |
|
| httpx | converting |
| mcp | the |
Requires Python ≥ 3.11. Windows, macOS and Linux are all supported — the CLI is Windows-first (UTF-8 output on cp936 consoles).
Usage
CLI
anymd convert <INPUT...> [-o PATH] [options]
INPUT file, glob, directory, http(s) URL, or - (stdin)anymd convert report.pdf # → stdout
anymd convert report.docx -o out.md # → file
anymd convert ./docs -o md/ # → directory tree → .md files
anymd convert "*.xlsx" --json # → NDJSON manifest
anymd convert scan.pdf --image-mode ocr # OCR scanned pages (needs [ocr])
anymd convert data.csv --max-rows 50 # cap table rows
anymd convert - --stdin-filename msg.eml # stdin with a filename hint
anymd convert report.pdf --page-range 3-7 # partial PDF
anymd --list-formats # everything supported + required extrasAgent contract
stdout carries only the product. Progress, warnings and diagnostics go to stderr.
Exit codes:
0all converted ·1partial failure ·2usage error ·3zero output (missing/unsupported inputs). A file you name explicitly that has no converter is an error (exit 1); the same file swept up by a glob or directory walk is skipped (exit 0).--strictpromotes those skips to errors too.--jsonemits one JSON object per line, each with atypefield so an agent can stream-parse:
{"type":"item","input":"a.docx","output":"a.md","status":"ok","converter":"docx","duration_ms":42,"bytes_in":12345,"bytes_out":678,"sha256_in":"…","warnings":[]}
{"type":"item","input":"legacy.doc","output":null,"status":"error","error":{"code":"UnsupportedFormatError","message":"legacy binary Office formats (.doc/.ppt) require LibreOffice…","hint":"run anymd --list-formats to see supported types"}}
{"type":"summary","total":2,"ok":1,"error":1,"skipped":0,"bytes_in":9000,"bytes_out":678,"duration_ms":120,"exit_code":1}Python API
from anymd import convert, convert_result, list_formats, ConvertContext
md = convert("report.docx")
result = convert_result("report.xlsx") # .metadata (sheets, pages, …) + .warnings
md = convert("https://example.com/post", # URL (needs anymd[web])
ctx=ConvertContext(frontmatter=True))MCP server
pip install "anymd[mcp]"
anymd-mcp # stdio transportTwo tools for any MCP client:
convert_document(source, max_rows, image_mode, frontmatter, max_chars, include_metadata, allow_remote)— one document → Markdown. Missing extras return an actionableERROR[...]string instead of a protocol error; output is truncated atmax_charswith a visible note.list_formats()— every supported extension and its extra.
Claude Desktop / Claude Code registration:
{"mcpServers": {"anymd": {"command": "anymd-mcp"}}}Supported formats
family | formats |
Office |
|
| |
Web |
|
E-books |
|
Data |
|
| |
Notebooks |
|
Images |
|
Archives |
|
Text/code |
|
No extension | content sniffing: PDF/zip-family/OLE/RTF/HTML/images/plain text |
Not supported (by design): legacy binary .doc / .ppt have no pip-installable reader —
anymd fails with a precise message telling you to re-save as .docx/.pptx.
(.xls is supported via xlrd.)
Notes
License: anymd is MIT. Two optional extras carry copyleft licenses that activate only for their own code paths:
pymupdf(AGPL-3.0, thepdfextra) andstriprtf(GPL-3.0, inside theofficeextra). They are separate packages installed at runtime — if you deploy anymd as a hosted service, review your obligations (a pdfminer.six-based fallback can replace thepdfextra).OCR engine: RapidOCR bundles its models in-wheel; nothing is downloaded at runtime. If model files fail to load in a restricted environment, OCR degrades to a warning, never an error.
LaTeX formulas (
[math]extra): math-dense pages are re-recognized with pix2text, which emits$...$LaTeX for formula regions (plain OCR flattensx²and displaces integral limits). pix2text downloads its models on first use — in China setHF_ENDPOINT=https://hf-mirror.com. Any failure degrades to plain OCR; this extra is heavy (pulls PyTorch) and is therefore not part ofanymd[all].Adobe Symbol maths: PDFs that embed the Adobe Symbol font frequently ship no
ToUnicodemap, so an extractor reports raw private-use codepoints and every operator vanishes. anymd decodes the common glyph set (= − + ∫ ∑ π ≤ ∞ ∂ ∈ …) and warns about whatever it could not resolve instead of dropping it quietly. Tall-delimiter fragments are deliberately left undecoded — seesrc/anymd/_symbolfont.pyfor why.Mixed pages (text layer + images): a page carrying both a real text layer and images that hold content gets the OCR reading appended below the text layer, minus any lines the text layer already contained. Nothing is dropped, nothing is repeated.
Remote inputs: http(s) URL inputs are fetched by default in the CLI — pass
--no-remoteto forbid them. Private, loopback and link-local addresses are rejected regardless, unless you explicitly pass--allow-private. The MCPconvert_documenttool is stricter: every URL needsallow_remote=True.Windows-first: the CLI reconfigures stdout to UTF-8 so piping Chinese text through cp936 consoles never raises
UnicodeEncodeError.
Development
git clone https://github.com/Ljf857/anymd.git
cd anymd
pip install -e ".[all,dev]"
pytest # slow OCR tests excluded by default
pytest -m slow # OCR end-to-end
ruff check src tests
ruff format src tests
mypy src/anymdReleasing is documented in docs/RELEASING.md.
Contributions are welcome — open an
issue or pull request. See
CONTRIBUTING.md for the ground rules.
License
MIT © 2026 Ljf857
Optional extras keep their own licenses (see Notes).
This server cannot be deployed
Maintenance
Related MCP Connectors
Document-to-Markdown MCP server — convert PDF, Office and HTML into LLM-ready Markdown.
Convert files, URLs, and documents to clean, AI-ready Markdown via MCP.
Document conversion MCP server: PDF to Markdown, image OCR, spreadsheet parsing.
Convert PDF, DOCX, HTML, and URLs to clean, LLM-ready markdown with tables preserved
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceConverts documents (PDF, DOCX, images, etc.) to Markdown using Microsoft's Markitdown library, with no local setup required. Integrates with AI agents via MCP for seamless document conversion.1-
- AlicenseNot gradedqualityAmaintenanceConverts documents and images to Markdown using Mistral AI's OCR, enabling AI-powered document processing via MCP-compatible clients like Claude Desktop.35 npm2MIT
- FlicenseAqualityDmaintenanceConverts files (PDF, DOCX, PPTX, XLSX, images via OCR) and URLs to Markdown, enabling AI clients to read them via a single MCP tool.1-
- AlicenseAqualityBmaintenanceLocal MCP server that converts PDF, Word, PowerPoint, Excel and more to Markdown on your machine, enabling coding agents like Claude Code and Cursor to read office files in the repo.331 npmMIT