Skip to main content
Glama

anymd

Any document → clean Markdown your LLM can actually read. Including the Chinese ones that break everything else.

A CLI, Python library and MCP server for AI agents — no cloud, no API key.

English | 简体中文

CI PyPI Python License: MIT Code style: ruff

$ anymd convert report.docx
# Quarterly Report
...

Why anymd

Most converters are tuned for clean, English, born-digital PDFs. anymd is tuned for the documents that actually break: a GBK-encoded CSV exported from a Chinese Windows desktop, a scanned exam paper, a .docx somebody renamed to .doc.

Three things it handles that general-purpose converters do not:

East-Asian encoding is a first-class case

UTF-8 / UTF-16 BOM / GB18030 / Big5 / HTML <meta charset> detection, plus a CLI that forces UTF-8 stdout — piping Chinese through a cp936 console never raises UnicodeEncodeError

Maths survives extraction

PDFs embedding the Adobe Symbol font usually extract with every =, − and ∫ silently missing (极大值 f 3 3 ln10). anymd decodes those glyphs, so formulas come out readable

Built for agents, not only humans

stdout carries only the product, --json gives an NDJSON manifest to stream-parse, exit codes mean something, and the OCR/maths layers degrade to a warning instead of a traceback

Plus the boring guarantees: truncation is always stated in the output, never silent; a .docx renamed .doc is still identified by content sniffing; hostile input (zip/tar bombs, SSRF, oversized files) is capped everywhere.

Where anymd is not the best choice

Saves you time:

  • Complex merged-cell tables. A timetable PDF with vertically merged cells comes out with columns misaligned — table_strategy makes no difference, it is a limitation of the upstream PDF table extractor. If tables are your whole problem, reach for docling or MinerU.

  • Formula-heavy scans. LaTeX output needs the optional [math] extra, which pulls PyTorch. Without it, scanned maths degrades to plain OCR.

  • Legacy binary .doc / .ppt. No pip-installable reader exists — re-save as .docx / .pptx first.

Related MCP server: DocMistral MCP Server

Installation

Not on PyPI yet. Packaging and the release workflow are ready (see docs/RELEASING.md); until the first upload lands, install from git:

pip install "anymd[all] @ git+https://github.com/Ljf857/anymd.git"

Once published, that collapses to:

pip install anymd            # core: text, code, csv/json/yaml/xml, html, eml, epub, markdown
pip install "anymd[all]"     # everything below, one shot

24 converters cover Office, PDF, e-mail, e-books, images, archives, data files and ~60 code formats. Heavy dependencies live in extras, and a missing one fails with the exact pip install "anymd[x]" command rather than a raw ImportError:

extra

adds

formats unlocked

pdf

pymupdf, pymupdf4llm

.pdf (headings, tables, page markers)

office

python-docx, openpyxl, xlrd, python-pptx, odfpy, striprtf

.docx .xlsx .xls .pptx .odt .ods .odp .rtf

ocr

rapidocr[onnxruntime], pillow

text recognition in images + scanned PDF pages

math

pix2text

LaTeX output ($...$) for math-dense scanned pages

email

extract-msg

.msg (Outlook); .eml is already core

web

httpx

converting http(s)://… URLs directly

mcp

mcp

the anymd-mcp MCP server

Requires Python ≥ 3.11. Windows, macOS and Linux are all supported — the CLI is Windows-first (UTF-8 output on cp936 consoles).

Usage

CLI

anymd convert <INPUT...> [-o PATH] [options]

INPUT   file, glob, directory, http(s) URL, or - (stdin)
anymd convert report.pdf                      # → stdout
anymd convert report.docx -o out.md           # → file
anymd convert ./docs -o md/                   # → directory tree → .md files
anymd convert "*.xlsx" --json                 # → NDJSON manifest
anymd convert scan.pdf --image-mode ocr       # OCR scanned pages (needs [ocr])
anymd convert data.csv --max-rows 50          # cap table rows
anymd convert - --stdin-filename msg.eml      # stdin with a filename hint
anymd convert report.pdf --page-range 3-7     # partial PDF
anymd --list-formats                          # everything supported + required extras

Agent contract

  • stdout carries only the product. Progress, warnings and diagnostics go to stderr.

  • Exit codes: 0 all converted · 1 partial failure · 2 usage error · 3 zero output (missing/unsupported inputs). A file you name explicitly that has no converter is an error (exit 1); the same file swept up by a glob or directory walk is skipped (exit 0). --strict promotes those skips to errors too.

  • --json emits one JSON object per line, each with a type field so an agent can stream-parse:

{"type":"item","input":"a.docx","output":"a.md","status":"ok","converter":"docx","duration_ms":42,"bytes_in":12345,"bytes_out":678,"sha256_in":"…","warnings":[]}
{"type":"item","input":"legacy.doc","output":null,"status":"error","error":{"code":"UnsupportedFormatError","message":"legacy binary Office formats (.doc/.ppt) require LibreOffice…","hint":"run anymd --list-formats to see supported types"}}
{"type":"summary","total":2,"ok":1,"error":1,"skipped":0,"bytes_in":9000,"bytes_out":678,"duration_ms":120,"exit_code":1}

Python API

from anymd import convert, convert_result, list_formats, ConvertContext

md = convert("report.docx")
result = convert_result("report.xlsx")     # .metadata (sheets, pages, …) + .warnings
md = convert("https://example.com/post",   # URL (needs anymd[web])
             ctx=ConvertContext(frontmatter=True))

MCP server

pip install "anymd[mcp]"
anymd-mcp        # stdio transport

Two tools for any MCP client:

  • convert_document(source, max_rows, image_mode, frontmatter, max_chars, include_metadata, allow_remote) — one document → Markdown. Missing extras return an actionable ERROR[...] string instead of a protocol error; output is truncated at max_chars with a visible note.

  • list_formats() — every supported extension and its extra.

Claude Desktop / Claude Code registration:

{"mcpServers": {"anymd": {"command": "anymd-mcp"}}}

Supported formats

family

formats

Office

.docx .xlsx .xlsm .xls .pptx .odt .ods .odp .rtf

PDF

.pdf (structure-aware; scanned or image-heavy pages → OCR when [ocr] installed, LaTeX formulas when [math] installed)

Web

.html .htm .xhtml .mht/.mhtml, http(s) URLs

E-books

.epub (spine-ordered)

Data

.csv .tsv → tables · .json .jsonl → tables or fenced · .yaml .yml .toml .xml .ini .cfg .conf .env .properties → fenced

Email

.eml (headers + body + attachment listing) · .msg

Notebooks

.ipynb (cells in order, outputs truncated)

Images

.png .jpg .jpeg .gif .bmp .tif .tiff .webp (OCR or metadata)

Archives

.zip .tar .tar.gz .tgz .tar.bz2 .tar.xz — members converted recursively

Text/code

.txt .log .md and ~60 code extensions → language-tagged fences

No extension

content sniffing: PDF/zip-family/OLE/RTF/HTML/images/plain text

Not supported (by design): legacy binary .doc / .ppt have no pip-installable reader — anymd fails with a precise message telling you to re-save as .docx/.pptx. (.xls is supported via xlrd.)

Notes

  • License: anymd is MIT. Two optional extras carry copyleft licenses that activate only for their own code paths: pymupdf (AGPL-3.0, the pdf extra) and striprtf (GPL-3.0, inside the office extra). They are separate packages installed at runtime — if you deploy anymd as a hosted service, review your obligations (a pdfminer.six-based fallback can replace the pdf extra).

  • OCR engine: RapidOCR bundles its models in-wheel; nothing is downloaded at runtime. If model files fail to load in a restricted environment, OCR degrades to a warning, never an error.

  • LaTeX formulas ([math] extra): math-dense pages are re-recognized with pix2text, which emits $...$ LaTeX for formula regions (plain OCR flattens x² and displaces integral limits). pix2text downloads its models on first use — in China set HF_ENDPOINT=https://hf-mirror.com. Any failure degrades to plain OCR; this extra is heavy (pulls PyTorch) and is therefore not part of anymd[all].

  • Adobe Symbol maths: PDFs that embed the Adobe Symbol font frequently ship no ToUnicode map, so an extractor reports raw private-use codepoints and every operator vanishes. anymd decodes the common glyph set (= − + ∫ ∑ π ≤ ∞ ∂ ∈ …) and warns about whatever it could not resolve instead of dropping it quietly. Tall-delimiter fragments are deliberately left undecoded — see src/anymd/_symbolfont.py for why.

  • Mixed pages (text layer + images): a page carrying both a real text layer and images that hold content gets the OCR reading appended below the text layer, minus any lines the text layer already contained. Nothing is dropped, nothing is repeated.

  • Remote inputs: http(s) URL inputs are fetched by default in the CLI — pass --no-remote to forbid them. Private, loopback and link-local addresses are rejected regardless, unless you explicitly pass --allow-private. The MCP convert_document tool is stricter: every URL needs allow_remote=True.

  • Windows-first: the CLI reconfigures stdout to UTF-8 so piping Chinese text through cp936 consoles never raises UnicodeEncodeError.

Development

git clone https://github.com/Ljf857/anymd.git
cd anymd
pip install -e ".[all,dev]"

pytest                 # slow OCR tests excluded by default
pytest -m slow         # OCR end-to-end
ruff check src tests
ruff format src tests
mypy src/anymd

Releasing is documented in docs/RELEASING.md.

Contributions are welcome — open an issue or pull request. See CONTRIBUTING.md for the ground rules.

License

MIT © 2026 Ljf857

Optional extras keep their own licenses (see Notes).

Related MCP Connectors

Related MCP Servers