Skip to main content
Glama
README.md
<div align="center">

# anymd

**Any document → clean Markdown your LLM can actually read.**
**Including the Chinese ones that break everything else.**

A CLI, Python library and MCP server for AI agents — no cloud, no API key.

[English](https://github.com/Ljf857/anymd/blob/main/README.md) | [简体中文](https://github.com/Ljf857/anymd/blob/main/README.zh-CN.md)

[![CI](https://github.com/Ljf857/anymd/actions/workflows/ci.yml/badge.svg)](https://github.com/Ljf857/anymd/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/anymd)](https://pypi.org/project/anymd/)
[![Python](https://img.shields.io/pypi/pyversions/anymd)](https://pypi.org/project/anymd/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://github.com/Ljf857/anymd/blob/main/LICENSE)
[![Code style: ruff](https://img.shields.io/badge/lint-ruff-261230.svg)](https://github.com/astral-sh/ruff)

```console
$ anymd convert report.docx
# Quarterly Report
...
```

</div>

---

## Why anymd

Most converters are tuned for clean, English, born-digital PDFs. anymd is tuned for
the documents that actually break: a GBK-encoded CSV exported from a Chinese Windows
desktop, a scanned exam paper, a `.docx` somebody renamed to `.doc`.

Three things it handles that general-purpose converters do not:

| | |
|---|---|
| **East-Asian encoding is a first-class case** | UTF-8 / UTF-16 BOM / GB18030 / Big5 / HTML `<meta charset>` detection, plus a CLI that forces UTF-8 stdout — piping Chinese through a cp936 console never raises `UnicodeEncodeError` |
| **Maths survives extraction** | PDFs embedding the Adobe Symbol font usually extract with every `=`, `−` and `∫` silently missing (`极大值 f  3  3 ln10`). anymd decodes those glyphs, so formulas come out readable |
| **Built for agents, not only humans** | stdout carries only the product, `--json` gives an NDJSON manifest to stream-parse, exit codes mean something, and the OCR/maths layers degrade to a warning instead of a traceback |

Plus the boring guarantees: truncation is always stated in the output, never silent;
a `.docx` renamed `.doc` is still identified by content sniffing; hostile input
(zip/tar bombs, SSRF, oversized files) is capped everywhere.

### Where anymd is *not* the best choice

Saves you time:

- **Complex merged-cell tables.** A timetable PDF with vertically merged cells comes
  out with columns misaligned — `table_strategy` makes no difference, it is a
  limitation of the upstream PDF table extractor. If tables are your whole problem,
  reach for [`docling`](https://github.com/docling-project/docling) or
  [`MinerU`](https://github.com/opendatalab/MinerU).
- **Formula-heavy scans.** LaTeX output needs the optional `[math]` extra, which pulls
  PyTorch. Without it, scanned maths degrades to plain OCR.
- **Legacy binary `.doc` / `.ppt`.** No pip-installable reader exists — re-save as
  `.docx` / `.pptx` first.

## Installation

> **Not on PyPI yet.** Packaging and the release workflow are ready
> (see [`docs/RELEASING.md`](https://github.com/Ljf857/anymd/blob/main/docs/RELEASING.md)); until the first upload lands,
> install from git:

```bash
pip install "anymd[all] @ git+https://github.com/Ljf857/anymd.git"
```

Once published, that collapses to:

```bash
pip install anymd            # core: text, code, csv/json/yaml/xml, html, eml, epub, markdown
pip install "anymd[all]"     # everything below, one shot
```

**24 converters** cover Office, PDF, e-mail, e-books, images, archives, data files and
~60 code formats. Heavy dependencies live in extras, and a missing one fails with the
exact `pip install "anymd[x]"` command rather than a raw `ImportError`:

| extra | adds | formats unlocked |
|---|---|---|
| `pdf` | pymupdf, pymupdf4llm | `.pdf` (headings, tables, page markers) |
| `office` | python-docx, openpyxl, xlrd, python-pptx, odfpy, striprtf | `.docx` `.xlsx` `.xls` `.pptx` `.odt` `.ods` `.odp` `.rtf` |
| `ocr` | rapidocr[onnxruntime], pillow | text recognition in images + scanned PDF pages |
| `math` | pix2text | LaTeX output (`$...$`) for math-dense scanned pages |
| `email` | extract-msg | `.msg` (Outlook); `.eml` is already core |
| `web` | httpx | converting `http(s)://…` URLs directly |
| `mcp` | mcp | the `anymd-mcp` MCP server |

> Requires Python ≥ 3.11. Windows, macOS and Linux are all supported — the CLI is
> Windows-first (UTF-8 output on cp936 consoles).

## Usage

### CLI

```bash
anymd convert <INPUT...> [-o PATH] [options]

INPUT   file, glob, directory, http(s) URL, or - (stdin)
```

```bash
anymd convert report.pdf                      # → stdout
anymd convert report.docx -o out.md           # → file
anymd convert ./docs -o md/                   # → directory tree → .md files
anymd convert "*.xlsx" --json                 # → NDJSON manifest
anymd convert scan.pdf --image-mode ocr       # OCR scanned pages (needs [ocr])
anymd convert data.csv --max-rows 50          # cap table rows
anymd convert - --stdin-filename msg.eml      # stdin with a filename hint
anymd convert report.pdf --page-range 3-7     # partial PDF
anymd --list-formats                          # everything supported + required extras
```

### Agent contract

- **stdout carries only the product.** Progress, warnings and diagnostics go to stderr.
- **Exit codes:** `0` all converted · `1` partial failure · `2` usage error · `3` zero output
  (missing/unsupported inputs). A file you name explicitly that has no converter is an
  **error** (exit 1); the same file swept up by a glob or directory walk is **skipped**
  (exit 0). `--strict` promotes those skips to errors too.
- **`--json`** emits one JSON object per line, each with a `type` field so an agent can stream-parse:

```json
{"type":"item","input":"a.docx","output":"a.md","status":"ok","converter":"docx","duration_ms":42,"bytes_in":12345,"bytes_out":678,"sha256_in":"…","warnings":[]}
{"type":"item","input":"legacy.doc","output":null,"status":"error","error":{"code":"UnsupportedFormatError","message":"legacy binary Office formats (.doc/.ppt) require LibreOffice…","hint":"run anymd --list-formats to see supported types"}}
{"type":"summary","total":2,"ok":1,"error":1,"skipped":0,"bytes_in":9000,"bytes_out":678,"duration_ms":120,"exit_code":1}
```

### Python API

```python
from anymd import convert, convert_result, list_formats, ConvertContext

md = convert("report.docx")
result = convert_result("report.xlsx")     # .metadata (sheets, pages, …) + .warnings
md = convert("https://example.com/post",   # URL (needs anymd[web])
             ctx=ConvertContext(frontmatter=True))
```

### MCP server

```bash
pip install "anymd[mcp]"
anymd-mcp        # stdio transport
```

Two tools for any MCP client:

- **`convert_document(source, max_rows, image_mode, frontmatter, max_chars, include_metadata, allow_remote)`** —
  one document → Markdown. Missing extras return an actionable `ERROR[...]` string instead of
  a protocol error; output is truncated at `max_chars` with a visible note.
- **`list_formats()`** — every supported extension and its extra.

Claude Desktop / Claude Code registration:

```json
{"mcpServers": {"anymd": {"command": "anymd-mcp"}}}
```

## Supported formats

| family | formats |
|---|---|
| Office | `.docx` `.xlsx` `.xlsm` `.xls` `.pptx` `.odt` `.ods` `.odp` `.rtf` |
| PDF | `.pdf` (structure-aware; scanned or image-heavy pages → OCR when `[ocr]` installed, LaTeX formulas when `[math]` installed) |
| Web | `.html` `.htm` `.xhtml` `.mht/.mhtml`, http(s) URLs |
| E-books | `.epub` (spine-ordered) |
| Data | `.csv` `.tsv` → tables · `.json` `.jsonl` → tables or fenced · `.yaml` `.yml` `.toml` `.xml` `.ini` `.cfg` `.conf` `.env` `.properties` → fenced |
| Email | `.eml` (headers + body + attachment listing) · `.msg` |
| Notebooks | `.ipynb` (cells in order, outputs truncated) |
| Images | `.png` `.jpg` `.jpeg` `.gif` `.bmp` `.tif` `.tiff` `.webp` (OCR or metadata) |
| Archives | `.zip` `.tar` `.tar.gz` `.tgz` `.tar.bz2` `.tar.xz` — members converted recursively |
| Text/code | `.txt` `.log` `.md` and ~60 code extensions → language-tagged fences |
| No extension | content sniffing: PDF/zip-family/OLE/RTF/HTML/images/plain text |

**Not supported (by design):** legacy binary `.doc` / `.ppt` have no pip-installable reader —
anymd fails with a precise message telling you to re-save as `.docx`/`.pptx`.
(`.xls` *is* supported via xlrd.)

## Notes

- **License:** anymd is MIT. Two optional extras carry copyleft licenses that activate only
  for their own code paths: `pymupdf` (AGPL-3.0, the `pdf` extra) and `striprtf`
  (GPL-3.0, inside the `office` extra). They are separate packages installed at runtime —
  if you deploy anymd as a hosted service, review your obligations (a pdfminer.six-based
  fallback can replace the `pdf` extra).
- **OCR engine:** [RapidOCR](https://github.com/RapidAI/RapidOCR) bundles its models
  in-wheel; nothing is downloaded at runtime. If model files fail to load in a restricted
  environment, OCR degrades to a warning, never an error.
- **LaTeX formulas (`[math]` extra):** math-dense pages are re-recognized with
  [pix2text](https://github.com/breezedeus/pix2text), which emits `$...$` LaTeX for
  formula regions (plain OCR flattens `x²` and displaces integral limits). pix2text
  downloads its models on first use — in China set `HF_ENDPOINT=https://hf-mirror.com`.
  Any failure degrades to plain OCR; this extra is heavy (pulls PyTorch) and is
  therefore not part of `anymd[all]`.
- **Adobe Symbol maths:** PDFs that embed the Adobe Symbol font frequently ship no
  `ToUnicode` map, so an extractor reports raw private-use codepoints and every operator
  vanishes. anymd decodes the common glyph set (`= − + ∫ ∑ π ≤ ∞ ∂ ∈ …`) and warns about
  whatever it could not resolve instead of dropping it quietly. Tall-delimiter fragments
  are deliberately left undecoded — see
  [`src/anymd/_symbolfont.py`](https://github.com/Ljf857/anymd/blob/main/src/anymd/_symbolfont.py) for why.
- **Mixed pages (text layer + images):** a page carrying both a real text layer and
  images that hold content gets the OCR reading appended *below* the text layer, minus
  any lines the text layer already contained. Nothing is dropped, nothing is repeated.
- **Remote inputs:** http(s) URL inputs are fetched by default in the CLI — pass
  `--no-remote` to forbid them. Private, loopback and link-local addresses are
  rejected regardless, unless you explicitly pass `--allow-private`. The MCP
  `convert_document` tool is stricter: every URL needs `allow_remote=True`.
- **Windows-first:** the CLI reconfigures stdout to UTF-8 so piping Chinese text through
  cp936 consoles never raises `UnicodeEncodeError`.

## Development

```bash
git clone https://github.com/Ljf857/anymd.git
cd anymd
pip install -e ".[all,dev]"

pytest                 # slow OCR tests excluded by default
pytest -m slow         # OCR end-to-end
ruff check src tests
ruff format src tests
mypy src/anymd
```

Releasing is documented in [`docs/RELEASING.md`](https://github.com/Ljf857/anymd/blob/main/docs/RELEASING.md).

Contributions are welcome — open an
[issue](https://github.com/Ljf857/anymd/issues) or pull request. See
[`CONTRIBUTING.md`](https://github.com/Ljf857/anymd/blob/main/CONTRIBUTING.md) for the ground rules.

## License

[MIT](https://github.com/Ljf857/anymd/blob/main/LICENSE) © 2026 Ljf857

Optional extras keep their own licenses (see [Notes](#notes)).