Skip to main content
Glama
README.md
# MCP_Documents

**A self-hosted MCP server for reading, extracting from and manipulating
documents — PDF first, but not PDF only.**

The seventh repo in the `MCP_*` fleet, and the one that closes the *research*
leg of the research → analytics → reporting path the fleet exists to serve.

> **Release [`v0.2.0`](https://github.com/azzindani/MCP_Documents/releases/tag/v0.2.0)
> — the second tagged release.** All 13 tools are implemented and deployed; CI is
> green on Ubuntu, macOS and Windows. Source only: no wheel and no container
> image are published, so build the image yourself from the `Dockerfile` here.
> The documents in `docs/` remain the contract the implementation satisfies, and
> `CLAUDE.md` is the rulebook for anyone changing it.

---

## Why this exists

The commercial PDF sites are upload-first. The documents people actually run
through them are contracts, invoices, payslips, medical and legal records.

**This does the same work and nothing leaves the machine.** No GPU, no cloud
API, no model weights, no subscription — and it works offline.

---

## What it does

**Extraction from documents too large to read.** A 500-page PDF is roughly
250,000 tokens; the agent driving it has about 10,000. So the server does not
return documents, it makes them addressable:

```
probe    what is this — pages, scanned or digital, where the structure is
find     WHERE something is — locations and counts, never the content
extract  one region you chose, cleaned, with a note on how it was obtained
```

That path is what lets an agent answer a question about a 500-page bundle inside
a small context. With a regex and named groups, `find` over 300 pages returns
**rows** rather than prose — which is what the sibling data server loads.

**Manipulation**, the operations a PDF site offers, done locally: assemble
(merge / split / reorder / rotate in one grammar), convert, compress, repair,
OCR, protect, redact.

**Any document, not just PDF.** One reader per format normalising into a single
internal model, so every tool works the same on PDF, HTML, `.docx`, `.xlsx`,
`.pptx`, `.eml`, `.epub`, `.xbrl`, markdown and plain text. With URL fetching
enabled, every path argument also accepts a link — the same call, whether the
HTML came from disk or the web.

**Bundles open too.** A `.zip` reads as its manifest, and a member is read by
naming it — `probe("filing.zip::instance.xbrl")` — so a filing that arrives as
an archive does not have to be unpacked by hand first. A member this server does
not read, or one wanted as a file, is saved as it is with
`convert(source="filing.zip::data.csv", to="file")`, after the same size guards.
Reading a member writes nothing where outputs go: it is unpacked into a cache
private to the server (`DOCS_MEMBER_CACHE`, default under the system temp folder).

**XBRL figures come back `native`.** Every other format's numbers are recovered
from layout and carry a confidence to match; an XBRL instance states its facts
in machine-readable fields, and the response says so. Values are reported
exactly as filed and never rescaled.

---

## The 13 tools

```
docs-read   probe · outline · find · extract · extract_tables · read_page · to_markdown
docs-edit   assemble · convert · optimize · ocr · protect · redact
```

Thirteen, not the twenty-five a PDF website shows, because that number is a
property of user interfaces — a button cannot take an argument and an agent's
verb can. `assemble` alone covers merge, split, extract pages, remove pages,
organise and rotate.

### One endpoint, two tools

`/mcp` serves all thirteen as two domain tools, `docs_read` and `docs_edit`,
each taking an `action` (one of the verbs above, by its own name) and an
`args` object whose every property says which actions take it:

```json
{"action": "outline", "args": {"source": "report.pdf"}}
```

Each action runs the tier's own tool, so validation and answers are identical.
An action asked of the wrong tool is pointed at the right one; an argument it
does not take is refused by name. `/read/mcp` and `/edit/mcp` keep serving.

---

## Documentation

| File | What is in it |
|---|---|
| `CLAUDE.md` | The rules. Read this first if you are an agent working here. |
| `docs/ARCHITECTURE.md` | The three-step path, the intermediate representation, provenance, budgets |
| `docs/SCHEMA.md` | Every tool's signature, response shape and refusals |
| `docs/TECH_STACK.md` | Libraries, licences, external binaries, the container budget |
| `docs/DECISIONS.md` | What was rejected and why — read before proposing a change |

---

## Two things worth knowing before you use it

**Reconstruction announces itself.** A PDF is glyphs at coordinates — paragraphs,
tables, reading order and headings are all *inferred*. Every extraction carries a
`basis` field saying how it was obtained: a table found from ruling lines and one
guessed from column gaps do not get the same confidence, and a page that is an
un-OCR'd scan says so instead of returning nothing.

**PDF → Word/PowerPoint is reconstruction, not conversion.** The commercial sites
use commercial engines and there is no CPU-only open-source path to that quality.
This ships it, labels it, and tells you when a document is a poor candidate.

---

## Install

Requires Python **3.14** and `uv`. Set `MCP_CONSTRAINED_MODE=1` on small
hardware to tighten every budget. A caller's regular expression, in `find` and
in `redact`, runs in a worker that is stopped after `DOCS_REGEX_SECONDS` of
matching (10, 5 constrained) and refused by name.

### Local, as a stdio server

```bash
uv sync
uv run python servers/docs_read/server.py      # 7 read tools
uv run python servers/docs_edit/server.py      # 6 edit tools
```

Two entries in your client's `mcp.json`, one per tier. Everything runs on the
CPU with no network; `convert(to='pdf')` needs LibreOffice and `ocr()` needs
Tesseract, and both say so by name when they are missing rather than failing
inside a subprocess.

### Docker, as a remote endpoint

One container, both tiers on one port, so the PDF stack loads once:

```bash
cp tokens.example.json tokens.json          # or use DOCS_API_KEY
mkdir -p oauth-state shared-files && sudo chown -R 999:999 oauth-state shared-files tokens.json
docker compose up -d --build

curl http://localhost:8850/health            # aggregate
curl http://localhost:8850/read/health       # per tier
# connect /mcp for the two domain tools, or /read/mcp and /edit/mcp for the tiers
```

The image carries LibreOffice and Tesseract. It does **not** carry Ghostscript
— that is a licence decision, not an omission, and `optimize()` reports the
capability it therefore lacks (see `docs/DECISIONS.md` §11). Build with
`--build-arg INSTALL_GHOSTSCRIPT=1` if you accept AGPL for your own deployment.

Mounts are `/read/mcp` and `/edit/mcp`. Auth is bearer-token, by precedence:
`DOCS_TOKENS_FILE` > `DOCS_TOKENS` > `DOCS_API_KEY` > open. **Open mode is for
localhost only** — a reachable deployment with no token set has no auth at all.
Set `DOCS_PUBLIC_URL` to the public origin, or OAuth discovery falls back to the
internal bind address and no remote client can complete it.

To give a caller a link rather than a path inside the container, point
`MCP_SHARED_DIR` at a directory your file server serves and set
`MCP_PUBLIC_BASE_URL` to its URL; every produced file then comes back with a
`public_url`. `MCP_FETCH_URLS=1` additionally lets any `source` argument be an
http(s) link — off by default, and private, loopback and cloud-metadata
addresses are refused even when it is on.
A Google Drive, Docs, Dropbox, GitHub or GitLab share link is read as the file
it points to, and a web page served where a file was asked for (a link that is
not public answers with a sign-in page) is refused, not parsed. A path from the
caller's side -- a chat's sandbox such as `/mnt/user-data/…` -- is refused by
name, with the ways to bring the file here.

Where a `source` goes, a file's bytes may go instead:
`data:application/pdf;name=report.pdf;base64,<bytes>` is saved to
`MCP_OUTPUT_DIR/inbox/report.pdf` before the tool runs. It is for a caller
whose file sits in its own sandbox (a claude.ai upload) with no link to give,
capped at `MCP_MAX_INLINE_MB` (default 10); the same bytes sent twice are one
file, and a taken name is never overwritten.
A bigger file goes in parts: add `part=2/5;sha256=<of the whole file>` to each.
The tool answers `tool_ran: false` with the parts still missing until the last
lands, then runs on the joined, checked file (`MCP_MAX_UPLOAD_MB`, default
100; an upload left unfinished for an hour is dropped).

Upload URLs are off by default. With `MCP_UPLOAD_URLS=1` and `MCP_UPLOAD_BASE_URL`
(this server's public origin), the refusal for a path on the caller's side
carries a URL minted for that file: `curl -T <file> '<url>'` from the sandbox
writes it to `MCP_OUTPUT_DIR/inbox/`, and the answer is the path to pass. The
bytes never pass through the model. A URL writes one file, once, within 15
minutes, up to `MCP_MAX_UPLOAD_MB`, under the name fixed when it was minted; a
forged, expired or spent one writes nothing. The route takes no API key -- its
signed token (`MCP_UPLOAD_SECRET`, else a key made per process) is the
credential -- so turning it on is the operator's decision.

### Checking a deployment

```bash
uv run python -m pytest tests/ -q                     # 497 offline tests
DOMAIN=http://localhost:8850 ./remote_smoke_test.sh   # all 13 tools over HTTP
```

The smoke test is the only thing that exercises LibreOffice and Tesseract, and
it is worth more than its size suggests: it found six defects the whole offline
suite did not, because it is the only check that hands these tools a document
real software produced. `DOMAIN` has no default on purpose — no hostname
appears anywhere in this repo.

TDQS

B3.3/5.0

Scored across 7 tools

Disambiguation4/5

The tools are mostly distinct: find locates snippets, read_page returns one page, extract handles a page range, extract_tables targets tables, outline exposes structure, probe identifies the document, and to_markdown does full conversion. The only mild ambiguity is between extract and to_markdown, since both return document text and could be selected for similar high-level tasks.

Naming Consistency2/5

Naming conventions are inconsistent: several bare verbs (find, extract, outline, probe), two verb_noun compounds (read_page, extract_tables), and one prepositional name (to_markdown). There is no shared prefix or pattern, though each name is readable on its own.

Tool Count5/5

Seven tools is appropriate for a document-reading server: each one addresses a distinct need such as search, page reading, range extraction, tables, outline, probe, and conversion. The count feels neither thin nor bloated.

Completeness4/5

The set covers the main document workflow: identify, outline, search, read single pages, extract text and tables, and convert to Markdown. Minor gaps remain: scanned documents have no explicit OCR path, and to_markdown's token-budget refusal may require manual page-range reconstruction for very large documents.

Maintenance

ActivityMaintained
ResponsivenessNo issues