MCP_Documents
Supports Markdown documents as a first-class format, allowing the server's reading, extraction, and manipulation tools to operate on Markdown files, and enabling conversion between Markdown and other document formats.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP_DocumentsProbe ~/Downloads/invoice.pdf and tell me if it's scanned."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP_Documents
A self-hosted MCP server for reading, extracting from and manipulating documents — PDF first, but not PDF only.
The seventh repo in the MCP_* fleet, and the one that closes the research
leg of the research → analytics → reporting path the fleet exists to serve.
Status: design complete, not yet implemented. See
CLAUDE.md§14 for the build order. The documents indocs/are the contract the implementation must satisfy.
Why this exists
The commercial PDF sites are upload-first. The documents people actually run through them are contracts, invoices, payslips, medical and legal records.
This does the same work and nothing leaves the machine. No GPU, no cloud API, no model weights, no subscription — and it works offline.
Related MCP server: document-intelligence-mcp
What it does
Extraction from documents too large to read. A 500-page PDF is roughly 250,000 tokens; the agent driving it has about 10,000. So the server does not return documents, it makes them addressable:
probe what is this — pages, scanned or digital, where the structure is
find WHERE something is — locations and counts, never the content
extract one region you chose, cleaned, with a note on how it was obtainedThat path is what lets an agent answer a question about a 500-page bundle inside
a small context. With a regex and named groups, find over 300 pages returns
rows rather than prose — which is what the sibling data server loads.
Manipulation, the operations a PDF site offers, done locally: assemble (merge / split / reorder / rotate in one grammar), convert, compress, repair, OCR, protect, redact.
Any document, not just PDF. One reader per format normalising into a single
internal model, so every tool works the same on PDF, HTML, .docx, .xlsx,
.pptx, .eml, .epub, .xbrl, markdown and plain text. With URL fetching
enabled, every path argument also accepts a link — the same call, whether the
HTML came from disk or the web.
Bundles open too. A .zip reads as its manifest, and a member is read by
naming it — probe("filing.zip::instance.xbrl") — so a filing that arrives as
an archive does not have to be unpacked by hand first.
XBRL figures come back native. Every other format's numbers are recovered
from layout and carry a confidence to match; an XBRL instance states its facts
in machine-readable fields, and the response says so. Values are reported
exactly as filed and never rescaled.
The 13 tools
docs-read probe · outline · find · extract · extract_tables · read_page · to_markdown
docs-edit assemble · convert · optimize · ocr · protect · redactThirteen, not the twenty-five a PDF website shows, because that number is a
property of user interfaces — a button cannot take an argument and an agent's
verb can. assemble alone covers merge, split, extract pages, remove pages,
organise and rotate.
Documentation
File | What is in it |
| The rules. Read this first if you are an agent working here. |
| The three-step path, the intermediate representation, provenance, budgets |
| Every tool's signature, response shape and refusals |
| Libraries, licences, external binaries, the container budget |
| What was rejected and why — read before proposing a change |
Two things worth knowing before you use it
Reconstruction announces itself. A PDF is glyphs at coordinates — paragraphs,
tables, reading order and headings are all inferred. Every extraction carries a
basis field saying how it was obtained: a table found from ruling lines and one
guessed from column gaps do not get the same confidence, and a page that is an
un-OCR'd scan says so instead of returning nothing.
PDF → Word/PowerPoint is reconstruction, not conversion. The commercial sites use commercial engines and there is no CPU-only open-source path to that quality. This ships it, labels it, and tells you when a document is a poor candidate.
Install
Requires Python 3.14 and uv. Set MCP_CONSTRAINED_MODE=1 on small
hardware to tighten every budget.
Local, as a stdio server
uv sync
uv run python servers/docs_read/server.py # 7 read tools
uv run python servers/docs_edit/server.py # 6 edit toolsTwo entries in your client's mcp.json, one per tier. Everything runs on the
CPU with no network; convert(to='pdf') needs LibreOffice and ocr() needs
Tesseract, and both say so by name when they are missing rather than failing
inside a subprocess.
Docker, as a remote endpoint
One container, both tiers on one port, so the PDF stack loads once:
cp tokens.example.json tokens.json # or use DOCS_API_KEY
mkdir -p oauth-state shared-files && sudo chown -R 999:999 oauth-state shared-files tokens.json
docker compose up -d --build
curl http://localhost:8850/health # aggregate
curl http://localhost:8850/read/health # per tierThe image carries LibreOffice and Tesseract. It does not carry Ghostscript
— that is a licence decision, not an omission, and optimize() reports the
capability it therefore lacks (see docs/DECISIONS.md §11). Build with
--build-arg INSTALL_GHOSTSCRIPT=1 if you accept AGPL for your own deployment.
Mounts are /read/mcp and /edit/mcp. Auth is bearer-token, by precedence:
DOCS_TOKENS_FILE > DOCS_TOKENS > DOCS_API_KEY > open. Open mode is for
localhost only — a reachable deployment with no token set has no auth at all.
Set DOCS_PUBLIC_URL to the public origin, or OAuth discovery falls back to the
internal bind address and no remote client can complete it.
To give a caller a link rather than a path inside the container, point
MCP_SHARED_DIR at a directory your file server serves and set
MCP_PUBLIC_BASE_URL to its URL; every produced file then comes back with a
public_url. MCP_FETCH_URLS=1 additionally lets any source argument be an
http(s) link — off by default, and private, loopback and cloud-metadata
addresses are refused even when it is on.
Checking a deployment
uv run python -m pytest tests/ -q # 205 offline tests
DOMAIN=http://localhost:8850 ./remote_smoke_test.sh # all 13 tools over HTTPThe smoke test is the only thing that exercises LibreOffice and Tesseract, and
it is worth more than its size suggests: it found six defects that 145 passing
tests did not, because it is the only check that hands these tools a document
real software produced. DOMAIN has no default on purpose — no hostname
appears anywhere in this repo.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI-driven PDF document processing including PDF to Markdown conversion, intelligent text and table extraction, image extraction, format conversion between PDF/Word/Markdown, batch processing, and fuzzy search - optimized for LLM context and RAG workflows.2MIT
- AlicenseNot gradedqualityDmaintenanceLocal document intelligence for AI agents — extract text, detect tables, read metadata, analyze structure, search keywords, and detect language from PDF and DOCX files. No cloud API required, no API key needed.MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to interact with local documents (PDF, Markdown, TXT) through tools for discovery, reading, extraction, summarization, comparison, keyword extraction, search, and analysis, ensuring privacy and offline capability.
- FlicenseNot gradedqualityAmaintenanceEnables AI agents to perform comprehensive PDF operations locally, including compression, text extraction, PII redaction, page organization, splitting, merging, watermarking, creation, and form filling, all without cloud uploads.
Related MCP Connectors
Markdown in, any format out. PDFs merged, split, watermarked. Runs on our own doc engines.
Turn any PDF into structured JSON via AI + OCR: invoices, bank statements, contracts.
Document API for AI-native software: render PDFs, e-sign, PAdES-seal, and verify.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/azzindani/MCP_Documents'
If you have feedback or need assistance with the MCP directory API, please join our Discord server