document-intelligence-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@document-intelligence-mcpExtract text from quarterly_report.pdf"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
document-intelligence-mcp
Local document intelligence for AI agents — extract text, detect tables, read metadata, analyze structure, search keywords, and detect language from PDF and DOCX files. No cloud API required, no API key needed.
Features
10 MCP Tools for PDF and DOCX processing
Local processing — no data leaves your machine
No API key required
Supports PDF (via PyMuPDF + pdfplumber) and Microsoft Word DOCX (via python-docx)
Language detection via langdetect (55+ languages)
Related MCP server: @paperjsx/mcp-server
Tools
Tool | Description |
| Extract all text from a PDF, page by page |
| Detect and extract tables from PDF |
| Read PDF metadata: title, author, dates, outline |
| Detect headings, font sizes, section structure |
| Search for keywords with context in PDF |
| Extract all text from a Word DOCX file |
| Extract all tables from a DOCX file |
| Analyze headings, styles, and structure of DOCX |
| Word count, sentence count, reading time, top words |
| Detect language of PDF or DOCX (55+ languages) |
Installation
pip install document-intelligence-mcpClaude Desktop Configuration
Add to your claude_desktop_config.json:
{
"mcpServers": {
"document-intelligence": {
"command": "document-intelligence-mcp"
}
}
}Usage Examples
Extract text from a PDF:
Extract the text from /path/to/report.pdfFind tables in a PDF:
Find all tables in /path/to/financial_report.pdfSearch for a keyword:
Search for "revenue" in /path/to/annual_report.pdfGet document stats:
Count the words and estimate reading time for /path/to/document.docxDetect language:
What language is /path/to/document.pdf written in?Requirements
Python 3.10+
PyMuPDF >= 1.24.0
pdfplumber >= 0.11.0
python-docx >= 1.1.0
langdetect >= 1.0.9
License
MIT License — free to use, modify, and distribute.
Built by AiAgentKarl | Part of the AI Agent Economy toolkit
This server cannot be deployed
Maintenance
Related MCP Connectors
Agent-native document parsing: PDF, scans and FR/EU invoices to structured JSON or Markdown.
PDF, image, video, OCR, screenshot, SQL, QR and text tools for agents. No API key, no signup.
Ingest, manage, and retrieve documents for RAG-powered AI applications
Turn any PDF into structured JSON via AI + OCR: invoices, bank statements, contracts.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceA local document processing toolkit for AI agents that extracts text, converts PDFs to Markdown, merges files, extracts tables, and summarizes documents without external API dependencies.4 npm98 PyPIMIT
- AlicenseNot gradedqualityFmaintenanceEnables AI agents to generate a variety of documents (PPTX presentations, DOCX reports, PDF invoices, XLSX spreadsheets) locally from JSON specs, without any API keys or network calls.28 npm1MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to process files locally — OCR images, extract text from PDFs and DOCX, and describe images using local vision models, all without sending data to external services.-
- AlicenseAqualityBmaintenanceIndexes local documents (PDF, Word, Markdown, text) into a SQLite database for AI agents to search and retrieve bounded, source-located passages. Runs fully locally with optional OCR, preserving privacy.5MIT