Docalyze
README.md
<!-- mcp-name: io.github.LunarPerovskite/docalyze -->
<p align="center">
<img src="logo.svg" alt="Docalyze Logo" width="150" height="150">
</p>
# Docalyze MCP Server
An MCP (Model Context Protocol) server that lets AI assistants read and visually analyze local documents — PDFs, Excel spreadsheets, CSV files, Word documents, PowerPoint presentations, and images.
No API keys required. The host AI (GitHub Copilot, Claude, etc.) does all the reasoning directly.
## Supported Formats
| Format | Extensions | Read | Visual |
|--------|-----------|:----:|:------:|
| PDF | `.pdf` | ✅ | ✅ |
| Excel | `.xlsx`, `.xls` | ✅ | ✅ |
| CSV / TSV | `.csv`, `.tsv` | ✅ | — |
| JSON | `.json` | ✅ | — |
| Word | `.docx` | ✅ | ✅ |
| PowerPoint | `.pptx` | ✅ | ✅ |
| Plain text | `.txt`, `.md` | ✅ | — |
| Images | `.png`, `.jpg`, `.jpeg`, `.gif`, `.bmp`, `.tiff`, `.webp` | — | ✅ |
## Tools
| Tool | Description |
|------|-------------|
| `list_documents` | List files under a directory, filtered by glob pattern |
| `document_info` | Get metadata (size, modified date, sheets) for a file |
| `read_document` | Extract text content from a document with pagination |
| `visual_evaluate_document` | Return page images inline so the AI can analyze charts, tables, and diagrams |
## Installation
### From VS Code (recommended)
Search for **docalyze** in the MCP server gallery (Extensions sidebar → MCP tab) and click Install.
### From PyPI
```bash
pip install docalyze-mcp-server
```
### From npm
```bash
npx docalyze-mcp-server
```
This requires [uv](https://docs.astral.sh/uv/) or pipx installed — the npm wrapper calls `uvx` to run the Python package automatically.
### Manual setup
Add to your VS Code `mcp.json` (or `settings.json`):
```jsonc
{
"servers": {
"docalyze": {
"type": "stdio",
"command": "python",
"args": ["-m", "docalyze_mcp_server"],
"env": {
"PYTHONIOENCODING": "utf-8"
}
}
}
}
```
Or, if you installed via pip and want to use the entry point:
```jsonc
{
"servers": {
"docalyze": {
"type": "stdio",
"command": "docalyze-mcp-server"
}
}
}
```
## Optional Dependencies
The base install handles PDF, Excel, CSV, JSON, and plain text. For additional formats:
```bash
# Word documents
pip install docalyze-mcp-server[docx]
# PowerPoint
pip install docalyze-mcp-server[pptx]
# OCR (requires Tesseract installed on your system)
pip install docalyze-mcp-server[ocr]
# Everything
pip install docalyze-mcp-server[all]
```
## Configuration
The server reads documents from a configurable root directory. Set the `DOCUMENTS_ROOT` environment variable to change it:
```jsonc
{
"servers": {
"docalyze": {
"type": "stdio",
"command": "docalyze-mcp-server",
"env": {
"DOCUMENTS_ROOT": "/path/to/your/documents"
}
}
}
}
```
If not set, it defaults to the directory containing the server script.
## License
MIT
TDQS
B3.2/5.0
Scored across 4 tools
Disambiguation5/5
Each tool has a distinct purpose: metadata retrieval, listing, text reading, and visual analysis. No two tools overlap in functionality.
Naming Consistency4/5
All tools use snake_case and end with 'document' or 'documents', but 'document_info' is a noun-noun pattern while others are verb-noun, causing minor inconsistency.
Tool Count5/5
Four tools is a reasonable scope for a document analysis server, covering essential read operations without being excessive.
Completeness4/5
The tools cover metadata, listing, text reading, and visual analysis. However, folder navigation or search is missing, which may be needed given the mention of directory structure.
Maintenance
ActivityInactive
ResponsivenessUnresponsive