Skip to main content
Glama
README.md
<!-- mcp-name: io.github.LunarPerovskite/docalyze -->
<p align="center">
  <img src="logo.svg" alt="Docalyze Logo" width="150" height="150">
</p>

# Docalyze MCP Server

An MCP (Model Context Protocol) server that lets AI assistants read and visually analyze local documents — PDFs, Excel spreadsheets, CSV files, Word documents, PowerPoint presentations, and images.

No API keys required. The host AI (GitHub Copilot, Claude, etc.) does all the reasoning directly.

## Supported Formats

| Format | Extensions | Read | Visual |
|--------|-----------|:----:|:------:|
| PDF | `.pdf` | ✅ | ✅ |
| Excel | `.xlsx`, `.xls` | ✅ | ✅ |
| CSV / TSV | `.csv`, `.tsv` | ✅ | — |
| JSON | `.json` | ✅ | — |
| Word | `.docx` | ✅ | ✅ |
| PowerPoint | `.pptx` | ✅ | ✅ |
| Plain text | `.txt`, `.md` | ✅ | — |
| Images | `.png`, `.jpg`, `.jpeg`, `.gif`, `.bmp`, `.tiff`, `.webp` | — | ✅ |

## Tools

| Tool | Description |
|------|-------------|
| `list_documents` | List files under a directory, filtered by glob pattern |
| `document_info` | Get metadata (size, modified date, sheets) for a file |
| `read_document` | Extract text content from a document with pagination |
| `visual_evaluate_document` | Return page images inline so the AI can analyze charts, tables, and diagrams |

## Installation

### From VS Code (recommended)

Search for **docalyze** in the MCP server gallery (Extensions sidebar → MCP tab) and click Install.

### From PyPI

```bash
pip install docalyze-mcp-server
```

### From npm

```bash
npx docalyze-mcp-server
```

This requires [uv](https://docs.astral.sh/uv/) or pipx installed — the npm wrapper calls `uvx` to run the Python package automatically.

### Manual setup

Add to your VS Code `mcp.json` (or `settings.json`):

```jsonc
{
  "servers": {
    "docalyze": {
      "type": "stdio",
      "command": "python",
      "args": ["-m", "docalyze_mcp_server"],
      "env": {
        "PYTHONIOENCODING": "utf-8"
      }
    }
  }
}
```

Or, if you installed via pip and want to use the entry point:

```jsonc
{
  "servers": {
    "docalyze": {
      "type": "stdio",
      "command": "docalyze-mcp-server"
    }
  }
}
```

## Optional Dependencies

The base install handles PDF, Excel, CSV, JSON, and plain text. For additional formats:

```bash
# Word documents
pip install docalyze-mcp-server[docx]

# PowerPoint
pip install docalyze-mcp-server[pptx]

# OCR (requires Tesseract installed on your system)
pip install docalyze-mcp-server[ocr]

# Everything
pip install docalyze-mcp-server[all]
```

## Configuration

The server reads documents from a configurable root directory. Set the `DOCUMENTS_ROOT` environment variable to change it:

```jsonc
{
  "servers": {
    "docalyze": {
      "type": "stdio",
      "command": "docalyze-mcp-server",
      "env": {
        "DOCUMENTS_ROOT": "/path/to/your/documents"
      }
    }
  }
}
```

If not set, it defaults to the directory containing the server script.

## License

MIT

TDQS

B3.2/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct purpose: metadata retrieval, listing, text reading, and visual analysis. No two tools overlap in functionality.

Naming Consistency4/5

All tools use snake_case and end with 'document' or 'documents', but 'document_info' is a noun-noun pattern while others are verb-noun, causing minor inconsistency.

Tool Count5/5

Four tools is a reasonable scope for a document analysis server, covering essential read operations without being excessive.

Completeness4/5

The tools cover metadata, listing, text reading, and visual analysis. However, folder navigation or search is missing, which may be needed given the mention of directory structure.

Maintenance

ActivityInactive
ResponsivenessUnresponsive