doc-mcp-server
by jiahuidegit
README.md
# đ Document Analyzer MCP Server
[English](README.md) | [įŽäŊ䏿](README.zh.md)
[](https://pypi.org/project/doc-mcp-server/)
[](https://opensource.org/licenses/MIT)
[](https://www.python.org/downloads/)
[](https://modelcontextprotocol.io)
> **Make AI understand complex documents** - MCP server solving AI context limitations
---
## đ¯ Key Features
- â
**Smart Document Analysis** - Auto-detect sections, handle merged cells
- â
**Multi-format Support** - Excel (.xlsx, .xls) | PDF/Word in development
- â
**Precise Field Mapping** - Field mapping table + section-level reading
- â
**High Performance** - Structured caching + lazy loading
## đ Quick Start
### Installation
**macOS / Linux (Recommended with pipx)**
```bash
# Install pipx
brew install pipx # macOS
# or sudo apt install pipx # Ubuntu/Debian
# Install doc-mcp-server
pipx install doc-mcp-server
```
**Windows**
```bash
pip install doc-mcp-server
```
For more installation options, see **[Full Installation Guide](docs/en/installation.md)**
### Configure Claude Code
Add to `~/.claude.json` or your project's config file:
```json
{
"mcpServers": {
"document-analyzer": {
"command": "doc-mcp-server"
}
}
}
```
For detailed configuration, see **[Quick Start Guide](docs/en/quickstart.md)**
## đ Full Documentation
- **[Installation Guide](docs/en/installation.md)** - Platform-specific installation steps
- **[Update Guide](docs/en/update.md)** - How to upgrade to the latest version
- **[Quick Start](docs/en/quickstart.md)** - Configuration and basic usage
- **[Usage Guide](docs/en/usage.md)** - Complete API and examples
- **[Troubleshooting](docs/en/troubleshooting.md)** - Common issues and solutions
## đĄ Usage Example
```python
# 1. Analyze document structure
analyze_document(file_path="/path/to/document.xlsx")
# 2. Read specific section
read_section(file_path="/path/to/document.xlsx", section_name="Section 1")
# 3. Read single field
read_field(file_path="/path/to/document.xlsx", field_key="Section1_CompanyName")
```
## đ ī¸ Available Tools
| Tool | Description |
|------|-------------|
| `analyze_document` | Analyze document structure and generate metadata |
| `get_structure` | Get cached document structure |
| `read_field` | Read specific field value |
| `read_section` | Read entire section data |
| `write_field` | Write field value (Excel only) |
| `list_sections` | List all sections |
| `list_fields` | List all fields |
| `export_structure` | Export document structure |
## đ¯ Why Use This?
**Problem**: Large Excel files consume massive tokens when directly read by AI
- â Traditional: Read entire 323-row Excel â 15000+ tokens â Often fails
- â
Using MCP: Structured reading â 2000 tokens â 90%+ success rate
**Performance Improvements**:
- đ Token consumption reduced by 87% (15000 â 2000)
- â
Success rate improved from 30% to 90%+
- ⥠Handles 323 rows à 24 columns with 4249 merged cells
## đ¤ Contributing & Feedback
- **Report Issues**: [GitHub Issues](https://github.com/jiahuidegit/doc-mcp-server/issues)
- **Contribute Code**: [CONTRIBUTING.md](CONTRIBUTING.md)
---
## đ License
MIT License - see [LICENSE](LICENSE) for details
---
**Made with â¤ī¸ by Yang Jiahui**
This server cannot be deployed
Maintenance
ActivityInactive
ResponsivenessNo issues