DOCX MCP Server
# DOCX MCP Server
Universal DOCX processing server implementing the Model Context Protocol (MCP) with full OOXML support.
## Features
- **Full OOXML Support**: Read/write all DOCX document parts (document.xml, styles, numbering, headers/footers, etc.)
- **Text Operations**: Extract, find, and replace text with literal or regex modes
- **Table Editing**: Insert/delete rows and columns, merge cells, set cell content
- **Structured Data Tags (SDT)**: Get/set content controls by tag or alias
- **Images**: List images, insert inline or anchored images with position/size control
- **Comments & Changes**: List comments, add/delete comments, accept all tracked changes
- **Document Properties**: Read/write metadata (title, author, subject, etc.)
- **LRU Caching**: Memory-efficient caching of document parts
- **Stdio Transport**: MCP communication via stdin/stdout
## Installation
```bash
npm install
npm run build
```
## Running
### Development
```bash
npm run dev
```
### Production
```bash
npm start
```
### With Claude Code
```bash
claude mcp add --scope user --transport stdio docx -- node /path/to/dist/index.js
```
## Tools
### Document Management
#### docx.open
Open a DOCX document from file or base64
```json
{
"docId": "uuid",
"parts": ["word/document.xml", "word/styles.xml", ...],
"partCount": 42,
"props": { "core": {...}, "app": {...} }
}
```
#### docx.close
Close and unload a document
#### docx.save
Save document to file or return as base64
### Part Management
#### docx.list_parts
List all parts in the document
#### docx.part_read
Read raw XML of a specific part
#### docx.part_write
Write/update XML content of a part
### Text Operations
#### docx.get_text
Extract all text from document
```bash
{
"docId": "uuid",
"scope": "document" | "headers" | "footers" | "all"
}
```
#### docx.find
Search for text with context
```bash
{
"docId": "uuid",
"query": "search term",
"mode": "literal" | "regex"
}
```
#### docx.replace_text
Replace text in document
```bash
{
"docId": "uuid",
"match": "old text",
"replace": "new text",
"mode": "literal" | "regex"
}
```
### Tables
#### docx.tables_list
List all tables with dimensions
```bash
{
"tables": [
{
"xpath": "/w:document/w:body/w:tbl[1]",
"rows": 3,
"colsApprox": 4
}
]
}
```
#### docx.table_edit
Modify table structure and content
```bash
{
"docId": "uuid",
"tableXPath": "/w:document/w:body/w:tbl[1]",
"op": {
"kind": "setCellText",
"row": 0,
"col": 1,
"text": "new value"
}
}
```
Operations:
- `setCellText(row, col, text)` - Set cell content
- `insertRow(at)` - Insert row at position
- `deleteRow(at)` - Delete row
- `insertCol(at)` - Insert column
- `deleteCol(at)` - Delete column
### Structured Data (SDT)
#### docx.sdt_get
Get content control content by tag or alias
#### docx.sdt_put
Update content control
### Images
#### docx.images_list
List all images with metadata
#### docx.image_add
Insert image inline or anchored
### Styles & Numbering
#### docx.styles_get / docx.styles_set
Read/write styles.xml
#### docx.numbering_get / docx.numbering_set
Read/write numbering.xml
### Headers/Footers
#### docx.headers_footers_list
List all header/footer parts
### Comments
#### docx.comments_list
List all comments
#### docx.comments_add
Add new comment
#### docx.changes_accept_all
Accept all tracked changes in document
### Metadata
#### docx.metadata_get
Get document properties (title, author, created, modified, etc.)
## Test Scenarios
### 1. Basic Read/Write
```bash
# Open document
docx.open: { "path": "/path/to/document.docx" }
# Get text
docx.get_text: { "docId": "returned-id" }
# Replace text
docx.replace_text: {
"docId": "returned-id",
"match": "old text",
"replace": "new text"
}
# Save
docx.save: { "docId": "returned-id", "returnBase64": true }
```
### 2. Table Manipulation
```bash
# List tables
docx.tables_list: { "docId": "id" }
# Edit cell
docx.table_edit: {
"docId": "id",
"tableXPath": "/w:document/w:body/w:tbl[1]",
"op": { "kind": "setCellText", "row": 0, "col": 0, "text": "Hello" }
}
```
### 3. Images
```bash
# List images
docx.images_list: { "docId": "id" }
```
### 4. Track Changes
```bash
# Accept all changes
docx.changes_accept_all: { "docId": "id" }
```
## Architecture
```
src/
index.ts # Entry point
errors.ts # Error definitions
logger.ts # Logging utilities
ooxml/
namespaces.ts # XML namespace definitions
emu.ts # EMU conversion utilities
dom.ts # XML DOM utilities (xmldom + fontoxpath)
xmlParser.ts # fast-xml-parser wrapper
parts.ts # DOCX ZIP part management
rels.ts # Relationships management
text.ts # Text extraction & replacement
tables.ts # Table operations
sdt.ts # Structured Data Tags
drawings.ts # Images & DrawingML
headersFooters.ts # Headers/Footers
styles.ts # Style operations
numbering.ts # Numbering operations
changes.ts # Track changes
comments.ts # Comments
store/
types.ts # Type definitions
docStore.ts # Document store + LRU cache
mcp/
schemas.ts # Tool input schemas
tools.ts # Tool implementations
server.ts # MCP server setup
```
## Dependencies
- `@modelcontextprotocol/sdk` - MCP implementation
- `jszip` - ZIP archive handling
- `fast-xml-parser` - Lossless XML parsing
- `@xmldom/xmldom` - DOM implementation
- `fontoxpath` - XPath queries
- `diff-match-patch` - Text diffing
- `lru-cache` - Memory-efficient caching
- `uuid` - Document ID generation
## Performance Notes
- Documents up to 10 MB supported
- LRU cache with 100-part limit and 1 GB memory cap
- Parts loaded on-demand, not fully into memory
- Dirty-part optimization: only modified parts saved to ZIP
- No deep copying of XML structures
## Limitations
- Headers/footers: basic support (complex section structures may need manual adjustment)
- Comments: basic list/add/delete (reply chains not fully supported)
- Track changes: accept-all available; detailed change inspection limited
- Styles: get/set full XML; no selective style merging
- EMU/sizing: calculated but rendered geometry depends on Word's layout engine
## License
MIT
TDQS
Scored across 34 tools
Many tools have distinct targets, but there is meaningful overlap between find and text_find_locations, get_text and text_runs_analyze, and find_empty_table_rows and get_table_content. The descriptions help clarify intent, but an agent could easily select the wrong tool for text or table inspection tasks.
All tools share the docx. prefix, which helps, but the naming pattern is inconsistent: some are verb_noun (list_parts, get_text), some are noun_verb (metadata_get, tables_list), and a few are bare verbs (open, save, find). Synonyms like get/list/read and set/put/write are also mixed.
With 34 tools, this exceeds the 25+ threshold and feels heavy. Several tools are highly granular and could be consolidated, such as find_empty_table_rows versus get_table_content, or find versus text_find_locations, without sacrificing clarity.
The server covers a wide range of DOCX operations: text, tables, comments, styles, headers/footers, SDT content, and validation. However, obvious gaps remain, including no create-from-scratch operation, no paragraph insertion, no table creation, and no image insertion or replacement.