verified-docx-mcp
Provides tools for working with .docx files stored on iCloud Drive, including PDF export via Word automation and status reporting for lock files and iCloud sync state.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@verified-docx-mcpExport my draft.docx to PDF and show the page count."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
verified-docx-mcp
An MCP server for local .docx files with verified writes. It is the
docx counterpart to
verified-googledocs-mcp:
the same evidence-first contract (every mutating tool re-reads the file
after writing and returns before/after evidence, never a bare "success"),
applied to a local Word document on a synced drive (OneDrive, iCloud Drive)
instead of a Google Doc.
Status
Early scaffold (issue #28 WP-02 of the build plan). Reading and writing a
.docx package needs no Word installation at all — this server manipulates
OOXML directly. Word is required only for export_pdf (rendering a page-
accurate PDF), and that additionally needs a macOS Automation grant for the
app hosting this server's process.
Tools implemented so far:
export_pdf(path, output_path)— render a.docxto PDF via Microsoft Word automation and report its page count.lock_status(path)— report Word/LibreOffice owner-file presence and sync-quiesce state, as data only. Never refuses.
More tools (document reads, markdown-based writes, tables, comments, tracked changes) land in later work packages of the same plan.
Related MCP server: DOCX-MCP
Install
Requires Python 3.12+ and uv.
uv syncRun
As an MCP server (stdio transport), typically registered by a client rather than run directly:
uv run --directory "<path-to-this-clone>" verified-docx-mcpDiagnose the Word render path on this machine (run from the SAME application that will host the MCP server — macOS grants Automation, and file access, per hosting app, not per terminal emulator in general):
uv run --directory "<path-to-this-clone>" verified-docx-mcp doctorTest
uv run pytest tests/unittests/unit is offline and runs against committed fixture .docx/.pdf
files — no Word installation required. tests/live needs a real Word
install and a macOS Automation grant; it is skipped unless --run-live is
passed.
Path safety
Every tool resolves its path argument through an allowlist
(VERIFIED_DOCX_MCP_ALLOWED_FILE_ROOTS, defaulting to the user's home
directory) and a denylist of well-known credential locations
(~/.ssh, ~/.aws, etc.) that is never overridable. See
src/verified_docx_mcp/paths.py.
License
MIT — see LICENSE.
Available Tools
2 toolsexport_pdfA
Render a local .docx to PDF via Microsoft Word and report its page count.
output_path must fall inside VERIFIED_DOCX_MCP_ALLOWED_FILE_ROOTS (defaults to the user's home directory), must not resolve to a credential path, and its parent directory must already exist. This is a read/render tool: the source .docx is never modified (Word opens a private staged copy — see render.py), so the return value has no "applied" key.
Returns pdf_path, sha256, page_count (best-effort; None — never a guessed 0 — when it cannot be determined), page_count_source ("word"|"pdfinfo"|"regex"|None), engine ("word"; the only engine, D2: no LibreOffice), and left_open_document (the rendered document's window title in Word — see render.py's module docstring for why it is never auto-closed).
Errors:
INVALID_INPUT - a bad path or output_path
RENDER_ENGINE_UNAVAILABLE - no render engine available (Word only)
AUTOMATION_NOT_GRANTED - macOS declined Automation control of Word
for this app; run verified-docx-mcp doctor for the fix, scoped to the app
hosting this MCP server's own process
WORD_SANDBOX_UNAVAILABLE - Word has never been launched on this
machine (its sandbox container does not
exist yet)
RENDER_FAILED - any other Word automation failure
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers: it discloses that the source .docx is never modified (Word opens a private staged copy), that page_count is best-effort and never a guessed 0, that left_open_document is never auto-closed (with a pointer to render.py for why), and enumerates all error conditions with their meanings. This is exemplary behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with clear section breaks for return values and errors. It is longer than typical, but every sentence earns its place by disclosing constraints, behavior, or error handling. It is front-loaded with the core purpose and then systematically covers edge cases. Slightly verbose in the error section, but justified for a tool with complex failure modes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (Word automation, sandbox, macOS permissions, best-effort page count), the description covers all critical aspects: input constraints, output format, error taxonomy, and behavioral quirks. The output schema exists, so return values are further structured. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not explicitly define 'path' beyond 'a local .docx', but it thoroughly defines output_path constraints and the return value semantics. The description adds substantial meaning to output_path and the overall behavior, though 'path' itself is only implicitly described as the source .docx. This is strong compensation for a 0% coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Render a local .docx to PDF via Microsoft Word and report its page count.' This clearly distinguishes it from the only sibling (lock_status), which is about lock state, not document conversion. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states constraints on output_path (must fall inside VERIFIED_DOCX_MCP_ALLOWED_FILE_ROOTS, must not resolve to a credential path, parent must exist), and clarifies this is a read/render tool, not a mutating one. It also names the only engine (Word) and explicitly says no LibreOffice, which is a clear exclusion. This is strong when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lock_statusA
Report Word/LibreOffice owner-file presence and sync-quiesce state.
Data only — never refuses. Owner-file detection recognizes Word's "~$" owner file (matched by shared filename suffix, not by constructing one exact expected name — see _word_owner_file_matches's docstring) and LibreOffice's ".~lock.#". Sync-quiesce takes two (size, mtime_ns) samples ~1.5s apart and also checks for ".tmp"/".~*" siblings in the same directory (OneDrive/iCloud staging artifacts); "sync_quiesced" is false if either signal suggests the file is still being written.
Returns path, owner_file ({present, path, format, owner_name}), sync_quiesced, sync_detail ({sample_1, sample_2, interval_seconds, tmp_siblings}).
Errors: INVALID_INPUT - path does not exist or is outside the allowed roots
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility. It explicitly states 'Data only — never refuses,' discloses the detection algorithms (owner-file patterns, two-sample timing, tmp-sibling checks), and even references an internal docstring. It also lists the exact error condition (INVALID_INPUT). This is extremely transparent about behavior beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is thorough but somewhat long and dense, including implementation details like reference to an internal docstring and naming sample intervals. It front-loads the purpose and then dives into specifics. While every sentence provides value, the length might exceed what is immediately necessary for an agent to decide to call the tool. A more compact version could retain the key facts.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (indicated by context), yet the description still enumerates the returned fields and their structure. It also explains behavior, timing, error cases, and the distinction between owner-file detection and sync-quiesce. For a tool of this complexity (multi-part return, nuanced detection), the description covers all necessary invocation context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter, 'path,' with no schema documentation (0% coverage). The description doesn't explicitly define 'path' but the context (checking a file) makes its meaning clear. The error condition 'path does not exist or is outside the allowed roots' adds semantic boundaries. While it doesn't explicitly say 'path is the file to check,' the surrounding text implies that strongly enough.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Report Word/LibreOffice owner-file presence and sync-quiesce state.' It clearly distinguishes the tool from the sibling 'export_pdf' by naming the exact monitoring function. The scope (owner files, sync quiescence) is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking file-writing stability ('sync_quiesced' state) and notes 'Data only — never refuses,' indicating a read-only check. It does not explicitly state when to use this versus export_pdf, but the sibling's purpose is obviously different. The phrase 'Data only' hints that it is safe to call without side effects, which is useful context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
export_pdf - First observed
lock_status
TDQS
Scored across 2 tools
export_pdf and lock_status have completely different purposes: one renders a docx to PDF, the other inspects owner/lock and sync state. There is no overlap in inputs, outputs, or side effects, so an agent should never confuse them.
Both names use snake_case and are readable, but export_pdf follows an imperative verb+object pattern while lock_status is a noun phrase without a verb. Adding a verb like get_lock_status or check_lock_status would make the naming pattern consistent.
Two tools is on the thin side and feels closer to a utility script than a full MCP server. However, the server appears intentionally scoped to safe docx rendering and preflight lock inspection, so the count is borderline rather than excessive.
The two tools cover the core workflow of checking whether a docx is locked/syncing and then rendering it to PDF. Minor gaps exist—such as no explicit docx validity check or wait-for-quiesce helper—but agents can work around them with repeated lock_status calls and careful error handling.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Document API for AI-native software: render PDFs, e-sign, PAdES-seal, and verify.
Markdown in, any format out. PDFs merged, split, watermarked. Runs on our own doc engines.
Document processing over MCP: merge, split and compress PDFs, run OCR, extract document text.
Composable APIs for document extraction, image transformation, and document & sheet generation.
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables creation, editing, and management of Word documents through JSON schema with support for rich content including text formatting, tables, images, code blocks, and lists. Provides comprehensive DOCX operations including opening existing documents, modifying content, and saving files to disk.12291MIT
- AlicenseNot gradedqualityDmaintenanceEnables comprehensive Microsoft Word document manipulation through the Model Context Protocol, with advanced table operations including creation, data management, formatting, and bulk operations. Supports document creation, editing, and saving with plans for full document content management.7MIT
- FlicenseAqualityDmaintenanceEnables natural language interaction with local .docx files, allowing users to find, read, search, and summarize Word documents using friendly names and location hints.53-
- AlicenseAqualityDmaintenanceProvides comprehensive read/write access to Word documents, including comments, track changes, and reply threads.91MIT