safe-docx
Safe DOCX Suite
English | Español | 简体中文 | Português (Brasil) | Deutsch
safe-docx by UseJunior — 코딩 에이전트를 활용한 문서 작업 자동화 도구입니다.
UseJunior 개발자 도구의 일부입니다.
Safe Docx는 기존 Microsoft Word .docx 파일을 정밀하게 편집하기 위한 오픈 소스 TypeScript 스택입니다. 에이전트가 변경 사항을 제안하고 사람이 서식을 유지하면서 안정적으로 문서를 편집해야 하는 워크플로우를 위해 구축되었습니다.
AI로 계약서를 검토할 때 가장 느린 단계는 Word에서 승인된 권장 사항을 적용하는 것입니다. Safe Docx는 이를 결정론적인 도구 호출로 변환합니다.
이 도구가 필요한 이유
AI 코딩 CLI는 코드 및 텍스트 파일에는 뛰어나지만 기존 .docx 편집에는 취약합니다. 비즈니스 및 법률 워크플로우는 여전히 Word 문서 기반으로 운영되므로, 다음과 같은 작업을 위한 네이티브 TypeScript 경로를 구축했습니다:
토큰 효율적인 형식으로 기존 문서 읽기 및 검색
서식을 파괴하지 않고 정밀하게 편집
깔끔한/변경 사항이 추적된 출력물 및 수정 사항 추출 아티팩트 생성
목표: 코딩 에이전트가 문서 작업도 수행할 수 있도록 지원합니다. Safe Docx는 자동화 과정에서도 서식과 검토 의미론이 유지되어야 하는 기존 Word 파일의 결정론적 편집에 중점을 둡니다.
Related MCP server: docx-mcp
포지셔닝
Safe Docx는 기존 .docx 파일에 대해 결정론적이고 로컬 우선 편집이 필요한 에이전트 워크플로우에 최적화되어 있습니다:
편집, 비교, 수정 사항 추출, 주석, 각주 및 레이아웃을 위한 타입 지정된 MCP 도구
테스트 증거 및 추적성 아티팩트를 통한 감사 가능한 동작
지원되는 사용 사례에서 Python이나 LibreOffice를 요구하지 않는 TypeScript 런타임 배포
Safe Docx는 생성 우선 .docx 라이브러리를 대체하기 위한 것이 아닙니다.
신뢰하는 기업
Am Law 상위 10대 로펌 — 다단계 계약서 번역 파이프라인
150명 규모의 지역 로펌 — 2,200만 개 이상의 토큰 계약서 마크업 처리
Gemini CLI — 호환 가능한 Word 편집 MCP 확장 기능
시작하기
npx -y @usejunior/safe-docx자세한 설정 및 도구 참조는 packages/docx-mcp/README.md를 참조하세요.
예시: 에이전트의 계약서 편집
Safe Docx가 설치된 코딩 에이전트(Claude Code, Cursor, Gemini CLI)에 프롬프트를 입력하면, 에이전트는 다음과 같은 MCP 도구 호출을 수행합니다:
User: Edit the NDA at ~/docs/NDA.docx — change the governing law
from "State of New York" to "State of Delaware" and save both
a clean copy and a tracked-changes copy.
Agent calls:
1. read_file(file_path="~/docs/NDA.docx", format="toon")
→ Returns paragraphs with stable IDs: _bk_1, _bk_2, ...
2. grep(file_path="~/docs/NDA.docx", pattern="State of New York")
→ Match in paragraph _bk_47
3. replace_text(
file_path="~/docs/NDA.docx",
target_paragraph_id="_bk_47",
old_string="State of New York",
new_string="State of Delaware",
instruction="Change governing law to Delaware"
)
4. save(
file_path="~/docs/NDA.docx",
save_to_local_path="~/docs/NDA-clean.docx",
tracked_save_to_local_path="~/docs/NDA-tracked.docx",
save_format="both"
)에이전트가 도구 호출을 자동으로 처리합니다. 사용자는 깔끔한 파일과 사람이 검토할 수 있는 변경 사항 추적 파일을 얻게 됩니다.
MCP 퀵스타트
Claude Code
claude mcp add safe-docx -- npx -y @usejunior/safe-docxClaude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json(macOS) 또는 %APPDATA%\Claude\claude_desktop_config.json(Windows)에 추가하세요:
{
"mcpServers": {
"safe-docx": {
"command": "npx",
"args": ["-y", "@usejunior/safe-docx"]
}
}
}Gemini CLI
{
"mcpServers": {
"safe-docx": {
"command": "npx",
"args": ["-y", "@usejunior/safe-docx"]
}
}
}모든 MCP 클라이언트
명령어:
npx인자:
["-y", "@usejunior/safe-docx"]전송: stdio
Safe Docx 최적화 대상
기존
.docx파일의 브라운필드 편집서식을 유지하는 텍스트 교체 및 단락 삽입
주석 및 각주 워크플로우
검토를 위한 변경 사항 추적 출력물(
download,compare_documents)구조화된 JSON으로 수정 사항 추출(
extract_revisions)
Safe Docx 최적화 대상 아님
Safe Docx는 처음부터 문서를 생성하는 툴킷이 아닙니다.
템플릿/프로그래밍 방식의 레이아웃에서 새로운 .docx 파일을 생성하는 것이 주 목적이라면 docx와 같은 패키지를 사용하세요.
또한 로컬 Safe Docx 런타임은 현재 의도적으로 Word 템플릿 파일(.dotx)을 거부합니다. 여기서 열기 전에 템플릿을 일반 .docx 문서로 변환하세요.
문서 제품군
이 저장소의 자동화된 픽스처 커버리지
Common Paper 스타일 상호 NDA 픽스처
Bonterms 상호 NDA 픽스처
의향서(LOI) 픽스처
ILPA 유한책임조합 계약서 레드라인 픽스처
복잡한 법률 및 비즈니스 .docx 클래스를 위해 설계됨
NVCA 자금 조달 양식
YC SAFE
투자 설명서
주문서 및 서비스 계약서
유한책임조합 계약서
패키지
@usejunior/docx-core: 기존.docx문서를 위한 기본 요소 + 비교 엔진@usejunior/docx-mcp: MCP 서버 구현 및 도구 인터페이스@usejunior/safe-docx: 표준 최종 사용자 설치 이름 (npx -y @usejunior/safe-docx)@usejunior/safedocx-mcpb: 비공개 MCP 번들 래퍼
신뢰성 및 보안
도구 스키마는
packages/docx-mcp/src/tool_catalog.ts에서 생성됩니다.OpenSpec 추적성 매트릭스:
packages/docx-mcp/src/testing/SAFE_DOCX_OPENSPEC_TRACEABILITY.md가정 매트릭스:
packages/docx-mcp/assumptions.md준수 가이드:
docs/safe-docx/sprint-3-conformance.md
FAQ
Safe Docx란 무엇인가요?
기존 Word 문서에 대해 결정론적이고 서식을 유지하는 편집이 필요한 코딩 에이전트 워크플로우를 위한 TypeScript 우선 DOCX 편집 스택입니다.
편집 중에 서식이 유지되나요?
그것이 핵심 설계 목표입니다. 도구 인터페이스는 문서 구조와 서식 의미론을 최대한 보존하는 정밀한 작업(replace_text, insert_paragraph, 레이아웃 제어)을 중심으로 구축되었습니다.
일반적인 런타임 사용 시 .NET, Python 또는 LibreOffice가 필요한가요?
아니요. 지원되는 런타임 사용 환경은 jszip + @xmldom/xmldom을 사용하는 JavaScript/TypeScript입니다.
처음부터 계약서를 생성할 수 있나요?
주요 목적이 아닙니다. 처음부터 생성하려면 docx와 같은 패키지를 사용하세요.
저장소 내 픽스처에서 어떤 문서 유형을 테스트했나요?
상호 NDA(Common Paper/Bonterms 스타일 픽스처 포함), 의향서, ILPA 유한책임조합 계약서 레드라인 픽스처 등을 테스트했습니다.
변호사만을 위한 도구인가요?
아니요. 동일한 브라운필드 .docx 편집 문제는 인사, 조달, 재무, 영업 운영 및 기타 문서 작업이 많은 워크플로우에서도 발생합니다.
MCP 사용자로서 어디서부터 시작해야 하나요?
npx를 통해 @usejunior/safe-docx를 사용하고 packages/docx-mcp/README.md의 설정 예시를 따르세요.
도구 스키마는 어디서 확인할 수 있나요?
packages/docx-mcp/docs/tool-reference.generated.md에서 생성된 참조를 확인하세요.
개발
npm ci
npm run build
npm run lint --workspaces --if-present
npm run test:run
npm run check:spec-coverage
npm run test:coverage:packages
npm run coverage:packages:check
npm run coverage:matrix참고 자료
Open Agreements — 코딩 에이전트로 표준 법률 템플릿(NDA, SAFE, NVCA) 채우기
UseJunior Developer Tools — 설치 옵션 및 도구 카탈로그가 포함된 제품 페이지
개인정보 보호
Safe Docx는 전적으로 로컬 컴퓨터에서 실행됩니다. 문서 내용은 외부 서버로 전송되지 않습니다. 자세한 내용은 개인정보 처리방침을 참조하세요.
거버넌스
Available Tools
30 toolsaccept_ai_editsADestructive
Selectively accept tracked changes by revision id or author in the in-memory session, leaving all other (e.g. third-party reviewer) revisions byte-untouched. This does not write file_path; call save to persist the mutation. Provide revision_ids (array of w:id values) to target specific revisions, or author to accept every revision by one actor. Sweeps document.xml and supported side-story parts (footnotes, endnotes, comments). An ambiguous overlap — a targeted revision structurally containing, or contained by, a non-targeted revision (nested ins/del/move) — hard-errors with code AMBIGUOUS_REVISION_OVERLAP and a structured overlaps list unless normalize_first is set (best-effort, no byte-identical promise).
| Name | Required | Description | Default |
|---|---|---|---|
| author | No | Accept every revision authored by this w:author. Convenience alternative to revision_ids. | |
| file_path | Yes | Path to the DOCX or ODT file. | |
| revision_ids | No | w:id values of the revisions to accept. Mutually preferred over author. | |
| normalize_first | No | Attempt best-effort resolution on an ambiguous (overlapping) revision graph instead of hard-erroring. No byte-identical guarantee. Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply destructiveHint=true; the description adds substantial behavior beyond that: the operation is in-memory and non-persisting, which document parts are swept, and the AMBIGUOUS_REVISION_OVERLAP hard-error with a structured overlaps list. This is exactly the context a mutation tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core action and the persistence caveat come first, then targeting, then scope, then the error case. Every sentence carries new information, though the error/overlap sentence is heavy enough to slow scanning slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers persistence implications, targeting modes, affected parts, and the failure mode. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3. The description still adds value by clarifying that revision_ids are w:id values, that author is a convenience alternative, and that normalize_first is best-effort with no byte-identical promise — semantics that go beyond the schema wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Selectively accept tracked changes by revision id or author in the in-memory session.' The 'selectively' and 'leaving all other revisions byte-untouched' scope makes it clearly distinct from the siblings accept_changes and reject_ai_edits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use revision_ids for specific revisions, author for a whole actor, and call save to persist since this tool does not write file_path. It also names the normalize_first fallback condition, so the agent knows when to reach for each option.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accept_changesBDestructive
Accept every supported tracked change in the document body, including inserted and deleted table-row revisions. Returns acceptance stats; unresolvedRowRevisions remains 0 for supported row markers.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the DOCX or ODT file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=false and destructiveHint=true, so the mutation risk is covered. The description usefully adds the scope ('document body') and the unsupported-marker caveat ('unresolvedRowRevisions remains 0 for supported row markers'), which is genuine extra context, but it does not state irreversibility or whether the result is persisted to disk or requires a separate save.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and scope, and the second sentence carries the return-value caveat. Dense and largely waste-free, though the 'acceptance stats' clause is slightly vague.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive mutation with no output schema, the description covers scope, the supported-change limitation, and a return-value detail. The main gap is persistence behavior and the boundary against accept_ai_edits, but the destructive annotation already conveys the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (file_path) with 100% schema description coverage, so the schema fully documents it. The description adds nothing about file_path semantics (format support wording mirrors the schema), so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Accept every supported tracked change in the document body') and narrows scope with 'supported' plus table-row revision handling. However, it never distinguishes itself from the sibling accept_ai_edits/reject_ai_edits tools, leaving the agent to infer which accept variant applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given — no prerequisites, no mention of when to prefer this over accept_ai_edits or how it pairs with extract_revisions/has_tracked_changes. The agent must guess the workflow context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_commentADestructive
Add a comment or threaded reply to a document. Provide target_paragraph_id + anchor_text for root comments, or parent_comment_id for replies. Supports DOCX and ODT (ODT backs comments with office:annotation; threaded replies are DOCX-only). Surface: revisionable + package-mutation — the body-story comment reference is tracked (w:ins), while comment text and author metadata are recorded in the save report non-revision change manifest.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Comment body text. | |
| author | Yes | Comment author name. | |
| initials | No | Author initials (defaults to first letter of author name). | |
| file_path | Yes | Path to the DOCX or ODT file. | |
| anchor_text | No | Text within the paragraph to anchor the comment to. If omitted, anchors to entire paragraph. | |
| parent_comment_id | No | Parent comment ID for threaded replies. | |
| target_paragraph_id | No | Paragraph ID to anchor the comment to (for root comments). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Explains beyond annotations: it is revisionable with tracked changes (w:ins) in body-story, and comment text/author metadata recorded in non-revision change manifest. Also notes ODT format behavior (office:annotation). Annotations only indicate destructiveHint=true, so this adds significant value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but well-structured: starts with main purpose, then usage patterns, format differences, and behavioral details. Every sentence adds value, though it could be slightly shortened without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, no output schema, and complexity of threaded comments vs root, the description covers purpose, parameter use cases, format limitations, and behavioral impact. It does not explain return values, but that is acceptable without an output schema. Slightly more could be done to clarify prerequisites or error conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning by explaining two usage patterns (root vs reply), the effect of omitting anchor_text (anchors to entire paragraph), and defaults for initials. This goes beyond the schema's descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool adds comments or threaded replies, distinguishes two modes (root vs reply), and specifies supported formats (DOCX, ODT) with DOCX-only for threaded replies. This distinguishes it from sibling tools like delete_comment or get_comments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear instructions on when to use target_paragraph_id+anchor_text versus parent_comment_id, and notes format-specific limitations. However, it does not explicitly state when not to use this tool or compare it to alternatives like batch_edit or insert_paragraph.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_footnoteADestructive
Add a footnote anchored to a paragraph. Optionally position the reference after specific text using after_text. Note: [^N] markers in read_file output are display-only and not part of the editable text used by replace_text. Surface: revisionable + package-mutation — the footnote reference and note text are tracked (w:ins), while footnote-part creation and registration are recorded in the save report non-revision change manifest.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Footnote body text. | |
| file_path | Yes | Path to the DOCX or ODT file. | |
| after_text | No | Text after which to insert the footnote reference. If omitted, appends at end of paragraph. | |
| target_paragraph_id | Yes | Paragraph ID to anchor the footnote to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations ('destructiveHint: true'), the description details revision tracking behavior: footnote reference and text are tracked as insertions, while footnote-part creation is recorded in the save report as non-revision changes. This provides valuable context for AI agent decision-making.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two sentences and a note. The first sentence plainly states purpose, the second adds optional parameter. The technical note about revision tracking, while dense, is relevant for transparency. Could be slightly streamlined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately covers what happens on invocation: tracked changes and save report details. It explains the effect of 'after_text' and the revision behavior, leaving little ambiguity for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minimal extra semantics, only clarifying that 'after_text' is optional and used for positioning. No additional depth beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Add') and resource ('footnote anchored to a paragraph'), distinguishing it from sibling tools like 'update_footnote' and 'delete_footnote'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly indicates when to use this tool (to add a footnote) and mentions optional positioning with 'after_text', but lacks explicit guidance on when not to use it or alternatives beyond the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_editADestructive
Single-agent front door for applying multiple edit steps (replace_text, insert_paragraph) to a document in one call. Validates all steps first, rejects conflicts before applying anything, then executes valid steps sequentially. Accepts inline steps or a plan_file_path JSON array. Surface: revisionable — every applied step emits native OOXML tracked changes.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | No | JSON array of edit steps. Each step needs step_id, operation, and operation-specific fields. | |
| file_path | Yes | Path to the DOCX or ODT file. | |
| plan_file_path | No | Path to a .json file containing an array of edit steps. Mutually exclusive with steps. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true and readOnlyHint=false. The description adds valuable behavioral details: validates all steps first, rejects conflicts, executes sequentially, and emits tracked changes, going beyond annotation-only information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, each adding essential information without redundancy. Purpose, behavior, and outcome are clearly communicated with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers key aspects: batch operation, validation, conflict handling, and output (tracked changes). However, lacks details on error handling for individual step failures or partial execution, leaving some gaps for complex scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, baseline 3. The description adds meaning by explaining that steps can be inline or via plan_file_path, and that each step needs step_id, operation, and operation-specific fields, compensating for the schema's empty items properties.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool is for applying multiple edit steps (replace_text, insert_paragraph) to a document in one call, distinguishing it from individual sibling tools like replace_text and insert_paragraph.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It identifies itself as a 'single-agent front door' for batch edits, implying usage for multiple steps. It mentions validation and conflict rejection but does not explicitly state when not to use it or compare to alternatives beyond the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clear_formattingADestructive
Clear specific run-level formatting (bold, italic, underline, highlight, color, font) from paragraphs. Surface: revisionable — clearing emits a native run-property-change revision (w:rPrChange).
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the DOCX or ODT file. | |
| clear_bold | No | Remove bold formatting. | |
| clear_font | No | Remove font family and size. | |
| clear_color | No | Remove font color. | |
| clear_italic | No | Remove italic formatting. | |
| paragraph_ids | No | Paragraph IDs to clear formatting from. If omitted, clears from all paragraphs. | |
| clear_highlight | No | Remove highlight formatting. | |
| clear_underline | No | Remove underline formatting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true, but the description adds important context: the operation is revisionable and emits a native run-property-change revision. This goes beyond annotations and helps the agent understand the side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no wasted words. The first sentence states the core function, the second adds behavioral context. Well-structured and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters (1 required) and no output schema, the description covers purpose, behavior (revisionable), and scope (paragraphs). It could mention error conditions or file modification details, but the combination with schema annotations is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage with descriptions for all parameters. The description repeats the list of formatting types but adds context about run-level and revisions. Since schema coverage is high, baseline is 3; the description offers moderate added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool clears specific run-level formatting (bold, italic, underline, highlight, color, font) from paragraphs, which is a specific verb+resource. It distinguishes from sibling tools like format_layout which likely deals with layout-level formatting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for clearing run-level formatting but does not explicitly state when to use this tool versus alternatives (e.g., format_layout). No exclusions or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_fileADestructive
Close an open file session, or close all sessions with explicit confirmation. Supports DOCX, ODT, and Google Docs.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | ||
| clear_all | No | ||
| file_path | No | Path to the DOCX or ODT file. | |
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true. The description adds supported file formats but does not disclose what happens to unsaved changes or other side effects of closing. Some additional behavioral context would be beneficial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the core purpose. It is front-loaded and contains no unnecessary words, though a bit more structure (e.g., separate lines for usage) could improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and the destructive nature, the description is minimal. It does not explain return values, side effects, or parameter interactions. For a closing tool, more detail on what happens after closing (e.g., saving) would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, but the description does not explain any parameter semantics. It does not clarify the roles of clear_all and confirm (which lack schema descriptions), nor does it add meaning beyond the schema for file_path and google_doc_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('close') and resource ('open file session'), and distinguishes from sibling tools like read_file, save, and batch_edit by clearly indicating this is about closing sessions. It also specifies supported file formats.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when an open file session exists and mentions explicit confirmation for closing all sessions, but provides no when-not guidance or alternatives. It lacks details on prerequisites or comparison to other closing scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_documentsARead-only
Compare two documents and produce a tracked-changes output document. Provide original_file_path + revised_file_path for standalone comparison, or file_path to compare session edits against the original. DOCX and ODF (.odt) support both modes. DOCX output always uses the revised archive as its package base and publishes tagged revisions; engine, strategy, reconstruction, premerge, and refinement selectors are not exposed. DOCX stats count insertions/deletions as contiguous ranges, expose tagged-token-v1 totals as insertedAtoms/deletedAtoms with atomMetricVersion, and report formatChanges separately from modifiedParagraphs. When a DOCX input difference is preserved in the output without tracked-change markup (for example a removed section's header or footer, or an unsupported header/footer topology), the response includes unrepresented_changes (objects with scope, kind, sectionIndex, for header/footer scopes role, and for the contentControl scope a contentControl identity: a changed w:sdtPr/w:sdtEndPr is published as the revised properties without markup, because CT_SdtPr admits no revision elements) plus one warnings string per entry, and the message carries a WARNING. Both fields are absent when no reportable unrepresented change was detected; the detector covers section properties, selected header/footer stories and content-control properties, so absence is not a guarantee about differences outside that scope. ODF compares at inline granularity (a modified paragraph is marked up in place — only the changed spans are struck or inserted).
| Name | Required | Description | Default |
|---|---|---|---|
| author | No | Author name for track changes. Default: 'Comparison' (DOCX) or the configured AI author (ODF). | |
| file_path | No | Path to the DOCX or ODT file. | |
| compare_moves | No | Detect moved content (DOCX only). Default: true. | |
| ignore_formatting | No | Ignore formatting differences (DOCX only). Default: false. | |
| revised_file_path | No | Path to the revised DOCX or .odt file. | |
| original_file_path | No | Path to the original DOCX or .odt file. | |
| save_to_local_path | Yes | Path to save the tracked-changes output (DOCX or .odt). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply readOnlyHint/destructiveHint; the description independently discloses deep behavior: unexposed selectors, revised-archive package base, insertion/deletion counting semantics with atomMetricVersion, formatChanges vs modifiedParagraphs, the unrepresented_changes/warnings contract, warning-message emission, and detector scope limits. This is exactly the beyond-annotations context the rubric rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but the definition is a single sprawling paragraph dense with implementation jargon (CT_SdtPr, w:sdtPr/w:sdtEndPr, tagged-token-v1) that would benefit from separation into mode, output, and warning sections. Every clause is informative, but the packing hurts scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-mode, seven-parameter tool with no output schema, the description covers both invocation modes, DOCX vs ODF granularity, output statistics fields, and the warning/unrepresented_changes return contract. An agent has enough to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so 3 is the baseline, but the description goes further by explaining how parameters interact (which two paths select standalone mode, which single path selects session mode) and clarifies that certain selectors are deliberately not exposed. That is semantic value the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a concrete verb+resource+artifact: 'Compare two documents and produce a tracked-changes output document.' That instantly distinguishes it from siblings like has_tracked_changes (detection only) and extract_revisions (reads existing markup), since this one generates a new revision-marked document.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit parameter-driven mode selection: original_file_path + revised_file_path for standalone comparison versus file_path to compare session edits against the original. That is clear contextual guidance, but it never names a sibling alternative or a when-not-to-use condition, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_to_odtA
Convert a DOCX document to OpenDocument Text (.odt) using the native model-to-model converter (no LibreOffice involved). Writes the .odt (default: source path with the .odt extension), validates ODF packaging safety before writing, and returns the output path plus a lossiness summary itemizing every downgraded construct. Conversion is semantic and intentionally lossy: text, headings, bold/italic/underline, hyperlinks, lists, and tables are mapped; richer styling, tracked changes, comments, and headers/footers are not. DOCX in, ODT out — Google Docs and .odt inputs are not supported.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | Path to the DOCX or ODT file. | |
| output_path | No | Where to write the .odt. Defaults to the source path with the .odt extension. | |
| allow_overwrite | No | Overwrite output_path if it already exists. Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description details the write operation, safety validation, lossy nature with specific mapped and unmapped constructs, and return values. This goes well beyond the annotations, which only indicate non-read-only and non-destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three sentences, each adding essential information. It is front-loaded with the core action and avoids redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fully covers inputs, output, processing details, limitations, and return value structure. Without an output schema, it adequately explains what the tool returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and description adds no extra parameter meaning beyond what the schema provides. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description precisely states the tool converts DOCX to ODT using a native converter, lists what is preserved and lost, and notes unsupported inputs. It clearly distinguishes from sibling tools like 'export' by specifying the conversion direction and limitations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the input requirement (DOCX) and output format (ODT), and lists unsupported features. However, it does not explicitly state when to avoid this tool or suggest alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_commentADestructive
Delete a comment and all its threaded replies from the document. Cascade-deletes all descendants. Surface: revisionable + package-mutation — the body-story comment reference removal is tracked (w:del), while comment/reply text cleanup is recorded in the save report non-revision change manifest.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the DOCX or ODT file. | |
| comment_id | Yes | Comment ID to delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true. The description adds valuable behavioral details: cascade-deletes all descendants and specifics about tracking (w:del, non-revision change manifest). This goes beyond the annotations, disclosing side effects and recording behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each serving a purpose. The first sentence is a clear action, the second adds technical behavioral context. Front-loaded but the second sentence may be dense; still efficient with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple delete tool with 2 parameters and no output schema, the description covers the core action and side effects. However, it lacks information on error conditions, prerequisites (e.g., file must be open), or what the response looks like, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both parameters. The tool description does not add extra meaning or syntax details beyond the schema, meeting the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Delete a comment and all its threaded replies') and identifies the resource (comment). It distinguishes this tool from siblings like add_comment or get_comments by specifying deletion of threaded replies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for deleting comments but does not explicitly state when to use or avoid this tool, nor does it mention alternatives (e.g., delete_footnote). Some context is given but lacks explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_footnoteADestructive
Delete a footnote and its reference from the document. Surface: revisionable — the reference and note text are removed as native OOXML tracked deletions (w:del).
| Name | Required | Description | Default |
|---|---|---|---|
| note_id | Yes | Footnote ID to delete. | |
| file_path | Yes | Path to the DOCX or ODT file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true. The description adds valuable detail: that the deletion is revisionable and performed as native OOXML tracked deletions (w:del), providing transparency beyond the annotations. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the primary action. Every sentence adds value: purpose and technical detail. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with 2 required parameters and no output schema, the description is complete. It explains the core behavior and tracked changes mechanism. Could optionally mention that note_id should come from get_footnotes, but not necessary for completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (both parameters described in schema). The description adds no additional meaning beyond the schema, but the schema itself is clear. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool deletes a footnote and its reference, distinguishing it from sibling tools like add_footnote and update_footnote. The verb 'delete' combined with the resource 'footnote' gives unambiguous purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is provided, nor are alternatives mentioned. However, the context from sibling tools makes it clear this is the deletion tool, providing some implicit guidance. Could be improved by noting prerequisites or limitations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exportA
Export a document to a portable rendering (Markdown, semantic HTML, or plain text). Writes an output file (default: source path with the format extension, e.g. .md, .html, or .txt) and returns its path, byte count, and the rendered content (under content). Intentionally lossy (no round-trip); HTML is the semantic tier, not pixel-faithful. DOCX only — Google Docs is not supported.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: 'markdown' (default, writes .md), 'html' (writes .html), or 'plaintext' (writes .txt). | |
| file_path | No | Path to the DOCX or ODT file. | |
| output_path | No | Where to write the rendering. Defaults to the source path with the format extension. | |
| allow_overwrite | No | Overwrite output_path if it already exists. Default: false. | |
| include_markdown | No | Include the rendered content (under `content`) in the response. Default: true; set false for large documents. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (both false), the description details output file writing, return values (path, byte count, content), lossy behavior, and semantic HTML nature. This fully discloses behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with main action, no wasted words. Every sentence provides essential information: purpose, output behavior, and limitations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters and no output schema, the description covers all needed context: formats, output file behavior, return values, and constraints (DOCX only, lossy). It is fully sufficient for agent understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value by explaining default output path behavior and the return of rendered content under 'content', which clarifies the include_markdown parameter's effect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool exports a document to Markdown, HTML, or plain text. It distinguishes from siblings like convert_to_odt by focusing on rendering formats, and specifies DOCX-only, providing a specific verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context (DOCX only, not Google Docs) but does not explicitly state when to use this tool over alternatives. However, siblings are distinct, so it is sufficient for guiding usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_revisionsARead-only
Extract tracked changes as structured JSON with before/after text per paragraph, revision details, and comments. Table rows inserted or deleted as a whole, and row property changes (w:trPr > w:ins / w:del / w:trPrChange), are reported as records with scope "row", keyed by the row's first paragraph, whose revisions are ROW_INSERTION / ROW_DELETION / FORMAT_CHANGE entries carrying the revision id, author and date. Supports pagination via offset and limit. Read-only - does not modify the document.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max entries per page (1-500). Default: 50. | |
| offset | No | 0-based offset for pagination. Default: 0. | |
| file_path | Yes | Path to the DOCX or ODT file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description confirms the read-only nature while adding real behavioral detail: how row-level insertions/deletions/format changes are surfaced with scope "row" and ROW_INSERTION/ROW_DELETION/FORMAT_CHANGE records, and that pagination is supported. This goes beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main purpose is front-loaded in the first sentence, followed by necessary detail on row-scope records and pagination. The middle sentences are dense and technical but each carries useful semantics rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values, and it does so well (before/after text, revision details, comments, row-scoped records). The only minor gap is that it doesn't restate accepted file types (DOCX/ODT, which the schema covers) or the author/date fields' format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so file_path, offset, and limit are already documented with ranges and defaults. The description only notes that pagination exists via offset and limit, adding no syntax or format detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Extract) and resource (tracked changes) plus the output shape (structured JSON with before/after text per paragraph, revision details, comments). Clear enough to distinguish from the sibling has_tracked_changes, which only checks presence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the purpose but never stated as when/when-not. It does not route the agent to siblings such as has_tracked_changes (to check first) or accept_changes/reject_ai_edits (to act on the revisions), so no explicit selection guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
format_layoutADestructive
Apply layout controls (paragraph spacing, table row height, cell padding). Google Docs supports paragraph spacing only. Surface: revisionable — DOCX geometry edits emit native property-change revisions (w:pPrChange/w:trPrChange/w:tcPrChange).
| Name | Required | Description | Default |
|---|---|---|---|
| strict | No | ||
| file_path | No | Path to the DOCX or ODT file. | |
| row_height | No | ||
| cell_padding | No | ||
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| paragraph_spacing | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by detailing that edits are revisionable (emit native property-change revisions) and notes platform-specific behavior. Annotations already indicate destructiveHint=true, and the description adds valuable context about tracked changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences that front-load the purpose and then add key behavioral context. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters with nested objects, no output schema), the description lacks detail on how to use parameters (e.g., row_indexes, cell_indexes) and the meaning of 'strict'. It does not cover return values or error conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only file_path and google_doc_id have descriptions). The description does not explain the semantics of parameters like 'strict', 'row_height', 'cell_padding', or 'paragraph_spacing' beyond naming the categories. An agent would need more detail to set these correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool applies layout controls (paragraph spacing, table row height, cell padding). It differentiates from sibling tools like 'clear_formatting' by focusing on specific layout properties.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mention of 'Google Docs supports paragraph spacing only' provides some context about when to use which parameter, but there is no explicit guidance on when to use this tool versus alternatives like 'clear_formatting' or 'insert_paragraph'. Usage context is implied but not fully elaborated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
format_numberingADestructive
Change one DOCX body paragraph’s direct numbering reference. Use remove=true to drop direct w:numPr, match_paragraph_id to adopt another paragraph’s explicit numbering, or num_id with ilvl to reference an existing numbering definition. This tool does not create numbering definitions or change style-inherited numbering. Effective edits emit a native w:pPrChange; identical requests are no-ops.
| Name | Required | Description | Default |
|---|---|---|---|
| ilvl | No | Existing numbering level for num_id; requires num_id. | |
| num_id | No | Existing positive decimal w:numId from this DOCX; requires ilvl. | |
| remove | No | Set true to remove the target paragraph’s direct w:numPr. | |
| file_path | No | Path to the DOCX or ODT file. | |
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| match_paragraph_id | No | Copy this paragraph’s complete direct num_id and ilvl to the target. | |
| target_paragraph_id | Yes | Target paragraph anchor returned by read_file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and readOnlyHint=false, but the description goes further by disclosing that effective edits emit a native w:pPrChange and that identical requests are no-ops. This tracked-change and idempotency behavior is exactly the kind of trait annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then modes, then constraints and behavioral footnote. No filler; every sentence carries distinct, useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-paragraph mutation tool with full annotation coverage and complete schema descriptions, the definition covers purpose, modes, and side effects well. It stops short of describing the return value, which is notable given there is no output schema, but the core call contract is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented, but the description adds genuine value by explaining how the parameters combine (three mutually distinct edit modes) and the semantic result of each. It doesn't add syntax detail beyond the schema, so it stays just above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Change) and precisely scoped resource (one DOCX body paragraph's direct numbering reference), immediately clarifying it operates on direct w:numPr rather than style-inherited numbering. An agent can distinguish this from generic formatting or numbering-creation tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit parameter-driven modes of use (remove=true, match_paragraph_id, num_id with ilvl) plus clear exclusions: it does not create numbering definitions or modify style-inherited numbering. This tells the agent both when and when-not to reach for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
format_sectionADestructive
Partially update one DOCX section’s page-number restart, page dimensions/orientation, or margins using a zero-based section_index from get_sections. Effective calls emit one native w:sectPrChange and preserve section topology, page-number format, columns, break type, and header/footer references. Orientation is literal and does not automatically swap dimensions. This tool does not create sections or edit header/footer content.
| Name | Required | Description | Default |
|---|---|---|---|
| margins | No | Partial margin update in twips. All seven values are required when w:pgMar is absent. | |
| file_path | No | Path to the DOCX or ODT file. | |
| page_size | No | Partial page-size update. Both dimensions are required when w:pgSz is absent. | |
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| section_index | Yes | Zero-based session-relative index returned by get_sections. | |
| page_number_start | No | Non-negative page number at which this section starts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation profile is known. The description adds substantive behavioral context beyond that: effective calls emit exactly one native w:sectPrChange and preserve section topology, page-number format, columns, break type, and header/footer references, plus the non-obvious caveat that orientation is literal and does not auto-swap dimensions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences with no filler; the updatable fields and the required index source are front-loaded, and the preservation guarantees and non-goals follow. Every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with nested margin/page_size objects and no output schema, the description covers the mutation semantics, preservation guarantees, and scope limits, which is nearly everything an agent needs. It does not restate return values, appropriately relying on the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by clarifying that section_index must be a zero-based index sourced from get_sections and that orientation is literal rather than dimension-swapping, which directly prevents a common invocation error.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (partially update) and a precise resource (one DOCX section's page-number restart, page dimensions/orientation, or margins). It names the sibling that supplies the required index (get_sections) and sets clear scope boundaries against insert_section_break and format_layout by ruling out section creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs the agent to obtain section_index from get_sections and states the negative conditions ('does not create sections or edit header/footer content'). It stops short of naming when to prefer format_layout or batch_edit for broader formatting, so it is clear context without full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_commentsARead-only
Get all comments from the document with IDs, authors, dates, text, and anchored paragraph IDs. Range-anchored DOCX comments also expose optional end_paragraph_id, start_run_index, start_char_offset, end_run_index, and end_char_offset fields describing the covered span. Includes threaded replies (DOCX). Supports DOCX and ODT. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the DOCX or ODT file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true and destructiveHint=false, so the description's 'Read-only' is consistent. It adds value by listing specific return fields and mentioning threaded replies and optional range-anchored fields, which go beyond what annotations provide. No contradictions are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two sentences that front-load the core purpose ('Get all comments from the document') and then expand with relevant details. Every sentence adds value without redundancy. It is efficiently structured for quick parsing by an AI agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a single parameter, 100% schema coverage, and no output schema, the description sufficiently covers the return values (IDs, authors, dates, text, anchored paragraph IDs, and optional span fields). The mention of threaded replies and supported formats completes the picture for a read-only comment retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage with a single 'file_path' parameter described as 'Path to the DOCX or ODT file.' The description does not add additional meaning beyond what the schema already provides, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a clear verb+resource combination ('Get all comments from the document') and lists specific fields (IDs, authors, dates, text, anchored paragraph IDs). It distinguishes itself from sibling tools like add_comment or delete_comment by focusing on retrieval. The mention of range-anchored DOCX comments and threaded replies adds specificity without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states it supports DOCX and ODT formats and is read-only, which implicitly guides when to use (read scenarios) versus write operations (e.g., add_comment, delete_comment). It does not explicitly exclude use cases or list alternatives, but the context of sibling tools provides sufficient differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_document_outlineARead-only
Get a compact structural map of a document's headings (DOCX only). Each entry is {paragraph_id, text, level, source}. Deterministic sources are word_style, list_metadata, and outline_level, selected in that precedence order and included by default. Heuristic sources are run_in_header, title_with_period, title_with_colon, title_caps_centered, and title_bare; set include_heuristic_headings=true to include them. JSON preserves levels 1-9; Markdown clamps visual ATX depth to 6. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: 'json' (default, structured outline array) or 'markdown' (indented ATX outline under `content`). | |
| file_path | No | Path to the DOCX file. | |
| include_heuristic_headings | No | When true, also include heuristic title/run-in/centered-caps headings alongside deterministic word_style, list_metadata, and outline_level headings. Default: false (all deterministic sources only). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive), so 'Read-only' is redundant. The description earns credit elsewhere: it discloses the deterministic-vs-heuristic source split, the precedence order word_style > list_metadata > outline_level, and that Markdown clamps ATX depth to 6 while JSON preserves 1-9.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and scope are front-loaded, and the sentence enumerating source types is dense but each element carries information. The source lists are long, but they define the actual selection contract, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by specifying the entry shape {paragraph_id, text, level, source} and the format-dependent level behavior. Minor gaps remain (no pagination/size hints for large documents), but it is largely complete for a read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema: the precedence order governing the deterministic sources and the level-preservation difference between json and markdown formats. This genuinely clarifies parameter effects rather than restating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a compact structural map of a document's headings') and scopes it ('DOCX only'), so an agent can distinguish it from siblings like get_sections or read_file. It stops short of explicitly naming an alternative, so it doesn't reach the top of the scale.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the DOCX-only constraint and explains when to flip include_heuristic_headings=true, which implies usage. However, it never states when NOT to use this tool or which sibling to reach for instead (e.g., get_sections for full content), leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_file_statusARead-only
Get file/session metadata including edit count, normalization stats, and cache info. Supports DOCX, ODT, and Google Docs.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | Path to the DOCX or ODT file. | |
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, indicating safe read operation. Description adds minimal context about return data (edit count, normalization stats, cache info) but does not elaborate on side effects or limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single, focused sentence with no extraneous information. Front-loaded with core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, description adequately hints at return values. Combined with strong annotations, it provides sufficient context for a simple metadata retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% coverage with clear descriptions for both parameters. Description mentions supported formats (DOCX, ODT, Google Docs) which aligns with schema, but adds no additional semantic detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'Get file/session metadata' with specific examples (edit count, normalization stats, cache info), and mentions supported formats. This distinguishes it from sibling tools like get_comments or get_footnotes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool vs alternatives. While it lists supported file types, it does not specify contexts where get_file_status is preferred over other metadata tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_footnotesARead-only
Get all footnotes from the document with IDs, display numbers, text, and anchored paragraph IDs. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the DOCX or ODT file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description adds 'Read-only' and lists returned data (IDs, display numbers, text, anchored paragraph IDs). Annotations already provide readOnlyHint=true and destructiveHint=false, so description adds minor context beyond annotations but does not contradict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence efficiently conveys purpose, scope, and data returned with no wasted words. Front-loaded with key action 'Get all footnotes'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one parameter and no output schema, description covers returned data and read-only nature. Lacks error handling or prerequisite info, but is sufficient for basic use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% coverage with file_path fully described. Description adds no additional parameter information beyond schema, meeting baseline expectation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description explicitly states 'Get all footnotes from the document' with specific attributes (IDs, display numbers, text, anchored paragraph IDs), clearly distinguishing it from sibling tools like add_footnote, delete_footnote, update_footnote.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Description states 'Read-only,' implying no modifications, but does not explicitly contrast with alternatives or provide when-not-to-use scenarios. Lacks explicit guidance relative to siblings, but the read-only nature is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sectionsARead-only
Read DOCX main-document sections in document order. Returns zero-based session-relative section_index values, paragraph/body boundary metadata, page numbering, page size, margins, and header/footer relationship references. Call again after any operation that changes section topology. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | Path to the DOCX or ODT file. | |
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true and destructiveHint=false already declared, the description's 'Read-only' adds nothing new, but it does disclose non-obvious behavior: the cache-staleness rule (re-call after topology changes) and that indices are session-relative rather than persistent. That is real behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core verb+resource before the return-value and re-call details. Slightly dense in enumerating return fields, but each clause carries information, and the trailing 'Read-only' is a minor redundancy against the annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description helpfully enumerates the return surface (section_index, boundary metadata, page numbering, size, margins, header/footer refs). Combined with the staleness rule, an agent has enough to call it correctly; only the choice between this tool and get_document_outline is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both file_path and google_doc_id are fully documented in the schema, including the URL-extraction hint. The description adds no parameter meaning at all, so the baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read DOCX main-document sections in document order,' and names the key return artifact (zero-based session-relative section_index values). It distinguishes itself from generic reads like read_file/grep by scoping to section topology, but never explicitly contrasts with the nearby get_document_outline sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call again after any operation that changes section topology' gives one concrete re-invocation trigger, which is genuinely useful. However, there is no guidance on when to prefer this over get_document_outline, read_file, or extract_revisions, leaving the agent to infer the choice among several read-oriented siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grepARead-only
Search paragraphs with regex. Use file_path for session-based search, file_paths for stateless multi-file search, or google_doc_id for Google Docs. ODT supported via file_path (single-file) only.
| Name | Required | Description | Default |
|---|---|---|---|
| pattern | No | ||
| patterns | No | ||
| file_path | No | Path to the DOCX or ODT file. | |
| file_paths | No | Multiple file paths for stateless multi-file search. No session created. | |
| search_xml | No | When true, search raw XML (word/document.xml) instead of paragraph text. | |
| whole_word | No | ||
| max_results | No | ||
| context_chars | No | ||
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| case_sensitive | No | ||
| include_context | No | When false, skip document view context (list labels, headers) for faster results. Default: true. | |
| dedupe_by_paragraph | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint and destructiveHint, confirming safety. The description adds behavioral context: it searches paragraphs with regex, specifies input modes, and notes ODT-only support via file_path. No contradictions. Additional details about regex flags or output format would improve transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise: two sentences that front-load the core action ('Search paragraphs with regex') and efficiently convey key usage distinctions. No redundant or unnecessary text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (12 parameters, no output schema, low schema coverage), the description is insufficient. It does not explain return values (e.g., matching paragraphs), result limits, or behavior of regex flags, leaving significant gaps for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low (42%), and the description adds meaning to only the file/file-path parameters (file_path, file_paths, google_doc_id) and ODT support. The remaining 7 parameters (e.g., pattern, case_sensitive, max_results) lack description in both schema and tool description, leaving their semantics unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search paragraphs with regex.' It specifies different input modes (file_path, file_paths, google_doc_id) and highlights the ODT limitation, effectively distinguishing it from sibling tools like read_file or replace_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use different input parameters (session-based vs stateless multi-file vs Google Docs) and notes ODT support limitations. However, it lacks explicit guidance on when not to use this tool or alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
has_tracked_changesARead-only
Check whether the document body contains tracked-change markers (insertions, deletions, moves, and property-change records). Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Path to the DOCX or ODT file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and destructiveHint=false. The description adds which marker types are checked but does not describe the return value or error behavior. Since no output schema exists, this is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with purpose, no redundant information. Highly concise and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple boolean check tool, the description is adequate but omits return type and error cases. Given no output schema, agents may not know if the result is a boolean or something else. However, the parameter is well-documented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage with a clear description for file_path. The tool description does not add anything beyond that, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks for tracked-change markers with specific types (insertions, deletions, moves, property-change records) and declares it read-only. This distinguishes it from sibling tools like accept_changes or extract_revisions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. For example, it doesn't suggest using it before accept_changes or that it's lighter than extract_revisions. The purpose implies usage but is not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
insert_paragraphADestructive
Insert a paragraph before/after an anchor paragraph by paragraph id. Supports DOCX, ODT, and Google Docs. (ODT paragraph ids are positional and shift after insertion — re-read before further edits.) Surface: revisionable — DOCX insertions emit native OOXML tracked changes.
| Name | Required | Description | Default |
|---|---|---|---|
| position | No | ||
| file_path | No | Path to the DOCX or ODT file. | |
| new_string | Yes | ||
| instruction | Yes | ||
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| style_source_id | No | Paragraph anchor to clone formatting (pPr and template run) from instead of the positional anchor. Accepts a `_bk_*` ID, or (DOCX only) any other bookmark name whose w:id-paired range covers exactly that one paragraph. Falls back to anchor with a warning if not found. | |
| positional_anchor_node_id | Yes | Anchor paragraph. Accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly this one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnlyHint=false and destructiveHint=true, so the mutation profile is known; the description adds genuinely useful context beyond that: DOCX insertions emit native OOXML tracked changes (revisionable surface) and ODT paragraph ids are positional and shift after insertion. It does not state permission/auth requirements or whether the operation can be undone outside tracked changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core action front-loaded, followed by supported surfaces and then caveats. The parenthetical ODT warning and tracked-changes note both earn their place; no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema and only a destructive hint in annotations, the description covers supported formats, the id-shifting hazard, and tracked-changes behavior. It leaves the required 'instruction' parameter and any auth requirements for Google Docs unexplained, so it is strong but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 57%, with the two most consequential parameters (positional_anchor_node_id and the enum position) already documented in the schema, and the description reinforces the anchor/position relationship. Two required parameters, new_string and instruction, have no explanation in either the schema or the description, so the description does not fully compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (insert) and resource (paragraph) and pins down the mechanism: before/after an anchor paragraph identified by paragraph id. This clearly separates it from sibling tools like replace_text or insert_section_break, which operate on different targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied by 'insert a paragraph before/after an anchor paragraph,' and the ODT caveat tells the agent to re-read before further edits. However, there is no explicit when-to-use/when-not guidance and no reference to alternatives such as replace_text or batch_edit for related edits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
insert_section_breakADestructive
Insert a tracked DOCX section break after a stable direct-body paragraph. The new boundary preserves the containing section’s page setup and header/footer relationship references. The following section inherits current properties by default; set inherit_properties=false to reset non-relationship properties, and optionally provide page-number/page-size/margin overrides in new_section. Call get_sections again after success because section indexes change.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | Path to the DOCX or ODT file. | |
| break_type | Yes | OOXML start behavior for the following section. | |
| new_section | No | Optional page-number and page-setup overrides for the following section. Complete page size/margins are required when reset removes those elements. | |
| paragraph_id | Yes | Stable paragraph id returned by read_file; must identify a direct main-body paragraph that does not already end a section. | |
| inherit_properties | No | Whether the following section retains current non-relationship properties. Default: true. Header/footer references are always preserved. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the mutation profile is known. The description adds real value beyond that: the break is tracked, existing page setup and header/footer relationship references are preserved, inheritance of non-relationship properties defaults on, and — most importantly — section indexes shift after success, requiring a re-read. Return format is unstated, but the actions needed to stay consistent are.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The verb+resource is front-loaded, and each subsequent sentence carries a distinct piece of information (preservation, inheritance, override, follow-up). It is dense and information-packed rather than padded, though the parameter detail could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with nested objects and no output schema, the definition covers the important ground: tracked-change behavior, preservation semantics, inheritance defaults, override options, and the index-invalidation side effect. Only error conditions and explicit sibling routing remain uncovered, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine interpretation: inherit_properties resets only non-relationship properties (header/footer refs always preserved), and new_section supplies optional page-number/page-size/margin overrides for the following section. This clarifies the interaction between the two parameters in ways the schema properties alone do not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Insert a tracked DOCX section break') with the anchoring condition ('after a stable direct-body paragraph'). This is clearly distinguishable from read-only siblings like get_sections and from format_section, which modifies an existing section rather than inserting a boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides operational context: default inheritance behavior, how to opt out via inherit_properties=false, and the explicit follow-up 'Call get_sections again after success because section indexes change.' It does not, however, contrast this tool with alternatives like insert_paragraph or state when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileARead-only
Read document content (DOCX, ODT, or Google Doc). Output is token-limited (~14k tokens) by default with pagination metadata (has_more, next_offset). Use offset/limit to paginate.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max paragraphs to return. When omitted, output is token-limited to ~14k tokens with pagination. | |
| format | No | ||
| offset | No | 1-based paragraph offset for pagination. Negative values count from end. | |
| node_ids | No | Paragraph selectors. Each accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused. Returned rows always report the paragraph's canonical `_bk_*` id, even when selected by another bookmark name; results are de-duplicated and returned in document order. | |
| file_path | No | Path to the DOCX or ODT file. | |
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| show_formatting | No | When true (default), shows inline formatting tags (<b>, <i>, <u>, <highlighting>, <a>). When false, emits plain text with no inline tags. | |
| comment_rendering | No | How to render comments in read_file output. Use "paragraph_notes" (default) for paragraph-local comment threads, "inline_markers" to add `[cm-start:N]`/`[cm-end:N]` milestones in TOON output (combined with the thread blocks), "endnotes" to collect threaded comments into a trailing #COMMENTS block in TOON output, or "none" for the legacy output with no comment rendering. | |
| include_footnotes | No | Single-call body + footnotes retrieval. When true and format="json", the response gains a document-wide TOP-LEVEL `footnotes` array — each entry is {id, display_number, ref_paragraph_ids (an ARRAY of the paragraph ids that reference it), paragraphs[] ({text, tagged_text with run-level formatting tags, style})} — preserving multi-paragraph bodies and footnote-internal bold/italic/citation formatting. This top-level array is NOT inlined into content[], so the 1:1 content[] index invariant is preserved. For backward compatibility a lightweight per-node `footnotes` array ({id, display_number, text}) is ALSO attached to each paragraph node it anchors, windowed to the returned slice. When true and format="toon", a trailing `#FOOTNOTES` sidecar block is appended (symmetric with `#COMMENTS`). Footnotes with an empty body or display_number 0 are excluded. No effect on simple output. Ignored for Google Docs and ODT. Default: false. | |
| include_fingerprint | No | When true and format="json", include a portable content_fingerprint ("sha256:nfkc:<32hex>") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept a `_bk_*` ID, or (DOCX only) any other bookmark name whose w:id-paired range covers exactly that one paragraph. No effect on TOON/simple output. Ignored for Google Docs and ODT. | |
| include_fingerprint_ordinal | No | When true together with include_fingerprint and format="json", add duplicate-disambiguation metadata to each paragraph: `content_fingerprint_ordinal` (1-based document-order position among paragraphs sharing the same content_fingerprint), `content_fingerprint_count_in_document` (total paragraphs sharing it, document-wide even under pagination), and `portable_paragraph_ref` ("<content_fingerprint>#<ordinal>"). Read-only disambiguator, NOT an edit anchor; reordering duplicates may change ordinals. No effect without include_fingerprint, and no effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds real behavioral value beyond that: the default output is token-limited to ~14k tokens and the response carries pagination metadata (has_more, next_offset), which an agent needs to avoid truncated reads.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, then output truncation behavior, then the pagination remedy. Every sentence carries distinct information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema, the description supplies the essential return-shape facts (14k token cap, has_more/next_offset) that the schema cannot. It does not describe the paragraph/ID structure of the content payload, but that gap is partly mitigated by the extensive per-parameter schema documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 91%, so individual parameters (including the deep footnote/fingerprint options and comment_rendering modes) are documented in the schema itself. The description restates offset/limit and the three input formats but adds no syntax or edge-case meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read document content') and enumerates the supported source types (DOCX, ODT, Google Doc), which lets an agent separate it from outline/section/grep siblings by intent. It never names a sibling to route against, so it stops short of full discrimination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage instruction is 'Use offset/limit to paginate,' which tells the agent how to handle the large-output case but not when to prefer read_file over get_document_outline, grep, or get_sections. Usage is implied by the pagination coaching rather than stated with alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reject_ai_editsADestructive
Selectively reject tracked changes by revision id or author in the in-memory session (restoring their pre-edit state), leaving all other revisions byte-untouched. This does not write file_path; call save to persist the mutation. Symmetric to accept_ai_edits: provide revision_ids or author, sweeps document.xml and supported side-story parts, and hard-errors on an ambiguous overlap (code AMBIGUOUS_REVISION_OVERLAP with a structured overlaps list) unless normalize_first is set.
| Name | Required | Description | Default |
|---|---|---|---|
| author | No | Reject every revision authored by this w:author. Convenience alternative to revision_ids. | |
| file_path | Yes | Path to the DOCX or ODT file. | |
| revision_ids | No | w:id values of the revisions to reject. Mutually preferred over author. | |
| normalize_first | No | Attempt best-effort resolution on an ambiguous (overlapping) revision graph instead of hard-erroring. No byte-identical guarantee. Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say destructiveHint=true; the description goes well beyond by disclosing that the mutation is in-memory and non-persistent, that it sweeps document.xml plus supported side-story parts, that all other revisions are byte-untouched, and that overlap triggers a hard error with code AMBIGUOUS_REVISION_OVERLAP and a structured overlaps list. This is exactly the behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the verb and scope, then persistence, then error behavior. Every clause carries new information, though the final sentence is packed with jargon (side-story parts, normalize_first) that slows parsing slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating, non-persisting tool with no output schema, the description covers the persistence contract, the affected document scope, and the failure mode with its error code and payload. Nothing an agent needs in order to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter meaning the schema lacks: revision_ids and author are alternative selectors, and normalize_first is a lossy escape hatch from the ambiguity error rather than a routine flag. The parenthetical on the error path is the main added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('reject tracked changes'), the selection keys (revision id or author), the scope ('in the in-memory session'), and the effect on untouched revisions. It also explicitly positions itself against the sibling accept_ai_edits and against save, so the agent can distinguish it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states this tool does not write file_path and that save must be called to persist — a when-to-use and what-to-do-next instruction that would otherwise be a common agent mistake. It names accept_ai_edits as the symmetric counterpart and gives the exact condition (normalize_first) for avoiding the ambiguous-overlap hard error.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replace_textADestructive
Replace text in a paragraph by provider paragraph id, preserving formatting where supported. Supports DOCX, ODT, and Google Docs. To delete an ordinary DOCX body paragraph, pass its complete visible text as old_string and an empty new_string; a clean save removes the paragraph and a tracked save keeps the deletion for review. Do not use this shortcut for paragraphs that carry section properties, are structurally required by a table cell, or own bookmark/comment anchors without inspecting the structure first. Surface: revisionable — DOCX edits emit native OOXML tracked changes (w:ins/w:del/w:rPrChange).
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | Path to the DOCX or ODT file. | |
| new_string | Yes | ||
| old_string | Yes | ||
| instruction | Yes | ||
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| normalize_first | No | Merge format-identical adjacent runs before searching. Useful when text is fragmented across runs. | |
| target_paragraph_id | Yes | Paragraph anchor. Accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly this one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds substantial context beyond them: revisionable surface with native OOXML tracked changes (w:ins/w:del/w:rPrChange), the clean-save-deletes vs tracked-save-retains distinction, and the structural hazards that make deletion unsafe. This is exactly the added value annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose and supported formats, then layers the deletion shortcut and its caveats, then the revisionable surface note. Dense but every sentence carries operational information; the warning sentence is long but justified by its safety content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description conveys the key post-call behavior (tracked vs clean save outcomes) and the formats supported. It does not describe the response payload shape, which is a minor gap for a mutation tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 57%, and the description adds semantics for old_string/new_string in the deletion shortcut (complete visible text) and mentions the paragraph-id acceptance. However instruction, normalize_first, file_path, and google_doc_id are left to the schema, so the description only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Replace text in a paragraph by provider paragraph id' — plus the scope (formatting preservation, supported formats DOCX/ODT/Google Docs). This clearly distinguishes it from insert_paragraph or batch_edit without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance (replace text by paragraph id), an alternative mechanism (empty new_string to delete a DOCX body paragraph), and explicit when-not-to warnings (section properties, table cell structural paragraphs, bookmark/comment anchors). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
saveADestructive
Persist the current in-memory document session. For DOCX: saves clean and/or tracked changes output. For ODT: saves an .odt package. For Google Docs: checkpoint (default) returns revisionId, or snapshot exports as DOCX. Surface: revisionable — the save report lists both the AI revisions applied and a non-revision change manifest of any package-level mutations (comment/footnote side parts, relationships) that have no tracked-change wrapper.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | No | Path to the DOCX or ODT file. | |
| save_format | No | ||
| google_doc_id | No | Google Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit | |
| allow_overwrite | No | ||
| clean_bookmarks | No | Controls removal of internal bookmarks from DOCX output. Behavior is intentionally three-way: OMIT (recommended for tracked/persistence saves) preserves the document's own bookmarks — only safe-docx paragraph anchors (`_bk_*`) are removed. Explicit `true` ALSO strips harness edit-span bookmarks (`edit-*`) to produce a clean deliverable; do NOT pass it when the tracked output feeds a redline pipeline, because that reproduces the pre-#609 loss of `edit-*` anchors. `false` keeps all bookmarks. Omitting is NOT equivalent to passing `true` — they differ precisely in whether original `edit-*` bookmarks survive. | |
| save_to_local_path | Yes | ||
| tracked_changes_author | No | ||
| tracked_changes_engine | No | Deprecated and ignored (#126). The redline is now the session's write-time tracked markup, serialized directly — there is no comparison engine to select. Use the compare_documents tool for comparison-based redlines. | |
| fail_on_rebuild_fallback | No | Deprecated and ignored (#126). The default save no longer runs the comparison reconstruction engine, so there is no rebuild fallback to guard against; accepted for backward compatibility only. | |
| tracked_save_to_local_path | No | ||
| allow_discard_preserved_revisions | No | Explicitly allow a clean artifact to auto-accept remaining revisions by the session AI author after accept_ai_edits/reject_ai_edits selectively left revisions unresolved. Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and destructiveHint=true; the description goes well beyond that by disclosing the "revisionable" surface and the contents of the save report (AI revisions plus a non-revision change manifest of package-level mutations with no tracked-change wrapper). That is genuine behavioral context an agent could not infer. It omits any warning about the destructive overwrite path despite allow_overwrite existing, which keeps it out of the top tier.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence, followed by format-specific behavior and then the save-report contract. It is dense and jargon-heavy ("non-revision change manifest of any package-level mutations") but every sentence carries distinct information with little redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 11 parameters, no output schema, no title, and only safety annotations, the description does useful work by describing the save report in place of a return schema and covering all three target formats. It still leaves the relationship among file_path, save_to_local_path, tracked_save_to_local_path, and allow_overwrite unresolved, which is a meaningful gap for a destructive write tool with this many inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 55%, so the schema already carries substantial parameter documentation (notably the very detailed clean_bookmarks field). The description adds only indirect meaning, mapping "clean and/or tracked" to the save_format options, and leaves key parameters like allow_overwrite, file_path vs save_to_local_path vs tracked_save_to_local_path unexplained. Baseline 3 is appropriate for this coverage level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ("Persist the current in-memory document session") and then enumerates concrete per-format outcomes for DOCX, ODT, and Google Docs. An agent can tell this writes the session to durable storage rather than merely closing it. It stops short of naming siblings (export, close_file) as alternatives, so it is clear but not sibling-differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Format-specific behavior is spelled out ("For DOCX: saves clean and/or tracked changes output", "For Google Docs: checkpoint (default) returns revisionId, or snapshot exports as DOCX"), which implies when each mode applies. However, there is no explicit when-to-use-vs-alternative guidance in the description itself — the pointer to compare_documents lives in a schema field, not here. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_footnoteADestructive
Update the text content of an existing footnote. Surface: revisionable — note-text changes emit native OOXML tracked changes (w:ins/w:del) inside the footnote body.
| Name | Required | Description | Default |
|---|---|---|---|
| note_id | Yes | Footnote ID to update. | |
| new_text | Yes | New footnote body text. | |
| file_path | Yes | Path to the DOCX or ODT file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds value beyond annotations by disclosing that changes emit native OOXML tracked changes (w:ins/w:del). Annotations already show readOnlyHint=false and destructiveHint=true, but the description clarifies the revisionable nature, enhancing transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundant information. First sentence states the core purpose; second sentence adds a critical behavioral detail. Every word is purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three simple parameters, no output schema, and clear annotations, the description provides sufficient context. It covers purpose and a key behavioral trait (tracked changes), making it complete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the description does not need to add much. It provides general context but does not elaborate on parameter formats or constraints beyond what is in the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool updates the text content of an existing footnote. It uses a specific verb ('Update') and resource ('footnote'), distinguishing it from siblings like 'add_footnote' and 'delete_footnote'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for updating footnote text but does not explicitly state when to use this tool versus alternatives or any prerequisites. The behavioral note about tracked changes is helpful but does not provide direct usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.21.2- Changed
compare_documents3 fields changed- added
Input schema / properties / compare_movesAdded value: +{ + "description": "Detect moved content (DOCX only). Default: true.", + "type": "boolean" +} - removed
Input schema / properties / engineRemoved value: -{ - "description": "Comparison engine (DOCX only). Default: 'auto'.", - "enum": [ - "auto", - "atomizer" - ], - "type": "string" -} - added
Input schema / properties / ignore_formattingAdded value: +{ + "description": "Ignore formatting differences (DOCX only). Default: false.", + "type": "boolean" +}
- Added
format_numbering - Added
format_section - Changed
get_document_outline1 field changed- changed
Input schema / properties / include_heuristic_headings / descriptionPrevious value: -"When true, also include heuristically-detected headings (manual title / run-in / centered-caps) alongside Word HeadingN styles. Default: false (style-based only)."New value: +"When true, also include heuristic title/run-in/centered-caps headings alongside deterministic word_style, list_metadata, and outline_level headings. Default: false (all deterministic sources only)."
- Added
get_sections - Changed
insert_paragraph2 fields changed- added
Input schema / properties / positional_anchor_node_id / descriptionAdded value: +"Anchor paragraph. Accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly this one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused." - changed
Input schema / properties / style_source_id / descriptionPrevious value: -"Paragraph _bk_* ID to clone formatting (pPr and template run) from instead of the positional anchor. Falls back to anchor with a warning if not found."New value: +"Paragraph anchor to clone formatting (pPr and template run) from instead of the positional anchor. Accepts a `_bk_*` ID, or (DOCX only) any other bookmark name whose w:id-paired range covers exactly that one paragraph. Falls back to anchor with a warning if not found."
- Added
insert_section_break - Changed
read_file3 fields changed- changed
Input schema / properties / include_fingerprint / descriptionPrevious value: -"When true and format=\"json\", include a portable content_fingerprint (\"sha256:nfkc:<32hex>\") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept only `_bk_*` IDs. No effect on TOON/simple output. Ignored for Google Docs and ODT."New value: +"When true and format=\"json\", include a portable content_fingerprint (\"sha256:nfkc:<32hex>\") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept a `_bk_*` ID, or (DOCX only) any other bookmark name whose w:id-paired range covers exactly that one paragraph. No effect on TOON/simple output. Ignored for Google Docs and ODT." - changed
Input schema / properties / include_footnotes / descriptionPrevious value: -"When true and format=\"json\", attach a `footnotes` array ({id, display_number, text}) to each paragraph node for the footnotes anchored to it. Windowed to the returned slice (a paginated walk returns each footnote exactly once) and counted toward the read token budget. Footnotes with an empty body or no anchored paragraph are excluded — use get_footnotes for the authoritative full enumeration. No effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false."New value: +"Single-call body + footnotes retrieval. When true and format=\"json\", the response gains a document-wide TOP-LEVEL `footnotes` array — each entry is {id, display_number, ref_paragraph_ids (an ARRAY of the paragraph ids that reference it), paragraphs[] ({text, tagged_text with run-level formatting tags, style})} — preserving multi-paragraph bodies and footnote-internal bold/italic/citation formatting. This top-level array is NOT inlined into content[], so the 1:1 content[] index invariant is preserved. For backward compatibility a lightweight per-node `footnotes` array ({id, display_number, text}) is ALSO attached to each paragraph node it anchors, windowed to the returned slice. When true and format=\"toon\", a trailing `#FOOTNOTES` sidecar block is appended (symmetric with `#COMMENTS`). Footnotes with an empty body or display_number 0 are excluded. No effect on simple output. Ignored for Google Docs and ODT. Default: false." - added
Input schema / properties / node_ids / descriptionAdded value: +"Paragraph selectors. Each accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused. Returned rows always report the paragraph's canonical `_bk_*` id, even when selected by another bookmark name; results are de-duplicated and returned in document order."
- Changed
replace_text1 field changed- added
Input schema / properties / target_paragraph_id / descriptionAdded value: +"Paragraph anchor. Accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly this one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused."
- Changed
save2 fields changed- added
Input schema / properties / allow_discard_preserved_revisionsAdded value: +{ + "description": "Explicitly allow a clean artifact to auto-accept remaining revisions by the session AI author after accept_ai_edits/reject_ai_edits selectively left revisions unresolved. Default: false.", + "type": "boolean" +} - added
Input schema / properties / clean_bookmarks / descriptionAdded value: +"Controls removal of internal bookmarks from DOCX output. Behavior is intentionally three-way: OMIT (recommended for tracked/persistence saves) preserves the document's own bookmarks — only safe-docx paragraph anchors (`_bk_*`) are removed. Explicit `true` ALSO strips harness edit-span bookmarks (`edit-*`) to produce a clean deliverable; do NOT pass it when the tracked output feeds a redline pipeline, because that reproduces the pre-#609 loss of `edit-*` anchors. `false` keeps all bookmarks. Omitting is NOT equivalent to passing `true` — they differ precisely in whether original `edit-*` bookmarks survive."
5 tool updates
v0.16.0- Added
accept_ai_edits - Added
get_document_outline - Changed
read_file1 field changed- added
Input schema / properties / include_fingerprint_ordinalAdded value: +{ + "description": "When true together with include_fingerprint and format=\"json\", add duplicate-disambiguation metadata to each paragraph: `content_fingerprint_ordinal` (1-based document-order position among paragraphs sharing the same content_fingerprint), `content_fingerprint_count_in_document` (total paragraphs sharing it, document-wide even under pagination), and `portable_paragraph_ref` (\"<content_fingerprint>#<ordinal>\"). Read-only disambiguator, NOT an edit anchor; reordering duplicates may change ordinals. No effect without include_fingerprint, and no effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false.", + "type": "boolean" +}
- Added
reject_ai_edits - Changed
save2 fields changed- changed
Input schema / properties / fail_on_rebuild_fallback / descriptionPrevious value: -"When true, return an error instead of a destructive output if the comparison engine falls back to rebuild mode (which destroys table structure). Default: false."New value: +"Deprecated and ignored (#126). The default save no longer runs the comparison reconstruction engine, so there is no rebuild fallback to guard against; accepted for backward compatibility only." - added
Input schema / properties / tracked_changes_engine / descriptionAdded value: +"Deprecated and ignored (#126). The redline is now the session's write-time tracked markup, serialized directly — there is no comparison engine to select. Use the compare_documents tool for comparison-based redlines."
26 tool updates
v0.12.1- Changed
accept_changes1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
add_comment1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
add_footnote1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Removed
apply_plan - Added
batch_edit - Changed
clear_formatting1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
close_file1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
compare_documents6 fields changed- changed
Input schema / properties / author / descriptionPrevious value: -"Author name for track changes. Default: 'Comparison'."New value: +"Author name for track changes. Default: 'Comparison' (DOCX) or the configured AI author (ODF)." - changed
Input schema / properties / engine / descriptionPrevious value: -"Comparison engine. Default: 'auto'."New value: +"Comparison engine (DOCX only). Default: 'auto'." - changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file." - changed
Input schema / properties / original_file_path / descriptionPrevious value: -"Path to the original DOCX file."New value: +"Path to the original DOCX or .odt file." - changed
Input schema / properties / revised_file_path / descriptionPrevious value: -"Path to the revised DOCX file."New value: +"Path to the revised DOCX or .odt file." - changed
Input schema / properties / save_to_local_path / descriptionPrevious value: -"Path to save the tracked-changes DOCX output."New value: +"Path to save the tracked-changes output (DOCX or .odt)."
- Added
convert_to_odt - Changed
delete_comment1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
delete_footnote1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Added
export - Changed
extract_revisions1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
format_layout1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
get_comments1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
get_file_status1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
get_footnotes1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
grep1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
has_tracked_changes1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Removed
init_plan - Changed
insert_paragraph1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Removed
merge_plans - Changed
read_file4 fields changed- added
Input schema / properties / comment_renderingAdded value: +{ + "description": "How to render comments in read_file output. Use \"paragraph_notes\" (default) for paragraph-local comment threads, \"inline_markers\" to add `[cm-start:N]`/`[cm-end:N]` milestones in TOON output (combined with the thread blocks), \"endnotes\" to collect threaded comments into a trailing #COMMENTS block in TOON output, or \"none\" for the legacy output with no comment rendering.", + "enum": [ + "none", + "paragraph_notes", + "endnotes", + "inline_markers" + ], + "type": "string" +} - changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file." - added
Input schema / properties / include_fingerprintAdded value: +{ + "description": "When true and format=\"json\", include a portable content_fingerprint (\"sha256:nfkc:<32hex>\") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept only `_bk_*` IDs. No effect on TOON/simple output. Ignored for Google Docs and ODT.", + "type": "boolean" +} - added
Input schema / properties / include_footnotesAdded value: +{ + "description": "When true and format=\"json\", attach a `footnotes` array ({id, display_number, text}) to each paragraph node for the footnotes anchored to it. Windowed to the returned slice (a paginated walk returns each footnote exactly once) and counted toward the read token budget. Footnotes with an empty body or no anchored paragraph are excluded — use get_footnotes for the authoritative full enumeration. No effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false.", + "type": "boolean" +}
- Changed
replace_text1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
save1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
- Changed
update_footnote1 field changed- changed
Input schema / properties / file_path / descriptionPrevious value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
23 tool updates
- First observed
accept_changes - First observed
add_comment - First observed
add_footnote - First observed
apply_plan - First observed
clear_formatting - First observed
close_file - First observed
compare_documents - First observed
delete_comment - First observed
delete_footnote - First observed
extract_revisions - First observed
format_layout - First observed
get_comments - First observed
get_file_status - First observed
get_footnotes - First observed
grep - First observed
has_tracked_changes - First observed
init_plan - First observed
insert_paragraph - First observed
merge_plans - First observed
read_file - First observed
replace_text - First observed
save - First observed
update_footnote
TDQS
Scored across 30 tools
Most tools map to a distinct resource+action (comments, footnotes, sections, revisions), so collisions are limited. However, accept_changes (accept all) vs accept_ai_edits (selective) vs reject_ai_edits, batch_edit vs replace_text/insert_paragraph, and format_layout vs format_section all have boundaries an agent must read carefully to separate.
Names are uniformly snake_case and overwhelmingly verb-first (get_comments, delete_footnote, insert_section_break). Minor deviations are single-word verbs (save, grep, export) and the predicate-style has_tracked_changes, but the pattern stays predictable.
At 30 tools this is above the comfortable 3–15 range, and overlapping families (three accept/reject tools, three footnote CRUD tools, three comment tools) add surface area. The domain is genuinely broad—multi-format editing, tracked changes, comments, sections, formatting—so most tools earn their place, but it is heavy and could be tightened.
Core read/edit/revision/comment/footnote/section lifecycles are well covered, with strong revision and formatting primitives. Notable gaps remain: no create-document tool, no table/row insertion despite table formatting controls, no way to apply run formatting (only clear it), and no comment-text update or endnote/image support.
Maintenance
Related MCP Connectors
Edit PDF text in place, fonts and layout kept. Compress, merge, split, convert and translate PDFs.
Real Word and Excel for agents: open your .docx/.xlsm, read, propose, apply, get it back intact.
Generate PDF/DOCX/XLSX/PPTX from templates+JSON. Convert Office/HTML/MD to PDF. Universal templating
Deterministic DOCX/PPTX/XLSX/PDF parser: track changes, comments, headers, footers, merged cells.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA powerful Word document processing service based on FastMCP, enabling AI assistants to create, edit, and manage docx files with full formatting support. Preserves original styles when editing content.191-
- AlicenseBqualityDmaintenanceEnables comprehensive management of Microsoft Word documents with 30+ tools for reading, writing, formatting, template merging, image extraction, equation extraction, and style application.241MIT
- AlicenseAqualityBmaintenanceFill standard legal agreement templates (NDAs, SAFEs, NVCA docs, employment, cloud terms) and produce DOCX files.311,593 npm59Apache 2.0
- AlicenseAqualityDmaintenanceProvides comprehensive read/write access to Word documents, including comments, track changes, and reply threads.91MIT