Skip to main content
Glama

read_file

Read-only

Retrieve content from DOCX, ODT, or Google Docs with token-limited pagination, offset/limit controls, and optional formatting, footnotes, and comments.

Instructions

Read document content (DOCX, ODT, or Google Doc). Output is token-limited (~14k tokens) by default with pagination metadata (has_more, next_offset). Use offset/limit to paginate.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoMax paragraphs to return. When omitted, output is token-limited to ~14k tokens with pagination.
formatNo
offsetNo1-based paragraph offset for pagination. Negative values count from end.
node_idsNoParagraph selectors. Each accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused. Returned rows always report the paragraph's canonical `_bk_*` id, even when selected by another bookmark name; results are de-duplicated and returned in document order.
file_pathNoPath to the DOCX or ODT file.
google_doc_idNoGoogle Doc ID or URL (alternative to file_path). Extract from URL: docs.google.com/document/d/{ID}/edit
show_formattingNoWhen true (default), shows inline formatting tags (<b>, <i>, <u>, <highlighting>, <a>). When false, emits plain text with no inline tags.
comment_renderingNoHow to render comments in read_file output. Use "paragraph_notes" (default) for paragraph-local comment threads, "inline_markers" to add `[cm-start:N]`/`[cm-end:N]` milestones in TOON output (combined with the thread blocks), "endnotes" to collect threaded comments into a trailing #COMMENTS block in TOON output, or "none" for the legacy output with no comment rendering.
include_footnotesNoSingle-call body + footnotes retrieval. When true and format="json", the response gains a document-wide TOP-LEVEL `footnotes` array — each entry is {id, display_number, ref_paragraph_ids (an ARRAY of the paragraph ids that reference it), paragraphs[] ({text, tagged_text with run-level formatting tags, style})} — preserving multi-paragraph bodies and footnote-internal bold/italic/citation formatting. This top-level array is NOT inlined into content[], so the 1:1 content[] index invariant is preserved. For backward compatibility a lightweight per-node `footnotes` array ({id, display_number, text}) is ALSO attached to each paragraph node it anchors, windowed to the returned slice. When true and format="toon", a trailing `#FOOTNOTES` sidecar block is appended (symmetric with `#COMMENTS`). Footnotes with an empty body or display_number 0 are excluded. No effect on simple output. Ignored for Google Docs and ODT. Default: false.
include_fingerprintNoWhen true and format="json", include a portable content_fingerprint ("sha256:nfkc:<32hex>") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept a `_bk_*` ID, or (DOCX only) any other bookmark name whose w:id-paired range covers exactly that one paragraph. No effect on TOON/simple output. Ignored for Google Docs and ODT.
include_fingerprint_ordinalNoWhen true together with include_fingerprint and format="json", add duplicate-disambiguation metadata to each paragraph: `content_fingerprint_ordinal` (1-based document-order position among paragraphs sharing the same content_fingerprint), `content_fingerprint_count_in_document` (total paragraphs sharing it, document-wide even under pagination), and `portable_paragraph_ref` ("<content_fingerprint>#<ordinal>"). Read-only disambiguator, NOT an edit anchor; reordering duplicates may change ordinals. No effect without include_fingerprint, and no effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv0.21.2
    • changedInput schema / properties / include_fingerprint / description
      Previous value: -"When true and format=\"json\", include a portable content_fingerprint (\"sha256:nfkc:<32hex>\") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept only `_bk_*` IDs. No effect on TOON/simple output. Ignored for Google Docs and ODT."New value: +"When true and format=\"json\", include a portable content_fingerprint (\"sha256:nfkc:<32hex>\") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept a `_bk_*` ID, or (DOCX only) any other bookmark name whose w:id-paired range covers exactly that one paragraph. No effect on TOON/simple output. Ignored for Google Docs and ODT."
    • changedInput schema / properties / include_footnotes / description
      Previous value: -"When true and format=\"json\", attach a `footnotes` array ({id, display_number, text}) to each paragraph node for the footnotes anchored to it. Windowed to the returned slice (a paginated walk returns each footnote exactly once) and counted toward the read token budget. Footnotes with an empty body or no anchored paragraph are excluded — use get_footnotes for the authoritative full enumeration. No effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false."New value: +"Single-call body + footnotes retrieval. When true and format=\"json\", the response gains a document-wide TOP-LEVEL `footnotes` array — each entry is {id, display_number, ref_paragraph_ids (an ARRAY of the paragraph ids that reference it), paragraphs[] ({text, tagged_text with run-level formatting tags, style})} — preserving multi-paragraph bodies and footnote-internal bold/italic/citation formatting. This top-level array is NOT inlined into content[], so the 1:1 content[] index invariant is preserved. For backward compatibility a lightweight per-node `footnotes` array ({id, display_number, text}) is ALSO attached to each paragraph node it anchors, windowed to the returned slice. When true and format=\"toon\", a trailing `#FOOTNOTES` sidecar block is appended (symmetric with `#COMMENTS`). Footnotes with an empty body or display_number 0 are excluded. No effect on simple output. Ignored for Google Docs and ODT. Default: false."
    • addedInput schema / properties / node_ids / description
      Added value: +"Paragraph selectors. Each accepts a safe-docx `_bk_*` id, or (DOCX only) any other bookmark name — e.g. a host application's own stable paragraph bookmark — whose w:id-paired range covers exactly one paragraph. Exact name match; a point bookmark or a multi-paragraph range is refused. Returned rows always report the paragraph's canonical `_bk_*` id, even when selected by another bookmark name; results are de-duplicated and returned in document order."
  2. Changed1 schema field changedv0.16.0
    • addedInput schema / properties / include_fingerprint_ordinal
      Added value: +{
      +  "description": "When true together with include_fingerprint and format=\"json\", add duplicate-disambiguation metadata to each paragraph: `content_fingerprint_ordinal` (1-based document-order position among paragraphs sharing the same content_fingerprint), `content_fingerprint_count_in_document` (total paragraphs sharing it, document-wide even under pagination), and `portable_paragraph_ref` (\"<content_fingerprint>#<ordinal>\"). Read-only disambiguator, NOT an edit anchor; reordering duplicates may change ordinals. No effect without include_fingerprint, and no effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false.",
      +  "type": "boolean"
      +}
  3. Changed4 schema fields changedv0.12.1
    • addedInput schema / properties / comment_rendering
      Added value: +{
      +  "description": "How to render comments in read_file output. Use \"paragraph_notes\" (default) for paragraph-local comment threads, \"inline_markers\" to add `[cm-start:N]`/`[cm-end:N]` milestones in TOON output (combined with the thread blocks), \"endnotes\" to collect threaded comments into a trailing #COMMENTS block in TOON output, or \"none\" for the legacy output with no comment rendering.",
      +  "enum": [
      +    "none",
      +    "paragraph_notes",
      +    "endnotes",
      +    "inline_markers"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / file_path / description
      Previous value: -"Path to the DOCX file."New value: +"Path to the DOCX or ODT file."
    • addedInput schema / properties / include_fingerprint
      Added value: +{
      +  "description": "When true and format=\"json\", include a portable content_fingerprint (\"sha256:nfkc:<32hex>\") on each paragraph. Read-only metadata derived from the paragraph's normalized visible text; NOT an edit anchor. Edit tools accept only `_bk_*` IDs. No effect on TOON/simple output. Ignored for Google Docs and ODT.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / include_footnotes
      Added value: +{
      +  "description": "When true and format=\"json\", attach a `footnotes` array ({id, display_number, text}) to each paragraph node for the footnotes anchored to it. Windowed to the returned slice (a paginated walk returns each footnote exactly once) and counted toward the read token budget. Footnotes with an empty body or no anchored paragraph are excluded — use get_footnotes for the authoritative full enumeration. No effect on TOON/simple output. Ignored for Google Docs and ODT. Default: false.",
      +  "type": "boolean"
      +}
  4. First observed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds real behavioral value beyond that: the default output is token-limited to ~14k tokens and the response carries pagination metadata (has_more, next_offset), which an agent needs to avoid truncated reads.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose, then output truncation behavior, then the pagination remedy. Every sentence carries distinct information with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with no output schema, the description supplies the essential return-shape facts (14k token cap, has_more/next_offset) that the schema cannot. It does not describe the paragraph/ID structure of the content payload, but that gap is partly mitigated by the extensive per-parameter schema documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 91%, so individual parameters (including the deep footnote/fingerprint options and comment_rendering modes) are documented in the schema itself. The description restates offset/limit and the three input formats but adds no syntax or edge-case meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read document content') and enumerates the supported source types (DOCX, ODT, Google Doc), which lets an agent separate it from outline/section/grep siblings by intent. It never names a sibling to route against, so it stops short of full discrimination.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage instruction is 'Use offset/limit to paginate,' which tells the agent how to handle the large-output case but not when to prefer read_file over get_document_outline, grep, or get_sections. Usage is implied by the pagination coaching rather than stated with alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.