Skip to main content
Glama
pvliesdonk
by pvliesdonk

Get Document Text

get_document_content
Read-onlyIdempotent

Return a document's OCR'd text, limited to 20,000 characters per request. Provide the offset from a truncation marker to read subsequent sections.

Instructions

Return the OCR'd text content of a document.

Documents such as books and technical standards can run to millions of characters, so each call is capped at 20,000. A partial result opens with a marker naming the character range returned, the document's full length, and the offset to pass to read the next section; text that fits under the cap is returned whole with no marker.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
offsetNoCharacter position to start reading from. Pass the value named in a truncation marker to continue from where it stopped.
max_charsNoMaximum number of characters to return, up to 20,000.
document_idYesID of the document to read.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv2.1.0
    • addedInput schema / properties / document_id / description
      Added value: +"ID of the document to read."
    • addedInput schema / properties / max_chars
      Added value: +{
      +  "default": 20000,
      +  "description": "Maximum number of characters to return, up to 20,000.",
      +  "exclusiveMinimum": 0,
      +  "maximum": 20000,
      +  "type": "integer"
      +}
    • addedInput schema / properties / offset
      Added value: +{
      +  "default": 0,
      +  "description": "Character position to start reading from.  Pass the value\nnamed in a truncation marker to continue from where it stopped.",
      +  "minimum": 0,
      +  "type": "integer"
      +}
  2. First observedv1.0.1

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the operation read-only, idempotent, and non-destructive, and the description adds substantial behavioral detail beyond that: the 20,000-character cap, the truncation marker containing character range and full length, and the offset continuation mechanism. It also clearly states when no marker appears, leaving no surprise about response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense: the first sentence states the core purpose, and the second paragraph explains the only non-obvious behavior. Every sentence earns its place, and the structure front-loads the most important information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only content-fetching tool, the description fully covers the pagination edge case that could confuse an agent, while the output schema handles return-value details. Given the 100% schema coverage and rich annotations, nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters document_id, offset, and max_chars are already well explained. The description adds a little contextual meaning by tying offset to the truncation marker, but this mostly repeats what the schema already says. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Return the OCR'd text content of a document.' This clearly distinguishes the tool from siblings like get_document, get_document_metadata, and get_document_thumbnail, which serve different purposes. No ambiguity remains about what the tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context by explaining that long documents require pagination and how the caller should continue reading via the returned offset. It does not name sibling alternatives or explicitly state when not to use this tool, but the scenario is obvious enough that an agent can determine appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.