Skip to main content
Glama

Read document content (text / scan)

clio_document_read

Read Clio documents without disk access: extract text from DOCX, PDF, TXT, EML, HTML with pagination, or return scanned pages as images for visual OCR.

Instructions

Returns the content of a Clio document directly in the response – no disk access needed. DOCX/PDF with a text layer/TXT/EML/HTML → text (paginate with offset/max_chars). Scanned PDFs without a text layer and images (JPG/PNG) → returns the pages as images that Claude reads (visual OCR); select pages with page_from/page_to (max 4 per call). mode: auto (default), text (text layer only), images (always page images – e.g. for stamps, signatures, tables). Accepts document_id or file_path.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNo
offsetNoCharacter offset to continue from
page_toNoFor page images: last page (max. 4 pages per call)
file_pathNoAlternative to document_id – absolute path to a local file
max_charsNoMax. characters of text in the response (default 40000)
page_fromNoFor page images: first page (default 1)
document_idNo
document_version_idNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0-beta.1

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses the return modality per format, that output is inline (no disk write), pagination via offset/max_chars, and a hard cap of max 4 image pages per call. It omits auth/permission prerequisites and read-only confirmation, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the key differentiator, then flows into format/mode/pagination details with an arrow-mapping style that is efficient. It is somewhat dense and would read better as short bullets, but nearly every clause carries useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a dual-mode (text vs. visual OCR) tool with 8 parameters and no output schema, the description covers format routing, mode semantics, pagination, page limits, and input source options. Remaining gaps (document_version_id meaning, permissions/inline-size implications) are modest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 63%, and the description compensates by explaining mode's three values, the pagination role of offset/max_chars, the page_from/page_to page-image selection with the 4-page limit, and that document_id or file_path is accepted. document_version_id and document_id remain undescribed in both schema and description, a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the content of a Clio document') and immediately scopes it with 'directly in the response – no disk access needed,' which is the key distinction from sibling tools like clio_document_download. An agent can tell exactly what it gets back versus a download or search tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for when each mode is appropriate, including a concrete example ('images – e.g. for stamps, signatures, tables') and format routing (text layer vs. scanned/image). It stops short of explicitly naming a sibling alternative to avoid, so it's strong context without full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.