Skip to main content
Glama

Read a remote (or sandboxed local) document

read_doc
Read-onlyIdempotent

Read remote PDFs, DOCX, and local files as Markdown. Specify an http(s) URL or sandboxed local path to extract text with pagination controls.

Instructions

Read an http(s) document (or a sandboxed local file) into Markdown.

Best for:
- Remote PDFs and DOCX from an http(s) URL (parsed locally, no remote API).
- Local PDF/DOCX/text/Markdown files, ONLY when local reads are enabled
  (see Security below).
- Paginating through a long document via `start` / `length`.

Not recommended for:
- Arbitrary HTML web pages -> `fetch` does reader-mode cleanup that this
  tool does not.
- Pages discovered through search -> `fetch` or `research`.

Security (local files are sandboxed and OFF by default):
- Local-file reads are DISABLED unless the server operator sets the
  SEARCH_MCP_DOCUMENT_ROOT env var to a directory. With it unset, a local
  path raises a "local file reads are disabled" error. Pass an http(s)
  URL instead, or ask the operator to enable the sandbox.
- When enabled, `source` must resolve INSIDE that root; relative paths
  resolve against the root (not the process CWD) and any `..` traversal
  that escapes the root is rejected. `file://` URLs are always rejected.
- Remote http(s) sources are unaffected by this setting.

Returns:
- markdown (default): rendered document text with a small header.
- json: {content, title, format, total_chars, start, returned_chars,
  truncated}. Use `total_chars` and `returned_chars` to drive pagination.

Common mistakes:
- Calling this on a normal article URL: you'll get raw HTML noise. Use
  `fetch` instead.
- Forgetting to advance `start` when paginating: next call should pass
  `start = previous_start + returned_chars`.
- Passing a negative `length` (raises an error) or a `start` past the end
  (clamped to EOF: you'll get `returned_chars == 0`, `start == total_chars`,
  and `truncated == False`, which is the signal you've paged off the end).

Args:
    source: http(s) URL, or a local path UNDER SEARCH_MCP_DOCUMENT_ROOT when
        local reads are enabled (disabled by default; see Security).
    start: Character offset to begin reading from. Default 0. Clamped into
        [0, total_chars]; a negative value is treated as 0.
    length: Max characters to return; None = read to end (still capped by
        the per-call max content size). Must be >= 0. A negative length
        is rejected with a ValueError.
    format: "markdown" or "json".

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
startNo
formatNomarkdown
lengthNo
sourceYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.12.0
    • changedOutput schema / (root)
      Previous value: -{
      -  "properties": {
      -    "result": {
      -      "anyOf": [
      -        {
      -          "type": "string"
      -        },
      -        {
      -          "additionalProperties": true,
      -          "type": "object"
      -        }
      -      ],
      -      "title": "Result"
      -    }
      -  },
      -  "required": [
      -    "result"
      -  ],
      -  "title": "read_docOutput",
      -  "type": "object"
      -}New value: +null
  2. First observedv0.1.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint, openWorldHint, and idempotentHint, but the description goes far beyond them: it details sandboxing rules (env var requirement, path traversal rejection, file:// rejection), error behavior (disabled error, negative length ValueError), clamping semantics, and the exact pagination signal (`returned_chars == 0`). This is rich behavioral context that annotations alone could never convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section serves a purpose: 'Best for' and 'Not recommended for' give usage guidance, 'Security' explains the critical env var behavior, 'Returns' details output formats, 'Common mistakes' prevents misusage, and 'Args' is a compact parameter reference. It front-loads the core purpose in the first line and organizes details logically, earning its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a genuinely complex tool (remote vs local, pagination, security, multiple output formats), yet the description covers every aspect needed to call it correctly: input expectations, edge cases (clamping, negative length), output structure (markdown header, json fields), and how to detect end-of-document. With no output schema present, the description fully substitutes that information. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description fully compensates. The Args section explains each parameter: source's format and security constraints, start's default and clamping behavior, length's range and rejection condition, and format's enum meaning. It also provides the json return fields to drive pagination, making each parameter's role crystal clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line states a specific verb and resource ('Read an http(s) document or sandboxed local file into Markdown'), and the 'Best for' / 'Not recommended for' sections explicitly differentiate it from siblings like fetch and research. An agent can immediately tell what this tool does and what it doesn't.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There are explicit usage recommendations with named alternatives: 'Not recommended for: Arbitrary HTML web pages -> `fetch` does reader-mode cleanup...' and 'Pages discovered through search -> `fetch` or `research`.' It also explains the crucial security prerequisite for local files. No ambiguity remains.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.