Skip to main content
Glama

Extraer texto de un PDF

read_pdf
Read-only

Extract text from a PDF at a direct URL, with optional page range selection; no OCR, so scanned PDFs return empty.

Instructions

Extract text from a PDF at a direct URL (max 5 MB, 20 pages per call; pages selects a range). No OCR: scanned PDFs come back empty.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesDirect URL to the PDF file
pagesNoPage range, 1-indexed (e.g. "3", "1-5", "1,4,9").
formatNotext summary (default) or json structured result.text

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv0.10.0
    • addedInput schema / properties / format / description
      Added value: +"text summary (default) or json structured result."
    • addedInput schema / properties / pages / description
      Added value: +"Page range, 1-indexed (e.g. \"3\", \"1-5\", \"1,4,9\")."
    • addedInput schema / properties / url / description
      Added value: +"Direct URL to the PDF file"
  2. Changed5 schema fields changedv0.8.14
    • removedInput schema / properties / format / title
      Removed value: -"Format"
    • removedInput schema / properties / pages / title
      Removed value: -"Pages"
    • removedInput schema / properties / url / title
      Removed value: -"Url"
    • removedInput schema / title
      Removed value: -"read_pdfArguments"
    • removedOutput schema / title
      Removed value: -"read_pdfDictOutput"
  3. First observedv0.8.12

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, openWorldHint), so the description's added value is the hard operational limits: 5 MB max, 20 pages per call, and the critical caveat that scanned PDFs return empty because there is no OCR. That last point is exactly the kind of failure-mode disclosure annotations cannot provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler. The core action and its size/page constraints are front-loaded, and the no-OCR limitation is placed last as a caution.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-format details are legitimately omitted, and the description covers size, page, and OCR boundaries. It stops short of noting whether non-PDF or non-direct URLs fail and how the text/json format choice affects downstream use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so url, pages and format are already documented, including the 1-indexed range syntax and the enum values. The description's "pages selects a range" restates the schema and it never mentions the format parameter, so it adds little beyond baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Extract text from a PDF") and immediately qualifies it with the required input form ("at a direct URL"). No sibling tool performs PDF text extraction, so an agent can select this unambiguously without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage conditions are implied rather than stated: the agent learns it needs a direct URL and that pages selects a range, but nothing says when to reach for this instead of download_resource or how to handle a scanned PDF once it returns empty. Adequate but no explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools