Skip to main content
Glama

PDF Text

pdf_text
Read-only

Extract clean text (per page + full) and metadata from any public PDF URL up to 25 MB / 200 pages. Pay per call with USDC or USDT on Base via x402. 3 free calls per wallet.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesPublic PDF URL.
pagesNo1-based page numbers to extract.
walletNoYour EVM wallet address (0x...). Unlocks the free tier (3 free calls per wallet).
api_keyNoOptional AMR enterprise API key.
max_pagesNo1-200, default 50.
x_paymentNoOptional signed x402 payment payload (base64 JSON) for USDC/USDT on Base.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
textYes
bytesYes
pagesYes
tokenNoToken actually used for payment (USDC, USDT or FREE).USDC
routingYesHow the request was routed (provider, fallback, cache, latency).
fetch_msYes
metadataYes
cost_usdcYesPrice charged in stablecoin units (USDC/USDT 1:1); 0 on the free tier.
final_urlYes
truncatedYes
char_countYes
fetched_atNoUTC timestamp when the data was fetched.
page_countYes
word_countYes
pages_extractedYes
freshness_secondsYesAge of the underlying data in seconds (0 = fetched live).
x_payment_responseNoBase64-encoded x402 settlement receipt (mirrors the X-PAYMENT-RESPONSE header). Only present when the call was paid with `x_payment`.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "description": "Response for ``POST /pdf-text``.",
      +  "properties": {
      +    "bytes": {
      +      "title": "Bytes",
      +      "type": "integer"
      +    },
      +    "char_count": {
      +      "title": "Char Count",
      +      "type": "integer"
      +    },
      +    "cost_usdc": {
      +      "description": "Price charged in stablecoin units (USDC/USDT 1:1); 0 on the free tier.",
      +      "title": "Cost Usdc",
      +      "type": "number"
      +    },
      +    "fetch_ms": {
      +      "title": "Fetch Ms",
      +      "type": "integer"
      +    },
      +    "fetched_at": {
      +      "description": "UTC timestamp when the data was fetched.",
      +      "format": "date-time",
      +      "title": "Fetched At",
      +      "type": "string"
      +    },
      +    "final_url": {
      +      "title": "Final Url",
      +      "type": "string"
      +    },
      +    "freshness_seconds": {
      +      "description": "Age of the underlying data in seconds (0 = fetched live).",
      +      "title": "Freshness Seconds",
      +      "type": "integer"
      +    },
      +    "metadata": {
      +      "additionalProperties": true,
      +      "title": "Metadata",
      +      "type": "object"
      +    },
      +    "page_count": {
      +      "title": "Page Count",
      +      "type": "integer"
      +    },
      +    "pages": {
      +      "items": {
      +        "additionalProperties": true,
      +        "type": "object"
      +      },
      +      "title": "Pages",
      +      "type": "array"
      +    },
      +    "pages_extracted": {
      +      "title": "Pages Extracted",
      +      "type": "integer"
      +    },
      +    "routing": {
      +      "description": "How the request was routed (provider, fallback, cache, latency).",
      +      "properties": {
      +        "cache_hit": {
      +          "default": false,
      +          "description": "True if the result was served from cache.",
      +          "title": "Cache Hit",
      +          "type": "boolean"
      +        },
      +        "fallback_used": {
      +          "default": false,
      +          "description": "True if the primary provider failed and a fallback served the result.",
      +          "title": "Fallback Used",
      +          "type": "boolean"
      +        },
      +        "latency_ms": {
      +          "default": 0,
      +          "description": "End-to-end routing latency in milliseconds.",
      +          "title": "Latency Ms",
      +          "type": "integer"
      +        },
      +        "provider": {
      +          "description": "Name of the upstream provider that served the result.",
      +          "title": "Provider",
      +          "type": "string"
      +        },
      +        "providers_tried": {
      +          "description": "Providers attempted, in order.",
      +          "items": {
      +            "type": "string"
      +          },
      +          "title": "Providers Tried",
      +          "type": "array"
      +        }
      +      },
      +      "required": [
      +        "provider"
      +      ],
      +      "title": "RoutingMeta",
      +      "type": "object"
      +    },
      +    "text": {
      +      "title": "Text",
      +      "type": "string"
      +    },
      +    "token": {
      +      "default": "USDC",
      +      "description": "Token actually used for payment (USDC, USDT or FREE).",
      +      "title": "Token",
      +      "type": "string"
      +    },
      +    "truncated": {
      +      "title": "Truncated",
      +      "type": "boolean"
      +    },
      +    "url": {
      +      "title": "Url",
      +      "type": "string"
      +    },
      +    "word_count": {
      +      "title": "Word Count",
      +      "type": "integer"
      +    },
      +    "x_payment_response": {
      +      "description": "Base64-encoded x402 settlement receipt (mirrors the X-PAYMENT-RESPONSE header). Only present when the call was paid with `x_payment`.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "freshness_seconds",
      +    "cost_usdc",
      +    "routing",
      +    "url",
      +    "final_url",
      +    "bytes",
      +    "fetch_ms",
      +    "page_count",
      +    "pages_extracted",
      +    "truncated",
      +    "char_count",
      +    "word_count",
      +    "metadata",
      +    "pages",
      +    "text"
      +  ],
      +  "type": "object"
      +}
  2. Changed4 schema fields changed
    • addedInput schema / properties / max_pages / default
      Added value: +50
    • addedInput schema / properties / max_pages / maximum
      Added value: +200
    • addedInput schema / properties / max_pages / minimum
      Added value: +1
    • changedInput schema / properties / wallet / description
      Previous value: -"Your EVM wallet address (0x...). Unlocks the free tier (3 calls per wallet)."New value: +"Your EVM wallet address (0x...). Unlocks the free tier (3 free calls per wallet)."
  3. Added

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint. The description adds critical behavioral context: it requires payment (x402 on Base), offers a free tier per wallet, and imposes size/page caps. These are not covered by annotations, making the description essential for safe invocation. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no fluff. The core function is front-loaded, and payment details are given succinctly. Every clause adds essential information. Ideal structure for an agent to quickly parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main function, constraints, payment, and free tier. With an output schema present, it doesn't need to explain return values. However, it doesn't mention error cases (e.g., exceeding limits or invalid URL) or explicitly distinguish from similar extraction tools. Still, given the annotations and schema, it's fairly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is documented. The description adds value by clarifying the output structure ('per page + full' text) and the 25MB/200-page constraint on the URL, which informs how to set max_pages. It also reinforces the wallet's role in the free tier. This goes slightly beyond the schema without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Extract'), a resource ('clean text and metadata from PDF'), and adds concrete constraints (public URL, size/page limits). It clearly distinguishes this PDF-specific tool from sibling tools like html_to_markdown or web_crawl by focusing on PDFs. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: works on public PDF URLs, with size and page limits, and requires payment via USDC/USDT unless using the free tier. However, it does not explicitly compare to alternatives or state when not to use it. The guidance is implied by the PDF focus but lacks explicit exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.