PDF text extraction (pdfjs-dist, no OCR/table/vendor)
pdf_extract[PAID — 0.04 USDC on Base via x402] Deterministic PDF -> per-page plain text + metadata, no OCR, no table extraction, no third-party PDF/OCR vendor — text pulled directly from the PDF's own embedded text objects via pdfjs-dist (Mozilla's own PDF.js, Apache-2.0) running in this process. Pass exactly one of url (public http(s) link, SSRF-hardened fetch) or pdf_base64 (raw base64 bytes). Max 10 MiB, max 40 pages processed per call. Scanned/image-only PDFs return empty page text with a warning, never a fabricated OCR guess. $0.04 USDC on Base, paid via x402 by the calling agent's own wallet. Call with no payment_header first to receive the payment requirements; your own wallet pays, never this server's.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Public http(s) link to a PDF. Mutually exclusive with pdf_base64. | |
| max_pages | No | Optional; defaults to 40 (the route's own cap). | |
| pdf_base64 | No | Raw PDF bytes, base64-encoded. Mutually exclusive with url. Max 10 MiB decoded. | |
| payment_header | No | Optional. Omit on your first call to receive the x402 payment challenge for free. After your own wallet/x402 client signs against that challenge, call this same tool again with the SAME business arguments plus this field set to the header your x402 client produced (typically { name: "PAYMENT-SIGNATURE", value: "<base64 payload>" }). This server never holds a wallet and never pays on your behalf. |