clean_pdf_research
Extract structured text, sections, and metadata from online research PDFs and whitepapers via direct HTTP/HTTPS URLs. Returns title, page counts, word count, and parsed content for analysis.
Instructions
Parses and extracts structured plain text, sections, and academic metadata from online PDF whitepapers and research papers.
Usage Guidelines:
Use this tool to ingest scientific papers (e.g., arXiv), technical documentation, or financial reports.
Constraint: Target document must be a direct HTTP/HTTPS URL pointing to a PDF file under 15MB.
Returns: Title, total/parsed page count, word count, and extracted text.
Do NOT use for general HTML web pages (use
clean_web_content).Do NOT use for YouTube videos (use
clean_youtube_transcript).Do NOT use for password-protected, DRM-encrypted, or scanned image-only PDFs without OCR.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Direct HTTP/HTTPS URL pointing to an online PDF document. | |
| max_pages | No | Maximum number of pages to parse (1 to 100, default: 30) to control token budget. | |
| auth_token_or_tx | No | Optional x402 micropayment authorization token or EVM transaction hash. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |