document
Extract text from PDF, DOCX, or PPTX links and convert it to markdown without a browser. Use it to get full content from shared documents or paper PDFs.
Instructions
Extract text from a PDF/DOCX/PPTX URL → markdown (no browser).
Use when the user shares a link to a document (PDF, Word, PowerPoint)
and wants its text content. Also useful after search_papers to get
full text from a paper's pdf_url.
Content-type sniffed and routed to pypdf / python-docx / python-pptx.
Optional deps — install with pip install 'pyrecrawl[docs]'.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| timeout | No | ||
| max_pages | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||