mineru_parse
Parses a document from a URL and returns a task ID to track extraction progress. Use it to convert PDFs, Office files, PPTs, and images into structured content.
Instructions
Parse a document URL. Returns task_id to check status.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| ocr | No | Enable OCR (pipeline only) | |
| url | Yes | Document URL (PDF, DOC, PPT, images) | |
| model | No | pipeline=fast, vlm=90% accuracy | |
| pages | No | Page range: 1-10,15 or 2--2 | |
| table | No | Table recognition | |
| formats | No | Extra export formats | |
| formula | No | Formula recognition | |
| language | No | Language code: ch, en, etc |