ocr_layout_paddleocr_vl
Extracts text with layout analysis using PaddleOCR-VL, returning text blocks with bounding boxes. Supports 109 languages and complex content like tables and formulas.
Instructions
Extract text with layout analysis. Returns blocks with bounding boxes.
Backend: paddleocr_vl. PaddleOCR-VL — 0.9B vision-language model on Apple Silicon (M1+). Most accurate, 109 languages, supports tables/formulas/charts. Requires paddleocr-vl Swift CLI.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | ||
| mode | No | base | |
| path | Yes |