ocr_text_paddleocr_vl
Extract text from images, returning full text, per-line confidence, and bounding boxes. Supports 109 languages, tables, formulas, and charts via PaddleOCR-VL.
Instructions
Extract text from an image. Returns full text, per-line confidence, and bounding boxes.
Backend: paddleocr_vl. PaddleOCR-VL — 0.9B vision-language model on Apple Silicon (M1+). Most accurate, 109 languages, supports tables/formulas/charts. Requires paddleocr-vl Swift CLI.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | ||
| mode | No | base | |
| path | Yes |