page_vision
Detect word bounding boxes in any element or viewport using local ML text detection. Returns pixel coordinates for captchas, canvas text, and scanned layouts.
Instructions
TIER 4 VISION — VISUAL CORTEX: detect word bounding boxes in an element (or viewport) via the built-in text-detection ML model. Returns JSON [{x, y, w, h}, ...] in image pixels. The first shipped visual-cortex model — see where text lives in ANY image (click-order captchas, canvas text, scanned layout). 100% local.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Ref of the element to analyze. Omit for the whole viewport. | |
| page_id | Yes | ||
| session_id | Yes |