vision_ocr
Extract readable text from screenshots and document images. Converts visual text into model-readable transcription for further processing.
Instructions
Read visible text from a screenshot/document image.
Returns the transcription as model text. Prefer detail="original".
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | ||
| detail | No | original | |
| max_tokens | No |