ocr_image
Extract readable text from one or more images—screenshots, documents, diagrams, error dialogs—with labeled sections for multi-image inputs.
Instructions
Extract readable text from one or more images (screenshots, documents, diagrams, error dialogs). Multi-image calls return a section per image label.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| image | No | Local file path, file:// URI, http(s) URL, data URL, or base64 image data | |
| images | No | One or more images. Prefer this for multi-image chats: ["path/a.png", "path/b.png"] or [{source, label: "1"}, {source, label: "2"}]. Labels default to "1", "2", ... | |
| prompt | No | Optional extra instruction for the vision model | |
| mimeType | No | Optional MIME type hint for a single bare-base64 `image` input, e.g. image/png |