Describe Image
describe_imageEnables text-only models to understand images: reads text, describes scenes, interprets charts, compares images, and returns results as text.
Instructions
Image-understanding tool for non-multimodal (text-only) models. Uses an external vision model to read text (OCR), describe scenes, interpret charts and diagrams, compare images, and return the results as text. Models that can understand images directly must not call this tool; use their own vision capability instead. Accepts image_path, image_url, image_base64, image_ref, or images[]. For faster, more useful answers, ask the specific question you need answered (e.g. "read the error text", "what does this chart show") instead of an open-ended "describe everything".
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| images | No | Ordered list of images to analyze. Use this for multiple images; do not combine it with top-level image fields. | |
| prompt | No | Specific question or instruction about the image(s). Focused questions (e.g. "Read the error message text", "List the visible menu items") are faster and more useful than open-ended "describe everything" prompts. Leave empty for a full description. | |
| image_ref | No | Short-lived opaque reference created by the VisionPower WebUI Inbox. Use this when a text-only host cannot pass an attachment through to the agent. | |
| image_url | No | Public http(s) URL of an image; support depends on the configured provider/model. Use image_base64 or image_ref when URL input is unavailable. | |
| image_path | No | Absolute path to a local raster image file. Use this when the image is available on disk. | |
| image_base64 | No | Base64-encoded image data without a data: URI prefix. | |
| output_format | No | Output shape. 'text' (default) returns a free-form description with an untrusted-source banner. 'structured' returns a JSON envelope: when formatValid is true, a single image has {answer, observations, extractedText?, limitations?} and multiple images have images[]; otherwise formatValid is false with formatError and rawResponse. | |
| image_mime_type | No | MIME type for image_base64. If omitted, VisionPower detects it from image bytes. |