Enables text-only AI agents to see images on demand by calling any OpenAI-compatible vision API for OCR, image analysis, structured extraction, image comparison, and GUI screenshot-to-accessibility-tree conversion.
Enables AI agents to analyze images through vision AI providers (Gemini, OpenAI, Claude), performing tasks like image description, object detection with bounding boxes, region-specific analysis, and precise color extraction without consuming context window with raw pixels.
Enables text-only LLMs to perceive images entirely on-device, providing vision capabilities like image description, OCR, table extraction, UI analysis, and region focusing without any cloud APIs or API keys.