Enables text-only LLMs to perceive images entirely on-device, providing vision capabilities like image description, OCR, table extraction, UI analysis, and region focusing without any cloud APIs or API keys.
Enables pure text LLMs to understand images by acting as a proxy to vision models via OpenAI-compatible APIs. Supports local files, URLs, and base64 inputs for image analysis.