Enables understanding local or network images via ZhipuAI's free vision model (glm-4.6v-flash), returning Chinese descriptions or answers to questions. Supports multiple images for comparison and optional model switching.
Enables text-only LLMs to see images/videos via cloud vision models, offering vision chat, OCR, grounding, and media info tools with free GLM fallback.
Enables vision capabilities for any AI model by routing image analysis requests through OpenRouter's vision models. It provides tools to analyze images from URLs, local file paths, or base64 data.