Provides a vision tool that converts images into structured text descriptions and OCR using the free GLM-4.6V-Flash model, enabling text-only LLMs like DeepSeek to understand images.
Enables text-only AI agents to see images on demand by calling any OpenAI-compatible vision API for OCR, image analysis, structured extraction, image comparison, and GUI screenshot-to-accessibility-tree conversion.
MCP server that gives text-only models vision capabilities via free GLM vision models, supporting image description, OCR, chart/document analysis, and grounding with automatic model fallback.