Provides image understanding and OCR via GLM-4.6V-Flash, supporting URL, base64, and local file inputs. Enables AI assistants to analyze images and extract text from screenshots, documents, and more.
Enables image analysis using the GLM-4.6V-Flash model, supporting image URLs and local file paths with custom prompts for tasks like OCR and chart analysis.
Enables image and video understanding plus audio transcription through natural language, using GLM-4.6V-Flash for visual analysis and faster-whisper for speech recognition.
Enables image analysis using GLM-4.5V's vision capabilities from Z.AI. Supports analyzing both local image files and URLs with customizable prompts and parameters.
Enables VLM-based image understanding through a unified API, supporting local llama.cpp and online Qwen3-VL backends, with tools for image analysis, OCR, chart analysis, translation, and multi-turn Q&A sessions.