Enables image and video understanding plus audio transcription through natural language, using GLM-4.6V-Flash for visual analysis and faster-whisper for speech recognition.
Enables Claude Code to understand images by transparently routing them to Qwen vision models, supporting both automatic gateway and MCP tool for file analysis.
Enables text-only AI coding agents to analyze images and videos via vision-capable models (Gemini, Grok, OpenRouter), returning text descriptions for reasoning.
Enables text-only models like Claude Code to recognize images by calling Qwen vision models from DashScope, returning text descriptions for continued reasoning.
Adds vision capabilities to Claude Code by leveraging Aliyun DashScope's Qwen3.7-Flash model to analyze local images or image URLs and return textual insights.