A
licenseNot graded
qualityB
maintenanceEnables text-only models to understand images by intercepting pasted images, saving them locally, and using a vision MCP server to analyze them with multimodal backends like Qwen-VL, Doubao, or GLM-4V.
1
MIT