Adds vision capabilities to text-only LLMs by integrating external vision models via MCP. It supports OCR, error screenshot reading, UI description, image comparison, and natural-language queries on images.
Bridges vision models to text-only coding models using Florence-2, enabling non-vision LLMs to describe images, extract text, and analyze screenshots via MCP tools.
Enables text-only LLMs to analyze images by bridging DeepSeek's web vision chat via MCP, supporting single/multiple image analysis, batch glob processing, and Windows screen capture for any MCP client.
Bridges text-only models like DeepSeek to 6 free multimodal vision APIs via MCP, enabling image understanding and analysis through automatic fallback and caching.