Provides vision capabilities to text-only LLMs by acting as a cloud vision adapter layer, enabling image understanding and OCR text extraction through a single MCP tool.
Enables any MCP client to perform image understanding and OCR via any OpenAI-compatible vision-language model. Supports local, private inference without images leaving the machine.
Bridges text-only AI models to Google Gemini for image analysis, providing structured visual descriptions, object detection, and answers to image-based questions via MCP.