ollama-vision-mcp
Related Servers
Alternatives to ollama-vision-mcp
No user-submitted related servers found.
Related Servers
- AlicenseAqualityBmaintenanceMCP server that adds vision capabilities to text-only AI models by sending images (local files, URLs, clipboard, screenshots) to a vision model and returning text descriptions.1208 npmMIT
- AlicenseNot gradedqualityBmaintenanceMCP server for local Ollama vision analysis, enabling text-only agents like Claude Code to inspect images via a single tool. Processes images locally with Ollama, keeping image bytes on the machine and returning text reports.2MIT
- FlicenseAqualityDmaintenanceMCP server for vision capabilities, enabling screenshot, camera, and image analysis using Ollama vision models.41-
- AlicenseNot gradedqualityCmaintenanceAn MCP server that enables any LLM to describe images from file paths, URLs, or base64 data by forwarding them to a supported vision provider such as OpenAI, Anthropic, or local Ollama models.969 npm10MIT
- AlicenseNot gradedqualityCmaintenanceMCP server that provides a 'borrowed eye' for text-only LLMs, enabling them to identify and describe local images via the Qwen VL vision model, including face recognition, scene description, OCR, and targeted visual questioning.2 npmApache 2.0
- AlicenseAqualityAmaintenanceAn MCP server that adds vision capability to any LLM by forwarding images to OpenRouter vision models, returning text analysis. Supports image analysis, model listing, and config diagnostics.314 npmMIT
TDQS
Scored across 4 tools
describe_image and ask_image both handle arbitrary image analysis, and process_clipboard_image explicitly supports describe/ocr/custom-question tasks, creating overlap. However, the input-source distinction (file vs clipboard) reduces ambiguity, and ocr_image is clearly distinct.
Three tools follow a consistent verb_noun pattern (describe_image, ocr_image, ask_image), and process_clipboard_image also uses verb_noun but adds a modifier. The naming is mostly consistent and predictable.
4 tools is a reasonable size for a vision-focused server, covering core operations without being bloated. It could have been higher if not for the redundancy between clipboard and file-based tools.
The toolset covers the main vision tasks (description, OCR, custom Q&A) and handles both file and clipboard inputs. Minor gaps like batch processing or image comparison are not essential for the stated purpose.