vision-mcp
Related Servers
Alternatives to vision-mcp
No user-submitted related servers found.
Related Servers
- AlicenseAqualityAmaintenanceEnables text-only AI coding agents to analyze images and videos via vision-capable models (Gemini, Grok, OpenRouter), returning text descriptions for reasoning.213 npmMIT
- AlicenseAqualityAmaintenanceGives text-only LLM coding agents vision by routing images to a multimodal model and returning detailed textual descriptions. Supports local files, URLs, clipboard, base64, raw bytes, and multiple providers like OpenAI, Anthropic, and Gemini.1124 npm12MIT
- AlicenseNot gradedqualityAmaintenanceEnables text-only AI agents to see images on demand by calling any OpenAI-compatible vision API for OCR, image analysis, structured extraction, image comparison, and GUI screenshot-to-accessibility-tree conversion.MIT
- AlicenseAqualityAmaintenanceEnables text-only coding agents to analyze local images using a dedicated vision provider, returning markdown and structured JSON evidence for screenshots, diagrams, UI mockups, and error captures.1121 npm10MIT
- AlicenseAqualityDmaintenanceEnables AI agents to analyze images, extract text, compare images, and analyze video through any OpenAI-compatible vision model.4166 npm20MIT
- FlicenseAqualityDmaintenanceProvides image understanding capabilities to coding models without vision support by automatically invoking a vision model and returning text descriptions, enabling seamless context-aware coding with images.12-
TDQS
Scored across 8 tools
Each tool targets a distinct visual analysis task (charts, errors, text extraction, general, UI diff, UI to code, diagrams, video), with clear usage guidance and a fallback for ambiguity, making selection straightforward.
Tool names are in snake_case but mix verb-first (e.g., analyze_data_visualization) and noun-first (e.g., image_analysis, video_analysis) patterns, plus unconventional names like ui_diff_check and ui_to_artifact, creating inconsistency.
With 8 specialized tools, the server is well-scoped for visual analysis, covering diverse needs without being excessive or sparse.
The tool set covers major visual understanding scenarios (charts, errors, text, UI, diagrams, video) comprehensively, with only a general fallback for edge cases, leaving no obvious gaps.