GLM-Vision MCP Server
Related Servers
Alternatives to GLM-Vision MCP Server
No user-submitted related servers found.
Related Servers
- AlicenseNot gradedqualityAmaintenanceProvides vision capabilities to text-only models (like DeepSeek) via Zhipu free vision models, enabling image analysis, OCR, and image comparison through natural language.MIT
- AlicenseNot gradedqualityBmaintenanceEnables text-only LLMs to see images/videos via cloud vision models, offering vision chat, OCR, grounding, and media info tools with free GLM fallback.Apache 2.0
- AlicenseAqualityDmaintenanceBridges a vision model to enable text-only models like DeepSeek to describe images, extract text, and compare images via MCP tools.528 npm9MIT
- FlicenseAqualityCmaintenanceEnables LLMs like DeepSeek to understand images by calling external vision models via OpenAI-compatible API. Provides tools to describe images or diagnose connectivity.2-
- AlicenseAqualityBmaintenanceProvides image recognition for text-only LLMs like DeepSeek by bridging to SenseNova multimodal model, enabling image description via the describe_image tool.121 npm3MIT
- AlicenseNot gradedqualityBmaintenanceEnables text-only LLMs to understand images by converting them into text descriptions, supporting multiple vision backends like cloud APIs, local models, and OCR engines.1MIT
TDQS
Scored across 1 tool
With only one tool, there is no possibility of overlap or misselection between tools. The single 'vision' tool has an unambiguous purpose (image recognition returning structured description and OCR text).
A single tool named 'vision' is trivially consistent since there are no siblings to clash with. It is a noun rather than a verb_noun pattern, but it is clear and readable.
One tool is thin in absolute terms, but the server has a tightly scoped single purpose (image-to-text), so a single capability fully matches its scope. It is appropriately minimal rather than under-built for the stated domain.
The tool covers the core need of converting an image to text with an optional custom prompt, which satisfies the surface for a vision utility. Minor gaps exist (no batch/multi-image handling or URL input), but agents can work around these by calling per-image.