Skip to main content
Glama

Related Servers

Alternatives to mcp-vision

No user-submitted related servers found.

    Related Servers

    • A
      license
      Not graded
      quality
      C
      maintenance
      An MCP server that enables any LLM to describe images from file paths, URLs, or base64 data by forwarding them to a supported vision provider such as OpenAI, Anthropic, or local Ollama models.
      1,150 npm
      9
      MIT
    • A
      license
      Not graded
      quality
      D
      maintenance
      A local MCP server that gives LLMs eyes for images by performing object detection (YOLOv8) and text recognition (EasyOCR), outputting descriptive statements about objects and text positions without any API key or cloud dependency.
      MIT
    • F
      license
      A
      quality
      C
      maintenance
      An MCP server that adds visual understanding to text-only LLMs via image understanding, OCR, and image comparison tools, with multi-provider fallback and context-aware Focus Hint for precise descriptions.
      3
      -
    • A
      license
      Not graded
      quality
      B
      maintenance
      An MCP server for image recognition and OCR via OpenAI-compatible vision APIs, supporting local files, URLs, and data URLs. Enables natural language image description and text extraction.
      10 npm
      2
      MIT

    TDQS

    A3.5/5.0

    Scored across 2 tools

    Disambiguation4/5

    Both tools involve object detection, but they serve distinct purposes: locate_objects returns detections while zoom_to_object crops to the best object. The overlap is minimal, and descriptions clarify the intended usage, though an agent might occasionally hesitate between the two.

    Naming Consistency5/5

    Both tool names follow a consistent verb_noun pattern (locate_objects, zoom_to_object), making them predictable and easy to distinguish at a glance.

    Tool Count3/5

    With only 2 tools, the server feels thin for a general vision MCP, but it is narrowly scoped to object detection and zooming. This borderline count is acceptable for a specialized utility, though it lacks breadth.

    Completeness3/5

    The surface covers the core workflow of detecting and cropping objects, but it omits common vision operations like classification or segmentation. For the narrow purpose stated, it is functional, but the 'vision' name implies a broader scope that is not fulfilled.

    Maintenance

    ActivityInactive
    ResponsivenessNo issues