Skip to main content
Glama

mcp-vision by

A Model Context Protocol (MCP) server exposing HuggingFace computer vision models such as zero-shot object detection as tools, enhancing the vision capabilities of large language or vision-language models.

This repo is in active development. See below for details of currently available tools.

Installation

Clone the repo:

git clone git@github.com:groundlight/mcp-vision.git

Build a local docker image:

cd mcp-vision
make build-docker

Related MCP server: llm-vision-mcp

Configuring Claude Desktop

Add this to your claude_desktop_config.json:

If your local environment has access to a NVIDIA GPU:

"mcpServers": {
  "mcp-vision": {
    "command": "docker",
    "args": ["run", "-i", "--rm", "--runtime=nvidia", "--gpus", "all", "mcp-vision"],
	"env": {}
  }
}

Or, CPU only:

"mcpServers": {
  "mcp-vision": {
    "command": "docker",
    "args": ["run", "-i", "--rm", "mcp-vision"],
	"env": {}
  }
}

When running on CPU, the default large-size object detection model make take a long time to laod and run inference. Consider using a smaller model as DEFAULT_OBJDET_MODEL (you can tell Claude directly to use a specific model too).

(Beta) It is possible to run the public docker image directly without building locally, however the download time may interfere with Claude's loading of the server.

"mcpServers": {
  "mcp-vision": {
    "command": "docker",
    "args": ["run", "-i", "--rm", "--runtime=nvidia", "--gpus", "all", "groundlight/mcp-vision:latest"],
	"env": {}
  }
}

Tools

The following tools are currently available through the mcp-vision server:

  1. locate_objects

  • Description: Detect and locate objects in an image using one of the zero-shot object detection pipelines available through HuggingFace (list for reference [https://huggingface.co/models?pipeline_tag=zero-shot-object-detection&sort=trending]).

  • Input: image_path (string) URL or file path, candidate_labels (list of strings) list of possible objects to detect, hf_model (optional string), will use "google/owlvit-large-patch14" by default, which could be slow on a non-GPU machine

  • Returns: List of dicts in HF object-detection format

  1. zoom_to_object

  • Description: Zoom into an object in the image, allowing you to analyze it more closely. Crop image to the object bounding box and return the cropped image. If many objects are present in the image, will return the 'best' one as represented by object score.

  • Input: image_path (string) URL or file path, label (string) object label to find and zoom and crop to, hf_model (optional), will use "google/owlvit-large-patch14" by default, which could be slow on a non-GPU machine

  • Returns: MCPImage or None

Example in blog post and video

Run Claude Desktop with Claude Sonnet 3.7 and mcp-vision configured as an MCP server in claude_desktop_config.json.

The prompt used in the example video and blog post was:

From the information on that advertising board, what is the type of this shop?
Options:
The shop is a yoga studio.
The shop is a cafe.
The shop is a seven-eleven.
The shop is a milk tea shop.

The image is the first image in the V*Bench/GPT4V-hard dataset and can be found here: https://huggingface.co/datasets/craigwu/vstar_bench/blob/main/GPT4V-hard/0.JPG (use the download link).

Note:

  • If you upload the image directly into the conversation with Claude instead of providing a download link, it will not be able to call the tools and will attempt to answer directly.

  • On accounts that have web search enabled, Claude will prefer to use web search over local MCP tools AFAIK. Disable web search for best results.

Development

Run locally using the uv package manager:

uv install
uv run python mcp_vision

Build the Docker image locally:

make build-docker

Run the Docker image locally:

make run-docker-cpu

or

make run-docker-gpu

[Groundlight Internal] Push the Docker image to Docker Hub (requires DockerHub credentials):

make push-docker

Troubleshooting

If Claude Desktop is failing to connect to mcp-vision:

  • Check the configuration is correct (CPU vs GPU)

  • Developer options may need to be enabled in Claude Desktop

  • Depending on the size of the model(s) used, give it a few minutes to download them from HuggingFace on first opening Claude Desktop. Once downloaded, the server will respond and Claude will connect.

On accounts that have web search enabled, Claude will prefer to use web search over local MCP tools AFAIK. Disable web search for best results.

TODO

  • Host best models online instead of requiring local download

  • Add more tools

Available Tools

2 tools
locate_objectsC

Detect, find and/or locate objects in the image found at image_path.

Args:
    image_path: path to the image
    candidate_labels: list of candidate object labels as strings
    hf_model (optional): huggingface zero-shot object detection model (default = "google/owlvit-large-patch14")
ParametersJSON Schema
NameRequiredDescriptionDefault
hf_modelNo
image_pathYes
candidate_labelsYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the default huggingface model, but does not describe return values (e.g., bounding boxes or confidence scores), side effects, network requirements, or error behavior. The agent is left guessing at what 'locate' means in terms of output and observable behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably sized but contains redundancy: 'Detect, find and/or locate' offers three verbs for the same action, adding little information. The Args list is useful but partially duplicates schema property names. It is front-loaded with the purpose sentence, but it could be tightened without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and no annotations, the description is incomplete. It does not explain what the tool returns (e.g., bounding boxes, coordinates, labels), which is essential for an agent to interpret the result. It also lacks details on supported image formats, error handling, or how output connects to the sibling 'zoom_to_object'. The agent can invoke the tool but cannot confidently use its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description's Args section adds meaningful parameter semantics. It explains image_path as 'path to the image', candidate_labels as 'list of candidate object labels as strings', and documents the hf_model default as 'google/owlvit-large-patch14', which the schema lacks. This is valuable, though it stops short of explaining formats, constraints, or advanced usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource combination: 'Detect, find and/or locate objects in the image found at image_path.' The purpose is immediately clear and distinguishable from a generic tool. However, it does not explicitly differentiate itself from the sibling tool 'zoom_to_object', relying on the tool name and phrasing for distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the sibling 'zoom_to_object'. It does not state conditions, prerequisites, or exclusions. Users must infer usage solely from the stated purpose, which is not enough for reliable tool selection in an agentic context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

zoom_to_objectA

Zoom into an object in the image, allowing you to analyze it more closely. Crop image to the object bounding box and return the cropped image. If many objects are present in the image, will return the 'best' one as represented by object score.

Args:
    image_path: path to the image
    label: object label to find and crop to
    hf_model (optional): huggingface zero-shot object detection model (default = "google/owlvit-large-patch14")
ParametersJSON Schema
NameRequiredDescriptionDefault
labelYes
hf_modelNo
image_pathYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that this is a crop operation, returns the cropped image, uses zero-shot detection with a default model, and picks the highest-scoring object. It does not cover failure modes, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main action is front-loaded and the Args section is cleanly organized. The first sentence is slightly redundant with the second, but there is little wasted text and no irrelevant detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple crop tool with no annotations and no output schema, it covers inputs, optional configuration, selection behavior, and the return type. The main missing piece is edge-case behavior, such as what happens when the requested label is not found.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the Args block fully compensates: image_path is defined, label is described as the object label to find and crop to, and hf_model is documented as optional with its default model. This is exactly the information an agent needs to call the tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is specific: it zooms into an object, crops to its bounding box, and returns the cropped image. Although it does not name locate_objects, the crop-and-return behavior clearly differentiates it from a tool that would locate objects and return coordinates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear intended use case ('allowing you to analyze it more closely') and clarifies selection behavior when multiple objects are present. It does not explicitly contrast with locate_objects or state when not to use it, but the context is clear enough for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.1
    • First observedlocate_objects
    • First observedzoom_to_object

TDQS

A3.5/5.0

Scored across 2 tools

Disambiguation4/5

Both tools involve object detection, but they serve distinct purposes: locate_objects returns detections while zoom_to_object crops to the best object. The overlap is minimal, and descriptions clarify the intended usage, though an agent might occasionally hesitate between the two.

Naming Consistency5/5

Both tool names follow a consistent verb_noun pattern (locate_objects, zoom_to_object), making them predictable and easy to distinguish at a glance.

Tool Count3/5

With only 2 tools, the server feels thin for a general vision MCP, but it is narrowly scoped to object detection and zooming. This borderline count is acceptable for a specialized utility, though it lacks breadth.

Completeness3/5

The surface covers the core workflow of detecting and cropping objects, but it omits common vision operations like classification or segmentation. For the narrow purpose stated, it is functional, but the 'vision' name implies a broader scope that is not fulfilled.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that enables any LLM to describe images from file paths, URLs, or base64 data by forwarding them to a supported vision provider such as OpenAI, Anthropic, or local Ollama models.
    1,150 npm
    9
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A local MCP server that gives LLMs eyes for images by performing object detection (YOLOv8) and text recognition (EasyOCR), outputting descriptive statements about objects and text positions without any API key or cloud dependency.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server for image recognition and OCR via OpenAI-compatible vision APIs, supporting local files, URLs, and data URLs. Enables natural language image description and text extraction.
    10 npm
    2
    MIT