mcp-vision
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-visionWhat objects are in this image? https://example.com/pic.jpg"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-vision by
A Model Context Protocol (MCP) server exposing HuggingFace computer vision models such as zero-shot object detection as tools, enhancing the vision capabilities of large language or vision-language models.
This repo is in active development. See below for details of currently available tools.
Installation
Clone the repo:
git clone git@github.com:groundlight/mcp-vision.gitBuild a local docker image:
cd mcp-vision
make build-dockerRelated MCP server: llm-vision-mcp
Configuring Claude Desktop
Add this to your claude_desktop_config.json:
If your local environment has access to a NVIDIA GPU:
"mcpServers": {
"mcp-vision": {
"command": "docker",
"args": ["run", "-i", "--rm", "--runtime=nvidia", "--gpus", "all", "mcp-vision"],
"env": {}
}
}Or, CPU only:
"mcpServers": {
"mcp-vision": {
"command": "docker",
"args": ["run", "-i", "--rm", "mcp-vision"],
"env": {}
}
}When running on CPU, the default large-size object detection model make take a long time to laod and run inference. Consider using a smaller model as DEFAULT_OBJDET_MODEL (you can tell Claude directly to use a specific model too).
(Beta) It is possible to run the public docker image directly without building locally, however the download time may interfere with Claude's loading of the server.
"mcpServers": {
"mcp-vision": {
"command": "docker",
"args": ["run", "-i", "--rm", "--runtime=nvidia", "--gpus", "all", "groundlight/mcp-vision:latest"],
"env": {}
}
}Tools
The following tools are currently available through the mcp-vision server:
locate_objects
Description: Detect and locate objects in an image using one of the zero-shot object detection pipelines available through HuggingFace (list for reference [https://huggingface.co/models?pipeline_tag=zero-shot-object-detection&sort=trending]).
Input:
image_path(string) URL or file path,candidate_labels(list of strings) list of possible objects to detect,hf_model(optional string), will use"google/owlvit-large-patch14"by default, which could be slow on a non-GPU machineReturns: List of dicts in HF object-detection format
zoom_to_object
Description: Zoom into an object in the image, allowing you to analyze it more closely. Crop image to the object bounding box and return the cropped image. If many objects are present in the image, will return the 'best' one as represented by object score.
Input:
image_path(string) URL or file path,label(string) object label to find and zoom and crop to,hf_model(optional), will use"google/owlvit-large-patch14"by default, which could be slow on a non-GPU machineReturns: MCPImage or None
Example in blog post and video
Run Claude Desktop with Claude Sonnet 3.7 and mcp-vision configured as an MCP server in claude_desktop_config.json.
The prompt used in the example video and blog post was:
From the information on that advertising board, what is the type of this shop?
Options:
The shop is a yoga studio.
The shop is a cafe.
The shop is a seven-eleven.
The shop is a milk tea shop.The image is the first image in the V*Bench/GPT4V-hard dataset and can be found here: https://huggingface.co/datasets/craigwu/vstar_bench/blob/main/GPT4V-hard/0.JPG (use the download link).
Note:
If you upload the image directly into the conversation with Claude instead of providing a download link, it will not be able to call the tools and will attempt to answer directly.
On accounts that have web search enabled, Claude will prefer to use web search over local MCP tools AFAIK. Disable web search for best results.
Development
Run locally using the uv package manager:
uv install
uv run python mcp_visionBuild the Docker image locally:
make build-dockerRun the Docker image locally:
make run-docker-cpuor
make run-docker-gpu[Groundlight Internal] Push the Docker image to Docker Hub (requires DockerHub credentials):
make push-dockerTroubleshooting
If Claude Desktop is failing to connect to mcp-vision:
Check the configuration is correct (CPU vs GPU)
Developer options may need to be enabled in Claude Desktop
Depending on the size of the model(s) used, give it a few minutes to download them from HuggingFace on first opening Claude Desktop. Once downloaded, the server will respond and Claude will connect.
On accounts that have web search enabled, Claude will prefer to use web search over local MCP tools AFAIK. Disable web search for best results.
TODO
Host best models online instead of requiring local download
Add more tools
Available Tools
2 toolslocate_objectsC
Detect, find and/or locate objects in the image found at image_path.
Args:
image_path: path to the image
candidate_labels: list of candidate object labels as strings
hf_model (optional): huggingface zero-shot object detection model (default = "google/owlvit-large-patch14")
| Name | Required | Description | Default |
|---|---|---|---|
| hf_model | No | ||
| image_path | Yes | ||
| candidate_labels | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the default huggingface model, but does not describe return values (e.g., bounding boxes or confidence scores), side effects, network requirements, or error behavior. The agent is left guessing at what 'locate' means in terms of output and observable behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably sized but contains redundancy: 'Detect, find and/or locate' offers three verbs for the same action, adding little information. The Args list is useful but partially duplicates schema property names. It is front-loaded with the purpose sentence, but it could be tightened without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema and no annotations, the description is incomplete. It does not explain what the tool returns (e.g., bounding boxes, coordinates, labels), which is essential for an agent to interpret the result. It also lacks details on supported image formats, error handling, or how output connects to the sibling 'zoom_to_object'. The agent can invoke the tool but cannot confidently use its result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description's Args section adds meaningful parameter semantics. It explains image_path as 'path to the image', candidate_labels as 'list of candidate object labels as strings', and documents the hf_model default as 'google/owlvit-large-patch14', which the schema lacks. This is valuable, though it stops short of explaining formats, constraints, or advanced usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource combination: 'Detect, find and/or locate objects in the image found at image_path.' The purpose is immediately clear and distinguishable from a generic tool. However, it does not explicitly differentiate itself from the sibling tool 'zoom_to_object', relying on the tool name and phrasing for distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the sibling 'zoom_to_object'. It does not state conditions, prerequisites, or exclusions. Users must infer usage solely from the stated purpose, which is not enough for reliable tool selection in an agentic context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zoom_to_objectA
Zoom into an object in the image, allowing you to analyze it more closely. Crop image to the object bounding box and return the cropped image. If many objects are present in the image, will return the 'best' one as represented by object score.
Args:
image_path: path to the image
label: object label to find and crop to
hf_model (optional): huggingface zero-shot object detection model (default = "google/owlvit-large-patch14")
| Name | Required | Description | Default |
|---|---|---|---|
| label | Yes | ||
| hf_model | No | ||
| image_path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that this is a crop operation, returns the cropped image, uses zero-shot detection with a default model, and picks the highest-scoring object. It does not cover failure modes, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main action is front-loaded and the Args section is cleanly organized. The first sentence is slightly redundant with the second, but there is little wasted text and no irrelevant detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple crop tool with no annotations and no output schema, it covers inputs, optional configuration, selection behavior, and the return type. The main missing piece is edge-case behavior, such as what happens when the requested label is not found.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the Args block fully compensates: image_path is defined, label is described as the object label to find and crop to, and hf_model is documented as optional with its default model. This is exactly the information an agent needs to call the tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is specific: it zooms into an object, crops to its bounding box, and returns the cropped image. Although it does not name locate_objects, the crop-and-return behavior clearly differentiates it from a tool that would locate objects and return coordinates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear intended use case ('allowing you to analyze it more closely') and clarifies selection behavior when multiple objects are present. It does not explicitly contrast with locate_objects or state when not to use it, but the context is clear enough for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.1- First observed
locate_objects - First observed
zoom_to_object
TDQS
Scored across 2 tools
Both tools involve object detection, but they serve distinct purposes: locate_objects returns detections while zoom_to_object crops to the best object. The overlap is minimal, and descriptions clarify the intended usage, though an agent might occasionally hesitate between the two.
Both tool names follow a consistent verb_noun pattern (locate_objects, zoom_to_object), making them predictable and easy to distinguish at a glance.
With only 2 tools, the server feels thin for a general vision MCP, but it is narrowly scoped to object detection and zooming. This borderline count is acceptable for a specialized utility, though it lacks breadth.
The surface covers the core workflow of detecting and cropping objects, but it omits common vision operations like classification or segmentation. For the narrow purpose stated, it is functional, but the 'vision' name implies a broader scope that is not fulfilled.
Maintenance
Related MCP Connectors
MCP server for OpenAI API (chat completions, image generation, embeddings) via AceDataCloud
MCP server for Qwen Image 3 AI image generation
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- AlicenseAqualityDmaintenanceAn MCP server that gives AI agents the ability to observe and understand images via multi-provider vision, object detection, hierarchical analysis, and color extraction.49 npm2MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that enables any LLM to describe images from file paths, URLs, or base64 data by forwarding them to a supported vision provider such as OpenAI, Anthropic, or local Ollama models.1,150 npm9MIT
- AlicenseNot gradedqualityDmaintenanceA local MCP server that gives LLMs eyes for images by performing object detection (YOLOv8) and text recognition (EasyOCR), outputting descriptive statements about objects and text positions without any API key or cloud dependency.MIT
- AlicenseNot gradedqualityBmaintenanceAn MCP server for image recognition and OCR via OpenAI-compatible vision APIs, supporting local files, URLs, and data URLs. Enables natural language image description and text extraction.10 npm2MIT