mcp-vision
OfficialServer Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| locate_objectsC | Detect, find and/or locate objects in the image found at image_path. |
| zoom_to_objectA | Zoom into an object in the image, allowing you to analyze it more closely. Crop image to the object bounding box and return the cropped image. If many objects are present in the image, will return the 'best' one as represented by object score. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
Both tools involve object detection, but they serve distinct purposes: locate_objects returns detections while zoom_to_object crops to the best object. The overlap is minimal, and descriptions clarify the intended usage, though an agent might occasionally hesitate between the two.
Both tool names follow a consistent verb_noun pattern (locate_objects, zoom_to_object), making them predictable and easy to distinguish at a glance.
With only 2 tools, the server feels thin for a general vision MCP, but it is narrowly scoped to object detection and zooming. This borderline count is acceptable for a specialized utility, though it lacks breadth.
The surface covers the core workflow of detecting and cropping objects, but it omits common vision operations like classification or segmentation. For the narrow purpose stated, it is functional, but the 'vision' name implies a broader scope that is not fulfilled.