Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
locate_objectsC

Detect, find and/or locate objects in the image found at image_path.

Args:
    image_path: path to the image
    candidate_labels: list of candidate object labels as strings
    hf_model (optional): huggingface zero-shot object detection model (default = "google/owlvit-large-patch14")
zoom_to_objectA

Zoom into an object in the image, allowing you to analyze it more closely. Crop image to the object bounding box and return the cropped image. If many objects are present in the image, will return the 'best' one as represented by object score.

Args:
    image_path: path to the image
    label: object label to find and crop to
    hf_model (optional): huggingface zero-shot object detection model (default = "google/owlvit-large-patch14")

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.5/5.0

Scored across 2 tools

Disambiguation4/5

Both tools involve object detection, but they serve distinct purposes: locate_objects returns detections while zoom_to_object crops to the best object. The overlap is minimal, and descriptions clarify the intended usage, though an agent might occasionally hesitate between the two.

Naming Consistency5/5

Both tool names follow a consistent verb_noun pattern (locate_objects, zoom_to_object), making them predictable and easy to distinguish at a glance.

Tool Count3/5

With only 2 tools, the server feels thin for a general vision MCP, but it is narrowly scoped to object detection and zooming. This borderline count is acceptable for a specialized utility, though it lacks breadth.

Completeness3/5

The surface covers the core workflow of detecting and cropping objects, but it omits common vision operations like classification or segmentation. For the narrow purpose stated, it is functional, but the 'vision' name implies a broader scope that is not fulfilled.

Maintenance

ActivityInactive
ResponsivenessNo issues