Skip to main content
Glama

locate_objects

Locate specified objects in an image by supplying its path and candidate labels. Uses zero-shot object detection to identify and return positions for each label.

Instructions

Detect, find and/or locate objects in the image found at image_path.

Args:
    image_path: path to the image
    candidate_labels: list of candidate object labels as strings
    hf_model (optional): huggingface zero-shot object detection model (default = "google/owlvit-large-patch14")

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
hf_modelNo
image_pathYes
candidate_labelsYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the default huggingface model, but does not describe return values (e.g., bounding boxes or confidence scores), side effects, network requirements, or error behavior. The agent is left guessing at what 'locate' means in terms of output and observable behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably sized but contains redundancy: 'Detect, find and/or locate' offers three verbs for the same action, adding little information. The Args list is useful but partially duplicates schema property names. It is front-loaded with the purpose sentence, but it could be tightened without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and no annotations, the description is incomplete. It does not explain what the tool returns (e.g., bounding boxes, coordinates, labels), which is essential for an agent to interpret the result. It also lacks details on supported image formats, error handling, or how output connects to the sibling 'zoom_to_object'. The agent can invoke the tool but cannot confidently use its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description's Args section adds meaningful parameter semantics. It explains image_path as 'path to the image', candidate_labels as 'list of candidate object labels as strings', and documents the hf_model default as 'google/owlvit-large-patch14', which the schema lacks. This is valuable, though it stops short of explaining formats, constraints, or advanced usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource combination: 'Detect, find and/or locate objects in the image found at image_path.' The purpose is immediately clear and distinguishable from a generic tool. However, it does not explicitly differentiate itself from the sibling tool 'zoom_to_object', relying on the tool name and phrasing for distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the sibling 'zoom_to_object'. It does not state conditions, prerequisites, or exclusions. Users must infer usage solely from the stated purpose, which is not enough for reliable tool selection in an agentic context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools