Skip to main content
Glama

Click

click
Destructive

Activate a desktop or browser UI element using an accessibility ref, visual mark, or image point; returns what changed and the next observation ID.

Instructions

Click an element: prefer ref, then mark (from a som observation), then point (image pixels, last resort). Pass the observation_id you act on. Returns what changed and the next observation_id.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
refNo
markNo
countNo
pointNo
buttonNoleft
modifiersNo
observation_idYesobservation_id of the observation you act on

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and non-idempotent, so the agent knows this mutates state. The description adds that it returns what changed and the next observation_id, which is behavioral context beyond the annotations. However, it doesn't elaborate on side effects such as whether the click might trigger navigation, downloads, or other destructive results beyond the generic hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with the action and target selection rule, followed by a required argument and return value. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A complex, destructive, multi-mode tool with an output schema and seven parameters. The description covers the essential selection logic and the crucial observation_id requirement, and defers return details to the output schema. It omits behavior for count/button/modifiers, but for a tool at this complexity level, most of the critical guidance is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 14%, so the description must compensate. It explains the priority and origin of ref, mark, and point ('from a som observation', 'image pixels, last resort'), and says to pass the observation_id you act on. It doesn't mention count, button, or modifiers, which is a gap, but the most critical mutually-exclusive parameters are covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Click) on a specific resource (an element) and immediately lists the three targeting modes with a priority order. It's distinguishable from siblings like type_text or drag because it names the element-targeting mechanisms.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent the preference order: 'prefer ref, then mark (from a som observation), then point (image pixels, last resort).' This is exactly the kind of when-to-use-which guidance that helps an agent choose among the mutually exclusive targeting parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.