Skip to main content
Glama

page_click

Destructive

Click a CSS selector or viewport coordinates in a real browser tab, with hit-testing, auto-scroll, and iframe support for reliable automation.

Instructions

Click a CSS/structured locator or viewport coordinates in a specific real browser tab using background CDP input. Coordinates are viewport-relative CSS pixels (the space getBoundingClientRect reports), NOT the device pixels capture_page_screenshot returns -- on a scaled display divide a screenshot pixel by devicePixelRatio first. Ambiguous or unreachable targets dispatch nothing; the tab is not activated and the desktop cursor does not move. Selector offsets are measured from the element's top-left corner; an omitted axis uses the element centre. In selector mode, duplicate CSS/structured matches are reduced to visible, interactable candidates so hidden modal templates do not win, then the point is hit-tested before anything is dispatched: an element below the fold is scrolled into view, and a point owned by another element returns status 'obscured' (with occluded_by) or 'outside_viewport' having clicked nothing. Coordinate mode is not hit-tested -- coordinates name a pixel, not an element. A selector click crossing a non-identity CSS-transformed iframe returns status 'unsupported_frame_transform' without dispatch; query/type paths remain available. A structured selector may use {'selector': '#pay', 'frame': [...]} as a CSS alias, or {'frame': [...], 'x': 20, 'y': 30} to click a point inside the final same-origin or cross-origin frame; frame-point mode is not hit-tested. Nested frame locators also support OOPIFs. Framed clicks do not scroll automatically and check every parent for obstruction. A binding invalidated during the call returns stale_frame; inspect input_dispatched before recovery and never replay partial or unknown input.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
buttonNoleft
clicksNo
timeoutNo
offset_xNo
offset_yNo
selectorNo
session_idNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even with destructiveHint=true in annotations, the description adds substantial behavioral detail well beyond annotations: background CDP dispatch, no tab activation, no desktop cursor movement, none dispatched for ambiguous/unreachable targets, hit-testing, scrolling, obscured/outside_viewport, unsupported_frame_transform, stale_frame, and input_dispatched guidance. This is far beyond the annotation metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but dense and logically organized, front-loading the core purpose and critical coordinate caveat before mode-specific details. Every sentence adds distinct information, though headings or bullet structure could improve scannability. It is appropriately detailed for the tool's complexity, not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with a long tail of edge cases, the description is remarkably complete: it covers coordinate conversion, selectors, offsets, frame and OOPIF behavior, hit-testing, statuses, scrolling, and stale-frame recovery with a warning to inspect input_dispatched. Since an output schema exists, return-value details are not required from the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries most of the semantic burden, and it delivers for key parameters: x/y as viewport CSS pixels requiring devicePixelRatio conversion, selector object forms as CSS aliases or frame-point mode, and offset_x/y measured from top-left with omitted axis centered. It does not explain button, clicks, timeout, or session_id, so it is not fully complete for every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: "Click a CSS/structured locator or viewport coordinates in a specific real browser tab using background CDP input." It clearly identifies two input modes, distinguishes itself from screenshot coordinate semantics by referencing capture_page_screenshot, and separates click behavior from type/press siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use coordinate mode versus selector mode, including explicit differences like "Coordinate mode is not hit-tested -- coordinates name a pixel, not an element." It also references fallback "query/type paths" in the transformed-iframe case, though it does not systematically enumerate sibling alternatives for all scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.