Skip to main content
Glama

phone_tap_element

Destructive

Taps a specified UI element on an Android device after re-reading the UI and confirming the snapshot matches, ensuring accurate interaction.

Instructions

Re-read UI and tap an enabled element only if the snapshot still matches.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
serialNo
snapshotYes
element_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the agent knows this mutates device state and is not repeatable. The description adds one genuine behavioral fact beyond that: it re-reads the UI and aborts the tap on a snapshot mismatch, which is a useful staleness guard. It does not say what happens on mismatch (error vs. silent no-op) or whether a fresh snapshot is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the precondition front-loaded and no filler. It is efficient, though its extreme brevity is partly the cause of the semantic gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but for a destructive tap tool with three parameters at 0% schema coverage the description leaves key facts unstated: the source and format of snapshot, the role of serial, and the failure behavior when the snapshot does not match. The guard concept is present but not operationalized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it only partially pays it: it clarifies that 'snapshot' is a staleness check token, but never explains where the snapshot comes from or its format, and 'serial' is not mentioned at all. 'element_id' is only implied by the word 'element'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (tap) and resource (element), plus a guard condition (only if the snapshot still matches), which is enough to separate it from the coordinate-based sibling phone_tap. It does not explicitly name the sibling it differs from, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'only if the snapshot still matches' implies usage context (act on a previously observed UI state rather than blindly tapping), but there is no explicit when-to-use vs. phone_tap / phone_find_elements guidance and no stated prerequisites. The agent must infer that a snapshot must first be obtained elsewhere.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.