Skip to main content
Glama
xzaleksey

open-mobile-mcp

by xzaleksey

tap_on_element

Tap any UI element on Android or iOS by text, test ID, or content description. Automatically locates the element and taps its center, with optional long press and post-tap screenshot verification.

Instructions

🥇 RECOMMENDED: Tap on a UI element by selector. Finds the element and taps its center automatically. Prefer over device_tap. If observing a log triggered by this tap, spawn the wait_for_log subagent BEFORE tapping. NOTE: text matching is exact — if an element renders with an emoji prefix (e.g. '🇫🇷 French A2'), passing 'French A2' will fail. Use get_semantic_hierarchy first to see the exact text, or use contentDescription/testId strategy instead.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
verifyNoIf true, capture and return a screenshot right after the tap — saves a separate get_viewport round-trip when you're about to check the result anyway.
deviceIdYes
durationNoOptional duration in ms. If > 0, performs a long press.
platformYes
selectorYesText, testId, or content description to find
strategyYesHow to find the element: 'text' (visible text), 'testId' (accessibility ID), or 'contentDescription'

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.1.4

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral disclosure burden. It does well by revealing that text matching is exact, that emoji prefixes can cause failures, that the tap targets the element center, that duration > 0 changes behavior to long press, and that verify captures a screenshot. A small gap is the lack of failure behavior or off-screen handling, but the disclosed caveats are highly relevant and actionable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place: recommendation, behavior, sibling comparison, log-spawning order, exact-match caveat, and mitigation strategy. The key advice is front-loaded with 'RECOMMENDED' and the most important behavioral caveat appears before the trailing example. Slightly dense, but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, this description is remarkably complete. It tells the agent which sibling to prefer, when to use a different sibling, what precondition to check (exact visible text via get_semantic_hierarchy), what parameter combination to use as a fallback, and how verify optimizes the workflow. An agent has enough context to select and invoke this tool correctly without opening other tool definitions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, with deviceId and platform left to inference. The description adds meaning beyond the schema by explaining exact-match semantics for the selector, warning about emoji prefixes, and clarifying the verify option as a way to avoid a separate get_viewport round-trip. This genuinely supports parameter selection, though it does not add detail for the two unnamed parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Tap on a UI element by selector' with automatic center targeting. It also distinguishes itself from device_tap by explicitly recommending this tool over that sibling, so an agent can immediately tell what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Prefer over device_tap' and provides precise routing guidance: use get_semantic_hierarchy first when text matching may be affected by prefixes, or fall back to contentDescription/testId strategy. It also instructs the agent to spawn wait_for_log before tapping when observing a log triggered by the tap. This is exemplary usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.