Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

tap

Tap or long-press screen elements by visible label, control number, or pixel coordinates to perform actions on an Android phone via adb.

Instructions

Tap, or long-press, something on the phone screen.

Args: target: The number from read_screen (for example 7), or the visible label of the control ("Send", "Sign in"). A label is matched against the control's text and description, and the best match wins. The screen is re-read first, so the target is resolved against what is on screen now; the reply names exactly what was tapped. x: Horizontal pixel coordinate. Use only when nothing on screen is nameable, such as a spot on a map or photo. y: Vertical pixel coordinate, used together with x. long_press: Hold instead of tapping, for context menus and drag handles.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
xNoHorizontal pixel coordinate, for tapping a point instead of a control.
yNoVertical pixel coordinate, for tapping a point instead of a control.
targetNoWhich control, by its number from read_screen or by its visible label.
long_pressNoHold instead of tapping, for a context menu or a drag handle.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses that the screen is re-read first, that target resolution uses text and description with best-match semantics, and that the reply names exactly what was tapped. This goes well beyond a generic 'tap' explanation and helps the agent predict runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a one-line summary followed by clearly labeled argument explanations. Every sentence adds useful information—examples, conditions, and behavioral notes—with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with four parameters and no annotations, the description covers all parameters, explains when to use each mode, describes the matching algorithm, and notes the response behavior. The presence of an output schema makes the lack of a detailed return-value explanation unnecessary; the definition is complete for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema coverage is 100%, the description adds substantial meaning beyond the schema: target may be an integer from read_screen or a visible label, matching is against text/description with best-match, x/y are coordinate fallbacks for non-nameable spots, and long_press triggers context-menu or drag-handle behavior. This richly compensates for the otherwise terse property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action and resource: 'Tap, or long-press, something on the phone screen.' This clearly identifies the tool as a screen-interaction primitive and differentiates it from siblings like swipe, scroll, and press_key, which involve different gestures or targets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage direction: target can be a read_screen number or visible label, and x/y should 'only' be used when 'nothing on screen is nameable.' It also explains long_press for context menus and drag handles. It stops short of explicitly contrasting tap with alternative sibling tools, but the guidance is sufficient for most cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.