Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

read_screen

List interactive controls currently on the Android screen, each with a number and label, to target taps or verify UI state. Optional text query filters results.

Instructions

List the controls currently on the phone's screen, by number and name.

This is the main way to see what is on screen, and it is far cheaper than a screenshot. Each control gets a number you can hand to tap, and a label you can match on.

Every action tool re-reads the screen before it acts, so both forms are resolved against the screen as it is at that moment. A number is precise -- it disambiguates two controls carrying the same text. A label is stable -- it survives the screen being redrawn. Use the number for a control you have just seen and are about to act on, and the label when the screen may have changed since you read it.

Args: query: Optional filter. Only controls whose text, description, or type contains this text are listed, best match first. Use it to ask "is the Send button up yet?" in one call.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryNoOnly list controls whose text or description contains this. Omit to list everything on the screen.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and excels. It discloses the return format (numbered controls with labels), the resolution semantics (numbers are precise, labels are stable), and that action tools re-read the screen, implying read_screen results may be stale after a redraw. This is rich behavioral context beyond what any annotation would add.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then layers value (cost comparison, number/label guidance) and finally parameter details. It is efficiently structured with clear paragraphs and zero filler. Every sentence earns its place, and the length is justified by the nuanced distinction it explains.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one optional parameter) and the presence of an output schema, the description covers all necessary aspects: what it does, how it differs from a screenshot, when to rely on numbers vs labels, when not to call it, and how to use the filter. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already describes the `query` parameter (100% coverage), the description enriches it: it specifies the matching criteria (text, description, or type), notes the ordering ('best match first'), and gives a practical example. This goes well beyond the schema's bare description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise statement: 'List the controls currently on the phone's screen, by number and name.' It explicitly differentiates from the sibling `take_screenshot` by noting it is 'far cheaper' and returns interactive controls rather than an image. This makes the tool's purpose unmistakable and distinct from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use this tool versus `take_screenshot` (cheaper for control info), and when not to call it at all: 'Every action tool re-reads the screen before it acts.' It also gives specific rules for choosing between number and label forms, and demonstrates a concrete use case for the query filter ('is the Send button up yet?'). These are explicit usage directives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.