Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

wait_for

Wait for a screen to settle or a specified text/app to appear after triggering an action, so the next step reads a stable screen instead of a mid-animation one.

Instructions

Wait until something becomes true, instead of sleeping a fixed time.

Use this after an action that triggers loading, so the next step reads a settled screen instead of one mid-animation. With no arguments it waits for the screen to stop changing.

Args: text: Wait until this text appears anywhere on screen. app: Wait until this app (a name or a package id) is in the foreground. timeout_seconds: Give up and report the timeout after this long.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
appNoWait until this app is in the foreground.
textNoThe text to type, or the text to replace the clipboard with.
timeout_secondsNoGive up and report the timeout after this long.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the disclosure burden. It explains that the tool waits for a condition rather than sleeping, that it can wait for screen stability, and that timeout_seconds causes it to give up and report the timeout. It does not detail polling mechanics, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core behavior, then gives one-sentence usage guidance, then a compact Args list. Every sentence adds value and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-argument polling tool with no annotations, it covers the main behavior, the recommended use case, all parameters, and timeout behavior. It does not say how multiple conditions combine or what the timeout response looks like, but the output schema can cover the return shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The Args section maps each parameter to its meaning and adds useful detail like app being 'a name or a package id' and text matching 'anywhere on screen.' This compensates for the input schema's text description, which appears copied from a type_text tool ('The text to type, or the text to replace the clipboard with.') and is misleading for wait_for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and outcome ('Wait until something becomes true, instead of sleeping a fixed time') and then gives three concrete wait conditions: text on screen, app in foreground, screen stable. This clearly distinguishes it from a blind sleep and from the sibling read/screen tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use it: 'Use this after an action that triggers loading, so the next step reads a settled screen instead of one mid-animation.' It does not provide formal when-not-to-use conditions or name alternatives, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.