Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

run_task

Execute a high-level goal on an Android phone by reading the screen, deciding actions, and iterating until done. Provide a plain-language goal; the tool handles the rest, reporting steps and outcome.

Instructions

Give the phone a goal and let it work the whole thing out itself.

Use this instead of driving the phone a step at a time. It reads the screen, decides what to do, does it, and reads again, until the goal is reported done or unreachable — so "open the email app" is one call here rather than ten round trips through you.

Every step reports the action, what came of it, and the two numbers Jev gave for the choice: probability is how likely that option was the best one and confidence is how sure it was of its own ranking. They are reported rather than gated on, so a run that succeeded on a thin lead says so, and a run that stopped for a reason other than reaching the goal says that instead.

achieved is what the run claimed. Read the steps for what actually happened. If a run needs looking at afterwards, set PHONE_CONTROL_RUNS to a directory and each run leaves a folder there with the full record of what it asked Jev and what came back; the reply then names the folder.

Prefer the step-by-step tools when you already know exactly what to tap, when the screen is visual rather than textual, or when you need to inspect something part-way through.

Args: goal: What to achieve, in plain language. For example "open the email app" or "turn on airplane mode". max_steps: How many actions to allow before giving up. Reaching it ends the run honestly as unfinished rather than as a failure of the goal.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
goalYesWhat you want done, in plain language. Phrase it as the outcome, not the steps.
max_stepsNoHow many actions to allow before giving up.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the iterative read-decide-act loop, the reporting of probability and confidence, the distinction between claimed achievement and actual steps, the optional run log directory, and the max_steps termination behavior. No contradictions with annotations since none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence adds value. It front-loads the core concept, then details behavior and parameter nuances. Slightly more verbose than necessary, but each paragraph earns its place. A 4 rather than 5 due to minor redundancy in the step-by-step preference explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an autonomous agent tool with multiple behavioral nuances, the description covers usage, alternatives, parameter semantics, and operational details. It even mentions the optional environment variable for logging. The output schema likely covers return values, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds significant extra meaning: it explains the goal should be phrased as an outcome, provides examples, and clarifies max_steps as a cap that results in an honest 'unfinished' state rather than failure. This goes beyond the schema's basic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it takes a goal and autonomously drives the phone to completion, contrasting with step-by-step tools. It distinguishes itself from siblings like 'tap' and 'open_app' by emphasizing the autonomous loop, making it unambiguous what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent when to use this tool ('Use this instead of driving the phone a step at a time') and when to prefer alternatives ('Prefer the step-by-step tools when you already know exactly what to tap, when the screen is visual rather than textual, or when you need to inspect something part-way through'). This is textbook usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.