Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

ask_jev

Ask typed questions about the phone screen to judge action success, error states, or visible choices. Include plain-language instructions and optional options/state for a decisive answer.

Instructions

Ask Jev a typed question about what is on the phone screen.

Use this for judgements that are not "what do I tap next": whether an action worked, whether the screen is an error state, or which of several visible items the user meant.

Jev reads text only — the control listing, plus whatever state you give it. It cannot see the screenshot, so keep the question about what the controls say and do, not about how they look.

Args: instructions: The question, in plain language. question_type: "choice" to pick one option, "yes_or_no" for the probability that something holds, or "scale" to place the screen on an ordered scale. options: For "choice", the alternatives. For "scale", the levels from lowest to highest. Write each as "short_key: what it means" to name the answer, or give just the meaning to key it by position. state: Extra context to judge against. The current screen reading is always included; use this for the goal or instruction being checked.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
stateNoExtra context to judge against. The current screen is always included.
optionsNoFor 'choice' and 'scale', the alternatives. For 'choice' write each as 'short_key: what it means'.
instructionsYesThe question, in plain language.
question_typeNoWhich shape of answer: 'choice' picks one option, 'yes_or_no' gives a probability, 'scale' places the state on a range.choice

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does so well: it discloses that Jev reads text only, cannot see the screenshot, and always receives the current screen reading plus any provided state. It does not explicitly state that the call is side-effect-free, but 'Ask... a question' strongly implies a read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but every section earns its place: purpose, exclusions, a critical limitation, and parameter guidance. The Args block is organized and front-loaded after the core usage statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations and an output schema present, the description is complete: it explains when to use it, what it can and cannot see, how to format each parameter, and what context to provide. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value by explaining the options format ('short_key: what it means' or positional keys), the meaning of question_type variants, and how state should be used as the goal or instruction being checked.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Ask Jev a typed question about what is on the phone screen.' It then distinguishes itself from the 'what do I tap next' class of action, which separates it from sibling decide_next_action without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit use cases: judging whether an action worked, recognizing error states, and disambiguating visible items. It also states what it is not for ('not what do I tap next'), though it does not name the sibling tool that should handle that case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.