Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

decide_next_action

Decides which on-screen control best advances your goal and returns a calibrated confidence to guide whether to act.

Instructions

Ask Jev which single control on screen best advances a goal.

Turns "send the message" or "log in" into one concrete action. Jev is a System One decision model: it answers with a typed choice drawn from the controls actually on screen, so it cannot invent a control that is not there, and it returns a calibrated confidence alongside the answer.

Jev reads text only. It is sent the numbered listing of on-screen controls plus your goal, and it never sees a screenshot — so ask it what to tap, not what the screen looks like. Anything about appearance (colour, layout, what a photo shows) is yours to judge from take_screenshot.

Read the confidence before acting. It describes how far the leading option stands from the others — it is not whether Jev could answer, and not whether the action is safe:

  • 0.70 and above: carry out the returned action.

  • 0.45 to 0.70: carry it out, then confirm with read_screen.

  • below 0.45: the leading options are near a coin flip. Do not act on it; call take_screenshot and decide from the image yourself.

Args: goal: What the user wants, in plain language. For example "reply to the most recent message saying I will be late". max_candidates: The most controls to offer in one choice. The default leaves room for the six built-in actions inside Jev's 255-option limit, so every control on screen is normally offered: measured over paired problems, accuracy held from 2 options to 255, and what costs accuracy is options that resemble each other, not how many there are. Lower this only to reduce token cost.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
goalYesWhat you want done, in plain language. Phrase it as the outcome, not the steps.
max_candidatesNoThe most controls to offer at once. Lower only to reduce token cost.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and succeeds: it discloses that Jev is text-only, never sees screenshots, cannot invent controls, and returns a calibrated confidence. It also carefully explains what confidence does and does not mean, adding a safety-relevant caveat that is easy for an agent to act on correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and organized into scannable confidence bandscars. It is somewhat long, and the paired-problems accuracy note could be trimmed, but every section serves some operational need, so the length is mostly justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter decision tool, the description covers invocation, output shape, confidence thresholds, uncertainty handling, and sibling-tool relationships. Since an output schema exists, the description does not need to detail return fields further; nothing an agent needs to call the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers both parameters, but the description goes well beyond it: it gives an example goal phrase, explains the six built-in actions within the 255-option limit, and clarifies that accuracy depends on option resemblance rather than count. This materially improves an agent's ability to set `max_candidates` appropriately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete verb–resource relation: 'Ask Jev which single control on screen best advances a goal,' and frames the output as a typed choice drawn from actual controls. It also distinguishes itself from `take_screenshot` and `read_screen` by clarifying that appearance is judged from screenshots, not this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use and alternative guidance: ask what to tap, not what the screen looks like, and leave appearance judgments to `take_screenshot`. It even provides confidence-based decision rules that route the agent to `read_screen` or `take_screenshot` when appropriate, leaving little room for misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.