Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
JEV_API_KEYNoThe Jev key, enabling the two decision tools. An OpenRouter sk-or-... key or a TypeSafe apikey_... one — its own shape routes it.
DSH_PHONE_SERIALNoWhich device to drive when several are attached.
PHONE_CONTROL_ADBNoPath to adb when it is not on PATH.
OPENROUTER_API_KEYNoThe same key under its older name, still read.
PHONE_CONTROL_RUNSNoDirectory where every run writes its report and step details; the tool's reply names it.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
statusA

Check the phone connection and see what it is currently doing.

Call this first when a command fails, and before starting anything multi-step. It reports the model, Android version, screen size, foreground app, whether the screen is on or locked, and the battery level. It changes nothing on the phone.

read_screenA

List the controls currently on the phone's screen, by number and name.

This is the main way to see what is on screen, and it is far cheaper than a screenshot. Each control gets a number you can hand to tap, and a label you can match on.

Every action tool re-reads the screen before it acts, so both forms are resolved against the screen as it is at that moment. A number is precise -- it disambiguates two controls carrying the same text. A label is stable -- it survives the screen being redrawn. Use the number for a control you have just seen and are about to act on, and the label when the screen may have changed since you read it.

Args: query: Optional filter. Only controls whose text, description, or type contains this text are listed, best match first. Use it to ask "is the Send button up yet?" in one call.

take_screenshotA

Capture the phone's screen and look at it.

Try read_screen first: it names things so you can act on the result, and it costs far less. Take a screenshot when the screen is genuinely visual -- a photo, a map, a game, a CAPTCHA, an icon-only toolbar -- when read_screen returns nothing useful, or when a decision came back with low confidence.

Args: max_width: Downscale the image to this width in pixels before returning it. Lower it to spend fewer tokens; 0 keeps full resolution.

tapA

Tap, or long-press, something on the phone screen.

Args: target: The number from read_screen (for example 7), or the visible label of the control ("Send", "Sign in"). A label is matched against the control's text and description, and the best match wins. The screen is re-read first, so the target is resolved against what is on screen now; the reply names exactly what was tapped. x: Horizontal pixel coordinate. Use only when nothing on screen is nameable, such as a spot on a map or photo. y: Vertical pixel coordinate, used together with x. long_press: Hold instead of tapping, for context menus and drag handles.

type_textA

Type text into a field on the phone.

Args: text: The text to type. Use "\n" for a line break. target: The field to type into, as a number or label from read_screen. When omitted, the text goes to whatever already has focus. replace: Clear the field first instead of appending to it. submit: Press Enter afterwards, for search bars and message boxes. verify: Re-read the screen afterwards and report whether the text actually landed, which catches a tap that missed the field.

scrollA

Scroll the content on screen.

Args: direction: Where you want to look, not which way the finger moves. "down" reveals content further down the page, "up" goes back toward the top, and "left" and "right" move sideways. distance: "small" nudges, "page" moves about one screenful (the default), "large" jumps most of the way. target: Optional number or label of the region to scroll, when the screen has more than one scrollable area.

swipeA

Drag between two points, for gestures no named control covers.

Use scroll for ordinary page movement. Reach for this to dismiss a card, pull to refresh, draw a pattern lock, or drag a slider.

Args: start_x: Horizontal pixel where the finger lands. start_y: Vertical pixel where the finger lands. end_x: Horizontal pixel where the finger lifts. end_y: Vertical pixel where the finger lifts. duration_ms: How long the drag takes. Longer is slower and more likely to be read as a drag rather than a fling.

press_keyB

Press a navigation, hardware, or system key.

Args: key: One of back, home, recents, enter, delete, escape, tab, space, move_home, move_end, volume_up, volume_down, power, wake, sleep, camera, play_pause, next_track, previous_track, copy, cut, paste, select_all, search, menu, page_up, page_down, notification_shade, quick_settings, collapse_shade.

open_appA

Bring an app to the front, by the name a person would say.

Accepts "YouTube", "youtube", "play store", or an exact package id. The name is checked against what is actually installed before anything is launched, and it waits for the app to reach the foreground before returning.

Args: name: The app to open. wait_seconds: How long to wait for it to reach the foreground.

open_urlA

Open a web address or deep link on the phone.

Args: url: A URL such as "https://example.com", or a deep link such as "geo:37.8,-122.4", "tel:+15551234567", or "sms:+15551234567".

list_appsA

List the apps installed on the phone, as a readable name and package id.

Use this when you are unsure what an app is called on this phone, or when open_app could not find the name you tried.

Args: query: Optional filter, matched against the package id.

wait_forA

Wait until something becomes true, instead of sleeping a fixed time.

Use this after an action that triggers loading, so the next step reads a settled screen instead of one mid-animation. With no arguments it waits for the screen to stop changing.

Args: text: Wait until this text appears anywhere on screen. app: Wait until this app (a name or a package id) is in the foreground. timeout_seconds: Give up and report the timeout after this long.

read_notificationsA

Open the notification shade, read what is in it, then close it again.

Args: close_after: Set False to leave the shade open, when you intend to tap one of the notifications.

clipboardA

Read or replace the phone's clipboard.

Useful for moving text between the phone and this conversation, and for putting non-ASCII text where a paste will pick it up, since adb can only type ASCII as keystrokes.

Args: action: "get" to read the clipboard, "set" to replace it. text: The text to place on the clipboard when setting.

run_shellA

Run a shell command on the phone and return its output.

The escape hatch for whatever the named tools do not cover: reading a setting, listing files, checking what is installed, granting a permission. Commands run as the shell user, which can read most of a stock phone but cannot change protected settings.

Args: command: The shell command line to run on the device.

transfer_fileA

Copy a file between this computer and the phone.

Args: direction: "to_phone" or "from_phone". computer_path: Path on this computer. "~" is expanded. phone_path: Path on the phone, for example /sdcard/Download/report.pdf.

decide_next_actionA

Ask Jev which single control on screen best advances a goal.

Turns "send the message" or "log in" into one concrete action. Jev is a System One decision model: it answers with a typed choice drawn from the controls actually on screen, so it cannot invent a control that is not there, and it returns a calibrated confidence alongside the answer.

Jev reads text only. It is sent the numbered listing of on-screen controls plus your goal, and it never sees a screenshot — so ask it what to tap, not what the screen looks like. Anything about appearance (colour, layout, what a photo shows) is yours to judge from take_screenshot.

Read the confidence before acting. It describes how far the leading option stands from the others — it is not whether Jev could answer, and not whether the action is safe:

  • 0.70 and above: carry out the returned action.

  • 0.45 to 0.70: carry it out, then confirm with read_screen.

  • below 0.45: the leading options are near a coin flip. Do not act on it; call take_screenshot and decide from the image yourself.

Args: goal: What the user wants, in plain language. For example "reply to the most recent message saying I will be late". max_candidates: The most controls to offer in one choice. The default leaves room for the six built-in actions inside Jev's 255-option limit, so every control on screen is normally offered: measured over paired problems, accuracy held from 2 options to 255, and what costs accuracy is options that resemble each other, not how many there are. Lower this only to reduce token cost.

run_taskA

Give the phone a goal and let it work the whole thing out itself.

Use this instead of driving the phone a step at a time. It reads the screen, decides what to do, does it, and reads again, until the goal is reported done or unreachable — so "open the email app" is one call here rather than ten round trips through you.

Every step reports the action, what came of it, and the two numbers Jev gave for the choice: probability is how likely that option was the best one and confidence is how sure it was of its own ranking. They are reported rather than gated on, so a run that succeeded on a thin lead says so, and a run that stopped for a reason other than reaching the goal says that instead.

achieved is what the run claimed. Read the steps for what actually happened. If a run needs looking at afterwards, set PHONE_CONTROL_RUNS to a directory and each run leaves a folder there with the full record of what it asked Jev and what came back; the reply then names the folder.

Prefer the step-by-step tools when you already know exactly what to tap, when the screen is visual rather than textual, or when you need to inspect something part-way through.

Args: goal: What to achieve, in plain language. For example "open the email app" or "turn on airplane mode". max_steps: How many actions to allow before giving up. Reaching it ends the run honestly as unfinished rather than as a failure of the goal.

ask_jevA

Ask Jev a typed question about what is on the phone screen.

Use this for judgements that are not "what do I tap next": whether an action worked, whether the screen is an error state, or which of several visible items the user meant.

Jev reads text only — the control listing, plus whatever state you give it. It cannot see the screenshot, so keep the question about what the controls say and do, not about how they look.

Args: instructions: The question, in plain language. question_type: "choice" to pick one option, "yes_or_no" for the probability that something holds, or "scale" to place the screen on an ordered scale. options: For "choice", the alternatives. For "scale", the levels from lowest to highest. Write each as "short_key: what it means" to name the answer, or give just the meaning to key it by position. state: Extra context to judge against. The current screen reading is always included; use this for the goal or instruction being checked.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.2/5.0

Scored across 19 tools

Disambiguation5/5

Each tool targets a distinct capability: screen reading (read_screen, take_screenshot), interaction (tap, type_text, scroll, swipe, press_key), app handling (list_apps, open_app), navigation (open_url), waiting (wait_for), notifications, clipboard, shell, file transfer, and three Jev-driven decision tools that are clearly separated by purpose (next action, autonomous task, general question). No two tools could be easily confused.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun/verb_phrase pattern (read_screen, take_screenshot, type_text, open_app, press_key, wait_for, run_shell, etc.). The style is uniform and predictable, making it easy for an agent to infer function from the name.

Tool Count4/5

With 19 tools, the server is slightly over the typical well-scoped range (3-15), but the scope of a phone-control server is broad, covering screen interaction, system controls, app management, and advanced decision-making. Each tool earns its place, though a few (e.g., separate Jev tools) could arguably be merged; still, the count is reasonable and not bloated.

Completeness5/5

The tool surface provides comprehensive coverage for phone control: reading the screen (text and visual), acting on screen content, navigating the system, managing apps and URLs, waiting for state changes, handling notifications, clipboard, shell escape hatch, file transfer, and high-level autonomous decision-making. No significant gaps are apparent for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues