android-phone-control
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| JEV_API_KEY | No | The Jev key, enabling the two decision tools. An OpenRouter sk-or-... key or a TypeSafe apikey_... one — its own shape routes it. | |
| DSH_PHONE_SERIAL | No | Which device to drive when several are attached. | |
| PHONE_CONTROL_ADB | No | Path to adb when it is not on PATH. | |
| OPENROUTER_API_KEY | No | The same key under its older name, still read. | |
| PHONE_CONTROL_RUNS | No | Directory where every run writes its report and step details; the tool's reply names it. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| statusA | Check the phone connection and see what it is currently doing. Call this first when a command fails, and before starting anything multi-step. It reports the model, Android version, screen size, foreground app, whether the screen is on or locked, and the battery level. It changes nothing on the phone. |
| read_screenA | List the controls currently on the phone's screen, by number and name. This is the main way to see what is on screen, and it is far cheaper than a
screenshot. Each control gets a number you can hand to Every action tool re-reads the screen before it acts, so both forms are resolved against the screen as it is at that moment. A number is precise -- it disambiguates two controls carrying the same text. A label is stable -- it survives the screen being redrawn. Use the number for a control you have just seen and are about to act on, and the label when the screen may have changed since you read it. Args: query: Optional filter. Only controls whose text, description, or type contains this text are listed, best match first. Use it to ask "is the Send button up yet?" in one call. |
| take_screenshotA | Capture the phone's screen and look at it. Try Args: max_width: Downscale the image to this width in pixels before returning it. Lower it to spend fewer tokens; 0 keeps full resolution. |
| tapA | Tap, or long-press, something on the phone screen. Args:
target: The number from |
| type_textA | Type text into a field on the phone. Args:
text: The text to type. Use "\n" for a line break.
target: The field to type into, as a number or label from |
| scrollA | Scroll the content on screen. Args: direction: Where you want to look, not which way the finger moves. "down" reveals content further down the page, "up" goes back toward the top, and "left" and "right" move sideways. distance: "small" nudges, "page" moves about one screenful (the default), "large" jumps most of the way. target: Optional number or label of the region to scroll, when the screen has more than one scrollable area. |
| swipeA | Drag between two points, for gestures no named control covers. Use Args: start_x: Horizontal pixel where the finger lands. start_y: Vertical pixel where the finger lands. end_x: Horizontal pixel where the finger lifts. end_y: Vertical pixel where the finger lifts. duration_ms: How long the drag takes. Longer is slower and more likely to be read as a drag rather than a fling. |
| press_keyB | Press a navigation, hardware, or system key. Args: key: One of back, home, recents, enter, delete, escape, tab, space, move_home, move_end, volume_up, volume_down, power, wake, sleep, camera, play_pause, next_track, previous_track, copy, cut, paste, select_all, search, menu, page_up, page_down, notification_shade, quick_settings, collapse_shade. |
| open_appA | Bring an app to the front, by the name a person would say. Accepts "YouTube", "youtube", "play store", or an exact package id. The name is checked against what is actually installed before anything is launched, and it waits for the app to reach the foreground before returning. Args: name: The app to open. wait_seconds: How long to wait for it to reach the foreground. |
| open_urlA | Open a web address or deep link on the phone. Args: url: A URL such as "https://example.com", or a deep link such as "geo:37.8,-122.4", "tel:+15551234567", or "sms:+15551234567". |
| list_appsA | List the apps installed on the phone, as a readable name and package id. Use this when you are unsure what an app is called on this phone, or when
Args: query: Optional filter, matched against the package id. |
| wait_forA | Wait until something becomes true, instead of sleeping a fixed time. Use this after an action that triggers loading, so the next step reads a settled screen instead of one mid-animation. With no arguments it waits for the screen to stop changing. Args: text: Wait until this text appears anywhere on screen. app: Wait until this app (a name or a package id) is in the foreground. timeout_seconds: Give up and report the timeout after this long. |
| read_notificationsA | Open the notification shade, read what is in it, then close it again. Args: close_after: Set False to leave the shade open, when you intend to tap one of the notifications. |
| clipboardA | Read or replace the phone's clipboard. Useful for moving text between the phone and this conversation, and for putting non-ASCII text where a paste will pick it up, since adb can only type ASCII as keystrokes. Args: action: "get" to read the clipboard, "set" to replace it. text: The text to place on the clipboard when setting. |
| run_shellA | Run a shell command on the phone and return its output. The escape hatch for whatever the named tools do not cover: reading a setting, listing files, checking what is installed, granting a permission. Commands run as the shell user, which can read most of a stock phone but cannot change protected settings. Args: command: The shell command line to run on the device. |
| transfer_fileA | Copy a file between this computer and the phone. Args: direction: "to_phone" or "from_phone". computer_path: Path on this computer. "~" is expanded. phone_path: Path on the phone, for example /sdcard/Download/report.pdf. |
| decide_next_actionA | Ask Jev which single control on screen best advances a goal. Turns "send the message" or "log in" into one concrete action. Jev is a System One decision model: it answers with a typed choice drawn from the controls actually on screen, so it cannot invent a control that is not there, and it returns a calibrated confidence alongside the answer. Jev reads text only. It is sent the numbered listing of on-screen
controls plus your goal, and it never sees a screenshot — so ask it what to
tap, not what the screen looks like. Anything about appearance (colour,
layout, what a photo shows) is yours to judge from Read the confidence before acting. It describes how far the leading option stands from the others — it is not whether Jev could answer, and not whether the action is safe:
Args: goal: What the user wants, in plain language. For example "reply to the most recent message saying I will be late". max_candidates: The most controls to offer in one choice. The default leaves room for the six built-in actions inside Jev's 255-option limit, so every control on screen is normally offered: measured over paired problems, accuracy held from 2 options to 255, and what costs accuracy is options that resemble each other, not how many there are. Lower this only to reduce token cost. |
| run_taskA | Give the phone a goal and let it work the whole thing out itself. Use this instead of driving the phone a step at a time. It reads the screen, decides what to do, does it, and reads again, until the goal is reported done or unreachable — so "open the email app" is one call here rather than ten round trips through you. Every step reports the action, what came of it, and the two numbers Jev gave
for the choice:
Prefer the step-by-step tools when you already know exactly what to tap, when the screen is visual rather than textual, or when you need to inspect something part-way through. Args: goal: What to achieve, in plain language. For example "open the email app" or "turn on airplane mode". max_steps: How many actions to allow before giving up. Reaching it ends the run honestly as unfinished rather than as a failure of the goal. |
| ask_jevA | Ask Jev a typed question about what is on the phone screen. Use this for judgements that are not "what do I tap next": whether an action worked, whether the screen is an error state, or which of several visible items the user meant. Jev reads text only — the control listing, plus whatever Args: instructions: The question, in plain language. question_type: "choice" to pick one option, "yes_or_no" for the probability that something holds, or "scale" to place the screen on an ordered scale. options: For "choice", the alternatives. For "scale", the levels from lowest to highest. Write each as "short_key: what it means" to name the answer, or give just the meaning to key it by position. state: Extra context to judge against. The current screen reading is always included; use this for the goal or instruction being checked. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 19 tools
Each tool targets a distinct capability: screen reading (read_screen, take_screenshot), interaction (tap, type_text, scroll, swipe, press_key), app handling (list_apps, open_app), navigation (open_url), waiting (wait_for), notifications, clipboard, shell, file transfer, and three Jev-driven decision tools that are clearly separated by purpose (next action, autonomous task, general question). No two tools could be easily confused.
All tool names follow a consistent snake_case verb_noun/verb_phrase pattern (read_screen, take_screenshot, type_text, open_app, press_key, wait_for, run_shell, etc.). The style is uniform and predictable, making it easy for an agent to infer function from the name.
With 19 tools, the server is slightly over the typical well-scoped range (3-15), but the scope of a phone-control server is broad, covering screen interaction, system controls, app management, and advanced decision-making. Each tool earns its place, though a few (e.g., separate Jev tools) could arguably be merged; still, the count is reasonable and not bloated.
The tool surface provides comprehensive coverage for phone control: reading the screen (text and visual), acting on screen content, navigating the system, managing apps and URLs, waiting for state changes, handling notifications, clipboard, shell escape hatch, file transfer, and high-level autonomous decision-making. No significant gaps are apparent for the stated purpose.