android-phone-control
Control an Android phone over ADB via MCP tools: inspect the screen, act on UI, manage apps/system, and automate goals with Jev.
Inspect phone/state:
statusreports model, Android version, foreground app, screen lock, battery;read_screenlists on-screen controls by number/label;take_screenshotcaptures visuals.Interact with UI:
tap,type_text,scroll,swipe,press_key.Launch and wait:
open_app,open_url,list_apps,wait_for.System/data access:
read_notifications,clipboard,run_shell,transfer_file.Jev decision helpers:
decide_next_actionfor next control,ask_jevfor typed screen judgments.Full automation:
run_tasktakes a plain-language goal, reads/decides/acts/checks until done or stopped.
Controls an Android phone over adb by reading the device screen and performing actions to complete tasks like opening apps, toggling settings, and taking photos.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@android-phone-controlOpen the Play Store and install Spotify"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
android-jev
Control your Android phone through Jev.
You say what you want done — open the Play Store, turn on aeroplane mode, take a
photo — and this does it. It plugs into any MCP client as a tool called run_task;
inside, it reads the phone's own screen, asks
Jev what to do next,
does it, checks the result, and tells you honestly whether it worked.
Everything runs over adb. Nothing is installed on the phone.
run_task("open the Play Store") → achieved: true done, 3 steps
run_task("search for YouTube") → achieved: true done, 5 steps
run_task("open YouTube") → achieved: true done, 3 stepsQuick start
What you need
A phone attached by USB, with a data-capable cable. It stays plugged in while it works.
USB debugging on — one setting, below.
macOS with Python 3.11+, and
uv. The picture route uses macOS Vision, which installs on macOS as a dependency; on Linux the loop works and reads the accessibility tree, and leaves the fallback out.Verified on a Pixel 8a running Android 17. The project has only been measured on that one, so treat other phones as untested rather than unsupported.
1. Connect the phone
Plug it in, and on the phone:
Settings → About phone → tap "Build number" seven times. Developer options appear.
Settings → System → Developer options → USB debugging → on.
Replug the cable. The phone shows Allow USB debugging? — tick Always allow and tap Allow. Nothing works until that tap happens, and it is deliberate: a computer cannot authorise itself.
Check the connection:
adb devices # the phone should say "device", not "unauthorized"2. Install the server
git clone https://github.com/FZ2000/android-jev ~/src/android-jev
cd ~/src/android-jev
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e .The distribution is android-jev, the import is phone_control and the command is
phone-control: three names, each read somewhere different. The command also answers
to android-jev, so that uvx android-jev — the spelling a registry client derives
from the distribution name — starts the server.
No adb? PHONE_CONTROL_ADB points at one, and the usual SDK locations are
searched:
curl -sSL -o /tmp/pt.zip \
https://dl.google.com/android/repository/platform-tools-latest-darwin.zip
unzip -q /tmp/pt.zip -d ~/.local/opt3. Give it a decision model
run_task works without one, using a keyword fallback, and is markedly better with
Jev: goals that do not name a control — turn off Bluetooth, open the settings
app — need it.
Either of two keys works, and the key's own shape says which gateway serves it: an
OpenRouter key beginning sk-or-, or a TypeSafe key beginning apikey_. Nothing in
the configuration says which you are using, because sending one gateway's model to the
other is a 400.
The server reads it from JEV_API_KEY — or OPENROUTER_API_KEY, if that is the name
already in use here — or from a file named jev-api-key or openrouter-api-key in the
secrets directory of whichever harness starts it. In DeepSeek Harness that is
$DSH_HOME/secrets/:
printf %s "$JEV_API_KEY" > "$DSH_HOME/secrets/jev-api-key"
chmod 600 "$DSH_HOME/secrets/jev-api-key"Whichever way it reaches the server, the key never appears in a state, a log, a run folder or a transcript.
4. Mount it in your agent
In Claude Code, this repository is also a plugin, which mounts the server and installs the skill in one step — steps 4 and 5 both:
claude plugin marketplace add FZ2000/android-jev
claude plugin install android-jev@android-jevIt asks for the Jev key once and keeps it in the system's credential store; leave it
empty to use JEV_API_KEY or OPENROUTER_API_KEY from your environment instead. The
server runs from the plugin's own copy of this repository through uv, so what it
does always matches the skill that describes it. It still needs uv and adb, and
the phone set up as in step 1.
For any other client, or to run from your own clone:
Any MCP client can mount this server; it speaks MCP over stdio. Send the agent this, with the path changed to wherever you cloned the repository — it says what the repository is and where its own documentation lives, because an agent that has to guess either of those will guess wrong. A prompt rather than a snippet per client: each client keeps its MCP configuration somewhere different, and changes the format between versions.
The repository at
~/src/android-jevis an MCP server that drives an Android phone attached by USB, together with the instructions for using it. Mount it for me, and tell me when you have.
Start it as a local stdio server: the absolute path
~/src/android-jev/.venv/bin/python, the single argument-m phone_control, run from that directory.Call it
android.It needs the Jev key in its environment, under
JEV_API_KEY— orOPENROUTER_API_KEYif that is the name already in use here. Either an OpenRoutersk-or-…key or a TypeSafeapikey_…one: the shape of the key decides which gateway serves it. If your MCP bridge scrubsKEY,PASSWORD,SECRETorTOKENout of a child process's environment, name the variable in the server's ownenvrather than expecting it to be inherited.Give it a generous per-call timeout, a couple of minutes: one of its tools waits for the phone by design.
Read
README.mdat the repository root first. Its quick start covers connecting the phone, enabling USB debugging, installing the server and supplying the key;docs/has the longer explanations. If the phone is not attached or the server will not start, read that rather than guessing.The instructions for using it are
skills/android-phone-control/SKILL.md, with every tool and argument inreferences/tools-reference.mdbeside it.When it is mounted, list its tools and tell me whether I need to start a new session before you can see them.
Three things catch people out, and the prompt is written the way it is to survive all three:
The interpreter path has to be absolute. A client starts the server itself and does not inherit this project's virtualenv, so a bare
python -m phone_controlresolves to the wrong interpreter or to none at all.The key cannot always be inherited. DeepSeek Harness, for one, strips every variable whose name matches
/KEY|PASSWORD|SECRET|TOKEN/ifrom a stdio child, so the key has to be named in the server's ownenvor it arrives empty.A mounted server is not a visible one. Most clients need a new session, or a restart, before a model sees the tools at all — which is why the prompt asks.
The tools surface under whatever namespace your client uses — mcp__android__run_task
in DeepSeek Harness, something else elsewhere. Nothing in these instructions depends on
the exact spelling, and neither should you.
5. Teach your agent how to use it
The skill is one body of text, skills/android-phone-control/SKILL.md, and this
repository already ships it rendered for every harness that reads instructions from the
working directory — which file each one reads is in
The skill, for any agent below. If you work in this
repository, your agent has it already and there is nothing to install.
For any other harness, send it this:
Install the phone-control skill in
~/src/android-jevfor yourself, so that you can drive my phone in any session. The instructions areskills/android-phone-control/SKILL.md, with one reference file beside it atreferences/tools-reference.md. Put both wherever you read skills from, and tell me whether you need to be restarted before it takes effect.
If your harness has nowhere to keep a skill — a hosted assistant with an instructions box and nothing else — paste this instead, which is the part that matters:
The phone is driven by an MCP server mounted as
android. To do anything on it, callrun_taskwith the outcome you want in plain language — "open the Play Store", "turn on aeroplane mode". Ask for one thing at a time and wait for each result. Readachievedandreasonfrom the reply, and relay them rather than retrying.
6. Sanity test: take a photograph
With the phone unlocked and the screen on, ask your agent:
run_task("open the camera app and take a photo")Then check the phone's own evidence rather than the report:
adb shell ls -t /sdcard/DCIM/Camera | head -3 # a new file at the topA new file and a report that it could not confirm the goal are both correct here, and this is the most useful thing the quick start can show you. A photograph leaves no trace on the screen — measured, the tree is byte-identical before and after the shutter — so the run can see that the camera opened and cannot see the picture. It says so instead of claiming a success it did not verify.
Run against the phone this was written on, through a fresh clone of this repository, that is what it looks like: five steps, three of them pressing the shutter, and then
"achieved": false, "outcome": "stalled",
"reason": "3 actions in a row left the screen exactly as it was, so the run was
going round rather than on"and three new files in DCIM/Camera. It took the photograph and could not know it
had, so it stopped rather than pressing the shutter forever. That is the honest
answer to a goal whose result is off-screen, and it is why the last line of this
test is one whose result is on the screen:
run_task("open the calculator app") # achieved: true, done, 3 steps, ~13sRelated MCP server: Android DevTools MCP
Where to go next
docs/engineering-decisions.md— why the code is the way it is, with the measurement that forced each decision.docs/code-map.md— every module, what it is for, and how they depend on each other.docs/android-apis.md— what this phone will and will not tell you, measured, including the reliable ways to prove an end state.docs/debugging.md— one run failing, and what to read.docs/tool-contract.md— the tool surface as a contract, including what is deliberately not built.docs/what-we-got-wrong.md— what was tried and did not work.CONTRIBUTING.md— the rules the tests live by.
Why it is built this way
Most adb wrappers expose tap(x, y), which is useless to an agent that cannot
see. This one exposes what is actually on screen:
#5 ImageButton labeled "Send message" at (970,2170) [tappable]
#4 EditText labeled "Type a message" at (450,2170) [text field, focused]Four decisions follow from that, and each cost something to learn.
The reading is structured, not a screenshot. read_screen reports every
window with its type, whether the keyboard is up, and a typed description of every
control — its category, label, bounds and a point to tap. An image is the fallback
for genuinely visual screens, not the default. See
docs/android-apis.md for what is extractable and what is
not.
It reports what has focus, not what is merely running. When a system window
such as the lock screen or the notification shade is in front, status says so
and reports no foreground app — rather than naming the app behind it, which is what
mFocusedApp would have said.
A control is addressed by name, then by number, and only then by coordinate. Every action tool re-reads the screen before it acts, so a name or a number is resolved against the screen as it is now; a coordinate is not. Read the reply — it names exactly what was acted on.
The agent states the goal; the server owns the loop. run_task("open the email app") is one call. The server reads its own screen, decides each step,
checks its own work, and repeats until the goal is reached or cannot be — then
reports what it did and why it stopped. An agent needs to know nothing about
accessibility trees, windows or coordinates to use it.
Each step is several typed questions, asked together. Inside that loop, Jev, TypeSafe's System One decision model, answers four questions in one request and only the answer belonging to the chosen operation is used:
question | what it settles |
| which kind of action, from what this screen can actually do |
| which control, if the operation taps one |
| which installed app, if the operation opens one |
| whether a fresh reading shows the goal met |
Splitting them is the part that mattered. One flat list holding operations, app
names and screen controls made the decider unsure on a launcher screen — it
answered give_up at 0.63 — and the reason is mechanical: confidence measures
how concentrated a choice is, so overlapping options always read as doubt,
whichever of them is right. Once the operations had a question of their own,
open_app won the same screen at 0.74.
Only what can be carried out is offered. Typing needs a focused field, scrolling needs something that scrolls, a website needs a value the device can resolve, a tap needs a control to tap. An option the loop cannot execute is a guaranteed dead step and worse than that — it reads as the model being unsure when the model was never at fault. A web address with no scheme was offered once; Android could not resolve it, the action did nothing, and the 0.41 confidence on that step was read as doubt for days.
The state is facts computed from the phone, not a dump to interpret. What the focused field holds, where each control sits in coarse terms, what tells two identical labels apart, what has already been tried on this screen. This is what makes a threshold unnecessary rather than merely wrong: asked whether a goal was complete, Jev scored 0.42 to 0.47 on a screen where it genuinely was and 0.51 on one where it was not, because in neither case did the state say whether the query had been sent or was still in the box. No number separates those two; a state that says what is on the screen does.
A low number is not a reason to stop, and the loop does not decide. Every rule that used to refuse an answer for scoring low is gone, because each one was measured ending a run that was working. What is left is a stop rule about the phone — three actions in a row that left the screen as it was, or two already taken on the same screen, means the run is going round rather than on — and both err toward running on, because a futile run costs a minute and a working run cut short costs the task.
One exception, and it is about safety rather than sureness. An action that would have a material effect — sending, deleting, paying, granting a permission — is not carried out unless the goal explicitly asks for it. That gate exists because a real run tapped Allow Chrome to record audio while pursuing an unrelated goal.
When the loop stops short, a reader takes over. With no reader configured a stop ends the run as it always did. With one, it reads the screen and says whether the goal is met, and may set one next move for the decider to aim at — never choosing the action itself. The reference implementation measured 0.39 becoming 0.92 on the same capture once a focus was set, which is the argument in one number: the fix belongs in what the decider is told, not in a threshold.
The phone is read once per step. Reading it was 84% of a step, half of that a second read whose only job was to see what the action had done. The post-action observation is now the next step's reading:
before step 7.29s reading 2.91s watching_the_result 3.18s deciding 0.26s
after step 3.75s reading 2.18s watching_the_result 0.00s deciding 0.24sWhat is sent is the source, not a summary of it. Jev's state carries the
accessibility tree's own field names and values, and each option is a word
(open_app) rather than our numbering. The
measured prompting evidence
is blunt: the same programs scored 0.894 as source, 0.598 as a syntax tree
and 0.530 as a control-flow graph, even though the graph carried more
information for 2.5× the tokens.
Running it directly
.venv/bin/python -m phone_controlIt speaks MCP over stdio and waits. Nothing to configure; the tools appear as
mcp__android__* once a client mounts it, as described in the quick start.
The tools
Seeing |
|
Acting |
|
System |
|
Doing the whole job |
|
Stepping by hand |
|
Full argument reference:
skills/android-phone-control/references/tools-reference.md.
The skill, for any agent
skills/android-phone-control/ teaches the
loop: start with the goal rather than the taps, when to read the tree rather than
take a screenshot, and how to read Jev's confidence. The body names no agent, so
the same instructions work anywhere.
Each tool reads a different file, so the skill is rendered into all of them from
one canonical source — skills/android-phone-control/SKILL.md:
File | Read by |
| Codex, Cursor, Copilot agent mode, and the wider AGENTS.md convention |
| Gemini CLI |
| GitHub Copilot |
| Cursor |
| Claude Code |
Edit the canonical skill and re-render; a test fails if any copy has drifted:
.venv/bin/python scripts/render_skill_for_agents.pyFor DeepSeek Harness, symlink or copy skills/android-phone-control/ into a
scanned skill root — $DSH_HOME/skills/ for every session, or .dsh/skills/ in a
project. Discovery is live; no restart.
Environment
Variable | Effect |
| Which device to drive when several are attached. |
| The Jev key, enabling the two decision tools. An OpenRouter |
| The same key under its older name, still read. |
| Path to |
Known limits
Whether a screen can be read at all varies moment to moment. This is the limit to know about, because it produces failures that look like the model's fault and are not. Measured on one page, read twice in a row: 2.9 seconds and 29 controls, then 73.8 seconds and a failure. The page cannot be dumped while it is doing something, and a page carrying a live figure is doing something most of the time and not all of it. So the same goal can succeed on one attempt and fail on the next, with nothing changed but the timing. Everything else below is a consequence or a neighbour of this one.
A screen hosting embedded Android Views inside Compose can read as almost empty. The interop wrapper is marked
NO_HIDE_DESCENDANTS, which hides its whole subtree from every accessibility client — no flag or setting overrides it, and Appium's workaround exists only because it ships its own dumper inside an installed APK. Measured, this is the common case on this device's Settings pages rather than the exception.When the tree will not describe a page, a photograph of it is read instead. That is the second route, and it reaches the tools as well as the loop: a page that cannot be dumped still comes back with its words. What it cannot do is give controls — the reading is lines of text with no tappable targets on them — so on such a page a run can conclude the goal and leave, and cannot act further. A page whose dump fails and whose photograph says nothing either is reported as a failure, not as an empty screen.
dumpsys activity topis not a foreground oracle. It reads as though it should be, and a probe built on it will silently read another app's state: asked while the Settings Battery page was in front, it reported fragments belonging to Photos.dumpsys windownames the page that is actually in front, and for a Settings page reached by tapping that is the genericSubSettings— the page's own title, read from the tree, is the reliable identity.A sleeping screen reads as empty. The display going off leaves one bare node.
read_screensays which of asleep, locked or canvas-only it is.Non-ASCII typing.
input textgoes through the virtual keyboard's character map, which returns null for anything it cannot generate — AOSP's own comment says "For robust text entry, do not use this function." Non-ASCII is routed through the clipboard and a paste instead. The proper fix isACTION_SET_TEXT.dumpsysoutput is internal and drifts between releases. Parsing is tested against dump shapes captured from real reports and exercised on a real device;docs/android-apis.mdseparates what was confirmed from what the device corrected.
Documentation
What each route to an Android phone offers and costs, the parsing traps, and what a real device confirmed | |
Where the tool surface is going, with the measurements — failures now raise, so | |
A diagnosis against a reference implementation of the same loop: the state and the questions are structurally wrong, and tuning will not fix them | |
Start here when a run goes wrong — what to read, in what order, and how to re-decide a recorded step without a phone | |
Every belief this code holds about Android, proven against the phone. A fake can only repeat what its author believed; these cannot | |
What protects |
Testing
scripts/check.sh --as-ci # what CI runs: no phone, no key
scripts/check.sh # everything, with the 95% coverage gate
.venv/bin/python -m pytest -q -m "not device" # everything that needs no phone
.venv/bin/python -m pytest -m device -q # everything that needs a real phone--as-ci is the one to run before pushing. It runs the fast suite on a machine that
has never been set up — no adb, no decision-model key, no home directory holding
either — because that is what the runner is, and the difference is not obvious from
a machine that has both. The suite passed here for weeks while main was red for
exactly that reason.
The default suite runs without a device; device tests skip themselves when none is attached, and when the phone is asleep or locked — that is a precondition, not a code property. The same is true of three other things the suite needs and cannot provide: a Jev key, a network that holds for the length of a run, and macOS's Vision framework, which is the only way to read a picture of a screen. A test that needs any of them stands itself aside with the reason rather than failing, because a dropped connection is not a fact about this code.
That last one is worth knowing when reading a build: on the Linux runner, every test
that photographs a screen stands aside, so the picture route is verified on a Mac
rather than in CI. scripts/check.sh --as-ci reproduces the missing adb and the
missing key, because it runs on a Mac it cannot reproduce the missing framework.
The counts are deliberately left out of this list. They change with every commit that adds a test, and a number in a README is stale the moment after it is written — these commands print the real ones. What does not change is the shape: the majority of the suite needs no phone, and the device suite is the part that cannot be faked.
Coverage has to be read from both suites together, against a gate of 95%: most of
this package only executes with a phone attached, so the fast suite alone fails a gate
on code it had no way to reach. scripts/check.sh runs both. CI runs the fast suite
only, because a hosted runner has no phone.
The usage scenarios
tests/scenarios/ is the specification: all 47 scenarios, each a thing a person asks
a phone to do, from one launch to goals no phone can reach. Each carries its own
starting state, its own ground truth read from the phone, and a screenshot at every
stage.
scripts/run_scenarios.py --tier 1 # a tier at a time
scripts/run_scenarios.py --id bluetooth-off # or one, by name
scripts/run_scenarios.py --all # about half an hourArtefacts land in artifacts/scenarios/<id>/. See
docs/debugging.md for how to read them.
The rules that keep the tests honest
scripts/install_git_hooks.sh # once: refuse a commit whose suite is red
.venv/bin/python -m pytest tests/test_no_xfail.py -q # no test may be allowed to failFour times in this project a fake agreed with its author and the belief was wrong:
that a bare domain opens a browser, that am start -d resolves it, that
cmd clipboard set places something, that a settings toggle is checkable. Each cost
real work. Hence two rules rather than good intentions — a test allowed to fail
reports green while the behaviour it describes does not work, and a belief about
Android belongs in
tests/test_what_the_phone_actually_does.py,
where the phone gets asked.
Watching a run afterwards
An MCP reply is one shot and a run leaves nothing behind unless you ask it to:
export PHONE_CONTROL_RUNS=~/.phone-control/runsEvery run then writes a folder — the report, the exact request and answer for every step, and where the seconds went — and the tool's reply names it. A recorded step can be re-decided offline, without the phone:
.venv/bin/python scripts/replay_run.py ~/.phone-control/runs/<folder> --step 0Status
Working and device-validated: opening apps, reading and acting on screens, scrolling, typing, window and keyboard state, screenshots.
Measured, on a Pixel 8a running Android 17
tests/test_jev_does_the_job.py asks for things a person would ask for and then
checks the phone rather than the report. Every goal below was independently
confirmed against the focused window, which the decision never sees:
Goal | Reported | Verified | Steps | Time |
"open the settings app" | achieved |
| 2 | 9.6 s |
"open the clock app" | achieved |
| 2 | 8.5 s |
"open the calculator app" | achieved |
| 2 | 8.3 s |
No false successes: nothing claimed a result the phone did not reach. The completion check agreed with the chosen action every time, at confidences of 0.97, 1.00 and 0.99.
One honest failure. Asked to "order a large pepperoni pizza", which no phone can do, it correctly reported the goal unmet — but it spent all ten steps exploring Chrome's interface first, including tapping "Allow Chrome to record audio". It explores before it concedes, and some of that exploration is consequential. Bounding that is open work.
Not yet done, and named rather than implied: the tool contract migration described
in docs/tool-contract.md, and any device run of the Jev
decision path, which needs a Jev key.
Available Tools
19 toolsask_jevA
Ask Jev a typed question about what is on the phone screen.
Use this for judgements that are not "what do I tap next": whether an action worked, whether the screen is an error state, or which of several visible items the user meant.
Jev reads text only — the control listing, plus whatever state you give
it. It cannot see the screenshot, so keep the question about what the
controls say and do, not about how they look.
Args: instructions: The question, in plain language. question_type: "choice" to pick one option, "yes_or_no" for the probability that something holds, or "scale" to place the screen on an ordered scale. options: For "choice", the alternatives. For "scale", the levels from lowest to highest. Write each as "short_key: what it means" to name the answer, or give just the meaning to key it by position. state: Extra context to judge against. The current screen reading is always included; use this for the goal or instruction being checked.
| Name | Required | Description | Default |
|---|---|---|---|
| state | No | Extra context to judge against. The current screen is always included. | |
| options | No | For 'choice' and 'scale', the alternatives. For 'choice' write each as 'short_key: what it means'. | |
| instructions | Yes | The question, in plain language. | |
| question_type | No | Which shape of answer: 'choice' picks one option, 'yes_or_no' gives a probability, 'scale' places the state on a range. | choice |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does so well: it discloses that Jev reads text only, cannot see the screenshot, and always receives the current screen reading plus any provided state. It does not explicitly state that the call is side-effect-free, but 'Ask... a question' strongly implies a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but every section earns its place: purpose, exclusions, a critical limitation, and parameter guidance. The Args block is organized and front-loaded after the core usage statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no annotations and an output schema present, the description is complete: it explains when to use it, what it can and cannot see, how to format each parameter, and what context to provide. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value by explaining the options format ('short_key: what it means' or positional keys), the meaning of question_type variants, and how state should be used as the goal or instruction being checked.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Ask Jev a typed question about what is on the phone screen.' It then distinguishes itself from the 'what do I tap next' class of action, which separates it from sibling decide_next_action without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit use cases: judging whether an action worked, recognizing error states, and disambiguating visible items. It also states what it is not for ('not what do I tap next'), though it does not name the sibling tool that should handle that case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clipboardA
Read or replace the phone's clipboard.
Useful for moving text between the phone and this conversation, and for putting non-ASCII text where a paste will pick it up, since adb can only type ASCII as keystrokes.
Args: action: "get" to read the clipboard, "set" to replace it. text: The text to place on the clipboard when setting.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | The text to type, or the text to replace the clipboard with. | |
| action | No | What to do: 'get' to read the clipboard, 'set' to replace it. | get |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses that 'set' replaces the clipboard and explains the adb ASCII limitation motivating non-ASCII use. It does not describe return behavior or side effects, but the presence of an output schema reduces the need to explain return values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The most important information is front-loaded in the first sentence. The use-case explanation earns its place by clarifying when the tool is valuable, and the Args section is compact and directly tied to the parameters. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with full schema coverage and an output schema, the description is nearly complete. It adds practical context about moving text and handling non-ASCII characters. The only minor gap is not describing what the 'get' action returns, but the output schema covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already explains action as 'get'/'set' and text as the clipboard content. The description's Args section mirrors the schema without adding significant new semantic detail, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Read or replace the phone's clipboard', which is a specific verb+resource statement. It clearly distinguishes this tool from siblings like type_text and read_screen by naming the clipboard as the target.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use the tool: moving text between the phone and the conversation, and handling non-ASCII text that adb cannot type. It does not explicitly name sibling alternatives or exclusions, but the use cases are concrete enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decide_next_actionA
Ask Jev which single control on screen best advances a goal.
Turns "send the message" or "log in" into one concrete action. Jev is a System One decision model: it answers with a typed choice drawn from the controls actually on screen, so it cannot invent a control that is not there, and it returns a calibrated confidence alongside the answer.
Jev reads text only. It is sent the numbered listing of on-screen
controls plus your goal, and it never sees a screenshot — so ask it what to
tap, not what the screen looks like. Anything about appearance (colour,
layout, what a photo shows) is yours to judge from take_screenshot.
Read the confidence before acting. It describes how far the leading option stands from the others — it is not whether Jev could answer, and not whether the action is safe:
0.70 and above: carry out the returned action.
0.45 to 0.70: carry it out, then confirm with
read_screen.below 0.45: the leading options are near a coin flip. Do not act on it; call
take_screenshotand decide from the image yourself.
Args: goal: What the user wants, in plain language. For example "reply to the most recent message saying I will be late". max_candidates: The most controls to offer in one choice. The default leaves room for the six built-in actions inside Jev's 255-option limit, so every control on screen is normally offered: measured over paired problems, accuracy held from 2 options to 255, and what costs accuracy is options that resemble each other, not how many there are. Lower this only to reduce token cost.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | What you want done, in plain language. Phrase it as the outcome, not the steps. | |
| max_candidates | No | The most controls to offer at once. Lower only to reduce token cost. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and succeeds: it discloses that Jev is text-only, never sees screenshots, cannot invent controls, and returns a calibrated confidence. It also carefully explains what confidence does and does not mean, adding a safety-relevant caveat that is easy for an agent to act on correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and organized into scannable confidence bandscars. It is somewhat long, and the paired-problems accuracy note could be trimmed, but every section serves some operational need, so the length is mostly justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter decision tool, the description covers invocation, output shape, confidence thresholds, uncertainty handling, and sibling-tool relationships. Since an output schema exists, the description does not need to detail return fields further; nothing an agent needs to call the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers both parameters, but the description goes well beyond it: it gives an example goal phrase, explains the six built-in actions within the 255-option limit, and clarifies that accuracy depends on option resemblance rather than count. This materially improves an agent's ability to set `max_candidates` appropriately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete verb–resource relation: 'Ask Jev which single control on screen best advances a goal,' and frames the output as a typed choice drawn from actual controls. It also distinguishes itself from `take_screenshot` and `read_screen` by clarifying that appearance is judged from screenshots, not this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use and alternative guidance: ask what to tap, not what the screen looks like, and leave appearance judgments to `take_screenshot`. It even provides confidence-based decision rules that route the agent to `read_screen` or `take_screenshot` when appropriate, leaving little room for misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_appsA
List the apps installed on the phone, as a readable name and package id.
Use this when you are unsure what an app is called on this phone, or when
open_app could not find the name you tried.
Args: query: Optional filter, matched against the package id.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Only list what matches this text. Omit to list everything. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses the output includes name and package ID and that query filters against package ID, but does not explicitly state read-only behavior or potential failure modes (e.g., no apps found, permission issues). For a simple list, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short, front-loaded paragraphs with a clear args section. Every sentence earns its place, and the primary purpose is stated first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional parameter and an existing output schema, the description covers purpose, usage, and parameter semantics sufficiently. No critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already describes the query parameter, but the description adds the crucial detail that matching occurs against the package ID, not the name. This adds meaning beyond the schema's generic text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists installed apps with readable name and package ID, and explicitly ties its purpose to resolving uncertainty about app names or failures of open_app. This distinguishes it from siblings like open_app and read_screen.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-to-use guidance: when unsure of an app's name or when open_app fails. This directly addresses the selection problem and gives a concrete use case, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_appA
Bring an app to the front, by the name a person would say.
Accepts "YouTube", "youtube", "play store", or an exact package id. The name is checked against what is actually installed before anything is launched, and it waits for the app to reach the foreground before returning.
Args: name: The app to open. wait_seconds: How long to wait for it to reach the foreground.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The app, by the name a person would say: 'settings', 'Play Store', 'com.android.chrome'. | |
| wait_seconds | No | How long to wait for it to reach the foreground before giving up. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It discloses that the app name is checked against installed apps before launching and that the tool waits for the foreground state, which is valuable behavioral context. It does not mention error handling when the app is not found, but the core behaviors are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with an opening sentence, a short paragraph on accepted formats, and an Args section. It is concise but somewhat redundant with the schema parameter descriptions. Front-loaded purpose and clear layout justify a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers the main flow: name resolution, validation against installed apps, and waiting for foreground. An output schema exists, so return values need not be explained. Missing details like failure modes are minor given the tool's scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters fully (100% coverage). The description repeats similar information but adds the explicit acceptance of variations like 'YouTube' and 'play store', slightly enhancing the semantic clarity. This is marginal beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Bring an app to the front') and the resource (an app), and specifies accepted input formats (natural names, package IDs). It does not explicitly contrast with sibling tools like open_url, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (when you need to open an installed app) but does not mention alternatives or exclusions. It lacks guidance on when to use open_url or list_apps instead, leaving the selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_urlA
Open a web address or deep link on the phone.
Args: url: A URL such as "https://example.com", or a deep link such as "geo:37.8,-122.4", "tel:+15551234567", or "sms:+15551234567".
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | A web address such as 'https://example.com', or a deep link such as 'geo:37.8,-122.4'. | |
| wait_seconds | No | How long to wait for it to reach the foreground before giving up. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations available, the description carries the disclosure burden; it does state the core side effect (opening a URL/deep link on the phone). It does not disclose failure/foreground behavior, potential app-switching, or other consequences, and the wait-for-foreground trait is only present in the schema rather than in the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded with the action, and contains no filler. The Args block is slightly redundant with the schema and incompletely lists the optional parameter, so it is not perfectly polished.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with an output schema, the description plus schema is largely sufficient: it covers the required URL parameter and accepted formats, while wait_seconds is documented in the schema. It could be more complete with explicit alternative routing, but an agent generally has enough to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is complete, so the baseline is 3, but the description adds tel: and sms: deep-link examples that go beyond the schema's geo: example. It does not mention wait_seconds in its Args block, though the schema already documents that parameter clearly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Open') and a clear resource ('a web address or deep link'), and the tel:/sms:/geo: examples make the intended scope concrete. This differentiates it from sibling open_app, which is about launching an app rather than resolving a URL/deep link.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied by the name and by the phrase 'web address or deep link', but the description never explicitly compares this tool to open_app or says when not to use it. There are no exclusions or routing hints beyond the action itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keyB
Press a navigation, hardware, or system key.
Args: key: One of back, home, recents, enter, delete, escape, tab, space, move_home, move_end, volume_up, volume_down, power, wake, sleep, camera, play_pause, next_track, previous_track, copy, cut, paste, select_all, search, menu, page_up, page_down, notification_shade, quick_settings, collapse_shade.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Which key, from the list in the description above. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full behavioral burden. It only names the action and valid keys; it does not disclose that keys like power, home, back, or volume can change device state, trigger navigation, or produce side effects. It also says nothing about return behavior or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A one-sentence definition is followed by a compact, scannable list of allowed keys. No filler words, and the verb and resource are front-loaded before the enumeration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter action tool with an output schema, the description is nearly complete: the agent knows the operation and all valid inputs. It falls short of 5 only because it omits behavioral context and alternative routing, but nothing essential is missing for making the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description fully enumerates all valid values for the key parameter, which is the core semantic content; the schema merely references that list. This adds real meaning beyond the schema, even though schema coverage is marked as 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Press') and a well-defined resource category ('navigation, hardware, or system key'), then enumerates the exact allowed keys. This makes it distinguishable from siblings like tap and type_text, though it does not explicitly contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use press_key instead of alternatives such as tap, type_text, scroll, or swipe. The key list implies some usage context, but there are no explicit conditions, exclusions, or comparisons to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_notificationsA
Open the notification shade, read what is in it, then close it again.
Args: close_after: Set False to leave the shade open, when you intend to tap one of the notifications.
| Name | Required | Description | Default |
|---|---|---|---|
| close_after | No | Set false to leave the shade open, when you intend to tap a notification. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the open/read/close sequence and notes that the shade can be left open via close_after=false. Since annotations are absent, this is helpful, but it still omits details such as whether notifications are cleared or whether the operation has any side effects beyond UI state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short, front-loaded with the core procedure, and followed by a single parameter note. There is no filler or unnecessary repetition beyond the duplicated Args text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one optional, well-documented parameter and an output schema present, the description is largely complete for a simple UI action. The remaining gaps are minor: it does not say what happens if the shade is already open or empty, nor does it explicitly route usage away from read_screen.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description's Args section mostly restates the schema's description of close_after. It adds no new semantic detail, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise procedure: open the notification shade, read its contents, then close it. This clearly targets the notification shade resource and distinguishes it from siblings like read_screen or take_screenshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to choose read_notifications over read_screen or other sibling tools. The only conditional advice concerns the close_after parameter, not tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_screenA
List the controls currently on the phone's screen, by number and name.
This is the main way to see what is on screen, and it is far cheaper than a
screenshot. Each control gets a number you can hand to tap, and a label
you can match on.
Every action tool re-reads the screen before it acts, so both forms are resolved against the screen as it is at that moment. A number is precise -- it disambiguates two controls carrying the same text. A label is stable -- it survives the screen being redrawn. Use the number for a control you have just seen and are about to act on, and the label when the screen may have changed since you read it.
Args: query: Optional filter. Only controls whose text, description, or type contains this text are listed, best match first. Use it to ask "is the Send button up yet?" in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Only list controls whose text or description contains this. Omit to list everything on the screen. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels. It discloses the return format (numbered controls with labels), the resolution semantics (numbers are precise, labels are stable), and that action tools re-read the screen, implying read_screen results may be stale after a redraw. This is rich behavioral context beyond what any annotation would add.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then layers value (cost comparison, number/label guidance) and finally parameter details. It is efficiently structured with clear paragraphs and zero filler. Every sentence earns its place, and the length is justified by the nuanced distinction it explains.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional parameter) and the presence of an output schema, the description covers all necessary aspects: what it does, how it differs from a screenshot, when to rely on numbers vs labels, when not to call it, and how to use the filter. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema already describes the `query` parameter (100% coverage), the description enriches it: it specifies the matching criteria (text, description, or type), notes the ordering ('best match first'), and gives a practical example. This goes well beyond the schema's bare description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise statement: 'List the controls currently on the phone's screen, by number and name.' It explicitly differentiates from the sibling `take_screenshot` by noting it is 'far cheaper' and returns interactive controls rather than an image. This makes the tool's purpose unmistakable and distinct from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use this tool versus `take_screenshot` (cheaper for control info), and when not to call it at all: 'Every action tool re-reads the screen before it acts.' It also gives specific rules for choosing between number and label forms, and demonstrates a concrete use case for the query filter ('is the Send button up yet?'). These are explicit usage directives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_shellA
Run a shell command on the phone and return its output.
The escape hatch for whatever the named tools do not cover: reading a setting, listing files, checking what is installed, granting a permission. Commands run as the shell user, which can read most of a stock phone but cannot change protected settings.
Args: command: The shell command line to run on the device.
| Name | Required | Description | Default |
|---|---|---|---|
| command | Yes | The shell command, run on the phone as the shell user. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It does add useful behavioral context: commands run as the shell user, can read most of a stock phone, and cannot change protected settings. However, it does not warn that arbitrary shell commands can have side effects, hang, or be destructive, which is a significant omission for an un-annotated generic shell tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized and front-loaded: one-sentence purpose, usage framing, privilege caveat, then an Arg block. The examples are useful and not wasteful; the only minor redundancy is the Arg line duplicating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The one parameter and output schema are present, and the core purpose/usage is covered. But because this is a generic arbitrary-shell tool with no annotations, a warning about destructive potential or hanging commands would be needed for the description to be fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the command parameter. The description adds only the phrase 'command line' and largely repeats the schema's meaning, with no additional shell syntax, quoting, timeout, or execution details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run a shell command on the phone') and immediately frames the tool as the escape hatch for cases named tools do not cover. This clearly distinguishes it from specialized siblings like list_apps or read_screen.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says the tool is for 'whatever the named tools do not cover' and gives concrete examples: reading a setting, listing files, checking what is installed, granting a permission. This tells an agent when to reach for it versus the specialized sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_taskA
Give the phone a goal and let it work the whole thing out itself.
Use this instead of driving the phone a step at a time. It reads the screen, decides what to do, does it, and reads again, until the goal is reported done or unreachable — so "open the email app" is one call here rather than ten round trips through you.
Every step reports the action, what came of it, and the two numbers Jev gave
for the choice: probability is how likely that option was the best one and
confidence is how sure it was of its own ranking. They are reported rather
than gated on, so a run that succeeded on a thin lead says so, and a run that
stopped for a reason other than reaching the goal says that instead.
achieved is what the run claimed. Read the steps for what actually
happened. If a run needs looking at afterwards, set PHONE_CONTROL_RUNS to a
directory and each run leaves a folder there with the full record of what it
asked Jev and what came back; the reply then names the folder.
Prefer the step-by-step tools when you already know exactly what to tap, when the screen is visual rather than textual, or when you need to inspect something part-way through.
Args: goal: What to achieve, in plain language. For example "open the email app" or "turn on airplane mode". max_steps: How many actions to allow before giving up. Reaching it ends the run honestly as unfinished rather than as a failure of the goal.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | What you want done, in plain language. Phrase it as the outcome, not the steps. | |
| max_steps | No | How many actions to allow before giving up. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the iterative read-decide-act loop, the reporting of probability and confidence, the distinction between claimed achievement and actual steps, the optional run log directory, and the max_steps termination behavior. No contradictions with annotations since none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence adds value. It front-loads the core concept, then details behavior and parameter nuances. Slightly more verbose than necessary, but each paragraph earns its place. A 4 rather than 5 due to minor redundancy in the step-by-step preference explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an autonomous agent tool with multiple behavioral nuances, the description covers usage, alternatives, parameter semantics, and operational details. It even mentions the optional environment variable for logging. The output schema likely covers return values, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds significant extra meaning: it explains the goal should be phrased as an outcome, provides examples, and clarifies max_steps as a cap that results in an honest 'unfinished' state rather than failure. This goes beyond the schema's basic descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it takes a goal and autonomously drives the phone to completion, contrasting with step-by-step tools. It distinguishes itself from siblings like 'tap' and 'open_app' by emphasizing the autonomous loop, making it unambiguous what this tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use this tool ('Use this instead of driving the phone a step at a time') and when to prefer alternatives ('Prefer the step-by-step tools when you already know exactly what to tap, when the screen is visual rather than textual, or when you need to inspect something part-way through'). This is textbook usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollA
Scroll the content on screen.
Args: direction: Where you want to look, not which way the finger moves. "down" reveals content further down the page, "up" goes back toward the top, and "left" and "right" move sideways. distance: "small" nudges, "page" moves about one screenful (the default), "large" jumps most of the way. target: Optional number or label of the region to scroll, when the screen has more than one scrollable area.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Which control, by its number from read_screen or by its visible label. | |
| distance | No | How far: a nudge, about one screenful, or most of the way. | page |
| direction | No | Which way to move, or which way the copy goes for a file. | down |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the key behavioral nuance (direction means where to look, not finger movement) and the distance semantics, but it does not mention error handling, what happens when scrolling is impossible, or side effects beyond the scroll itself. The output schema exists but is not described in the description, so return behavior is left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a well-structured docstring with a concise intro and a bulleted Args section. Every sentence earns its place, and the most critical clarification (direction semantics) is front-loaded. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple scroll tool, the description covers the parameters thoroughly and the presence of an output schema covers return values. It lacks edge-case details (e.g., what happens when reaching the end of scrollable content or invalid targets), but these are minor and likely inferable from the tool's purpose. Overall, it is sufficiently complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds substantial value: it explains that 'down' reveals content further down, clarifies 'distance' options (small/page/large), and specifies that 'target' is an optional number or label. This goes beyond the schema's terse descriptions, especially clarifying direction semantics, which is often misunderstood.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (scroll) and the resource (content on screen), and it immediately clarifies the direction semantics ('where you want to look, not which way the finger moves'), which distinguishes it from gesture-based tools like swipe. The purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides detailed guidance on parameter usage (direction, distance, target) and implicitly explains when to use it by contrasting with other tools, but it does not explicitly name alternatives or state when not to use it. It is clear enough for an agent to select it for scrolling, though explicit exclusions would improve it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusA
Check the phone connection and see what it is currently doing.
Call this first when a command fails, and before starting anything multi-step. It reports the model, Android version, screen size, foreground app, whether the screen is on or locked, and the battery level. It changes nothing on the phone.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It explicitly states 'It changes nothing on the phone', providing a crucial non-destructive guarantee. It also enumerates outputs, aiding expectation setting. Missing minor details like behavior on connection failure, but sufficient for a read-only status check.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, each serving a distinct purpose: purpose, usage guidance, and reported data plus safety. No filler, information is front-loaded and well organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description covers all essentials: what it does, when to call it, what it returns, and that it's side-effect free. The existing output schema handles detailed return structure, so nothing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description adds no parameter meaning. Following the baseline for 0 parameters, a score of 4 is appropriate; the description doesn't need to elaborate on inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks the phone connection and provides current device state. It lists specific information reported (model, Android version, screen size, foreground app, screen state, battery), distinguishing it from sibling tools like read_screen or take_screenshot which focus on UI content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to call this first when a command fails and before any multi-step operation. This gives clear, actionable guidance on when to use the tool, and the non-destructive nature further encourages its use as a diagnostic first step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
swipeA
Drag between two points, for gestures no named control covers.
Use scroll for ordinary page movement. Reach for this to dismiss a card,
pull to refresh, draw a pattern lock, or drag a slider.
Args: start_x: Horizontal pixel where the finger lands. start_y: Vertical pixel where the finger lands. end_x: Horizontal pixel where the finger lifts. end_y: Vertical pixel where the finger lifts. duration_ms: How long the drag takes. Longer is slower and more likely to be read as a drag rather than a fling.
| Name | Required | Description | Default |
|---|---|---|---|
| end_x | Yes | Horizontal pixel where the drag ends. | |
| end_y | Yes | Vertical pixel where the drag ends. | |
| start_x | Yes | Horizontal pixel where the drag starts. | |
| start_y | Yes | Vertical pixel where the drag starts. | |
| duration_ms | No | How long the drag takes. Slower is more likely to be read as a drag than a fling. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals touch semantics ('finger lands/lifts'), and explains the effect of duration ('Longer is slower and more likely to be read as a drag rather than a fling'). While it doesn't cover every side effect, the examples and duration behavior provide solid transparency above the baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: one-sentence summary, then usage guidance, then a brief args list. Every section adds value, and the repeated arg descriptions are short and not excessive. Well-structured for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 params, output schema present), the description covers the core context: what it does, when to use it, and behavior nuances. It omits coordinate origin details, but that is likely standard environment knowledge. Overall, an agent can call this tool correctly without further documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the description largely repeats the schema's parameter descriptions. It adds minimal new meaning beyond the schema, such as the phrasing 'finger lands' vs 'drag ends', but no substantive new information. Baseline 3 is appropriate when the schema already fully documents parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a 'drag between two points' and provides specific use cases (dismiss card, pull to refresh, pattern lock, slider). It explicitly differentiates from the sibling `scroll` tool, leaving no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use this vs. alternatives: 'Use `scroll` for ordinary page movement. Reach for this to dismiss a card, pull to refresh, draw a pattern lock, or drag a slider.' This gives clear when-to-use guidance and a direct exclusion for ordinary scrolling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_screenshotA
Capture the phone's screen and look at it.
Try read_screen first: it names things so you can act on the result, and
it costs far less. Take a screenshot when the screen is genuinely visual --
a photo, a map, a game, a CAPTCHA, an icon-only toolbar -- when read_screen
returns nothing useful, or when a decision came back with low confidence.
Args: max_width: Downscale the image to this width in pixels before returning it. Lower it to spend fewer tokens; 0 keeps full resolution.
| Name | Required | Description | Default |
|---|---|---|---|
| max_width | No | Downscale the image to this width in pixels. Lower costs fewer tokens to look at. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It accurately explains that the screenshot is returned, can be downscaled, and incurs a token cost, which is meaningful operational behavior. It does not mention potential limitations like image format or screen capture permissions, but those are not critical for this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a one-sentence purpose, a focused usage paragraph that earns its place by routing the agent correctly, and a compact args explanation. No filler or redundant content; the structure supports fast parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool with no annotations or output schema, this description covers what an agent needs: when to use it, how to control image size, and what the result is (an image to look at). It also names the sibling tool and the tradeoff, making it fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents `max_width` at 100% coverage, so baseline is 3. The description adds a useful semantic detail not in the schema: `0 keeps full resolution`, and reiterates the token-cost tradeoff in a more decision-oriented way. This exceeds baseline by providing practical guidance on how to choose the value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool captures the phone screen and looks at it, using a specific verb and resource. It explicitly differentiates itself from the sibling `read_screen` by naming what it is not and when the alternative is preferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives direct, actionable guidance: try `read_screen` first, use a screenshot only for genuinely visual content or when `read_screen` is unhelpful or low-confidence. This explicitly defines when to use the tool versus its sibling, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tapA
Tap, or long-press, something on the phone screen.
Args:
target: The number from read_screen (for example 7), or the visible
label of the control ("Send", "Sign in"). A label is matched against
the control's text and description, and the best match wins. The
screen is re-read first, so the target is resolved against what is
on screen now; the reply names exactly what was tapped.
x: Horizontal pixel coordinate. Use only when nothing on screen is
nameable, such as a spot on a map or photo.
y: Vertical pixel coordinate, used together with x.
long_press: Hold instead of tapping, for context menus and drag handles.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | Horizontal pixel coordinate, for tapping a point instead of a control. | |
| y | No | Vertical pixel coordinate, for tapping a point instead of a control. | |
| target | No | Which control, by its number from read_screen or by its visible label. | |
| long_press | No | Hold instead of tapping, for a context menu or a drag handle. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it delivers: it discloses that the screen is re-read first, that target resolution uses text and description with best-match semantics, and that the reply names exactly what was tapped. This goes well beyond a generic 'tap' explanation and helps the agent predict runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a one-line summary followed by clearly labeled argument explanations. Every sentence adds useful information—examples, conditions, and behavioral notes—with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four parameters and no annotations, the description covers all parameters, explains when to use each mode, describes the matching algorithm, and notes the response behavior. The presence of an output schema makes the lack of a detailed return-value explanation unnecessary; the definition is complete for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds substantial meaning beyond the schema: target may be an integer from read_screen or a visible label, matching is against text/description with best-match, x/y are coordinate fallbacks for non-nameable spots, and long_press triggers context-menu or drag-handle behavior. This richly compensates for the otherwise terse property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Tap, or long-press, something on the phone screen.' This clearly identifies the tool as a screen-interaction primitive and differentiates it from siblings like swipe, scroll, and press_key, which involve different gestures or targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage direction: target can be a read_screen number or visible label, and x/y should 'only' be used when 'nothing on screen is nameable.' It also explains long_press for context menus and drag handles. It stops short of explicitly contrasting tap with alternative sibling tools, but the guidance is sufficient for most cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transfer_fileA
Copy a file between this computer and the phone.
Args: direction: "to_phone" or "from_phone". computer_path: Path on this computer. "~" is expanded. phone_path: Path on the phone, for example /sdcard/Download/report.pdf.
| Name | Required | Description | Default |
|---|---|---|---|
| direction | Yes | Which way to move, or which way the copy goes for a file. | |
| phone_path | Yes | Path on the phone, for example /sdcard/Download/report.pdf. | |
| computer_path | Yes | Path on this computer. '~' is expanded. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full behavioral burden. It states the operation (copy), supports bidirectional transfer via 'direction', and notes tilde expansion for computer_path. However, it does not disclose overwrite behavior, error handling, or whether directories are created, which are important for a file copy operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise: a one-sentence purpose followed by a terse parameter list. It is front-loaded with the main function, and every line adds information without redundancy. There is no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three required parameters and a straightforward operation, the description covers the essential usage: direction and both paths. It omits behavioral details like overwrite policies, but the presence of an output schema (though its content is unknown) may address return values. Given the simplicity and unrelated siblings, this is fairly complete, though a note on conflict handling would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all parameters. The description adds value by explicitly listing the valid values for 'direction' (to_phone/from_phone) and clarifying that '~' is expanded in computer_path, which are not fully specified in the schema. This is a meaningful enhancement beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool copies a file between a computer and a phone, with a specific verb ('copy') and resource ('file'). It is distinct from all sibling tools, none of which handle file transfer, so there is no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly indicates usage by explaining the purpose and the parameters, but it does not explicitly mention when to use it versus alternatives. However, since no sibling tool serves the same function, the intended usage is clear from context alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
type_textA
Type text into a field on the phone.
Args:
text: The text to type. Use "\n" for a line break.
target: The field to type into, as a number or label from read_screen.
When omitted, the text goes to whatever already has focus.
replace: Clear the field first instead of appending to it.
submit: Press Enter afterwards, for search bars and message boxes.
verify: Re-read the screen afterwards and report whether the text
actually landed, which catches a tap that missed the field.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text to type, or the text to replace the clipboard with. | |
| submit | No | Press Enter afterwards, for search bars and message boxes. | |
| target | No | Which control, by its number from read_screen or by its visible label. | |
| verify | No | Read the screen back afterwards and report whether the text landed. | |
| replace | No | Clear the field first instead of adding to what it holds. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden well: it explains the default append behavior, the replace option to clear first, the submit action, and verify re-reading the screen to confirm input landed. It does not fully disclose every edge case, but the key side effects are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a one-line purpose followed by a tight Args list. Every sentence earns its place by explaining behavior rather than repeating generic boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all five parameters with practical guidance, explains side effects clearly, and works with the output schema to give the agent enough to invoke the tool correctly. No critical invocation details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: the '\n' line-break convention, focus behavior when target is omitted, and the purpose of verify as a safety net for missed taps. Each parameter is given practical semantics not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Type text into a field on the phone') and immediately clarifies the tool's core action. It is clearly distinct from siblings such as tap, press_key, and clipboard, even without naming them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives solid parameter-level usage guidance, such as using submit 'for search bars and message boxes' and verify to catch a missed field. However, it never explicitly says when to choose this tool over alternatives like tap, press_key, or clipboard; that selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forA
Wait until something becomes true, instead of sleeping a fixed time.
Use this after an action that triggers loading, so the next step reads a settled screen instead of one mid-animation. With no arguments it waits for the screen to stop changing.
Args: text: Wait until this text appears anywhere on screen. app: Wait until this app (a name or a package id) is in the foreground. timeout_seconds: Give up and report the timeout after this long.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | Wait until this app is in the foreground. | |
| text | No | The text to type, or the text to replace the clipboard with. | |
| timeout_seconds | No | Give up and report the timeout after this long. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It explains that the tool waits for a condition rather than sleeping, that it can wait for screen stability, and that timeout_seconds causes it to give up and report the timeout. It does not detail polling mechanics, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core behavior, then gives one-sentence usage guidance, then a compact Args list. Every sentence adds value and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-argument polling tool with no annotations, it covers the main behavior, the recommended use case, all parameters, and timeout behavior. It does not say how multiple conditions combine or what the timeout response looks like, but the output schema can cover the return shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The Args section maps each parameter to its meaning and adds useful detail like app being 'a name or a package id' and text matching 'anywhere on screen.' This compensates for the input schema's text description, which appears copied from a type_text tool ('The text to type, or the text to replace the clipboard with.') and is misleading for wait_for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and outcome ('Wait until something becomes true, instead of sleeping a fixed time') and then gives three concrete wait conditions: text on screen, app in foreground, screen stable. This clearly distinguishes it from a blind sleep and from the sibling read/screen tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use it: 'Use this after an action that triggers loading, so the next step reads a settled screen instead of one mid-animation.' It does not provide formal when-not-to-use conditions or name alternatives, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
19 tool updates
v0.1.0- First observed
ask_jev - First observed
clipboard - First observed
decide_next_action - First observed
list_apps - First observed
open_app - First observed
open_url - First observed
press_key - First observed
read_notifications - First observed
read_screen - First observed
run_shell - First observed
run_task - First observed
scroll - First observed
status - First observed
swipe - First observed
take_screenshot - First observed
tap - First observed
transfer_file - First observed
type_text - First observed
wait_for
TDQS
Scored across 19 tools
Each tool targets a distinct capability: screen reading (read_screen, take_screenshot), interaction (tap, type_text, scroll, swipe, press_key), app handling (list_apps, open_app), navigation (open_url), waiting (wait_for), notifications, clipboard, shell, file transfer, and three Jev-driven decision tools that are clearly separated by purpose (next action, autonomous task, general question). No two tools could be easily confused.
All tool names follow a consistent snake_case verb_noun/verb_phrase pattern (read_screen, take_screenshot, type_text, open_app, press_key, wait_for, run_shell, etc.). The style is uniform and predictable, making it easy for an agent to infer function from the name.
With 19 tools, the server is slightly over the typical well-scoped range (3-15), but the scope of a phone-control server is broad, covering screen interaction, system controls, app management, and advanced decision-making. Each tool earns its place, though a few (e.g., separate Jev tools) could arguably be merged; still, the count is reasonable and not bloated.
The tool surface provides comprehensive coverage for phone control: reading the screen (text and visual), acting on screen content, navigating the system, managing apps and URLs, waiting for state changes, handling notifications, clipboard, shell escape hatch, file transfer, and high-level autonomous decision-making. No significant gaps are apparent for the stated purpose.
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Drive real devices from your AI Coding tool. Embed a client SDK (Unity, Godot, Flutter, iOS/macOS, Android, React Native, Web) in your app, then capture screenshots, traverse the UI tree, inject taps and key events, and run automated test tasks on the physical device over a secure relay.
Disposable cloud Android emulators for coding agents: run an APK or PR build, tap, type, screenshot.
Melaya is a remote MCP server. It gives an assistant hands on your own Android phone and browser: it reads the screen through the accessibility tree, then taps, types and navigates inside the apps and sites you allow-list, with no per-app API. It also builds, schedules and runs agent pipelines across 6k+ connected tools. OAuth 2.1, nothing to install.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI models to control Android devices via ADB through natural language commands, supporting screen analysis and automated actions.339 npmMIT
- AlicenseAqualityDmaintenanceEnables an agent to inspect and interact with Android emulators or physical devices via ADB, capturing UI snapshots, tapping nodes, typing text, and reading app logs.10MIT
- AlicenseBqualityDmaintenanceEnables AI agents to control Android devices via ADB, supporting gestures, input, screenshots, UI analysis, and app management.1916 npmISC
- AlicenseNot gradedqualityDmaintenanceEnables coding agents to observe and operate Android devices through ADB and uiautomator2, supporting screenshot capture, UI hierarchy access, and actions like tap, swipe, and text input.1Apache 2.0