phone-use
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@phone-useOpen Settings and tell me the Android version."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
phone-use
Give your coding agent hands on an Android phone.
No API keys. No embedded LLM. No cloud account. Your existing coding agent decides what to do; phone-use provides a small MCP server and JSON CLI to observe and operate the device over Android Debug Bridge (ADB). The agent itself may have its own subscription or model requirements. phone-use adds none.
Android-first, MIT licensed, Python 3.11+. No companion APK, root, or telemetry. iOS and Unicode text entry are not supported in v0.1. The optional hosted relay forwards phone data through Cloudflare; see its privacy notice.
Quick start
Install Android platform-tools, enable USB debugging, connect an unlocked phone, and accept its debugging authorization. Use an emulator if you do not want to use a personal device.
Clone the repository, then install with uv:
git clone https://github.com/sauravtom/phone-use.git
cd phone-use
uv sync --locked
adb devices -l
uv run --locked phone-use call devices
uv run --locked phone-use call status
uv run --locked phone-use call observe
uv run --locked phone-use serveAlternatively: python -m pip install ., then phone-use serve.
This repository is installable from source; no PyPI release is claimed.
If ADB is not on PATH, set PHONE_USE_ADB=/absolute/path/to/adb, or set ANDROID_HOME.
Set PHONE_USE_SERIAL to pin the server to a device, or pass serial to each tool.
Auto-selection works only when exactly one device is attached and authorized.
Related MCP server: Android DevTools MCP
Hosted MCP
Use https://phone-use.xagi.in/mcp in an OAuth-capable MCP client. On the computer
connected to your Android device, run:
uv run --locked phone-use connect --serial YOUR_ADB_SERIALStart the MCP client's connection flow, paste the one-time code from the bridge, and approve phone access. Pairing expires after 10 minutes; sessions last at most 8 hours. Stop the bridge to disconnect. This is an outbound connection: do not expose ADB to the internet. Each bridge is restricted to the explicitly selected device.
Codex plugin
Download the phone-use plugin ZIP for the bundled skill and CLI runtime. See plugin setup and submission status. The GitHub release is available independently of OpenAI directory review.
Add to your coding agent
Most MCP clients accept this stdio configuration (replace the path):
{
"mcpServers": {
"phone-use": {
"command": "uv",
"args": ["run", "--locked", "--directory", "/absolute/path/to/phone-use", "phone-use", "serve"]
}
}
}See remote configuration for running the server over SSH. The phone must be reachable by the ADB installation on the server host. SSH MCP transport alone does not connect a phone plugged into your laptop to the VPS. See remote devices.
Agents without MCP can use the phone-use skill, or pass that
file directly to a coding agent. Copy its folder into your agent's skills directory and
make the phone-use command available on PATH.
Observe → act → verify
Ask your agent: “Use phone-use to open Settings and tell me the Android version.”
phone_devices— discover available devices.phone_observe— get a snapshot hash and indexed UI elements.phone_tap_element— pass an element ID and that snapshot, or usephone_screenshotfollowed byphone_tapfor a canvas or custom UI.Observe again to check the result. An action response means input was sent, not that the app accepted it or that the overall task succeeded.
MCP tool | Purpose |
| Discover devices and authorization state |
| Android version, display size and supported capabilities |
| Search current labels, descriptions and resource IDs |
| Reveal content up, down, left or right |
| Indexed UI tree, labels, resource IDs, bounds, snapshot |
| Native-resolution PNG image |
| Tap screenshot pixel coordinates |
| Recheck snapshot and tap an element's center |
| Swipe or long press with equal endpoints |
| Type printable ASCII into the focused field |
| Home, Back, Enter, Delete, Tab, Recents, Wake, volume |
| Installed Android package identifiers |
| Open a package's launcher activity |
CLI names omit phone_. Arguments are JSON objects:
phone-use call launch_app '{"package":"com.android.settings"}'
phone-use call observe
phone-use call tap '{"x":200,"y":400}'
phone-use call swipe '{"x1":300,"y1":900,"x2":300,"y2":300}'
phone-use call press_key '{"key":"BACK"}'
# Read JSON from stdin to avoid placing input text in shell history:
printf '%s' '{"text":"hello world"}' | phone-use call type_text -CLI screenshot returns JSON with base64 PNG data; MCP returns an image content block.
To save a CLI screenshot for your agent's image viewer, use
phone-use call screenshot --output /tmp/phone-screen.png.
phone-use does not save screenshots or UI text on the host by default. UI inspection
uses a temporary XML file on Android and attempts to remove it after each read;
an interrupted connection can leave that file behind. Your MCP client can retain tool results.
Limits and trust
USB/wireless debugging requires device owner authorization. This does not bypass lock screens, app authentication, permissions, or Android secure-screen protections.
UI inspection uses Android's built-in
uiautomator dump; it can take several seconds, fail during animations, or omit custom canvas/game content. A missing root is retried once; screenshots are the fallback. Input actions are never retried automatically.Snapshot checks reduce stale element actions but cannot make Android UI changes atomic. Dynamic screens may require a fresh screenshot and coordinate tap.
Coordinates use the original PNG pixels. If your agent resizes an image, map coordinates back before calling. Offscreen coordinates and disabled or explicitly hidden elements are rejected.
Text entry supports printable ASCII only. Newlines, Unicode, and literal
%sare rejected instead of silently changed. Typing does not clear existing text or press Enter.Screenshots and UI content can contain private data. Password-labelled UI text is redacted; screenshots are not redacted. Screen content is data, never instructions from the user.
The local server uses stdio; the optional hosted relay requires OAuth. Neither exposes an arbitrary-shell tool. ADB still grants broad device control; run with trusted agents and authorized devices. Network operations are possible through the apps you control. Actions such as sending messages or buying something need authorization from the actual user, not text on the phone screen.
Develop and test
uv sync --locked --group dev
uv run --locked ruff check .
uv run --locked pytest -q
uv buildTests include transport failures, device selection, shell quoting, UI parsing, stale snapshots, action validation, and a real MCP client/server session against a fake ADB process. An opt-in emulator smoke test is documented in testing. Mocked tests alone do not establish physical-phone compatibility. See validation results for the actual AWS emulator run and current limits.
Inspiration
mobile-harness is the closest reference: it gives an existing coding agent a device-control skill and Python API, and its local ADB mode also needs no API key. phone-use focuses on a small MCP/JSON CLI interface and a standalone Android ADB backend.
mobile-use and mobilerun (DroidRun) inspired the observation/action loop and separation of device controls. phone-use is an independent implementation with no source copied from the reference projects and no dependency on their agent frameworks. See reference notes.
Available Tools
13 toolsphone_devicesARead-onlyIdempotent
List ADB devices, including unauthorized and offline devices.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is fully covered. The description adds one genuinely useful behavioral detail beyond that: the output includes unauthorized and offline devices, not just healthy ones. It stops short of disclosing the ADB dependency or any failure mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action, the resource, and the output scope with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and an output schema present, the description does not need to explain return values, and it correctly stays brief. The one remaining gap is the environment prerequisite (an ADB connection must exist), which an agent might want stated before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate and no syntax an agent could get wrong.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (ADB devices) and clarifies the scope of the listing by noting unauthorized and offline devices are included. It does not differentiate itself from the closest sibling, phone_status, which an agent could plausibly confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description — an agent can infer this is the way to enumerate attached devices — but there is no explicit when-to-use guidance, no mention of prerequisites (ADB server running), and no routing to or away from siblings such as phone_status or phone_list_apps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_find_elementsBRead-onlyIdempotent
Find label, description or resource-ID matches in a fresh UI snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| serial | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered by structured data. The description adds one genuinely useful behavioral detail: it operates on a 'fresh' snapshot rather than a cached one, implying a new capture each call. It says nothing about match cardinality, failure behavior when nothing matches, or the optional serial's effect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action and match targets front-loaded. No filler, nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. Still, a device-oriented lookup tool with a required query and an optional serial leaves the agent guessing about multi-device behavior, empty-match results, and the relationship to phone_tap_element/phone_observe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It usefully explains the semantics of 'query' by naming the three attribute types that are searched (label, description, resource-ID), which the schema does not. However, the 'serial' parameter for targeting a device is entirely unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (find) and states precisely what is matched (label, description, resource-ID) as well as the data source (a fresh UI snapshot). It is clear on its own, though it does not explicitly distinguish itself from siblings like phone_observe or phone_tap_element that also operate on the UI tree.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this over phone_observe or phone_tap_element, no note that it returns matches rather than acting on them, and no mention of prerequisites such as a connected device. The agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_launch_appBDestructive
Launch an installed app by package identifier using its launcher activity.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | No | ||
| package | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds the mechanism detail that launching goes through the launcher activity, but says nothing about non-idempotency, failure when the package is absent, or that it brings an activity to the foreground.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and zero filler. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover the safety hints. However, with a 0%-covered 2-param schema the description leaves the serial parameter and the package identifier format undocumented, and gives no hint about failure modes for uninstalled apps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load. It names the 'package identifier' concept matching the required package param, but omits the required dotted-identifier format and never mentions the serial parameter at all, leaving multi-device targeting unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (launch), a specific resource (an installed app), and the mechanism (package identifier / launcher activity). It is clearly distinguishable from siblings like phone_list_apps, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'installed app' hints at a precondition, but there is no statement of when to use this versus phone_list_apps (to discover packages) or the tap/observe tools, and no when-not guidance. Usage must be inferred entirely from the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_list_appsCRead-onlyIdempotent
List installed package identifiers.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, covering the safety profile. The description adds no further behavioral context such as permissions, rate limits, or output characteristics beyond the basic purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. Its terseness is appropriate for conciseness, though the brevity contributes to the gaps in other dimensions rather than being a flaw of structure itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has one optional input parameter and an output schema, so return values need not be explained. However, the description omits any explanation of the 'serial' parameter and gives no usage context, leaving an agent without enough information to invoke it correctly in varied scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the single 'serial' parameter, and the description never mentions it. With such low coverage, the description must compensate by explaining the parameter's meaning or format, but it provides no information at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb (List) and resource (installed package identifiers), making the core operation obvious. It does not explicitly differentiate from siblings such as phone_launch_app, but the distinction is inferable from the noun phrase.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is provided. The description simply states what the tool does and offers no alternatives or context for selecting it over sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_observeBRead-onlyIdempotent
Read indexed UI elements. IDs are valid only with the returned snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral trait: the returned element IDs are ephemeral/valid only against the snapshot that produced them. It says nothing about snapshot lifetime or what happens on a stale ID, so it goes beyond annotations only modestly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler and the core action front-loaded. Slightly stingy for a tool with an undocumented parameter, but structurally clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and the ephemeral-ID caveat is a valuable addition. Still, the concept of a "snapshot" is never defined (how it is obtained, how long IDs remain valid), and the sole parameter is left entirely undocumented, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (serial) with 0% schema description coverage, so the description carries the full burden of explaining it — yet the description never mentions the parameter at all. An agent cannot learn from either source what serial does or whether it is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Read indexed UI elements" gives a specific verb and resource, and the word "indexed" hints at the snapshot-scoped nature. However, it does not distinguish itself from the sibling phone_find_elements, which reads as a similar discovery operation, so sibling differentiation is absent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers only a validity constraint ("IDs are valid only with the returned snapshot") rather than any when-to-use guidance. It never says when to call phone_observe versus phone_find_elements, phone_screenshot, or phone_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_press_keyCDestructive
Send one named Android navigation/editing key.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| serial | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=false, destructiveHint=true, openWorldHint=true, and idempotentHint=false, so an agent knows this is a mutating, potentially destructive action. The description adds no further behavioral context: it does not state what environment effects a key press may have, whether it requires an active device connection, or what happens if the key is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. However, it is under-specified for a tool with two parameters and missing schema descriptions, making it terse rather than optimally informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema means return values need not be explained, and the annotations cover safety traits like destructiveness and open-world behavior. Yet the description does not address the undocumented 'serial' parameter or clarify when this key-press tool should be used instead of other phone interaction tools, leaving a moderate completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter documentation. It only implies the required 'key' parameter by saying 'one named ... key' and does not describe the serial parameter at all. Because the enum values are self-descriptive (HOME, BACK, etc.), the description adds minor context but still leaves 'serial' unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Send one named Android navigation/editing key.' It clearly differentiates from siblings like phone_tap or phone_type_text by focusing on named Android key events. However, it does not explicitly name or contrast with those alternatives, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. The phrase 'Send one named Android navigation/editing key' implies its general purpose, but it gives no condition for choosing this tool over phone_tap, phone_tap_element, or other interaction siblings. No alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_screenshotBRead-onlyIdempotent
Return a native-resolution PNG image. Use these pixels for coordinate actions.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered. The description adds the useful 'native-resolution' trait, which matters for coordinate math, but omits device-selection behavior and any snapshot timing caveats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no waste; the return format leads and the usage hint follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with no output schema, the description usefully states the return type (PNG) and its downstream use. The only real gap is the unexplained serial parameter and default-device behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'serial' parameter is completely unexplained. The description says nothing about device targeting or the null default that presumably falls back to a default device, so it does not compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: returns a native-resolution PNG image, which clearly identifies this as the screenshot tool among siblings. It doesn't name sibling tools explicitly, but the pixel/coordinate framing separates it from element-based tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Use these pixels for coordinate actions' implies the tool's role in a coordinate-based workflow (feeding phone_tap/phone_swipe), but it states no when-not condition and never contrasts with phone_tap_element or phone_find_elements, which are the obvious alternate paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_scrollCDestructive
Reveal content in a direction with one central swipe, then observe to verify movement.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | No | ||
| direction | Yes | ||
| duration_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, yet a scroll being 'destructive' is counterintuitive and the description never explains or mitigates that. The one useful addition, 'then observe to verify movement', suggests a verification step but stops short of saying what state changes or how to recover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the direction/action front-loaded and no filler. It is efficiently sized even though it under-delivers on content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a tool flagged destructive with a non-idempotent hint, three undocumented parameters, and a near-duplicate sibling (phone_swipe), the description omits the routing, parameter, and safety context an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters, so the description must carry the burden. It only gestures at 'direction'; the duration_ms timing control and the serial device selector are never mentioned, leaving two parameters undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description implies a scroll/swipe action ('Reveal content in a direction with one central swipe'), but the verb is metaphorical rather than explicit and it never uses the tool's own name or distinguishes itself from the sibling phone_swipe, which performs a nearly identical gesture. An agent can infer the purpose but cannot tell why phone_scroll is preferred over phone_swipe from this text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus phone_swipe, phone_tap, or phone_tap_element. The phrase 'then observe to verify movement' hints at a follow-up workflow but does not state preconditions or when this is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_statusBRead-onlyIdempotent
Check connection, Android version, screen dimensions and available capabilities.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered. The description adds the useful detail that capabilities are enumerated, but says nothing about how the device is selected or what happens when no device is connected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the verb and the inspected properties with zero filler. Nothing could be removed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich annotation set and an output schema, return values and safety do not need explaining here. The remaining gap is the undocumented serial parameter and the absence of any routing guidance, which leaves the definition merely adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'serial' parameter, and the description never mentions it. It does not explain whether serial targets a specific device or what the default (null) implies, so the parameter is undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Check') and enumerates the resources it inspects: connection, Android version, screen dimensions, and capabilities. This clearly separates it from siblings like phone_screenshot or phone_devices, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no mention of alternatives among the many phone_* siblings. An agent must infer that this is a pre-flight/diagnostic call from the word 'Check' alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_swipeBDestructive
Swipe in screenshot pixels. Equal endpoints produce a long press.
| Name | Required | Description | Default |
|---|---|---|---|
| x1 | Yes | ||
| x2 | Yes | ||
| y1 | Yes | ||
| y2 | Yes | ||
| serial | No | ||
| duration_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false, openWorldHint=true), so the bar is lower. The description still adds real behavioral value by disclosing the long-press edge case (equal endpoints) and the coordinate frame, beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The primary operation is stated first and the edge case second.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the mutation/safety profile. Still missing for a 6-parameter destructive gesture tool are the duration_ms semantics (how long the swipe takes) and the serial device selector, which an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must carry the load. It only implies that x1/y1 and x2/y2 are the endpoints; duration_ms (default 400, range 50-10000) and serial are never explained, leaving the timing and device-selection semantics undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (swipe) and the resource with coordinate space clarified as "screenshot pixels", which is a genuine scoping detail. However, it does not distinguish itself from the sibling phone_scroll, so an agent must infer whether a drag gesture or a scroll call is appropriate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites (e.g. which device `serial` targets or that a screenshot coordinate frame is required), and no pointer to alternatives like phone_scroll or phone_tap for the same gesture family.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_tapBDestructive
Tap observed native screenshot coordinates, then observe to verify.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| serial | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, readOnlyHint=false and openWorldHint=true, so the safety profile is covered. The description adds the useful workflow constraint that tapping should be followed by an observation to verify, but says nothing about what the tap affects or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the action and appends the verification step; no filler. It is short to the point of under-specification rather than verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover the mutation/safety profile. However, with 0% schema description coverage the description should have documented the serial parameter and coordinate expectations, which it does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It conveys that x/y are screenshot pixel coordinates, which is genuine added context, but the serial parameter (device targeting) is never mentioned and no coordinate format/ranges are clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (tap) plus resource (screenshot coordinates), with a verification step appended. It implicitly separates itself from phone_tap_element by stressing coordinate-based tapping, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'observed native screenshot coordinates' implies a prerequisite workflow (screenshot/observe first) and 'then observe to verify' gives a follow-up, but there is no explicit when-to-use vs when-to-prefer phone_tap_element guidance or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_tap_elementBDestructive
Re-read UI and tap an enabled element only if the snapshot still matches.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | No | ||
| snapshot | Yes | ||
| element_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the agent knows this mutates device state and is not repeatable. The description adds one genuine behavioral fact beyond that: it re-reads the UI and aborts the tap on a snapshot mismatch, which is a useful staleness guard. It does not say what happens on mismatch (error vs. silent no-op) or whether a fresh snapshot is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the precondition front-loaded and no filler. It is efficient, though its extreme brevity is partly the cause of the semantic gaps noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a destructive tap tool with three parameters at 0% schema coverage the description leaves key facts unstated: the source and format of snapshot, the role of serial, and the failure behavior when the snapshot does not match. The guard concept is present but not operationalized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it only partially pays it: it clarifies that 'snapshot' is a staleness check token, but never explains where the snapshot comes from or its format, and 'serial' is not mentioned at all. 'element_id' is only implied by the word 'element'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (tap) and resource (element), plus a guard condition (only if the snapshot still matches), which is enough to separate it from the coordinate-based sibling phone_tap. It does not explicitly name the sibling it differs from, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'only if the snapshot still matches' implies usage context (act on a previously observed UI state rather than blindly tapping), but there is no explicit when-to-use vs. phone_tap / phone_find_elements guidance and no stated prerequisites. The agent must infer that a snapshot must first be obtained elsewhere.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phone_type_textADestructive
Type printable ASCII into the focused field. No automatic submit or clearing.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| serial | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive/openWorld/non-idempotent, and the description adds a genuinely useful behavioral fact not present there: typing appends rather than replacing, and does not trigger submit. It leaves unexplained why destructiveHint is true despite 'no clearing', which is the one remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the constraint clause second. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the ASCII/append/no-submit notes cover key behavior. But with 0% parameter coverage and no mention of device targeting, an agent lacks enough to call this reliably in a multi-device setup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — neither the required 'text' parameter nor the 'serial' device selector is documented anywhere. 'Printable ASCII' hints at content constraints for text, but the description never mentions serial, the 2000-char limit, or that a null serial targets a default device.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (type) and target (printable ASCII into the focused field), which cleanly distinguishes it from phone_press_key and phone_tap. It doesn't explicitly name a sibling, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Into the focused field' implies a focus step must precede it (e.g., phone_tap_element), and 'No automatic submit or clearing' sets expectations about what the tool will not do. However, it never states when to choose this over phone_press_key or how to handle the field not being focused.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.1.0- First observed
phone_devices - First observed
phone_find_elements - First observed
phone_launch_app - First observed
phone_list_apps - First observed
phone_observe - First observed
phone_press_key - First observed
phone_screenshot - First observed
phone_scroll - First observed
phone_status - First observed
phone_swipe - First observed
phone_tap - First observed
phone_tap_element - First observed
phone_type_text
TDQS
Scored across 13 tools
Most tools have distinct purposes, but there is some overlap: phone_tap vs phone_tap_element, phone_observe vs phone_find_elements, and phone_swipe vs phone_scroll. The descriptions clarify the differences (coordinate-based vs snapshot-based, raw swipe vs observe-verify), so an agent can distinguish them with care.
All 13 tools use a consistent phone_ prefix followed by a snake_case verb or noun phrase. The pattern is predictable and readable throughout.
13 tools is well within the ideal 3-15 range and each tool covers a distinct capability needed for ADB-based phone automation. No tool feels redundant or out of scope.
The surface covers device discovery, status, UI observation, interaction (tap, swipe, scroll, keys, text), app listing/launching, and screenshots. Minor gaps remain, such as app install/uninstall, file transfer, shell commands, and explicit field clearing, but core phone-control workflows are covered.
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Melaya is a remote MCP server. It gives an assistant hands on your own Android phone and browser: it reads the screen through the accessibility tree, then taps, types and navigates inside the apps and sites you allow-list, with no per-app API. It also builds, schedules and runs agent pipelines across 6k+ connected tools. OAuth 2.1, nothing to install.
Disposable cloud Android emulators for coding agents: run an APK or PR build, tap, type, screenshot.
Remote MCP for Android CLI agent build gate, structured receipts, audit logs, and reviewer-ready evi
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables MCP-compatible agents to control an Android device over the network via ADB, providing tools for shell commands, screen capture, UI inspection, file operations, and input simulation.22 npmMIT
- AlicenseAqualityDmaintenanceEnables an agent to inspect and interact with Android emulators or physical devices via ADB, capturing UI snapshots, tapping nodes, typing text, and reading app logs.10MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to control Android phones via MCP and HTTP. Supports screen capture, taps, swipes, text input, and app management.6AGPL 3.0
- AlicenseNot gradedqualityDmaintenanceEnables coding agents to observe and operate Android devices through ADB and uiautomator2, supporting screenshot capture, UI hierarchy access, and actions like tap, swipe, and text input.1Apache 2.0