Phonebox
Server Details
Cloud Android phones for AI agents: create a phone, read the screen, tap, type, install APKs, park.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- combined-ai/phonebox-android-mcp
- GitHub Stars
- 0
- Server Listing
- Phonebox Android MCP server
TDQS
Score is being calculated.
Available Tools
13 toolsactactDestructiveInspect
Run up to 20 actions in order (tap, type, key, swipe, scroll, back, home, open_app, open_url, wait_for, …). The batch stops at the first failure, and the result shows the screen where it stopped, unless the screen couldn't be read, or not in time. Never repeat an action whose outcome is unknown; observe first.
| Name | Required | Description | Default |
|---|---|---|---|
| actions | Yes | From 1 to 20 actions, run in order. Durations are in milliseconds, and the waits in one batch may add up to 60 seconds. Coordinates are device pixels. | |
| observe | No | What the result shows of the screen afterwards: ui (the element list, the default), screenshot, both, or none. A batch that stops early shows the element list, unless the screen couldn't be read, or not in time. | ui |
| phone_id | Yes | Phone ID from create_phone or list_phones. |
complete_app_uploadcomplete app uploadAIdempotentInspect
Finish an upload once its file has arrived, and wait while Phonebox reads and stores the APK. It answers with the upload: ready to install with install_app, or ended with an error saying why. If it is still processing when this returns, call it again with the same arguments: a repeat only waits.
| Name | Required | Description | Default |
|---|---|---|---|
| upload_id | Yes | The upload's ID, from create_app_upload. | |
| storage_id | Yes | The storageId that upload_url answered with, as {"storageId": "…"}. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false/readOnlyHint=false; the description corroborates and enriches this by explaining the blocking wait, the two terminal outcomes (ready vs. ended with an error reason), and that a repeat call merely waits rather than re-uploading. That retry/polling behavior is genuine value beyond the structured fields, though auth or timeout details are not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then the possible outcomes, then the retry rule — no wasted preamble. The middle sentence ('It answers with the upload: ready to install with install_app, or ended with an error saying why') is slightly stilted in construction but still earns its place by covering the return states.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully covers the return states and the polling contract, and annotations carry the safety profile. For a two-parameter mutation tool this is nearly complete; only timeout/retry-limit expectations are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both upload_id and storage_id documented including their formats and provenance, so the schema carries the parameter burden. The description adds only the implicit point that retries must reuse the same arguments; it offers no additional syntax or format guidance, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Finish an upload'), names the precondition (file has arrived), and the side effect (Phonebox reads and stores the APK). It routes the agent to the correct sibling at each stage — create_app_upload before, install_app after — so it is clearly distinguishable from the other upload/install tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: call after the file has arrived, and call again with the same arguments if processing is still in progress. It names install_app as the follow-on step. It stops short of stating any when-not-to-use case or a maximum retry/backoff policy, so it is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_app_uploadcreate app uploadAInspect
Start an upload of your own APK, such as the debug build you just made, to install it on a phone. MCP can't carry a file, so this returns a one-off upload_url and the exact curl command to send the APK there from your shell; then call complete_app_upload with the storageId it answers. Send the URL one file of up to 500 MiB, within an hour.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing the return payload (one-off upload_url plus an exact curl command), the handoff contract (call complete_app_upload with the storageId), and hard limits (one file, up to 500 MiB, URL valid for one hour). These are exactly the operational constraints an agent must know for a non-idempotent, open-world upload tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the action and rationale, then the mechanics, then the limits. Every clause carries information an agent needs; nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by describing the return value and the follow-up call. For a zero-parameter tool whose real payload moves through a shell command, the description covers purpose, limits, and completion path completely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero parameters, so there is nothing for the description to disambiguate; baseline is 4. The description correctly implies no arguments are needed and instead explains the out-of-band data flow.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start an upload of your own APK') with immediate scope ('to install it on a phone'). It distinguishes itself from the sibling complete_app_upload by naming it as the required follow-up step, so an agent can place both tools in the workflow without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use ('your own APK, such as the debug build you just made') and explicitly routes to complete_app_upload with the returned storageId. There is no explicit when-not condition or mention of the alternative install path (install_app), but the trigger and the sequencing are unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_phonecreate phoneAInspect
Create a cloud Android phone and wait until it is ready (usually under two minutes). It costs $0.06 a minute while running. Always park it when you're done. If it is still creating or starting when this returns, call get_phone with its ID and wait "ready" to keep waiting: it only reads, while start_phone would renew the session.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | A name to tell the phone apart, up to 80 characters. The default is its ID. | |
| country | No | Where the phone appears to be: one of the listed two-letter ISO 3166-1 country codes, such as US or DE. Any other code is refused before a phone is made. Leave it out for any country. | |
| metadata | No | Up to 20 string pairs of your own, such as a customer ID. Keys have up to 40 characters and values up to 256. | |
| idle_timeout | No | Seconds after the phone's last activity, its last_active_at, at which it parks itself, from 60 to 3600. Calls move last_active_at at most once every 15 seconds. The default is 300. | |
| max_duration | No | The most seconds this session may run before the phone parks, from 60 to 10800. The default is up to 900, shortened automatically to available credit and the spending limit. | |
| idempotency_key | No | A unique key for this phone. If a call times out, repeat it with the same key and settings to get the same phone instead of a second one. If that phone has failed, create the next one with a new key. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark this as a non-read-only, open-world, non-destructive write; the description adds the material behavior an agent cannot infer: provisioning takes under two minutes, it bills $0.06 per minute while running, and it may return before the phone is ready. That cost and timing disclosure is genuinely decision-relevant for a spend-incurring tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with what the tool does and the cost, then the operational rules. The final sentence about get_phone versus start_phone is slightly dense and requires re-reading, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the necessary work of explaining the return state (may still be creating or starting) and the follow-up call with the returned ID. It stops short of describing the full returned object, but for a 6-param, all-optional tool with full schema coverage this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so name, country, metadata, idle_timeout, max_duration, and idempotency_key are all already documented in the schema, and the description adds no parameter-level detail beyond the returned ID. Baseline 3 is appropriate when the schema carries the full parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Create a cloud Android phone") plus the completion semantics ("wait until it is ready"), which immediately separates it from siblings like start_phone, park_phone, and get_phone. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit lifecycle guidance: always park when done, and if the phone is still creating/starting, use get_phone with its ID to keep waiting rather than start_phone, which would renew the session. It names the alternative tool and the condition that selects it, which is exactly the routing an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_phoneget phoneARead-onlyIdempotentInspect
Read one phone: its status, session and when it parks. With wait, wait until it is ready or parked. It never starts, renews or bills anything, so it is the way to wait for a phone that is still creating or starting.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | "ready" or "parked": wait until the phone gets there. Leave it out to read the phone as it is. | |
| phone_id | Yes | Phone ID from create_phone or list_phones. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds real behavioral value on top: it explains that the call can block when 'wait' is set, and reassures that it never starts, renews or bills. It does not say what happens if the wait never reaches the target state (timeout/error behavior), which is the one notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core read action, then the wait behavior, then the negative scope clause. No filler and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the right thing by enumerating the returned fields (status, session, park time). Annotations carry the safety semantics. The only omission is the timeout/failure behavior of a 'wait' call that never reaches ready or parked.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum parameter is already documented ('ready' or 'parked': wait until the phone gets there). The description's 'With wait, wait until it is ready or parked' restates the schema rather than adding format or edge-case meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Read one phone' plus what it returns (status, session, park time). This distinguishes it from list_phones and the mutating siblings (start_phone, park_phone) without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete use condition: 'the way to wait for a phone that is still creating or starting', and an implicit exclusion ('It never starts, renews or bills anything'). It stops short of naming the alternative tools (start_phone/act) explicitly, so a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
install_appinstall appInspect
Install an app from Phonebox's app library by package name, in the background. It doesn't reach Google Play: to get a Play Store app, open the Play Store app on the phone with act and install it there. To check on an install later, open the app with act: open_app answers app_not_found until the install has finished. With upload instead of package, it installs your own APK (see create_app_upload) and waits until the phone lists the app, or says why it failed. While that upload is still installing on the phone, and for 10 minutes after that install succeeds, calling install_app again with it starts nothing: a repeat only waits for it, or reports it.
| Name | Required | Description | Default |
|---|---|---|---|
| upload | No | An upload of your own APK whose status is ready, from complete_app_upload. Give this or package. | |
| package | No | The app's Android package name, such as com.whatsapp: an app from Phonebox's app library. Give this or upload. | |
| replace | No | With upload: uninstall the app first when it is on the phone, which deletes its data. Use it for a build signed with another key, or an older version_code. | |
| phone_id | Yes | Phone ID from create_phone or list_phones. |
list_appslist appsRead-onlyIdempotentInspect
List the apps installed on a phone: package, label, version and whether it came with the phone. Open one with act and open_app.
| Name | Required | Description | Default |
|---|---|---|---|
| phone_id | Yes | Phone ID from create_phone or list_phones. |
list_phoneslist phonesRead-onlyIdempotentInspect
List this project's phones, newest first, with their status and when each parks: up to 50 a page, and next_cursor for the next page. Parked phones keep their apps and sign-ins and cost nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | next_cursor from the page before, to read the next page. | |
| status | No | List only the phones in this status. Without it, failed and deleted phones are left out: ask for failed to list those. |
live_view_urllive view urlInspect
Create a link a human can open to watch and control the phone, for example to sign in or solve a check. The link expires (default one hour). Anyone with the link controls the phone until it expires; share it only with your user and never paste it into logs or shared chats.
| Name | Required | Description | Default |
|---|---|---|---|
| phone_id | Yes | Phone ID from create_phone or list_phones. | |
| expires_in | No | Seconds until the link stops working, from 60 to 86400. The default is 3600. |
observeobserveRead-onlyIdempotentInspect
Read the screen: every visible element with a ref number, text and center, plus a screenshot when asked. Tap elements by ref with the act tool.
| Name | Required | Description | Default |
|---|---|---|---|
| phone_id | Yes | Phone ID from create_phone or list_phones. | |
| screenshot | No | Whether to include a screenshot, 720 pixels wide. The default is true. |
park_phonepark phoneIdempotentInspect
Park a phone to stop billing. Do this whenever you're done.
| Name | Required | Description | Default |
|---|---|---|---|
| phone_id | Yes | Phone ID from create_phone or list_phones. |
reboot_phonereboot phoneInspect
Restart a phone that is stuck, keeping everything on it, and wait until it is ready again. A phone can reboot once every 10 minutes; a reboot the phone service didn't confirm is never sent twice: wait for ready with get_phone instead. If it is still creating or starting when this returns, call get_phone with its ID and wait "ready" to keep waiting: it only reads, while start_phone would renew the session.
| Name | Required | Description | Default |
|---|---|---|---|
| phone_id | Yes | Phone ID from create_phone or list_phones. |
start_phonestart phoneInspect
Start a parked phone and wait until it is ready. Its apps, files and sign-ins are kept. On a phone that is already running, it renews the session: a new deadline, with credit reserved for it. If it is still creating or starting when this returns, call get_phone with its ID and wait "ready" to keep waiting: it only reads, while start_phone would renew the session.
| Name | Required | Description | Default |
|---|---|---|---|
| phone_id | Yes | Phone ID from create_phone or list_phones. | |
| idle_timeout | No | Seconds after the phone's last activity, its last_active_at, at which it parks itself, from 60 to 3600. Calls move last_active_at at most once every 15 seconds. Leave it out to keep the phone's current setting. | |
| max_duration | No | The most seconds this session may run before the phone parks, from 60 to 10800. Leave it out to keep the phone's current setting. |
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
- First observed
act - First observed
complete_app_upload - First observed
create_app_upload - First observed
create_phone - First observed
get_phone - First observed
install_app - First observed
list_apps - First observed
list_phones - First observed
live_view_url - First observed
observe - First observed
park_phone - First observed
reboot_phone - First observed
start_phone
Related MCP Connectors
Disposable cloud Android emulators for coding agents: run an APK or PR build, tap, type, screenshot.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
- LimrunOAuthcom.limrun
Cloud iOS simulators and Android emulators your agent can create, drive, and throw away.
The cloud for agents. Tools for AI agents to register, build, and deploy other agents. Zero human required.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI assistants to remotely control Android devices in the cloud via P2P WebRTC, providing tools for screenshots, touch input, file management, and automation scripts.26 npmMIT
- AlicenseCqualityCmaintenanceA lightweight bridge enabling AI agents to perform real-world tasks on Android devices such as app navigation, UI interaction, and automated QA testing without requiring computer-vision pipelines or preprogrammed scripts.142,158 PyPI880MIT
- AlicenseBqualityAmaintenanceGive any LLM agent a real Android or iPhone. 62 MCP tools: tap, swipe, type, screenshot, screen-tree reading, app launch, camera, TTS, crash reports, batched execution. Android via ADB, iPhone via WebDriverAgent, on-device inference, Docker+KVM emulators. Works with Claude Code, Cursor, LangChain, LlamaIndex, and any MCP client. MIT.66377MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to automate Android and iOS devices, including emulators and physical devices, by capturing screenshots, reading UI accessibility trees, performing touch and text input, and managing app lifecycle and logs.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.