Skip to main content
Glama

Phonebox

Server Details

Cloud Android phones for AI agents: create a phone, read the screen, tap, type, install APKs, park.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
combined-ai/phonebox-android-mcp
GitHub Stars
0
Server Listing
Phonebox Android MCP server

TDQS

Score is being calculated.

Available Tools

13 tools
actact
Destructive
Inspect

Run up to 20 actions in order (tap, type, key, swipe, scroll, back, home, open_app, open_url, wait_for, …). The batch stops at the first failure, and the result shows the screen where it stopped, unless the screen couldn't be read, or not in time. Never repeat an action whose outcome is unknown; observe first.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsYesFrom 1 to 20 actions, run in order. Durations are in milliseconds, and the waits in one batch may add up to 60 seconds. Coordinates are device pixels.
observeNoWhat the result shows of the screen afterwards: ui (the element list, the default), screenshot, both, or none. A batch that stops early shows the element list, unless the screen couldn't be read, or not in time.ui
phone_idYesPhone ID from create_phone or list_phones.
complete_app_uploadcomplete app uploadA
Idempotent
Inspect

Finish an upload once its file has arrived, and wait while Phonebox reads and stores the APK. It answers with the upload: ready to install with install_app, or ended with an error saying why. If it is still processing when this returns, call it again with the same arguments: a repeat only waits.

ParametersJSON Schema
NameRequiredDescriptionDefault
upload_idYesThe upload's ID, from create_app_upload.
storage_idYesThe storageId that upload_url answered with, as {"storageId": "…"}.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false/readOnlyHint=false; the description corroborates and enriches this by explaining the blocking wait, the two terminal outcomes (ready vs. ended with an error reason), and that a repeat call merely waits rather than re-uploading. That retry/polling behavior is genuine value beyond the structured fields, though auth or timeout details are not addressed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action, then the possible outcomes, then the retry rule — no wasted preamble. The middle sentence ('It answers with the upload: ready to install with install_app, or ended with an error saying why') is slightly stilted in construction but still earns its place by covering the return states.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully covers the return states and the polling contract, and annotations carry the safety profile. For a two-parameter mutation tool this is nearly complete; only timeout/retry-limit expectations are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both upload_id and storage_id documented including their formats and provenance, so the schema carries the parameter burden. The description adds only the implicit point that retries must reuse the same arguments; it offers no additional syntax or format guidance, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Finish an upload'), names the precondition (file has arrived), and the side effect (Phonebox reads and stores the APK). It routes the agent to the correct sibling at each stage — create_app_upload before, install_app after — so it is clearly distinguishable from the other upload/install tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: call after the file has arrived, and call again with the same arguments if processing is still in progress. It names install_app as the follow-on step. It stops short of stating any when-not-to-use case or a maximum retry/backoff policy, so it is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_app_uploadcreate app uploadAInspect

Start an upload of your own APK, such as the debug build you just made, to install it on a phone. MCP can't carry a file, so this returns a one-off upload_url and the exact curl command to send the APK there from your shell; then call complete_app_upload with the storageId it answers. Send the URL one file of up to 500 MiB, within an hour.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing the return payload (one-off upload_url plus an exact curl command), the handoff contract (call complete_app_upload with the storageId), and hard limits (one file, up to 500 MiB, URL valid for one hour). These are exactly the operational constraints an agent must know for a non-idempotent, open-world upload tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the action and rationale, then the mechanics, then the limits. Every clause carries information an agent needs; nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by describing the return value and the follow-up call. For a zero-parameter tool whose real payload moves through a shell command, the description covers purpose, limits, and completion path completely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero parameters, so there is nothing for the description to disambiguate; baseline is 4. The description correctly implies no arguments are needed and instead explains the out-of-band data flow.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start an upload of your own APK') with immediate scope ('to install it on a phone'). It distinguishes itself from the sibling complete_app_upload by naming it as the required follow-up step, so an agent can place both tools in the workflow without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use ('your own APK, such as the debug build you just made') and explicitly routes to complete_app_upload with the returned storageId. There is no explicit when-not condition or mention of the alternative install path (install_app), but the trigger and the sequencing are unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_phonecreate phoneAInspect

Create a cloud Android phone and wait until it is ready (usually under two minutes). It costs $0.06 a minute while running. Always park it when you're done. If it is still creating or starting when this returns, call get_phone with its ID and wait "ready" to keep waiting: it only reads, while start_phone would renew the session.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoA name to tell the phone apart, up to 80 characters. The default is its ID.
countryNoWhere the phone appears to be: one of the listed two-letter ISO 3166-1 country codes, such as US or DE. Any other code is refused before a phone is made. Leave it out for any country.
metadataNoUp to 20 string pairs of your own, such as a customer ID. Keys have up to 40 characters and values up to 256.
idle_timeoutNoSeconds after the phone's last activity, its last_active_at, at which it parks itself, from 60 to 3600. Calls move last_active_at at most once every 15 seconds. The default is 300.
max_durationNoThe most seconds this session may run before the phone parks, from 60 to 10800. The default is up to 900, shortened automatically to available credit and the spending limit.
idempotency_keyNoA unique key for this phone. If a call times out, repeat it with the same key and settings to get the same phone instead of a second one. If that phone has failed, create the next one with a new key.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark this as a non-read-only, open-world, non-destructive write; the description adds the material behavior an agent cannot infer: provisioning takes under two minutes, it bills $0.06 per minute while running, and it may return before the phone is ready. That cost and timing disclosure is genuinely decision-relevant for a spend-incurring tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with what the tool does and the cost, then the operational rules. The final sentence about get_phone versus start_phone is slightly dense and requires re-reading, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the necessary work of explaining the return state (may still be creating or starting) and the follow-up call with the returned ID. It stops short of describing the full returned object, but for a 6-param, all-optional tool with full schema coverage this is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so name, country, metadata, idle_timeout, max_duration, and idempotency_key are all already documented in the schema, and the description adds no parameter-level detail beyond the returned ID. Baseline 3 is appropriate when the schema carries the full parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Create a cloud Android phone") plus the completion semantics ("wait until it is ready"), which immediately separates it from siblings like start_phone, park_phone, and get_phone. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit lifecycle guidance: always park when done, and if the phone is still creating/starting, use get_phone with its ID to keep waiting rather than start_phone, which would renew the session. It names the alternative tool and the condition that selects it, which is exactly the routing an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_phoneget phoneA
Read-onlyIdempotent
Inspect

Read one phone: its status, session and when it parks. With wait, wait until it is ready or parked. It never starts, renews or bills anything, so it is the way to wait for a phone that is still creating or starting.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNo"ready" or "parked": wait until the phone gets there. Leave it out to read the phone as it is.
phone_idYesPhone ID from create_phone or list_phones.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds real behavioral value on top: it explains that the call can block when 'wait' is set, and reassures that it never starts, renews or bills. It does not say what happens if the wait never reaches the target state (timeout/error behavior), which is the one notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the core read action, then the wait behavior, then the negative scope clause. No filler and every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the right thing by enumerating the returned fields (status, session, park time). Annotations carry the safety semantics. The only omission is the timeout/failure behavior of a 'wait' call that never reaches ready or parked.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the enum parameter is already documented ('ready' or 'parked': wait until the phone gets there). The description's 'With wait, wait until it is ready or parked' restates the schema rather than adding format or edge-case meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Read one phone' plus what it returns (status, session, park time). This distinguishes it from list_phones and the mutating siblings (start_phone, park_phone) without needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete use condition: 'the way to wait for a phone that is still creating or starting', and an implicit exclusion ('It never starts, renews or bills anything'). It stops short of naming the alternative tools (start_phone/act) explicitly, so a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_appinstall appInspect

Install an app from Phonebox's app library by package name, in the background. It doesn't reach Google Play: to get a Play Store app, open the Play Store app on the phone with act and install it there. To check on an install later, open the app with act: open_app answers app_not_found until the install has finished. With upload instead of package, it installs your own APK (see create_app_upload) and waits until the phone lists the app, or says why it failed. While that upload is still installing on the phone, and for 10 minutes after that install succeeds, calling install_app again with it starts nothing: a repeat only waits for it, or reports it.

ParametersJSON Schema
NameRequiredDescriptionDefault
uploadNoAn upload of your own APK whose status is ready, from complete_app_upload. Give this or package.
packageNoThe app's Android package name, such as com.whatsapp: an app from Phonebox's app library. Give this or upload.
replaceNoWith upload: uninstall the app first when it is on the phone, which deletes its data. Use it for a build signed with another key, or an older version_code.
phone_idYesPhone ID from create_phone or list_phones.
list_appslist apps
Read-onlyIdempotent
Inspect

List the apps installed on a phone: package, label, version and whether it came with the phone. Open one with act and open_app.

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_idYesPhone ID from create_phone or list_phones.
list_phoneslist phones
Read-onlyIdempotent
Inspect

List this project's phones, newest first, with their status and when each parks: up to 50 a page, and next_cursor for the next page. Parked phones keep their apps and sign-ins and cost nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNonext_cursor from the page before, to read the next page.
statusNoList only the phones in this status. Without it, failed and deleted phones are left out: ask for failed to list those.
live_view_urllive view urlInspect

Create a link a human can open to watch and control the phone, for example to sign in or solve a check. The link expires (default one hour). Anyone with the link controls the phone until it expires; share it only with your user and never paste it into logs or shared chats.

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_idYesPhone ID from create_phone or list_phones.
expires_inNoSeconds until the link stops working, from 60 to 86400. The default is 3600.
observeobserve
Read-onlyIdempotent
Inspect

Read the screen: every visible element with a ref number, text and center, plus a screenshot when asked. Tap elements by ref with the act tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_idYesPhone ID from create_phone or list_phones.
screenshotNoWhether to include a screenshot, 720 pixels wide. The default is true.
park_phonepark phone
Idempotent
Inspect

Park a phone to stop billing. Do this whenever you're done.

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_idYesPhone ID from create_phone or list_phones.
reboot_phonereboot phoneInspect

Restart a phone that is stuck, keeping everything on it, and wait until it is ready again. A phone can reboot once every 10 minutes; a reboot the phone service didn't confirm is never sent twice: wait for ready with get_phone instead. If it is still creating or starting when this returns, call get_phone with its ID and wait "ready" to keep waiting: it only reads, while start_phone would renew the session.

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_idYesPhone ID from create_phone or list_phones.
start_phonestart phoneInspect

Start a parked phone and wait until it is ready. Its apps, files and sign-ins are kept. On a phone that is already running, it renews the session: a new deadline, with credit reserved for it. If it is still creating or starting when this returns, call get_phone with its ID and wait "ready" to keep waiting: it only reads, while start_phone would renew the session.

ParametersJSON Schema
NameRequiredDescriptionDefault
phone_idYesPhone ID from create_phone or list_phones.
idle_timeoutNoSeconds after the phone's last activity, its last_active_at, at which it parks itself, from 60 to 3600. Calls move last_active_at at most once every 15 seconds. Leave it out to keep the phone's current setting.
max_durationNoThe most seconds this session may run before the phone parks, from 60 to 10800. Leave it out to keep the phone's current setting.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updates
    • First observedact
    • First observedcomplete_app_upload
    • First observedcreate_app_upload
    • First observedcreate_phone
    • First observedget_phone
    • First observedinstall_app
    • First observedlist_apps
    • First observedlist_phones
    • First observedlive_view_url
    • First observedobserve
    • First observedpark_phone
    • First observedreboot_phone
    • First observedstart_phone

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    C
    maintenance
    A lightweight bridge enabling AI agents to perform real-world tasks on Android devices such as app navigation, UI interaction, and automated QA testing without requiring computer-vision pipelines or preprogrammed scripts.
    14
    2,158 PyPI
    880
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Give any LLM agent a real Android or iPhone. 62 MCP tools: tap, swipe, type, screenshot, screen-tree reading, app launch, camera, TTS, crash reports, batched execution. Android via ADB, iPhone via WebDriverAgent, on-device inference, Docker+KVM emulators. Works with Claude Code, Cursor, LangChain, LlamaIndex, and any MCP client. MIT.
    66
    377
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to automate Android and iOS devices, including emulators and physical devices, by capturing screenshots, reading UI accessibility trees, performing touch and text input, and managing app lifecycle and logs.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.