MobileRun MCP Server
Enables AI agents to drive real Android devices through MobileRun, including reading the screen and UI state, tapping, swiping, scrolling, typing, pressing Android system keys, launching apps, and capturing device-side recordings of actions.
Enables AI agents to drive real iOS devices (iPhones) through MobileRun, including reading the screen and UI state, tapping, swiping, scrolling, typing, pressing iPhone system keys, launching apps, and capturing device-side recordings of actions.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MobileRun MCP Serveropen Settings, turn on Wi-Fi, then send me the recording of it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MobileRun for AgentEnv
AI agents drive real phones. This plugin turns a MobileRun cloud phone, a real iPhone or Android device in a data centre, into an environment an agent drives through eighteen MCP tools: it reads the screen, taps, swipes, types and opens apps, and the device itself records video and a timed trajectory of every action. Environments are not limited to services; any interface an agent can act on can be one, and a phone is about as real an interface as there is. This repository is an environment plugin for the AgentEnv Framework, Scale AI's open-source framework for building RL environments.

Claude Sonnet 5.5, given one sentence, drives a real iPhone 15 Pro on MobileRun to walking directions from the
Brandenburg Gate to Museum Island: it opens Maps from the home screen, searches, notices its query was cut to "Mus",
clears and retypes it, types the start point by hand (Location Services is off on the phone) and switches to walking.
23 tool calls in 2 minutes 23 seconds, shown at 4×. On the right, what the model wrote before each action and every
mobilerun_* call with its result; on the left, the device's own server-side recording, with each tap ringed where and
when the device registered it. The run ends on an end-state check read from the phone's ui_state after the model said
it was done: Museum Island on screen, walking mode, a duration.
Contents: Run it yourself · Built on the AgentEnv Framework · The environment's tools · Recording · What has been verified · Known limits · Troubleshooting · Repository layout · Contributing · Development
Run it yourself
You need Docker, Python 3.11 or newer, a MobileRun account with an API key (dr_sk_..., from their API Keys page) and
a device on it. Android and bring-your-own devices are self-serve; a rented iPhone currently means contacting
MobileRun's team (pricing).
pip install agentenv-framework
agent-env plugin add 'agentenv-mobilerun @ git+https://github.com/kaihackney-scale/agentenv-mobilerun'
export MOBILERUN_API_KEY=dr_sk_...
agent-env mobilerun doctor --skip-agent --platform ios # key, devices, capabilities, a screenshot health probe
agent-env mobilerun setup --id mobilerun --platform ios # build the MCP server image, register the env
agent-env env deploy --id mobilerun # resolve an idle phone and deploy itThat is the whole installation: the env type, the task step, the bundle and the agent-env mobilerun commands
register themselves through entry points, so there is nothing to add to .agentenv/config.toml. setup builds the
image locally and registers it in agent-env's own image store, so there is no registry to pull from and nothing to
docker login to. It builds linux/amd64 by default, for remote sandbox VMs; on Apple Silicon with the local sandbox,
--build-platform linux/arm64 builds a native image instead of an emulated one. The deploy ends like this:
Deployed!
Instance ID: mobilerun-...
Env MCP Url: http://127.0.0.1:62694/mcpHand that MCP URL to an agent, or to any MCP client, and it has a phone. setup --device ID pins one device; without
it each deploy resolves an idle one that passes its health probe. Use the same --platform for doctor and setup:
doctor defaults to Android, and checks for a device of that platform.
The bundled task
The package also ships a bundle, mobilerun: deploy the env, hand the phone to an agent, ask it to open Settings and
the display screen, and collect the recording. The recording only exists if the env was set up with --record;
without it the collect step finds nothing and says so.
agent-env mobilerun setup --id mobilerun --platform ios --record
agent-env run # lists the installed bundles; mobilerun is one
agent-env mobilerun doctor --platform ios # this time without --skip-agent
agent-env run mobilerunIt needs an A2A agent, and agent-env ships none. The task's deploy_agent step names no agent, so it deploys
your configured default: [agents] default_a2a_agent_id in .agentenv/config.toml, else a2a-default, which on a
clean install does not exist. Point the setting at any A2A agent that drives MCP tools, or register one under that id
with agent-env a2a-agent put --id a2a-default --dockerfile path/to/your/agent/Dockerfile; the protocol package's
a2a_agent SDK (pip install 'agentenv-framework-protocol[agent]') is one way to build it. The task names no agent so
that your default wins, which means agent-env's own run preflight cannot see a missing one; doctor resolves it the
way deploy_agent does and fails first, before any phone or gateway is spent.
Related MCP server: mobile-mcp
Built on the AgentEnv Framework
This plugin is built on the AgentEnv Framework
(GitHub, pip install agentenv-framework): Scale AI's open-source
framework for building RL environments, with composable environments behind an MCP gateway, any agent in any
sandbox, and tasks as DAGs. The framework does the heavy lifting; this repository adds the phone.
┌────────────────────────────┐ ┌──────────────────────┐ ┌───────────┐
│ agent-env gateway │ │ api.mobilerun.ai │ │ a real │
│ ┌──────────────────────┐ │ HTTPS │ │ │ phone │
│ │ mobilerun MCP server │──┼───────▶│ device plane (REST) │───────▶│ │
│ └──────────────────────┘ │ └──────────────────────┘ └───────────┘
└────────────────────────────┘One container. No host machine, no USB, no WebDriverAgent, no device pool: a cloud phone needs outbound HTTPS and nothing else. Each piece maps to a framework concept:
AgentEnv concept | Here |
Environment: MCP tools and an env card in one container |
|
Plugin: a pip package with entry points |
|
| |
| |
Agents: A2A agents handed the env's MCP server | none shipped: any A2A agent that drives MCP tools; the task deploys your default |
Registry: versioned images, envs and runs |
|
Start with the framework's getting started and core concepts to build an environment of your own.
The environment's tools
Eighteen tools, capability-gated per device: the env reads the phone's capability map at deploy time and registers only what that phone can honour, so an agent never sees a tool that is guaranteed to fail. Two phones on one account can legitimately differ. Every coordinate is in screenshot pixels.
Tool | What it does |
| Capture the screen; coordinates read off it are what the gesture tools take, unscaled |
| The screenshot's pixel size, the device's point size and the ratio between them |
| The labelled on-screen elements with tappable centres, the app and keyboard state; compacted from the raw tree |
| Sleep, then let the caller take a fresh screenshot |
| Block until the device plane calls the phone ready |
| Tap a point |
| Tap a point twice |
| Press and hold a point |
| Drag from one point to another |
| Scroll by swiping across the middle of the screen |
| Type into the focused field, waiting for the text to land |
| Clear the focused field |
| A system key by name: |
| The installed apps |
| Bring an app to the foreground by package name (Android) or bundle id (iOS) |
| Open a URL or deep link directly |
| Read the clipboard |
| Write the clipboard |
The coordinate contract
You work in screenshot pixels. Read a coordinate off the image and pass it in unchanged; the conversion happens server-side, once. It is not cosmetic. On a live iPhone 15 Pro:
| 1179 × 2556 physical pixels |
| 393 × 852 logical points |
ratio | exactly 3.0 |
The device plane taking points rather than pixels is undocumented (the schema is a bare integer pair), so it was
established from a recorded trajectory of taps sent through that same endpoint: they land at x=196, the centre of a
393-point screen, where the pixel centre would be 589. Passing pixels straight through misses by 3×, which puts most
of the screen out of bounds. The plugin resolves the ratio once per device and routes every gesture through one
conversion; every gesture result echoes both spaces, e.g. {"x": 589, "y": 1278, "device_point": [196, 426]}, so a
trace shows what was aimed at and what was sent. If the ratio cannot be established, coordinates pass through
unchanged and the fact is logged loudly.
ui_state is compacted, deliberately
The raw accessibility tree is not usable as a tool result: on a live iPhone it is 183 KB, about 46,000 tokens across
412 nodes, of which 85 carry a label. mobilerun_ui_state returns the labelled elements with their centres, the
notable flags and the app and keyboard state, about 1,700 tokens, and declares elements_total and truncated so a
caller can see what was filtered. isClickable is false on every node of a real iPhone tree (it is an Android
field), so it is neither reported nor used as a filter; filtering on it would drop the whole screen.
Recording
MobileRun records server-side, which is strictly better than anything a client can do. Opt in at setup:
agent-env mobilerun setup --id mobilerun --recordEach deploy then starts a video and trajectory recording. Add the collect step after the step that prompts your agent:
{ "id": "collect", "type": "mobilerun_collect_recording", "key_prefix": "recordings" }The trajectory carries exact coordinates, seq, at_ms, gesture durations, the display scale and rotation, a
valid/reasons contract, and the part a client-side recorder cannot know: video.timeline.action_zero_to_video_ms
with an uncertainty_ms, the action-clock to video-clock offset measured by the device that captured both. That is
how the gif above was cut: each tap's at_ms says where it is in the video. The step copies both artifacts into
your configured object store, because the video route answers a 302 to a presigned URL that expires in 900
seconds, and each recording has a retentionDays after which MobileRun returns 410 Gone.
⚠️ A recording captures whatever is on screen, not just the task: notifications, whatever app was open, any signed-in account, and whatever screen the last user left an app on. Settings reopens where it was left, which can be General > About with the phone's serial number and IMEI. Recording is off by default, nothing here publishes anywhere, and anything you publish from a recording should be checked frame by frame.
What has been verified
Live, 2026-10-05, on rented iPhone 15 Pros, published agentenv-framework 0.9.1267 and the local sandbox:
setup, env deploy, the env card the gateway composes (with this server's own card as its mobilerun child), all
18 tools listed through the gateway, screenshot, ui_state, tap, scroll, type_text (plain and with clear)
and press_key home executing and observed on the device, the server-side recording and its trajectory, teardown
through the env's reattach and close(), and a model (Claude Sonnet 5.5, by function calling against the env's MCP
tools) completing a 23-step task end to end: the gif above.
Earlier, on 2026-09-21: device listing and resolution, the capability map, the health probe, doctor, the
pixel-to-point conversion, screenshots, ui_state compaction, list_apps and the full recording download path.
Not verified yet: anything on Android (the account's only Android device is in maintenance), so treat the
Android path, including press_key's Android codes, as unexercised; the bundle's agent-env run mobilerun, which goes
through an A2A agent rather than a model calling the tools directly; and clear_text, double_tap, long_press,
open_deep_link and the clipboard tools on a device. See
Known limits for launch_app.
Known limits
Stated rather than papered over.
No exclusivity. MobileRun's device plane has no lease or reservation primitive, and neither does this plugin. Two deploys against one account can resolve the same phone and both drive it. Pin distinct devices with
--device, or run one job per device.stateis not health. A phone can reportstate: "ready"while being completely undrivable. Everything here health-probes with a screenshot instead, anddevices --probeexposes the same check.launch_appon iOS needs care. On one iPhone, launching Settings by bundle id returnedok, Settings never opened, and the phone stopped answering screenshots and gestures while still reportingready. Tapping the app's icon instead worked on another phone. One case, so not a rule, but prefer the icon on iOS until it is understood.Location Services is off on MobileRun's phones. Maps and anything else that needs "my location" asks for a start point instead, so give tasks explicit places.
A badged icon is labelled by its badge. In an iPhone's accessibility tree, the Settings icon with a red badge is labelled
1 new item, notSettings. Find icons by what the screenshot shows, not by label alone.press_keycodes are Android's.POST /devices/{id}/globaltakes a bare integer and documents no enum. On an iPhone,homeis verified andbackis expected; useagent-env mobilerun probe-globalto check the rest.ui_stateis declared asAndroidStatein MobileRun's OpenAPI spec even on an iPhone, where the contents are XCUIElement-shaped. Don't write a parser that trusts the declared schema.long_pressanddouble_tapare composed from the tap and swipe primitives; the device plane has neither.No factory reset in the run loop.
Env.reset()presses home and nothing else. MobileRun'sPOST /devices/{id}/resetclears installed apps and user data; it is reachable asMobileRunClient.factory_resetfor an operator who means it, and is never wired into a task.
Troubleshooting
no MobileRun API key found: export MOBILERUN_API_KEY, or add it to the secret store your config.toml
configures under that name. An exported but empty variable counts as absent.
device X reports state='ready' but failed its screenshot health probe: the phone is degraded below anything this
plugin controls. Reboot it from the MobileRun dashboard; if that does not clear it (a reboot can make things worse),
it is a MobileRun support item.
no idle ready device: the message lists every device and its state. A device with an active task is not idle.
A tool is missing from the agent's list: the phone's capability map says it is unsupported. agent-env mobilerun doctor prints the supported and unsupported sets.
agent-env run mobilerun fails at deploy_agent: there is no default agent. See
The bundled task; agent-env mobilerun doctor reports it.
Commands
Command | What it does |
| The four checks that explain every failure (key, devices, capabilities, health probe) and the agent |
| List devices; |
| Build the MCP image and register the env ( |
| Discover what the undocumented |
Repository layout
src/agentenv_mobilerun/
env.py MobileRunEnv: resolve and health-probe a phone, deploy the MCP server behind a gateway
client.py the device-plane client, on MobileRun's official mobilerun-sdk
server/ the MCP server image: main.py (tools, env card), Dockerfile, requirements.txt
steps/ the mobilerun_collect_recording task step
cli.py agent-env mobilerun: doctor, devices, setup, probe-global
capabilities.py the capability map's wire encoding, shared by the env and the server
api_key.py key resolution, free of agent-env so doctor runs without it
bundles/mobilerun/ the open-settings task and the smoke eval
examples/ the same task as a standalone JSON, with notes on supplying an agent
tests/ client, server, env, CLI, recording and env-card tests; no device or network needed
docs/media/ the demo gifContributing
Contributions are welcome: new tools, Android verification, tasks and docs. Open an issue to discuss a larger change first, then send a pull request; CI must pass, and a maintainer reviews every pull request before it merges.
Development
uv venv && uv pip install -e '.[dev]'
.venv/bin/pytest # 177 tests
.venv/bin/ruff check --select F821 src tests[dev] installs agentenv-framework on Python 3.11 and newer, so the whole suite runs. On 3.10 the half the container
ships (the client and the MCP server) still runs and the framework-dependent modules skip themselves. The image
installs neither extra (see src/agentenv_mobilerun/server/requirements.txt): keeping the framework out of it is what
lets the container and the env class be upgraded independently. CI (.github/workflows/tests.yml) runs the suite on
Python 3.10 and 3.12 and gates undefined names.
Tests drive the client through httpx.MockTransport, so request construction (URLs, bodies, redirect following) is
exercised for real with nothing stubbed above the socket, and the deploy tests drive the framework's real gateway
provider class with only its network call replaced. No device, API key or network access is needed. Every async
test carries @pytest.mark.asyncio, because asyncio_mode is strict.
Licence and credits
This repository is licensed under the Apache License 2.0 (LICENSE).
AgentEnv Framework (scaleapi/agentenv-framework, Apache-2.0) runs the env, the gateway, the tasks and the registry this plugin plugs into;
agentenv-framework-protocol(Apache-2.0) serves the env card.MobileRun provides the phones; the client is built on its official
mobilerun-sdk(Apache-2.0).The MCP server uses the MCP Python SDK (MIT), httpx (BSD-3-Clause) and Pillow (MIT-CMU).
This server cannot be deployed
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
- LimrunOAuthcom.limrun
Cloud iOS simulators and Android emulators your agent can create, drive, and throw away.
Drive real Android & iOS devices and web browsers from natural language for mobile + web QA. 290+ tools across device control, app management, automation sessions, browser automation, and flow recording / replay. Bearer-auth — get a token at robotactions.com → Profile → API Tokens.
Drive real devices from your AI Coding tool. Embed a client SDK (Unity, Godot, Flutter, iOS/macOS, Android, React Native, Web) in your app, then capture screenshots, traverse the UI tree, inject taps and key events, and run automated test tasks on the physical device over a secure relay.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables AI agents to drive an iPhone or iOS Simulator through 30 tools and 4 resources, using an observe/act/verify loop with built-in safety approvals.311MIT
- AlicenseAqualityCmaintenanceEnables agents and LLMs to interact with native iOS and Android applications and devices via structured accessibility snapshots or coordinate-based taps, supporting simulators, emulators, and real devices.3968,067 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceGives AI agents uniform perception and control over Android, iOS, and mock mobile devices through MCP, including taps, typing, app navigation, deep links, SMS/OTP retrieval, and human-in-the-loop approvals.Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to automate Android and iOS devices, including emulators and physical devices, by capturing screenshots, reading UI accessibility trees, performing touch and text input, and managing app lifecycle and logs.MIT