Mobile Next MCP Server
OfficialMobile Next - MCP server for Mobile Development and Automation | iOS, Android, Simulator, Emulator, and Real Devices
This is a Model Context Protocol (MCP) server that enables scalable mobile automation, development through a platform-agnostic interface, eliminating the need for distinct iOS or Android knowledge. You can run it on emulators, simulators, and real devices (iOS and Android). This server allows Agents and LLMs to interact with native iOS/Android applications and devices through structured accessibility snapshots or coordinate-based taps based on screenshots.
https://github.com/user-attachments/assets/bb084777-beb3-4930-ae6f-8d3fe694ddde
🚀 Mobile MCP Roadmap: Building the Future of Mobile
Join us on our journey as we continuously enhance Mobile MCP! Check out our detailed roadmap to see upcoming features, improvements, and milestones. Your feedback is invaluable in shaping the future of mobile automation.
Main use cases
How we help to scale mobile automation:
📲 Native app automation (iOS and Android) for testing or data-entry scenarios.
📝 Scripted flows and form interactions without manually controlling simulators/emulators or real devices (iPhone, Samsung, Google Pixel etc)
🧭 Automating multi-step user journeys driven by an LLM
👆 General-purpose mobile application interaction for agent-based frameworks
🤖 Enables agent-to-agent communication for mobile automation usecases, data extraction
Main Features
🚀 Fast and lightweight: Uses native accessibility trees for most interactions, or screenshot based coordinates where a11y labels are not available.
🤖 LLM-friendly: No computer vision model required in Accessibility (Snapshot).
🧿 Visual Sense: Evaluates and analyses what's actually rendered on screen to decide the next action. If accessibility data or view-hierarchy coordinates are unavailable, it falls back to screenshot-based analysis.
📊 Deterministic tool application: Reduces ambiguity found in purely screenshot-based approaches by relying on structured data whenever possible.
📺 Extract structured data: Enables you to extract structred data from anything visible on screen.
🎯 Platform Support
Platform | Supported |
iOS Real Device | ✅ |
iOS Simulator | ✅ |
Android Real Device | ✅ |
Android Emulator | ✅ |
Related MCP server: Mobile Next MCP
🔧 Available MCP Tools
For detailed implementation and parameter specifications, see
src/server.ts
Device Management
mobile_list_available_devices- List all available devices (simulators, emulators, and real devices)mobile_get_screen_size- Get the screen size of the mobile device in pixelsmobile_get_orientation- Get the current screen orientation of the devicemobile_set_orientation- Change the screen orientation (portrait/landscape)
App Management
mobile_list_apps- List all installed apps on the devicemobile_launch_app- Launch an app using its package namemobile_terminate_app- Stop and terminate a running appmobile_install_app- Install an app from file (.apk, .ipa, .app, .zip)mobile_uninstall_app- Uninstall an app using bundle ID or package name
Screen Interaction
mobile_take_screenshot- Take a screenshot to understand what's on screenmobile_save_screenshot- Save a screenshot to a filemobile_list_elements_on_screen- List UI elements with their coordinates and propertiesmobile_click_on_screen_at_coordinates- Click at specific x,y coordinatesmobile_double_tap_on_screen- Double-tap at specific coordinatesmobile_long_press_on_screen_at_coordinates- Long press at specific coordinatesmobile_swipe_on_screen- Swipe in any direction (up, down, left, right)
Input & Navigation
mobile_type_keys- Type text into focused elements with optional submitmobile_press_button- Press device buttons (HOME, BACK, VOLUME_UP/DOWN, ENTER, etc.)mobile_open_url- Open URLs in the device browser
Platform Support
iOS: Simulators and real devices via native accessibility and WebDriverAgent
Android: Emulators and real devices via ADB and UI Automator
Cross-platform: Unified API works across both iOS and Android
🏗️ Mobile MCP Architecture
📚 Wiki page
More details in our wiki page for setup, configuration and debugging related questions.
Installation and configuration
Standard config works in most of the tools:
{
"mcpServers": {
"mobile-mcp": {
"command": "npx",
"args": ["-y", "@mobilenext/mobile-mcp@latest"]
}
}
}Add via the Amp VS Code extension settings screen or by updating your settings.json file:
"amp.mcpServers": {
"mobile-mcp": {
"command": "npx",
"args": [
"@mobilenext/mobile-mcp@latest"
]
}
}Amp CLI:
Run the following command in your terminal:
amp mcp add mobile-mcp -- npx @mobilenext/mobile-mcp@latestTo setup Cline, just add the json above to your MCP settings file.
Use the Claude Code CLI to add the Mobile MCP server:
claude mcp add mobile-mcp -- npx -y @mobilenext/mobile-mcp@latestFollow the MCP install guide, use json configuration above.
Use the Codex CLI to add the Mobile MCP server:
codex mcp add mobile-mcp npx "@mobilenext/mobile-mcp@latest"Alternatively, create or edit the configuration file ~/.codex/config.toml and add:
[mcp_servers.mobile-mcp]
command = "npx"
args = ["@mobilenext/mobile-mcp@latest"]For more information, see the Codex MCP documentation.
Use the Copilot CLI to interactively add the Mobile MCP server:
/mcp addYou can edit the configuration file ~/.copilot/mcp-config.json and add:
{
"mcpServers": {
"mobile-mcp": {
"type": "local",
"command": "npx",
"tools": [
"*"
],
"args": [
"@mobilenext/mobile-mcp@latest"
]
}
}
}For more information, see the Copilot CLI documentation.
Click the button to install:
Or install manually:
Go to Cursor Settings -> MCP -> Add new MCP Server. Name to your liking, use command type with the command npx -y @mobilenext/mobile-mcp@latest. You can also verify config or add command like arguments via clicking Edit.
Use the Gemini CLI to add the Mobile MCP server:
gemini mcp add mobile-mcp npx -y @mobilenext/mobile-mcp@latestClick the button to install:
Or install manually:
Go to Advanced settings -> Extensions -> Add custom extension. Name to your liking, use type STDIO, and set the command to npx -y @mobilenext/mobile-mcp@latest. Click "Add Extension".
Follow the MCP Servers documentation. For example in .kiro/settings/mcp.json:
{
"mcpServers": {
"mobile-mcp": {
"command": "npx",
"args": [
"@mobilenext/mobile-mcp@latest"
]
}
}
}Follow the MCP Servers documentation. For example in ~/.config/opencode/opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"mobile-mcp": {
"type": "local",
"command": [
"npx",
"@mobilenext/mobile-mcp@latest"
],
"enabled": true
}
}
}Open Qodo Gen chat panel in VSCode or IntelliJ → Connect more tools → + Add new MCP → Paste the standard config above.
Click Save.
Open Windsurf settings, navigate to MCP servers, and add a new server using the command type with:
npx @mobilenext/mobile-mcp@latestOr add the standard config under mcpServers in your settings as shown above.
SSE Server Mode
By default, Mobile MCP runs over stdio. To start an SSE server instead, use the --listen flag:
npx @mobilenext/mobile-mcp@latest --listen 3000This binds to localhost:3000. To bind to a specific interface:
npx @mobilenext/mobile-mcp@latest --listen 0.0.0.0:3000Then configure your MCP client to connect to http://<host>:3000/mcp.
Authorization
To require Bearer token authorization on the SSE server, set the MOBILEMCP_AUTH environment variable:
MOBILEMCP_AUTH=my-secret-token npx @mobilenext/mobile-mcp@latest --listen 3000When set, all requests must include the header Authorization: Bearer my-secret-token.
🛠️ How to Use 📝
After adding the MCP server to your IDE/Client, you can instruct your AI assistant to use the available tools. For example, in Cursor's agent mode, you could use the prompts below to quickly validate, test and iterate on UI intereactions, read information from screen, go through complex workflows. Be descriptive, straight to the point.
✨ Example Prompts
Workflows
You can specifiy detailed workflows in a single prompt, verify business logic, setup automations. You can go crazy:
Search for a video, comment, like and share it.
Find the video called " Beginner Recipe for Tonkotsu Ramen" by Way of
Ramen, click on like video, after liking write a comment " this was
delicious, will make it next Friday", share the video with the first
contact in your whatsapp list.Download a successful step counter app, register, setup workout and 5-star the app
Find and Download a free "Pomodoro" app that has more than 1k stars.
Launch the app, register with my email, after registration find how to
start a pomodoro timer. When the pomodoro timer started, go back to the
app store and rate the app 5 stars, and leave a comment how useful the
app is.Search in Substack, read, highlight, comment and save an article
Open Substack website, search for "Latest trends in AI automation 2025",
open the first article, highlight the section titled "Emerging AI trends",
and save article to reading list for later review, comment a random
paragraph summary.Reserve a workout class, set timer
Open ClassPass, search for yoga classes tomorrow morning within 2 miles,
book the highest-rated class at 7 AM, confirm reservation,
setup a timer for the booked slot in the phoneFind a local event, setup calendar event
Open Eventbrite, search for AI startup meetup events happening this
weekend in "Austin, TX", select the most popular one, register and RSVP
yes to the event, setup a calendar event as a reminder.Check weather forecast and send a Whatsapp/Telegram/Slack message
Open Weather app, check tomorrow's weather forecast for "Berlin", and
send the summary via Whatsapp/Telegram/Slack to contact "Lauren Trown",
thumbs up their response.Schedule a meeting in Zoom and share invite via email
Open Zoom app, schedule a meeting titled "AI Hackathon" for tomorrow at
10AM with a duration of 1 hour, copy the invitation link, and send it via
Gmail to contacts "team@example.com".More prompt examples can be found here.
Prerequisites
What you will need to connect MCP with your agent and mobile devices:
node.js v22+
MCP supported foundational models or agents, like Claude MCP, OpenAI Agent SDK, Copilot Studio
Simulators, Emulators, and Real Devices
When launched, Mobile MCP can connect to:
iOS Simulators on macOS/Linux
Android Emulators on Linux/Windows/macOS
iOS or Android real devices (requires proper platform tools and drivers)
Make sure you have your mobile platform SDKs (Xcode, Android SDK) installed and configured properly before running Mobile Next Mobile MCP.
Telemetry
Mobile MCP collects anonymous usage telemetry via PostHog. To disable it, set the MOBILEMCP_DISABLE_TELEMETRY environment variable:
MOBILEMCP_DISABLE_TELEMETRY=1 npx @mobilenext/mobile-mcp@latestFor json configurations:
{
"mcpServers": {
"mobile-mcp": {
"command": "npx",
"args": ["-y", "@mobilenext/mobile-mcp@latest"],
"env": {
"MOBILEMCP_DISABLE_TELEMETRY": "1"
}
}
}
}Running in "headless" mode on Simulators/Emulators
When you do not have a real device connected to your machine, you can run Mobile MCP with an emulator or simulator in the background.
For example, on Android:
Start an emulator (avdmanager / emulator command).
Run Mobile MCP with the desired flags
On iOS, you'll need Xcode and to run the Simulator before using Mobile MCP with that simulator instance.
xcrun simctl listxcrun simctl boot "iPhone 16"
Thanks to all contributors ❤️
We appreciate everyone who has helped improve this project.
Available Tools
33 toolsmobile_allocate_remote_deviceAllocate Remote DeviceA
Reserve a physical device from the remote cloud fleet for exclusive use, returning a device identifier usable with the other mobile_* tools. Unlike local devices, a remote device is a shared and billed resource borrowed for the session - only call this after the user has explicitly asked to use a remote/cloud device, never speculatively or as a fallback when a local device isn't found. Requires mobile_login_to_cloud_provider to have been called first; if this fails with an authentication error, call that tool then retry. Use mobile_list_remote_devices first to see which names and versions actually exist in the fleet before filtering by them. Release the device with mobile_release_remote_device once the whole task is finished - releasing wipes the device's state, so do not release and reallocate between steps of the same task just to be tidy.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Filter by device name/model. Supports a trailing * for prefix match (e.g. "iPhone*"), or an exact name (e.g. "iPhone 16"). | |
| type | No | Device type filter. Currently only "real" (physical devices) is supported by the fleet. | |
| wait | No | If true, block until the device has finished allocating and is ready to use, up to timeoutSeconds. If false/omitted, this returns as soon as the reservation is made, but the device may not be immediately ready. | |
| version | No | Filter by OS version. Supports comparison prefixes >=, >, <=, < (e.g. ">=18"), or an exact version (e.g. "18.6.2"). Multiple values are ANDed together. | |
| platform | Yes | The platform to allocate a device for | |
| timeoutSeconds | No | Seconds to wait for allocation when wait is true. Defaults to 900 (15 minutes). Only relevant when wait is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the read/mutate/open-world profile; the description adds real operational context beyond them: the device is shared and billed, release wipes device state, and auth failures should route to login-then-retry. It also warns against release/reallocate churn mid-task.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then routes to alternatives and caveats in tightly packed sentences with no filler. It is somewhat dense, but every sentence carries actionable guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating allocation tool with no output schema, the description covers the return value (device identifier), prerequisites, lifecycle (release when done), and cost/sharing implications. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (name, type, wait, version, platform, timeoutSeconds) are already documented in the schema, including prefix/comparison syntax. The description adds no parameter-level detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Reserve') and resource ('a physical device from the remote cloud fleet') and clarifies it returns a device identifier usable with other mobile_* tools. It explicitly distinguishes itself from local-device tools ('Unlike local devices...').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when/when-not: only after the user explicitly asks for remote/cloud, never speculatively or as a fallback. Names prerequisites (mobile_login_to_cloud_provider first, mobile_list_remote_devices to discover names) and the counterpart tool (mobile_release_remote_device) with guidance on when to release.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_batch_commandsBatch CommandsADestructive
Run multiple tools in sequence in a single call, e.g. click, type, click, type. Use this to fill forms or perform multi-step flows without round-trips. The device argument is applied to every step unless a step provides its own.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | Yes | Tools to run, in order | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| stopOnError | No | Stop at the first failing step. Defaults to true | |
| listElementsAtEnd | No | Run mobile_list_elements_on_screen after the last step and include its result. Defaults to false |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, so the agent knows steps can mutate device state. The description adds the device-inheritance rule, which is useful, but says nothing about partial-failure state, whether nested batch calls are permitted, or what the aggregate result looks like. With annotations covering the safety profile, this is an adequate but not rich disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with purpose before the usage example and the device-inheritance caveat. Every sentence carries information; nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a batch orchestrator with nested step objects and no output schema, the description covers purpose, usage and device scoping well. It leaves the return shape and partial-failure semantics implicit, though stopOnError and listElementsAtEnd are documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds a rule the schema does not state: the device argument applies to every step unless a step supplies its own. That inheritance behavior is genuinely useful for constructing a steps array correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — run multiple tools in sequence in one call — and immediately clarifies its scope with a concrete example (click, type, click, type). No sibling tool does batching, so the agent can distinguish it at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use context ('fill forms or perform multi-step flows without round-trips'), which tells the agent when this tool beats issuing individual calls. It stops short of explicit exclusions or guidance on when a single-step call is preferable, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_click_on_screen_at_coordinatesClick ScreenA
Click on the screen, either at x,y coordinates or on an element by its ref (e.g. "@e5") from the latest mobile_list_elements_on_screen result. Prefer ref when the element is listed.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | The x coordinate to click on the screen, in pixels. Required unless ref is given | |
| y | No | The y coordinate to click on the screen, in pixels. Required unless ref is given | |
| ref | No | Element ref from mobile_list_elements_on_screen, e.g. "@e5". Takes precedence over x,y | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds the freshness constraint that refs come from the latest list_elements result, but says nothing about failure behavior, coordinate origins, or what happens if neither x,y nor ref resolve.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and immediately followed by the ref-vs-coordinates preference. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, no-output-schema action tool, the description plus fully documented schema cover what an agent needs to invoke it correctly. Only minor gaps remain around the coordinate system and failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so x, y, ref and device are all already documented, including the ref-over-x,y precedence. The description reinforces the mutual exclusivity and the provenance of ref, which is mild added value over the structured fields, so baseline 3 fits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (click) and resource (screen) with two distinct modes: coordinates or element ref. It is clear and self-contained, but does not explicitly contrast itself with near siblings such as mobile_double_tap_on_screen or mobile_long_press_on_screen_at_coordinates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear selection rule: 'Prefer ref when the element is listed,' and points to mobile_list_elements_on_screen as the source of refs. It stops short of stating when coordinates are mandatory (e.g. unlisted/custom targets) or what to do when no ref exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_clipboardClipboardA
Read or replace the device clipboard. Pass text to set the clipboard, omit it to read the current clipboard.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Text to place on the clipboard. Omit to read the clipboard instead. | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false and openWorldHint=false; the description is consistent with that (it can mutate, and 'replace' clarifies the write is an overwrite, not an append). It does not say what the read mode returns or whether the prior clipboard content is recoverable, so it adds only moderate value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the read/write rule front-loaded immediately after the purpose statement. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only two parameters and full schema coverage, the description covers what an agent needs to call the tool. The one gap is the return behavior in read mode, which is not described and there is no output schema to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, including the omit-to-read semantics for 'text' and the pointer to mobile_list_available_devices for 'device'. The description restates rather than extends this, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (the device clipboard) and both verbs it supports (read or replace), which is exactly the dual-mode behavior an agent needs to recognize. No sibling tool touches the clipboard, so no differentiation is required.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the selection rule between the two modes: 'Pass text to set the clipboard, omit it to read the current clipboard.' That is clear operational context, though it names no alternatives or preconditions (e.g. device availability).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_double_tap_on_screenDouble Tap ScreenC
Double-tap on the screen at given x,y coordinates.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | The x coordinate to double-tap, in pixels | |
| y | Yes | The y coordinate to double-tap, in pixels | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the safety posture is covered. The description adds no behavioral context beyond the action itself — nothing about coordinate space, device state requirements, or that it targets the currently foregrounded app.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the action and inputs front-loaded and no filler. It is terse to the point of being sparse, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter input tool with full schema coverage and annotations, this is minimally sufficient. It is missing any routing against the similar click/long-press siblings, which is the main gap given how many gesture tools exist.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including pixel units for x/y and the pointer to mobile_list_available_devices for device. The description adds nothing beyond what the schema already documents, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (double-tap) and resource (screen) with the coordinate inputs named. It clearly distinguishes from mobile_click_on_screen_at_coordinates and mobile_long_press_on_screen_at_coordinates by gesture type, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to prefer a double-tap over the sibling click, long-press, or swipe gestures, and no preconditions stated. The agent must infer usage entirely from the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_fold_deviceFold DeviceA
Fold or unfold a foldable device, such as iPhone Duo simulators and foldable Android emulators. Provide either posture or angle.
| Name | Required | Description | Default |
|---|---|---|---|
| angle | No | The desired hinge angle in degrees, between 0 (folded) and 180 (open). iOS only, e.g. iPhone Duo. Use instead of posture. | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| posture | No | The desired hinge posture: folded (0 degrees), half-open (90 degrees) or open (180 degrees) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose that this is a non-read-only, non-destructive, closed-world operation, so the description need not repeat the safety profile. It adds useful context about foldable device types and the posture/angle input requirement, but stops short of describing failure modes, prerequisites such as device allocation, or the effect on device state beyond the fold itself.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the operation and followed by the key input constraint. The examples and instruction all earn their place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity, rich parameter schema, and annotations covering the mutation profile, the description supplies enough context to call the tool correctly. It could be slightly stronger by clarifying whether one of posture/angle is mandatory, but the schema and description together cover the essentials.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters in detail. The description adds a meaningful constraint not encoded in the required list: the caller must provide either posture or angle, which helps resolve the otherwise ambiguous optional parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb pair (fold/unfold) and resource (foldable device), with examples that make the scope immediately clear. It is distinguishable from siblings like mobile_set_orientation because it targets foldable hardware and hinge posture/angle rather than general screen orientation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context by naming the device class (foldable simulators/emulators) and instructs the caller to provide either posture or angle. It does not explicitly name an alternative tool or state when not to use it, but the context is sufficient to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_get_crashGet Crash ReportARead-only
Get the full content of a crash report by its ID. Use mobile_list_crashes to find available crash IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The crash report ID to retrieve | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safe-read profile is covered without description help. The description adds the 'full content' nuance (as opposed to the summary implied by listing), but says nothing about report size, permissions, or whether content can be truncated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the dependency stated second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read with no output schema, the description plus annotations give enough to invoke correctly, and 'full content' hints at the return payload. It stops short of describing the report's form (structured vs raw text), leaving a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'id' and 'device' are already documented in the schema. The phrase 'by its ID' only echoes the schema's id description, adding no new syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Get the full content of a crash report') and scopes it by ID, which cleanly separates it from the sibling mobile_list_crashes that enumerates crash IDs. An agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the prerequisite path explicitly: 'Use mobile_list_crashes to find available crash IDs.' That tells the agent both when to call this tool (once an ID is known) and where to get the input, mirroring the routing pattern of well-scored definitions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_get_device_logsGet Device LogsARead-only
Collect live device logs (logcat on Android, unified log on iOS) as one JSON object per line. Only logs emitted after the call starts are captured, so trigger the behavior you want to observe right before or during the call. Stops after limit entries or after 30 seconds of silence. Each line of a stack trace counts as one entry, so filter by process or tag to avoid filling the limit with noise. Start with a small limit (20-50) for a quick look. For large captures, pass saveTo to write logs to a file and only the entry count is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Stop after this many log entries | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| filter | No | Filters, ANDed together. key=value includes, key!=value excludes. Keys: pid, process, tag, level, subsystem, category, message. Example: ["tag=ActivityManager", "level=Error", "process!=SpringBoard"] | |
| saveTo | No | Path to write the logs to instead of returning them. Filename must end with .log, .txt, or .jsonl |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint, but the description adds substantial behavioral detail beyond them: only post-call logs are captured, termination on limit or 30s of silence, stack-trace lines counting as separate entries, and the saveTo side effect where only the entry count is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool captures, then the timing constraint, then termination behavior, then practical guidance. Six tight sentences with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description explains the return shape (one JSON object per line, or only an entry count with saveTo), which an agent needs to interpret results. Nothing required for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value beyond the schema: practical guidance on limit sizing (20-50), the consequence of saveTo (file written, only count returned), and filter-by-process/tag noise avoidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Collect live device logs') and even names the platform-specific mechanisms (logcat, unified log). It is clearly distinguishable from siblings like mobile_get_crash and mobile_list_crashes by being a live-streaming capture tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable context: trigger the behavior right before/during the call, start with a small limit (20-50), filter by process or tag, and use saveTo for large captures. It does not explicitly name an alternative sibling tool or state when-not to use it, keeping it just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_get_foreground_appGet Foreground AppARead-only
Get the app currently in the foreground on the device. Use this to verify which app or screen you are on before interacting with it.
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the pre-interaction verification use case, but says nothing about what identifying information is returned (package name, activity, screen label), which matters for a 'get' tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler; the core action is front-loaded and the usage hint follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless-output 'get' tool with no output schema, the description should indicate the shape of the result (app identifier vs. screen). It conveys purpose and usage but leaves the return value undefined, which is a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single device parameter is fully documented in the schema, including a pointer to mobile_list_available_devices. The description adds no parameter meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (foreground app) with clear scope: the app currently in the foreground on the device. This is readily distinguishable from siblings such as mobile_list_apps (all installed apps) or mobile_launch_app.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the usage context: verify which app or screen you are on before interacting with it. That gives an agent a clear trigger condition, though it names no alternative tool or exclusion case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_get_orientationGet OrientationBRead-only
Get the current screen orientation of the device
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds no behavioral context such as what orientation values are returned or whether it reflects device vs app orientation, so it contributes little beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It is exactly as long as it needs to be for a simple read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with no output schema, the description omits any indication of the return form (e.g., portrait/landscape values). An agent knows what to call but not what to expect back, leaving a small but real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the device parameter is fully documented, including a pointer to mobile_list_available_devices. The description adds no parameter detail, so the baseline 3 for schema-covered parameters is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Get the current screen orientation') clearly. It is distinguishable from the opposite-direction sibling mobile_set_orientation by the read verb, though it never names or contrasts with that sibling explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as mobile_set_orientation or mobile_get_screen_size. The read verb only implies the use case, leaving the agent to infer everything.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_get_screen_sizeGet Screen SizeARead-only
Get the screen size of the mobile device in pixels
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, capping the need for safety disclosure. The description adds the output unit (pixels) and scope ('mobile device'), but says nothing about return format, orientation caveats, or whether values change with rotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero waste. The unit, which is the most important qualifier, is placed at the end where it naturally qualifies 'size'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with full schema coverage and annotations, the definition is sufficient to invoke correctly. The only missing piece is whether the returned size is in physical or logical pixels and whether it reflects current orientation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the device parameter is fully documented (including a pointer to mobile_list_available_devices). The description adds no parameter meaning beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (screen size) with the unit ('in pixels'), distinguishing it from siblings like mobile_get_orientation. An agent knows exactly what this returns without checking the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no mention of prerequisites, and no reference to alternatives. The schema's device description points to mobile_list_available_devices, but the prose itself offers nothing about when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_install_appInstall AppC
Install an app on mobile device
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | The path to the app file to install. For iOS simulators, provide a .zip file or a .app directory. For Android provide an .apk file. For iOS real devices provide an .ipa file | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered structurally. The description adds no behavioral context of its own: it does not say the device must be booted, whether it overwrites an existing install, how long the operation takes, or what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single clause with no filler, and the verb+resource lead makes it front-loaded. It is efficient but so terse that it reads as under-specified rather than deliberately minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 required params, no output schema) and the schema plus annotations cover safety and parameter details, so an agent can invoke it. Still, the description omits any mention of device readiness or platform-specific install behavior that a caller would want before running it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters in detail, including per-platform file formats for 'path' and the device-listing hint for 'device'. The description adds nothing beyond the schema, which is the expected baseline of 3 when structured fields do the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Install) and resource (app on mobile device), which is enough to distinguish it from siblings like mobile_uninstall_app and mobile_launch_app. It does not, however, explicitly name or contrast those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives such as mobile_launch_app or mobile_uninstall_app, and no prerequisites stated (e.g., device must be running/allocated). The only useful routing hint, listing available devices, lives in the schema rather than the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_launch_appLaunch AppB
Launch an app on mobile device. Use this to open a specific app. You can find the package name of the app by calling list_apps_on_device.
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| locale | No | Comma-separated BCP 47 locale tags to launch the app with (e.g., fr-FR,en-GB) | |
| packageName | Yes | The package name of the app to launch |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds nothing behavioral beyond that: no statement about failure when the package is absent, whether the call blocks until the app is foregrounded, or how the optional locale affects the launched process.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and no padding. The first two sentences partially restate each other ('Launch an app...' / 'Use this to open a specific app'), which is the only real waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity launch tool with full schema coverage and no output schema, this is close to sufficient. It still leaves the agent without guidance on the primary failure mode (app not installed) and on whether the launch is synchronous, which an agent orchestrating subsequent screen interactions would want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (device, locale, packageName) are already documented in the schema; the baseline is 3. The description reinforces where packageName comes from but adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Launch an app on mobile device') that an agent can act on immediately. It does not, however, explicitly differentiate itself from nearby siblings such as mobile_terminate_app or mobile_install_app beyond the obvious antonym/adjacency.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives implied usage ('Use this to open a specific app') and a helpful pointer for sourcing the required package name, but offers no when-not guidance (e.g., what to do if the app is not installed, or when to use install vs launch). The referenced tool name 'list_apps_on_device' does not match the actual sibling 'mobile_list_apps', weakening the routing value.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_list_appsList AppsCRead-only
List all the installed apps on the device
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds nothing beyond that — no note about return format, filtering, or whether system/user apps are included, which would have been the value-add for a listing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no wasted words. It is efficient, though so terse that it omits the scoping detail (user vs system apps) that would make it more actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a trivial read-only listing tool with no output schema and a fully documented single parameter, this is minimally adequate. It omits what the listing contains (package names, labels) and there are no annotations-independent behavioral notes, leaving a small but real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter with 100% schema description coverage, so the schema fully documents 'device' including how to discover valid identifiers. The description adds no parameter meaning beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (installed apps), scoped to 'on the device'. It is clearly distinguishable from sibling list/get tools like mobile_get_foreground_app or mobile_list_remote_devices by the word 'installed', though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, prerequisites, or alternatives. The only routing hint ('Use mobile_list_available_devices to find which devices are available') lives in the schema's parameter description, not the description itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_list_available_devicesList DevicesARead-only
List all available devices. This includes both physical mobile devices and mobile simulators and emulators. It returns both Android and iOS devices. These are local devices already connected to this machine, ready to use immediately at no cost - for devices from the shared remote cloud fleet, use mobile_list_remote_devices instead.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safe-read profile is covered. The description adds real context beyond that: the devices are local, already connected, cost nothing, and include both emulators/simulators and physical Android/iOS hardware. It does not describe the return shape (e.g. device IDs or ordering), which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences with the core capability front-loaded and the alternative routing placed last. Every sentence carries distinct information (scope, device types, platforms, locality/cost, alternative).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with no output schema, the description covers scope, platform coverage, locality, cost, and the sibling alternative. Nothing required to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is no parameter behavior an agent could misconstrue. No parameter-level detail is needed or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource (list available devices) and immediately scopes it to local physical devices plus simulators/emulators across Android and iOS. It explicitly distinguishes itself from the similarly named sibling mobile_list_remote_devices, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use (local devices 'ready to use immediately at no cost') and an explicit when-not plus the named alternative ('for devices from the shared remote cloud fleet, use mobile_list_remote_devices instead'). This is the routing information an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_list_crashesList Crash ReportsCRead-only
List crash reports available on the device
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety and environment scope are covered by structured data. The description adds no behavioral context beyond the annotation payload -- nothing about whether reports persist, how many are returned, or whether the device must be active.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that is front-loaded and wastes no words. It is concise, though arguably too terse for the information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should describe what a crash report contains (bundle id, timestamp, stack trace?) and whether results are paginated or limited. Neither is present, so an agent cannot anticipate the return shape from this definition alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'device' parameter already documents itself, including the pointer to mobile_list_available_devices. Per the high-coverage baseline this sits at 3; the description adds no format or discovery detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('crash reports') with device scope, which cleanly separates it from the singular mobile_get_crash sibling. It doesn't explicitly name that sibling or explain the distinction, but the list-vs-get contrast is reasonably inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to use this tool versus alternatives such as mobile_get_crash (fetch one report) or mobile_get_device_logs. The only contextual pointer lives in the parameter schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_list_elements_on_screenList Screen ElementsARead-only
List elements on screen with their ref, coordinates, and display text or accessibility label. Use the ref with mobile_click_on_screen_at_coordinates. Refs and coordinates stay valid as long as the screen does not change; re-list only after navigation or a layout change.
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| format | No | Output format. "text" (default) is one compact line per element, "json" is a json array |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only and open-world, but the description adds important behavioral context: refs and coordinates are valid only while the screen does not change. It also clarifies the returned fields, though it does not discuss filtering, result size, or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three tight sentences, front-loaded with the core action and followed by integration and freshness caveats. Every sentence earns its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully summarizes return contents and ref validity. The input schema covers parameters, and annotations cover safety, leaving no significant gap for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the device and format parameters are already fully documented in the input schema. The description adds no additional meaning about parameter syntax, defaults, or effects, making 3 the appropriate baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List elements on screen,' and enumerates what is returned (ref, coordinates, display text or accessibility label). This clearly distinguishes the tool from siblings like mobile_get_screen_size or mobile_take_screenshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to use the output with mobile_click_on_screen_at_coordinates and gives a clear condition for re-listing after navigation or layout change. It lacks explicit guidance on when to prefer this over alternatives such as screenshots or other screen-inspection tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_list_remote_devicesList Remote DevicesARead-only
List the catalog of device models (make, platform, OS version) available to reserve from the remote cloud device fleet. This is different from mobile_list_available_devices, which lists real devices and simulators/emulators already connected to this local machine and ready to use immediately at no cost. Remote devices live in a shared cloud fleet: they are not usable until reserved with mobile_allocate_remote_device, and reserving one may be a limited/billed resource. Requires mobile_login_to_cloud_provider to have been called first; if this fails with an authentication error, call that tool then retry.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint; the description adds meaningfully more: devices are unusable until reserved, reservation may be a limited/billed resource, and cloud login must have run first with a retry path on auth failure. This is exactly the beyond-annotation context the dimension rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what is listed, then contrast, then prerequisites/cost and error handling. Multi-sentence but every sentence carries routing, dependency, or cost information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no params, no output schema, and only two annotations, the description supplies the full operational picture: workflow position (login -> list -> allocate), cost/availability caveat, and the difference from the local-device sibling. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description correctly implies no filtering/arguments are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the catalog of device models ... from the remote cloud device fleet') and explicitly distinguishes itself from mobile_list_available_devices by contrasting cloud fleet vs local connected devices. An agent can pick between the two without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (reservable cloud fleet) vs when-not (local devices via mobile_list_available_devices), plus the required prerequisite (mobile_login_to_cloud_provider) and the follow-up action (mobile_allocate_remote_device). Failure handling for auth errors is spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_login_to_cloud_providerLogin to Cloud ProviderA
Start authenticating this machine with the remote device cloud provider. This is required once before mobile_list_remote_devices or mobile_allocate_remote_device will work; if either of those fails with an authentication error, call this tool and then retry. This starts a browser-based device-code login and returns quickly with a URL and a one-time code - it does NOT wait for the login to complete. Show the URL and code to the user verbatim and ask them to open the URL and enter the code in their own browser. The login keeps running in the background after this tool returns; once the user confirms they've completed it, retry the remote devices tool that originally failed. Only call this after the user has explicitly asked to connect to, log into, or use remote/cloud devices - never call it speculatively, since it interrupts the user to act in their browser.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say openWorldHint=true, not destructive, not read-only. The description adds the critical behavioral facts the annotations cannot convey: it is asynchronous, returns quickly with a URL and one-time code, does NOT wait for completion, keeps running in the background, and requires the user to act in their own browser. This is exactly the extra context needed for an agent to act correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a long paragraph, but every sentence carries distinct operational information: purpose, prerequisite, trigger, mechanics, user instruction, and restriction. Nothing is redundant and the trigger condition is front-loaded after the purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description explains what the tool returns (URL and one-time code), how to present it (verbatim), and what to do afterwards (retry the failed tool). Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so there is nothing to mis-specify; the baseline for a no-param tool applies. The description even clarifies the interaction payload is returned, not passed in.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start authenticating this machine with the remote device cloud provider') and explicitly names the two sibling tools it unblocks. An agent can distinguish it from the many device-management siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Spells out prerequisites (required once before list/allocate will work), the trigger condition (either fails with an auth error), the sequence (call this, then retry), and an explicit prohibition ('never call it speculatively'). This is about as complete as usage guidance gets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_long_press_on_screen_at_coordinatesLong Press ScreenA
Long press on the screen at given x,y coordinates. If long pressing on an element, use the mobile_list_elements_on_screen tool to find the coordinates.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | The x coordinate to long press on the screen, in pixels | |
| y | Yes | The y coordinate to long press on the screen, in pixels | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| duration | No | Duration of the long press in milliseconds. Defaults to 500ms. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the agent knows this interacts with an external device without being destructive. The description adds nothing beyond that — no mention of the 500ms default duration, whether the gesture triggers context menus, or any app-state side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, the core action is front-loaded, and the second sentence is a genuinely useful routing hint. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, an agent doesn't need return-value documentation, and annotations plus a fully documented schema cover the rest. The description is nearly complete for a simple gesture tool, missing only any note on gesture timing/behavioral effect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (x, y, device, duration all documented, including the 500ms default), so the schema carries the parameter burden. The description only restates x/y and points at a discovery tool for coordinates, adding marginal value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Long press on the screen at given x,y coordinates'), clearly distinct from sibling gestures like mobile_click_on_screen_at_coordinates, mobile_double_tap_on_screen, and mobile_swipe_on_screen. It does not explicitly name those siblings, so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description routes the agent to mobile_list_elements_on_screen for finding coordinates on an element, which is real prerequisite guidance. However, it gives no guidance on when to choose a long press over a click, double tap, or button press, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_open_urlOpen URLC
Open a URL in browser on device
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to open | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and destructiveHint=false, so the safety profile is covered by structured data. The description adds nothing beyond that — no mention that this launches an external browser, makes network requests, or affects foreground state (which mobile_get_foreground_app would observe).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the action front-loaded and no filler. Efficient, though it leans on terseness rather than structured completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter, well-documented tool with no output schema, the minimal description is survivable, but it omits the openWorld/network side effects and any interaction with sibling navigation tools, leaving the agent to infer behavior from annotations alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters are documented in-schema, including the device-discovery tip. The description adds no format, constraint, or example detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb (open), resource (URL), and context (in browser on device), which distinguishes it from the other mobile_* action tools. It does not explicitly contrast with siblings like mobile_launch_app, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives such as mobile_launch_app or the clicking/typing tools. The only procedural hint (use mobile_list_available_devices) lives in the schema parameter description, not the tool description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_press_buttonPress ButtonC
Press a button on device
| Name | Required | Description | Default |
|---|---|---|---|
| button | Yes | The button to press. Supported buttons: BACK (android only), HOME, VOLUME_UP, VOLUME_DOWN, ENTER, DPAD_CENTER (android tv only), DPAD_UP (android tv only), DPAD_DOWN (android tv only), DPAD_LEFT (android tv only), DPAD_RIGHT (android tv only) | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and destructiveHint=false, covering the safety profile. The description adds nothing beyond that baseline — no indication of what a press affects, whether state changes persist, or any auth/device requirements. With annotations present the bar is lower, but this text contributes no behavioral context at all.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. It is efficient, though arguably under-specified rather than optimally concise; still, no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With full schema coverage and annotation-provided safety hints, and no output schema, the definition is minimally viable. It omits any context about what the button press does downstream or where it applies, which for an input tool leaves a modest gap without being inadequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the schema fully documents 'button' with its supported values and 'device' with a pointer to mobile_list_available_devices. The description adds no additional parameter meaning, so the baseline 3 applies when the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Press a button on device'), so an agent immediately knows this performs a key/button input. It does not, however, differentiate itself from siblings like mobile_click_on_screen_at_coordinates or mobile_long_press_on_screen_at_coordinates, which would require naming what makes a 'button' press distinct from a coordinate tap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the many other input tools (click, double tap, long press, type keys). No preconditions such as needing an app in the foreground or a device selected are stated. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_release_remote_deviceRelease Remote DeviceADestructive
Release a device previously reserved with mobile_allocate_remote_device back to the remote cloud fleet so it becomes available to others. Releasing is destructive to the device's state: apps installed, files pushed, and any other changes made during this session are lost, and a later mobile_allocate_remote_device call may take time and could return a different physical unit. Only release once the whole task is finished - if there is more work to do on the same device shortly, keep holding it rather than releasing and reallocating.
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to release back to the remote fleet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=true, but the description goes well beyond them by enumerating exactly what is lost (installed apps, pushed files, session changes) and warning that a later allocation may be slow and return a different physical unit. It does not cover auth/permission requirements or return behavior, but the added state-loss and reallocation context is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and sibling link, then consequences and the workflow rule. Four sentences, all earning their place, though the final sentence could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter release tool with no output schema, the definition covers purpose, consequences, reallocation caveats, and the timing rule an agent needs to avoid premature release. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, so the schema already documents the device identifier. The description adds only indirect meaning by tying the identifier to a previously allocated reservation; baseline 3 applies when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('release') and resource ('a device previously reserved') and explicitly names the counterpart tool mobile_allocate_remote_device, so the agent can distinguish this from siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('once the whole task is finished') and when-not ('if there is more work to do on the same device shortly, keep holding it'), plus the alternative action to take in that case. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_save_screenshotSave ScreenshotC
Save a screenshot of the mobile device to a file
| Name | Required | Description | Default |
|---|---|---|---|
| scale | No | Scale factor (0.0-1.0). Ignored if maxSize is provided. | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| saveTo | Yes | The path to save the screenshot to. Filename must end with .png, .jpg, or .jpeg | |
| maxSize | No | Maximum width/height in pixels, keeping aspect ratio. Omit for full size. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the agent already knows this is a non-destructive local write. The description's addition that it saves to a file is consistent and mildly useful, but it says nothing about overwriting existing files, required permissions, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler; the action and destination are front-loaded. It is efficient, though arguably too terse to be maximally useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter mutation tool with full schema coverage and no output schema, the description is adequate but thin: it omits sibling differentiation against mobile_take_screenshot and does not state whether the saved path is returned or whether existing files are overwritten.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (scale, maxSize, device, saveTo) are fully documented in the schema, including the scale/maxSize interaction and the .png/.jpg/.jpeg filename requirement. The description adds no parameter detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a clear verb and resource (save a screenshot to a file), which is more specific than the title. However, it never distinguishes itself from the sibling mobile_take_screenshot, leaving the agent to guess which one writes to disk versus returns image data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus mobile_take_screenshot, nor any mention of prerequisites such as needing a device identifier. The only implied usage is the presence of a saveTo parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_set_locationSet LocationA
Override the GPS location reported by the device, or clear the override to restore the real location. Omit latitude and longitude to clear.
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| latitude | No | Latitude in decimal degrees, e.g. 37.7749. Omit together with longitude to clear the override. | |
| longitude | No | Longitude in decimal degrees, e.g. -122.4194. Omit together with latitude to clear the override. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, but not what the write actually does; the description fills that gap by explaining it overrides the reported GPS and that clearing restores the real location, which is meaningful beyond the annotations. It omits whether the override persists, whether it requires a connected device, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the primary mutation stated first and the clearing behavior second; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given full schema coverage, a small parameter set, and no output schema, the description covers the essential operation and the reset path. It stops short of clarifying behavior when only one of latitude/longitude is supplied or whether the override survives device reconnect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both latitude/longitude descriptions already contain the omit-together-to-clear rule, so the description largely restates schema content. Baseline 3 applies since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (override/clear) and resource (GPS location reported by the device), plus the inverse operation. No sibling tool manipulates location, so an agent can identify it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes the clearing behavior (omit latitude and longitude) and the restoration effect, which is the key usage condition. No alternatives exist among siblings, so no exclusion guidance is needed, but it never states prerequisites about being connected to a device.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_set_orientationSet OrientationC
Change the screen orientation of the device
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| orientation | Yes | The desired orientation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false (a mutation) and destructiveHint=false, so the safety profile is covered. The description adds nothing beyond the annotations — it does not say whether the change persists, what happens if the device is locked, or how failures surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. It is appropriately sized for a two-parameter tool, though it is so terse it leaves the usage and behavioral gaps noted above.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 required params, no nested objects, no output schema) and the schema carries most of the load, so the description is close to sufficient. It still omits any behavior on invocation and any reference to the paired get-orientation sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters, including the device pointer to mobile_list_available_devices and the orientation enum. The description adds no additional parameter meaning, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('change the screen orientation') and the 'change' framing implicitly distinguishes it from the sibling mobile_get_orientation. However, it does not explicitly name the sibling or the scope (persistent vs temporary), so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no mention of the sibling mobile_get_orientation or any alternative. Usage is only implied by the tool name; an agent gets no exclusions or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_start_screen_recordingStart Screen RecordingA
Start recording the screen of a mobile device. The recording runs in the background until stopped with mobile_stop_screen_recording. Returns the path where the recording will be saved.
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| output | No | The file path to save the recording to. Filename must end with .mp4. If not provided, a temporary path will be used. | |
| timeLimit | No | Maximum recording duration in seconds. The recording will stop automatically after this time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, non-destructive, non-open-world, so the safety profile is covered. The description adds genuine behavioral context beyond that: the recording persists in the background, must be explicitly stopped, and the return value is a save path rather than the recording itself. It omits edge cases such as behavior when a recording is already active or any storage/permission constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the lifecycle, then the return value. No filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains the return value (the path where the recording will be saved) and the stop counterpart, which is what an agent needs to call it correctly. Minor gaps remain around concurrent recordings and failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so device, output, and timeLimit are fully documented in the schema, including the .mp4 requirement and auto-stop semantics. The description adds no parameter detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Start recording the screen of a mobile device") and names the paired sibling mobile_stop_screen_recording, distinguishing it from screenshot tools like mobile_take_screenshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes the lifecycle: it "runs in the background until stopped with mobile_stop_screen_recording", which tells the agent this is a paired start/stop operation. It does not, however, say when to prefer recording over mobile_take_screenshot or mobile_save_screenshot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_stop_screen_recordingStop Screen RecordingA
Stop an active screen recording on a mobile device. Returns the file path, size, and approximate duration of the recording.
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is already covered. The description usefully adds the return payload (file path, size, duration), but says nothing about what happens if no recording is active or whether the file persists afterward.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: purpose first, return values second. No filler, no redundancy with the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly compensates by naming the returned fields. For a single-parameter tool with full annotation coverage, this is close to complete, though the failure mode (no active recording) is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, so the schema fully documents the device identifier and even points to mobile_list_available_devices. The description adds nothing beyond the schema, which is the expected baseline when coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (stop an active screen recording) and scopes it to a mobile device. The contrast with the sibling mobile_start_screen_recording is obvious from the verb, but the description never names or references the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the phrase 'an active screen recording' — the agent can infer this applies when a recording is in progress, but the description gives no explicit when-to-use condition, no error case for when nothing is recording, and no routing to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_swipe_on_screenSwipe ScreenC
Swipe on the screen
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | The x coordinate to start the swipe from, in pixels. If not provided, uses center of screen | |
| y | No | The y coordinate to start the swipe from, in pixels. If not provided, uses center of screen | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| distance | No | The distance to swipe in pixels. Defaults to 400 pixels for iOS or 30% of screen dimension for Android | |
| direction | Yes | The direction to swipe |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the description carries a lower burden, but it adds literally nothing beyond them. It does not mention that a device must be allocated/connected, nor that swipes may trigger navigation or scrolling side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but brevity here reflects under-specification rather than efficiency. The single phrase is front-loaded but earns no informational place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, 2 required, and no output schema, the description should at minimum convey the gesture's effect and any device-session prerequisite. It leaves everything to the schema and annotations, which is inadequate for a device-interaction tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents device, direction, x, y, and distance with defaults and enum values. Per the baseline rule, a 3 is appropriate since the description adds no parameter meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description "Swipe on the screen" merely restates the tool name and title, adding no scope, target, or distinguishing detail. It does not differentiate from siblings like mobile_click_on_screen_at_coordinates or mobile_long_press_on_screen_at_coordinates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance whatsoever on when to use this versus gesture alternatives (tap, double tap, long press). No prerequisites (such as needing an active device session) or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_take_screenshotTake ScreenshotARead-only
Take a screenshot of the mobile device. Use this to understand what's on screen, if you need to press an element that is available through view hierarchy then you must list elements on screen instead. The screenshot is usually smaller than the screen, so when the device reports its screen size the result also states how to convert positions in the screenshot into screen coordinates before tapping. Do not cache this result.
| Name | Required | Description | Default |
|---|---|---|---|
| scale | No | Scale factor (0.0-1.0). Ignored if maxSize is provided. | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| maxSize | No | Maximum width/height in pixels, keeping aspect ratio. Defaults to 1024. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the bar is lower, yet the description still adds real behavioral context the annotations cannot convey: the image is usually smaller than the screen, the result states how to convert screenshot positions into screen coordinates before tapping, and the result must not be cached. It stops short of noting anything about size/quality tradeoffs (scale, maxSize) affecting the returned image.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler, and the primary purpose is front-loaded before the alternative-tool guidance and the coordinate-conversion caveat. Slightly dense in the middle sentence but every clause carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by explaining that the result includes coordinate conversion metadata, which is the key thing an agent must know to act on the returned image. Nothing critical is missing for a read-only capture tool, though the interaction with mobile_save_screenshot remains unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already fully self-documented by the schema (device, scale, maxSize with ranges and defaults). The description adds nothing about how scale or maxSize change the output, so the baseline 3 applies rather than anything higher.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ("Take a screenshot of the mobile device") and explicitly names the condition that routes the agent elsewhere (view hierarchy -> list elements). It distinguishes itself from mobile_list_elements_on_screen but never clarifies its relationship to the sibling mobile_save_screenshot, so the boundary between capture-and-return vs capture-and-persist is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use ("understand what's on screen") and an explicit when-not-to-use with the named alternative ("if you need to press an element that is available through view hierarchy then you must list elements on screen instead"). This is exactly the routing an agent needs between the two screen-inspection tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_terminate_appTerminate AppC
Stop and terminate an app on mobile device
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| packageName | Yes | The package name of the app to terminate |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so safety profile is covered. The description adds no further context such as whether termination requires the app to be running, whether it is reversible, or what happens to unsaved state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no wasted clauses. 'Stop and terminate' is mildly redundant but not harmful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter, non-destructive action tool with full schema coverage and no output schema, this is minimally adequate. It stops short of clarifying how termination differs from uninstalling, which is the main ambiguity given the sibling set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the 'device' parameter even points the agent to mobile_list_available_devices. The description adds nothing beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Stop and terminate an app'), so the agent knows the action targets an app on a device. It does not differentiate from siblings like mobile_uninstall_app or mobile_launch_app, which would require knowing terminate vs. remove vs. start.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to use this rather than mobile_uninstall_app (removes the app) or mobile_launch_app (starts it). The agent must infer usage purely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_type_keysType TextC
Type text into the focused element
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text to type | |
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| submit | Yes | Whether to submit the text. If true, the text will be submitted as if the user pressed the enter key. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-destructive, open-world operation, so the safety profile is covered. The description adds essentially nothing beyond that: it does not say what happens when no element is focused, whether typing appends or replaces, or how submission/errors behave.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler, which is efficient. The terseness borders on under-specification rather than verbosity, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
All three parameters are documented in the schema and annotations cover the safety profile, so the core is complete. Still missing is the focus precondition and error behavior, which matter for a tool that acts on whatever happens to be focused.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with all three parameters documented in the schema, so the schema carries the semantic load and the baseline of 3 applies. The description restates the 'text' idea but adds no format, encoding, or device-selection detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and target: 'Type text into the focused element', which distinguishes it from sibling input tools like mobile_press_button and mobile_click_on_screen_at_coordinates. It does not, however, name those siblings or otherwise route the agent among them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives. It does not tell the agent that an element must first be focused (e.g., via a tap), nor when to prefer mobile_press_button for special keys instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mobile_uninstall_appUninstall AppCDestructive
Uninstall an app from mobile device
| Name | Required | Description | Default |
|---|---|---|---|
| device | Yes | The device identifier to use. Use mobile_list_available_devices to find which devices are available to you. | |
| bundle_id | Yes | Bundle identifier (iOS) or package name (Android) of the app to be uninstalled |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=false, so the safety profile is covered and the description need not restate it. Beyond that, the description adds nothing: no mention of lost app data, no confirmation that the action removes the app permanently, and no device-state requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler or repetition. It is efficient, though arguably under-specified rather than optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity two-parameter mutation tool with full schema coverage and a clear destructiveHint annotation, and no output schema, the essential information is present via structured fields. The description itself contributes little but is not missing anything critical for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both 'device' and 'bundle_id' fully documented in the schema, including a pointer to mobile_list_available_devices and platform-specific identifier guidance. The description adds no parameter detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Uninstall an app'), clearly distinguishing it from install/launch siblings by name. However, it does not differentiate from the closely related mobile_terminate_app, which an agent could plausibly confuse with a full uninstall.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites (e.g. that the app should be installed, or that terminate_app is a lighter alternative), and no exclusions. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
33 tool updates
v1.0.6- Changed
mobile_allocate_remote_device1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_batch_commands1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_click_on_screen_at_coordinates1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_clipboard1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_double_tap_on_screen1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Added
mobile_fold_device - Changed
mobile_get_crash1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_get_device_logs1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_get_foreground_app1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_get_orientation1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_get_screen_size1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_install_app1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_launch_app1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_list_apps1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_list_available_devices1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_list_crashes1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_list_elements_on_screen1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_list_remote_devices1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_login_to_cloud_provider1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_long_press_on_screen_at_coordinates1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_open_url1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_press_button1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_release_remote_device1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_save_screenshot1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_set_location1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_set_orientation1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_start_screen_recording1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_stop_screen_recording1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_swipe_on_screen1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_take_screenshot1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_terminate_app1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_type_keys1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
mobile_uninstall_app1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
16 tool updates
v1.0.3- Added
mobile_allocate_remote_device - Added
mobile_batch_commands - Changed
mobile_click_on_screen_at_coordinates6 fields changed- added
Input schema / properties / refAdded value: +{ + "description": "Element ref from mobile_list_elements_on_screen, e.g. \"@e5\". Takes precedence over x,y", + "type": "string" +} - changed
Input schema / properties / x / descriptionPrevious value: -"The x coordinate to click on the screen, in pixels"New value: +"The x coordinate to click on the screen, in pixels. Required unless ref is given" - added
Input schema / properties / x / minimumAdded value: +0 - changed
Input schema / properties / y / descriptionPrevious value: -"The y coordinate to click on the screen, in pixels"New value: +"The y coordinate to click on the screen, in pixels. Required unless ref is given" - added
Input schema / properties / y / minimumAdded value: +0 - changed
Input schema / requiredPrevious value: -[ - "device", - "x", - "y" -]New value: +[ + "device" +]
- Added
mobile_clipboard - Changed
mobile_double_tap_on_screen2 fields changed- added
Input schema / properties / x / minimumAdded value: +0 - added
Input schema / properties / y / minimumAdded value: +0
- Added
mobile_get_device_logs - Added
mobile_get_foreground_app - Changed
mobile_list_elements_on_screen1 field changed- added
Input schema / properties / formatAdded value: +{ + "description": "Output format. \"text\" (default) is one compact line per element, \"json\" is a json array", + "enum": [ + "text", + "json" + ], + "type": "string" +}
- Added
mobile_list_remote_devices - Added
mobile_login_to_cloud_provider - Changed
mobile_long_press_on_screen_at_coordinates2 fields changed- added
Input schema / properties / x / minimumAdded value: +0 - added
Input schema / properties / y / minimumAdded value: +0
- Added
mobile_release_remote_device - Changed
mobile_save_screenshot2 fields changed- added
Input schema / properties / maxSizeAdded value: +{ + "description": "Maximum width/height in pixels, keeping aspect ratio. Omit for full size.", + "exclusiveMinimum": 0, + "maximum": 9007199254740991, + "type": "integer" +} - added
Input schema / properties / scaleAdded value: +{ + "description": "Scale factor (0.0-1.0). Ignored if maxSize is provided.", + "exclusiveMinimum": 0, + "maximum": 1, + "type": "number" +}
- Added
mobile_set_location - Changed
mobile_swipe_on_screen2 fields changed- added
Input schema / properties / x / minimumAdded value: +0 - added
Input schema / properties / y / minimumAdded value: +0
- Changed
mobile_take_screenshot2 fields changed- added
Input schema / properties / maxSizeAdded value: +{ + "description": "Maximum width/height in pixels, keeping aspect ratio. Defaults to 1024.", + "exclusiveMinimum": 0, + "maximum": 9007199254740991, + "type": "integer" +} - added
Input schema / properties / scaleAdded value: +{ + "description": "Scale factor (0.0-1.0). Ignored if maxSize is provided.", + "exclusiveMinimum": 0, + "maximum": 1, + "type": "number" +}
7 tool updates
v0.0.53- Added
mobile_get_crash - Changed
mobile_launch_app1 field changed- added
Input schema / properties / localeAdded value: +{ + "description": "Comma-separated BCP 47 locale tags to launch the app with (e.g., fr-FR,en-GB)", + "type": "string" +}
- Changed
mobile_list_available_devices2 fields changed- removed
Input schema / properties / noParamsRemoved value: -{ - "properties": {}, - "type": "object" -} - removed
Input schema / requiredRemoved value: -[ - "noParams" -]
- Added
mobile_list_crashes - Changed
mobile_save_screenshot1 field changed- changed
Input schema / properties / saveTo / descriptionPrevious value: -"The path to save the screenshot to"New value: +"The path to save the screenshot to. Filename must end with .png, .jpg, or .jpeg"
- Added
mobile_start_screen_recording - Added
mobile_stop_screen_recording
22 tool updates
v1.0.0- Changed
mobile_click_on_screen_at_coordinates3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "x", - "y" -]New value: +[ + "device", + "x", + "y" +]
- Added
mobile_double_tap_on_screen - Changed
mobile_get_orientation4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - removed
Input schema / properties / noParamsRemoved value: -{ - "additionalProperties": false, - "properties": {}, - "type": "object" -} - changed
Input schema / requiredPrevious value: -[ - "noParams" -]New value: +[ + "device" +]
- Changed
mobile_get_screen_size4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - removed
Input schema / properties / noParamsRemoved value: -{ - "additionalProperties": false, - "properties": {}, - "type": "object" -} - changed
Input schema / requiredPrevious value: -[ - "noParams" -]New value: +[ + "device" +]
- Added
mobile_install_app - Changed
mobile_launch_app3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "packageName" -]New value: +[ + "device", + "packageName" +]
- Changed
mobile_list_apps4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - removed
Input schema / properties / noParamsRemoved value: -{ - "additionalProperties": false, - "properties": {}, - "type": "object" -} - changed
Input schema / requiredPrevious value: -[ - "noParams" -]New value: +[ + "device" +]
- Changed
mobile_list_available_devices2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - removed
Input schema / properties / noParams / additionalPropertiesRemoved value: -false
- Changed
mobile_list_elements_on_screen4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - removed
Input schema / properties / noParamsRemoved value: -{ - "additionalProperties": false, - "properties": {}, - "type": "object" -} - changed
Input schema / requiredPrevious value: -[ - "noParams" -]New value: +[ + "device" +]
- Added
mobile_long_press_on_screen_at_coordinates - Changed
mobile_open_url3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "url" -]New value: +[ + "device", + "url" +]
- Changed
mobile_press_button3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "button" -]New value: +[ + "device", + "button" +]
- Changed
mobile_save_screenshot3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "saveTo" -]New value: +[ + "device", + "saveTo" +]
- Changed
mobile_set_orientation3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "orientation" -]New value: +[ + "device", + "orientation" +]
- Added
mobile_swipe_on_screen - Changed
mobile_take_screenshot4 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - removed
Input schema / properties / noParamsRemoved value: -{ - "additionalProperties": false, - "properties": {}, - "type": "object" -} - changed
Input schema / requiredPrevious value: -[ - "noParams" -]New value: +[ + "device" +]
- Changed
mobile_terminate_app3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "packageName" -]New value: +[ + "device", + "packageName" +]
- Changed
mobile_type_keys3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / deviceAdded value: +{ + "description": "The device identifier to use. Use mobile_list_available_devices to find which devices are available to you.", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "text", - "submit" -]New value: +[ + "device", + "text", + "submit" +]
- Added
mobile_uninstall_app - Removed
mobile_use_default_device - Removed
mobile_use_device - Removed
swipe_on_screen
17 tool updates
- First observed
mobile_click_on_screen_at_coordinates - First observed
mobile_get_orientation - First observed
mobile_get_screen_size - First observed
mobile_launch_app - First observed
mobile_list_apps - First observed
mobile_list_available_devices - First observed
mobile_list_elements_on_screen - First observed
mobile_open_url - First observed
mobile_press_button - First observed
mobile_save_screenshot - First observed
mobile_set_orientation - First observed
mobile_take_screenshot - First observed
mobile_terminate_app - First observed
mobile_type_keys - First observed
mobile_use_default_device - First observed
mobile_use_device - First observed
swipe_on_screen
TDQS
Scored across 33 tools
Most tools target a clearly distinct action+resource (launch vs terminate vs install vs uninstall, click vs double_tap vs long_press, local vs remote device listing). The main ambiguity is the mobile_take_screenshot / mobile_save_screenshot pair, and gesture tools overlap slightly, but descriptions mostly disambiguate them.
Every tool uses the mobile_ prefix with consistent snake_case verb_noun phrasing (mobile_launch_app, mobile_get_screen_size, mobile_list_elements_on_screen). No camelCase or stylistic deviations anywhere in the set.
33 tools is well above the practical ceiling for a single coherent surface, even for rich mobile automation. Several tools are closely related (take/save screenshot, three gesture variants) and could be consolidated, making the set heavier than it needs to be.
The surface covers an impressive lifecycle: device discovery/allocation/release, app install/launch/terminate/uninstall, input, screenshots, recording, logs, crashes, location, orientation, and clipboard. Only minor gaps remain (e.g. file push/pull, network condition control), which are workable.
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Melaya is a remote MCP server. It gives an assistant hands on your own Android phone and browser: it reads the screen through the accessibility tree, then taps, types and navigates inside the apps and sites you allow-list, with no per-app API. It also builds, schedules and runs agent pipelines across 6k+ connected tools. OAuth 2.1, nothing to install.
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
A Model Context Protocol server for Wix AI tools
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables AI agents to control and automate Android devices through natural language, supporting actions like app management, UI interactions, and device monitoring.59MIT
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables scalable mobile automation for iOS and Android through a platform-agnostic interface, allowing LLMs to interact with mobile applications via accessibility snapshots or screenshot-based inputs.1967,856 npm2Apache 2.0
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables AI assistants to interact with iOS simulators, perform accessibility testing, manage apps, and automate complex iOS workflows.32Apache 2.0
- AlicenseAqualityCmaintenanceA Model Context Protocol server for ad-hoc UI testing of Android and iOS apps, enabling LLM agents to interact with mobile app UIs and react to observations.4013 npm3MIT