open-mobile-mcp
Automates Android devices: capture screenshots, UI hierarchy, OCR, perform taps, swipes, typing, read logs, manage app lifecycle, and more.
Manages Expo dependencies by running npx expo install and provides bundler management for React Native/Expo projects.
Automates iOS devices via Maestro: perception (screenshots, UI hierarchy), interaction (tap, swipe, type), logging, and app lifecycle management.
Manages the Metro JavaScript bundler: start, stop, and restart Metro for React Native app development.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@open-mobile-mcptake a screenshot of the current screen"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Open Mobile MCP Server 📱
An open-source Model Context Protocol (MCP) server for mobile automation. Give any LLM eyes and hands on a real Android or iOS device — screenshot, tap, swipe, read logs, and verify your app without writing test code.
Works with Claude Code, Claude Desktop, Cursor, and any other MCP-compatible client.
Features
Perception: Screenshots, semantic UI hierarchy, OCR, element finder, layout health analysis.
Interaction: Tap, swipe, type, pinch, rotate, long-press, hardware key presses.
Logging: Per-app Android log filtering via PID (pass
deviceId+packageIdto eliminate system noise). Background log watching viawait_for_log.Environment: Metro bundler management, app lifecycle, deep links, screen recording, locale switching.
Text Input: Unicode/Cyrillic/CJK/Emoji support via ADB Keyboard with automatic keyboard restore.
Related MCP server: airi-android
Prerequisites
Node.js (v18+)
ADB installed and in PATH (for Android).
Maestro (required for iOS; fallback for Android input).
Mac/Linux:
curl -Ls "https://get.maestro.mobile.dev" | bashWindows:
powershell -Command "iwr -useb https://get.maestro.mobile.dev | iex"
(Optional) ADB Keyboard — only needed for non-ASCII input (Unicode, Cyrillic, Emoji).
Download from GitHub and install:
adb install ADBKeyboard.apk.
Configuration
macOS / Linux
{
"mcpServers": {
"open-mobile-mcp": {
"command": "npx",
"args": ["open-mobile-mcp"]
}
}
}Windows
{
"mcpServers": {
"open-mobile-mcp": {
"command": "npx",
"args": ["open-mobile-mcp"],
"env": {
"MAESTRO_HOME": "C:\\Users\\YOUR_USER\\.maestro",
"PATH": "C:\\Users\\YOUR_USER\\.maestro\\maestro\\bin;C:\\Windows\\system32;C:\\Windows;..."
}
}
}
}Note: On Windows, explicitly setting
MAESTRO_HOMEandPATHis often required formaestroto be found.
git clone https://github.com/xzaleksey/open-mobile-mcp.git
cd open-mobile-mcp
npm install && npm run buildThen use "command": "node", "args": ["/path/to/open-mobile-mcp/build/index.js"] in your MCP config.
Tools
Perception
Tool | Platform | Description |
| Android/iOS | List connected emulators and simulators |
| Android/iOS | Screenshot (~800px wide). Use |
| Android/iOS | Pruned UI tree as JSON |
| Android/iOS | OCR via Tesseract.js (default |
| Android/iOS | Set default OCR language (e.g. |
| Android/iOS | Find elements by |
| Android/iOS | Poll until element appears (default 20s) |
| Android/iOS | Cropped screenshot of a specific element |
| — | Compare two base64 screenshots, returns diff % |
| Android/iOS | Detect deep nesting and layout performance issues |
Interaction
Tool | Platform | Description |
| Android/iOS | Recommended — find + tap by selector. Note: text matching is exact; emoji prefixes (e.g. |
| Android/iOS | Raw coordinate tap. Must use original device pixels, not screenshot pixels. |
| Android/iOS | Swipe by coordinates |
| Android/iOS | Type text (handles Unicode) |
| Android | Two-finger pinch/zoom |
| Android | Two-finger rotation |
| Android/iOS | Hardware keys: |
Environment & Logs
Tool | Platform | Description |
| Android/iOS | Start/stop/restart Metro. Pass |
| Android/iOS | Manual control over |
| Android/iOS | Recent Metro/Android/iOS logs. Returns |
| Android/iOS | Recent error/exception lines across all sources |
| Android/iOS | Network logcat lines. For iOS, filters the internal log capture buffer (enable via |
| Android/iOS | Block until a log pattern matches. See subagent pattern below. |
| Android/iOS | Launch, stop, install, or uninstall apps |
| Android/iOS | Open a URL or deep link |
| Android/iOS | Reset app to fresh-install state |
| Android | Version, permissions, install date |
| Android/iOS | Screen recording to |
| Android/iOS | Run a Maestro YAML flow |
| — | Run |
| — | Run |
wait_for_log — Background Subagent Pattern
wait_for_log blocks until a pattern appears in the log buffer. Calling it directly in the main agent freezes the conversation. Always delegate it to a background subagent in Claude Code:
// Step 1 — spawn the watcher BEFORE the action that will trigger the log
// (In Claude Code, use Agent tool with run_in_background: true)
// Subagent prompt: "Call wait_for_log with pattern 'route: /home', timeout 60000. Report the result."
// Step 2 — perform the action in the main agent
tap_on_element({ selector: "Home", strategy: "text" });
// Step 3 — main agent continues freely; gets notified when subagent finishesLicense
MIT
Available Tools
32 toolsanalyze_layout_healthC
Analyze the UI layout for performance or health issues (e.g. deep nesting).
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It implies a read-only analysis but does not state what it returns, whether it has side effects, or what 'health issues' specifically covers. The tool's behavior remains largely opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose and provides a concrete example. There is no wasted wording or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema and no annotations, the description should explain what the tool returns and any relevant constraints. It does not, leaving the agent without essential context for interpreting results or understanding the tool's role in a workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, and the description adds no meaning for deviceId or platform. While the parameter names and the platform enum are somewhat self-explanatory, the description does not compensate for the lack of schema field descriptions, leaving the agent to guess about deviceId format and valid values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('analyze') and resource ('UI layout') with an illustrative example ('deep nesting'). It distinguishes itself from interaction tools, though it doesn't explicitly contrast with sibling diagnostic tools like get_semantic_hierarchy or capture_diff.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives. It doesn't mention context, prerequisites, or what differentiates it from sibling tools such as get_semantic_hierarchy or get_viewport.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_diffA
Compare two base64 images and return diff percentage.
| Name | Required | Description | Default |
|---|---|---|---|
| currentBase64 | Yes | ||
| baselineBase64 | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. The word 'Compare' and the phrase 'return diff percentage' imply a non-mutating, computation-only operation, but the description does not address constraints such as equal image dimensions, invalid base64 handling, or whether the operation is side-effect-free in an explicit way.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler: it front-loads the action and resource, then states the output. Every clause contributes information, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with two required parameters and no output schema, the description covers the input domain and the output concept. However, it omits the exact return shape, the definition of 'diff percentage' (e.g., percentage of differing pixels), and constraints such as matching image dimensions. It is adequate but leaves room for ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the schema only provides parameter names and string types, so the description must add meaning. It does this by identifying both parameters as base64-encoded images and clarifying their role in a comparison. It does not specify raw base64 versus data-URI format or size limits, but the core meaning of both parameters is clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Compare'), precise inputs ('two base64 images'), and the result ('return diff percentage'). It is clearly distinct from the sibling device-control and image-capture tools, so an agent can immediately identify its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, and it does not mention any exclusions or prerequisites. An agent cannot tell from the description whether this is the right choice for visual regression checking or how it relates to tools like get_element_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clear_app_dataA
Clear all data and cache for an app (equivalent to Settings → App → Clear Data). Resets the app to a fresh-install state. Useful for testing onboarding or reproducing first-launch bugs.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes | ||
| packageId | Yes | Android package ID or iOS bundle ID (e.g. 'com.example.app') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It communicates the destructive effect ('Clear all data and cache') and the consequence ('Resets the app to a fresh-install state'). However, it does not explicitly warn about irreversibility, permissions, or platform-specific behavior (especially on iOS, where 'Clear Data' is not a standard concept), leaving meaningful gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, with the action in the first sentence, followed by the reset consequence and then use cases. No fluff or redundancy; appropriately front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has no output schema, so return values need no explanation. Still, the description omits important context: it mentions an Android-specific settings path even though platform includes iOS, and it does not mention prerequisites like the app being installed. This is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only packageId is documented). The description adds no parameter-specific guidance: nothing about how deviceId is used or how platform affects behavior. It fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Clear all data and cache') and resource (an app), clarifies it as equivalent to Settings → App → Clear Data, and adds the fresh-install reset outcome. This clearly distinguishes it from sibling tools like manage_app_lifecycle or get_app_info.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a direct use case: 'Useful for testing onboarding or reproducing first-launch bugs,' which tells an agent when to use it. However, it does not explicitly name alternatives or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
configure_ocrA
Set the default OCR language for the session.
| Name | Required | Description | Default |
|---|---|---|---|
| language | Yes | Language code(s), e.g., 'eng', 'eng+fra', 'jpa'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds helpful context beyond the input schema by stating that this is the default OCR language and that the setting is session-scoped. It does not disclose invalid-language behavior, whether the change affects in-flight operations, or whether it can be reset, and there are no annotations to carry that burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence communicates exactly what the tool does with zero filler. Every word contributes to purpose or scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter setter, the description and schema together are sufficient for an agent to make the call. The only gap is absence of guidance about edge cases like invalid language codes or default reset behavior, which is minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the input schema already fully explains the language parameter with examples. The tool description itself adds no parameter-level meaning, which matches the baseline of 3 when the schema handles the semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Set, resource, scope ('the default OCR language for the session). This leaves no ambiguity about what operation is performed. It also helps distinguish configure_ocr from the many interaction/read siblings since none of them do configuration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: an agent can infer to call this tool when it needs to set the OCR language for the current session. However, there are no explicit when-to-use or alternative guidance, and no exclusions such as when OCR defaults are unnecessary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_listA
List connected active Android emulators and iOS simulators.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. 'List' implies a read-only operation and 'connected active' scopes the result, but the description does not disclose return format, empty-device behavior, or any side-effect guarantee beyond the verb.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short sentence with no filler. Every word adds meaning: 'List', 'connected active', 'Android emulators and iOS simulators' all contribute to the tool's scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter enumeration tool, the one-line description defines the core contract well. The main gaps are the unspecified return structure and absence of usage guidance, but these are minor for such a simple operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so the baseline is 4. With no parameters to describe, the description does not need to add parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific action verb ('List') and a precise resource ('connected active Android emulators and iOS simulators'). This clearly separates it from sibling tools that operate on individual devices or views, such as device_tap or get_viewport.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives like device_type or get_app_info, and no mention of exclusions. The intended use is only implied by the verb 'List', not explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_pinchA
Perform a pinch gesture (two-finger zoom) on the device. Use 'out' to zoom in (fingers spread apart) and 'in' to zoom out (fingers come together). Works on real Android phones (no root needed) via UIAutomation MotionEvent injection.
| Name | Required | Description | Default |
|---|---|---|---|
| spread | No | Max distance in logical pixels each finger travels from center (default 200) | |
| centerX | Yes | X coordinate of the pinch center in original screen pixels | |
| centerY | Yes | Y coordinate of the pinch center in original screen pixels | |
| deviceId | Yes | ||
| duration | No | Gesture duration in ms (default 500) | |
| platform | Yes | ||
| direction | Yes | 'out' = zoom in (spread fingers apart), 'in' = zoom out (fingers come together) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals the Android-specific UIAutomation injection method and that no root is needed, but it states 'Works on real Android phones' while the schema allows 'ios' as a platform enum value, creating a misleading inconsistency. It also doesn't mention whether it works on emulators or limitations when coordinates are invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is two sentences with no filler. The primary purpose is front-loaded, and the unusual direction semantics are clarified immediately in the second sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters and no output schema, the description covers core behavior and platform specifics but leaves gaps. The platform inconsistency is an important missing clarification, and it does not mention failure behavior, coordinate system edge cases, or how the tool reports results (even though no output schema exists).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71%, so most parameters are already described. The description repeats the direction mapping that is already in the schema (e.g., 'out' = zoom in) but adds no new information about parameters such as spread, duration, or coordinate units. It does not compensate for the ~29% of parameters lacking schema descriptions (deviceId, platform).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: performs a pinch gesture (two-finger zoom), and uniquely clarifies the non-intuitive direction mapping ('out' = zoom in, 'in' = zoom out). This effectively distinguishes it from sibling tools like device_swipe, device_tap, and device_rotate_gesture.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (for pinch/zoom gestures) through its explicit purpose, but does not explicitly name alternatives or provide exclusion criteria. It is clear enough given the tool name and sibling context, but lacks explicit routing guidance such as 'use this instead of device_swipe when you need a two-finger gesture'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_press_keyB
Press a hardware or system key. Also accepts raw Android keycodes as numbers.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Key name or raw Android keycode number | |
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It adds the keycode-acceptance behavior, but that is already restated in the schema's key parameter description. It omits important behavioral context such as platform-specific key support (e.g., 'back'/'home'/'recents' are Android concepts, and 'enter' may behave differently on iOS) and what happens for unsupported key/platform combinations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core action front-loaded; there is no filler. The second sentence about keycodes earns its place as a critical acceptance criterion, though it slightly duplicates the schema's key description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 required parametersFootnote, no output schema, and no annotations, the description captures the core action but leaves notable gaps: no guidance on platform/key compatibility, no behavior description for invalid or unsupported keys, and no indication of what the tool returns. An agent could easily invoke 'home' or 'back' on an iOS platform without being warned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'key' is described). The description supplements the 'key' parameter by reiterating that raw numeric Android keycodes are accepted, which adds semantic clarity. However, it offers nothing for the undocumented 'deviceId' and 'platform' parameters, so the low coverage gap is only partially compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Press') with a clear resource ('hardware or system key'), and the keycode sentence adds precision about what inputs are accepted. It functionally distinguishes from siblings like device_tap, device_type, and device_swipe, though it doesn't explicitly name any alternative or contrast itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'hardware or system key' implies this is for physical/system buttons rather than on-screen UI elements, which loosely signals when to use it over device_tap or tap_on_element. However, no explicit when/when-not conditions or alternative tools are mentioned; the guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_rotate_gestureA
Perform a two-finger rotation gesture (e.g. to rotate a map or image). Positive degrees = clockwise. Android only (uses UIAutomation, no root needed).
| Name | Required | Description | Default |
|---|---|---|---|
| radius | No | Distance of each finger from center in pixels (default 120) | |
| centerX | Yes | X center of rotation in screen pixels | |
| centerY | Yes | Y center of rotation in screen pixels | |
| degrees | Yes | Degrees to rotate. Positive = clockwise. | |
| deviceId | Yes | ||
| duration | No | Gesture duration in ms (default 500) | |
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool is Android-only, uses UIAutomation, requires no root, and defines rotation direction. This is meaningful behavioral context for an action tool, though it does not mention failure behavior or return value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three tight sentences with no filler. The core action is front-loaded, followed by the most decision-relevant details: direction and platform constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple gesture tool, the description and schema together cover the essential information: what it does, platform, direction, and required parameters. It could be more complete by mentioning the expected outcome or how failures surface, but nothing critical is missing for tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71%, so most parameters are already described. The description adds the Android-only constraint relevant to the platform parameter, and 'two-finger' clarifies the role of center and radius. However, it mostly repeats the degrees direction already present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Perform a two-finger rotation gesture' with concrete examples (rotate a map or image). This clearly distinguishes it from sibling gesture tools like device_pinch and device_swipe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: use for two-finger rotation, with examples, and it explicitly restricts the tool to Android ('Android only'). It does not name alternatives or say when to prefer another gesture tool, but the platform and gesture type are well specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_swipeA
⚠️ Low-level: Swipe from (x1,y1) to (x2,y2). Use this for custom gestures or when you need precise swipe control. Use coordinates from the physical screenshot (originalWidth/originalHeight). On Android, this tool automatically scales coordinates if a display override (logical resolution) is detected.
| Name | Required | Description | Default |
|---|---|---|---|
| x1 | Yes | Start X coordinate in original screen pixels | |
| x2 | Yes | End X coordinate in original screen pixels | |
| y1 | Yes | Start Y coordinate in original screen pixels | |
| y2 | Yes | End Y coordinate in original screen pixels | |
| deviceId | Yes | ||
| duration | No | Optional duration in ms. Default is 300. | |
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It provides important behavioral details: coordinates must come from the physical screenshot, and on Android it automatically scales coordinates when a display override is detected. This reveals meaningful platform-specific behavior beyond the raw schema, though it doesn't mention return values or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, with the core action front-loaded ('Swipe from (x1,y1) to (x2,y2)'), followed by the use case and a single coordinated note on coordinates and scaling. Every sentence earns its place, and there is no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-level gesture tool with no output schema and no annotations, the description covers the core action, coordinate source, platform scaling, and use cases. It doesn't explain return values or error handling, but those are less critical for a raw coordinate swipe. The context provided is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to the coordinate parameters by tying them to the physical screenshot's originalWidth/originalHeight, and explains Android's automatic coordinate scaling. This supplements the schema's 'original screen pixels' descriptions, and the coverage was already at 71%. The remaining parameters (deviceId, platform, duration) are either self-explanatory or have schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Swipe') and resource (from (x1,y1) to (x2,y2)), and distinguishes itself from sibling tools by explicitly labeling it 'Low-level' and targeting 'custom gestures or when you need precise swipe control'. This makes it clear how it differs from higher-level tools like device_tap or device_pinch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives direct usage guidance: 'Use this for custom gestures or when you need precise swipe control.' It implies this is not for ordinary interactions that higher-level tools handle, though it does not explicitly name alternatives or state when not to use it. The positive condition is clear, so this earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_tapA
⚠️ Low-level: Tap at raw screen coordinates. Prefer tap_on_element for reliability. Use coordinates from the physical screenshot (originalWidth/originalHeight). On Android, this tool automatically scales coordinates if a display override (logical resolution) is detected.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate in original screen pixels | |
| y | Yes | Y coordinate in original screen pixels | |
| verify | No | If true, capture and return a screenshot right after the tap. | |
| deviceId | Yes | ||
| duration | No | Optional duration in ms. If > 0, performs a long press. | |
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the low-level risk and the Android auto-scaling behavior, which are useful. It does not describe return values, failure modes, or side effects beyond the tap itself, so some behavioral context is still missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the warning and the primary alternative, then pack coordinate guidance and Android-specific behavior into the rest. Every sentence earns its place with no repetition or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-level tap tool, the description gives the crucial coordinate-source rule and platform scaling detail. It does not mention what the tool returns or how verify behaves, but the schema covers verify, and a simple tap tool does not need a longer narrative to be callable correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description adds essential meaning to x and y by specifying they come from the physical screenshot's originalWidth/originalHeight. It also adds platform-specific scaling context for Android. The remaining parameters are adequately explained in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Tap at raw screen coordinates'. It also distinguishes itself from the sibling tap_on_element by explicitly labeling itself as low-level and less reliable, so an agent can tell them apart immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly steers agents toward tap_on_element when reliability matters, and supplies the coordinate source ('physical screenshot originalWidth/originalHeight'). The guidance is practical and directional, though it does not list exhaustive when-not-to-use scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_typeC
Type text into the device.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| verify | No | If true, capture and return a screenshot right after typing. | |
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the full burden of behavioral disclosure. It only says 'Type text into the device' and does not explain whether text is appended or replaced, whether an element must be focused, or whether platform-specific behaviors exist. It is minimally honest but lacks meaningful behavioral depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: a single five-word sentence with no filler and the verb front-loaded. It loses points only because the brevity crosses into under-specification, but on a pure conciseness and structure basis it is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three required parameters, no output schema, and no annotations, the description is incomplete. It omits target semantics, return behavior, platform caveats, and the role of the optional verify flag. An agent would need to infer or inspect other tools to know how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%; only the optional 'verify' parameter is documented. The description adds no meaning to 'text', 'deviceId', or 'platform', and does not clarify how these parameters relate to the typing operation. Given the low coverage, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Type') and resource ('the device'), making the core action clear. It is distinct from sibling tools like tap_on_element and device_tap because it specifically indicates text input. However, it does not mention any target or modality nuances, so it is clear but not fully differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any mention of prerequisites such as a focused text field. The single sentence provides no exclusions, conditions, or relation to similar tools like get_screen_text or device_press_key.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_elementA
Find UI elements by selector. Returns elements with pre-parsed coordinates (centerX, centerY, left, top, right, bottom, width, height) ready for use. For tapping, prefer tap_on_element which does find+tap in one step.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes | ||
| selector | Yes | ||
| strategy | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It usefully states that the tool returns elements with pre-parsed coordinates 'ready for use,' but it does not mention whether it waits for the element, returns one or many matches, or how failures are signaled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: action first, return format second, and a routing recommendation third. The coordinate details are dense but purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and the description covers its core purpose and return format, which is important given there is no output schema. Still, the complete lack of parameter guidance and limited behavioral detail leave gaps that an agent must infer from parameter names and enum values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It only references the selector concept and does not explain deviceId, platform values, or strategy semantics such as testId, text, and contentDescription.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Find UI elements by selector,' and then clarifies the return value (pre-parsed coordinates). It also explicitly distinguishes itself from tap_on_element by noting that tool performs find+tap in one step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear guidance to prefer tap_on_element when the goal is tapping, which prevents misuse. However, it does not contrast this tool with other related siblings like wait_for_element or get_semantic_hierarchy, leaving some usage boundaries implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_app_infoA
Get version, install date, SDK target, data directory, and granted/denied permissions for an installed app. Android only.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes | ||
| packageId | Yes | Android package ID (e.g. 'com.google.android.apps.maps') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It clearly conveys a read-only inspection action ('Get') and lists the data it returns, plus an Android-only constraint. It does not mention output format or failure behavior, but those are secondary for a non-destructive info tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, no filler, with the most important scope constraint ('Android only') and the return fields front-loaded. Every word adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description tells the caller what data will be returned and the platform constraint, and all three parameters are required. However, it omits any description of deviceId and the exact output shape, and with no output schema or annotations this leaves some ambiguity for a first-time agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: only packageId has a description. The description adds 'Android only', which helps clarify the platform parameter, but it does not explain deviceId at all and adds nothing about packageId beyond the schema example. For a low-coverage schema the description should compensate more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Get') and resource ('app info') and enumerates the exact returned fields: version, install date, SDK target, data directory, and permissions. The 'Android only' qualifier clearly distinguishes this from iOS-oriented or generic device-inspection siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear context: use when you need metadata about an installed Android app. It also provides an exclusion ('Android only'), implying this tool is not for iOS. It does not explicitly name alternative tools for other cases, so it falls just short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bundler_logsA
Get recent logs from Metro bundler, Android, or iOS. To wait for a specific line, use wait_for_log in a background subagent instead of polling.
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | Log source: 'metro', 'android', 'ios', or 'all' (default 'all') | |
| deviceId | No | Restrict android/ios sources to this device's own log buffer. Useful when capturing logs for multiple devices at once; omit to merge all devices' buffers. | |
| tailLength | No | Number of lines to return (default 100) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It clearly conveys a read-only snapshot behavior ('get recent logs') and warns against polling, but it does not disclose whether the call blocks, how logs are ordered, or any source-specific behavioral nuances. This is adequate for a simple log-fetching tool but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The first sentence states the purpose, and the second adds the key usage caveat. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple log-retrieval tool with fully documented parameters and no output schema, the description covers the core use case and points to the relevant alternative. It does not describe output shape or pagination, but that is a minor gap given the straightforward nature of the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are already fully documented in the schema. The description repeats the source list but adds no additional parameter-level meaning beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is explicit: it names the action ('Get'), the resource ('recent logs'), and the exact sources ('Metro bundler, Android, or iOS'). It also differentiates itself from the wait_for_log sibling by noting that waiting for a specific line is that tool's job, not this one's.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives explicit routing guidance: if the goal is to wait for a specific line, use wait_for_log in a background subagent instead of polling this tool. This clearly states when not to use this tool and names the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_element_imageC
Get a cropped screenshot of a specific UI element.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes | ||
| selector | Yes | ||
| strategy | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden, but it only states the action. It does not disclose whether the element must already be present, whether the tool waits or retries, how failures are reported, or how the cropped image is returned, which are important behavioral traits for a screenshot tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundancy. Every word adds meaning and the structure is appropriate for a simple fetch operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too sparse for a 4-parameter tool with no annotations and no output schema. It omits the return representation, failure conditions, and any relationship to the sibling UI-inspection tools, so the agent lacks enough context to invoke it reliably in non-obvious situations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage and the description adds no parameter detail beyond referencing a 'specific UI element'. It does not clarify how selector and strategy combine, what deviceId/platform mean in context, or how the cropping is derived, leaving the agent to guess at parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Get a cropped screenshot' of a UI element, which clearly conveys the core operation and distinguishes it from sibling tools like get_viewport or capture_diff. It could be stronger by naming alternatives or noting the output type, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus sibling tools such as get_viewport, capture_diff, or find_element, and no prerequisites are mentioned (e.g., element visibility or discovery). The agent is left to infer the appropriate context from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_network_logsA
Pull network-related logs from the device. For Android, filters logcat. For iOS, filters the internal log capture buffer (must be running via manage_platform_logs or manage_bundler). For Expo / React Native apps, all console.log output (including fetchApi network traces) is emitted under the 'ReactNativeJS' logcat tag — use filter 'ReactNativeJS' or leave default. For native Android HTTP clients use 'OkHttp'. Note: a recurring warning 'ReconnectingWebSocket: Couldn't connect to ws://:8081/inspector/network' is harmless — it means Metro bundler's DevTools WebSocket is not reachable from the device (Metro not running or port 8081 blocked). The app still works; start Metro via 'npx expo start' or 'npx react-native start' on the same network to silence it.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Regex matched against logcat lines (case-insensitive). Expo/RN: 'ReactNativeJS' (all console.log including fetchApi traces). Native Android: 'OkHttp', 'Volley', 'CRONET'. Default: 'ReactNativeJS|OkHttp|Volley|CRONET' | |
| deviceId | Yes | ||
| platform | No | android | |
| tailLength | No | Number of raw log lines to scan before filtering (default 1000). Increase if recent logs have rolled out of the buffer. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the disclosure burden and does so well: it warns about the harmless ReconnectingWebSocket error, explains why it occurs, and clarifies that the app still works. It also discloses the prerequisite that iOS log capture must already be running. It does not discuss return format or buffering limits, but the operational caveats are substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but not padded: every sentence adds platform behavior, a filter recommendation, or a troubleshooting warning. It is front-loaded with the core purpose and then branches into operational detail. A light restructuring into bullets could tighten it, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter log tool with no output schema, the description covers platform differences, necessary preconditions, default filter behavior, and a common false-positive error. It could also state the output format or empty-result behavior, but the essentials an agent needs to select and call the tool correctly are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions cover filter and tailLength, and the description reinforces and expands on both: it gives concrete filter aliases ('ReactNativeJS', 'OkHttp', 'Volley', 'CRONET') and the buffer-rollover hint for tailLength. deviceId and platform remain undocumented in prose but are self-explanatory from the schema, so the description compensates for the 50% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Pull network-related logs') and then gives platform-specific meaning, which clearly separates it from sibling tools like get_bundler_logs or stream_errors. It leaves no doubt about what is returned: log lines tied to network activity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete usage context: iOS requires manage_platform_logs/manage_bundler to be running, Android reads logcat, and it recommends specific filter values per framework. It does not explicitly list exclusions or name alternative sibling tools, but the platform-scoped guidance is enough for an agent to know when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_screen_textB
Get all text visible on screen using OCR.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| language | No | OCR language code (e.g. 'eng', 'fra', 'deu'). Default: 'eng' | |
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals that OCR is used, which is a meaningful method detail, but it does not mention return structure, whether coordinates are included, behavior when no text is found, or any prerequisites such as the device screen being on or OCR being configured.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. It communicates the core purpose and method efficiently, earning every word.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only OCR tool, the description is adequate but has clear gaps. There is no output schema, and the description does not clarify the return format or the meaning of deviceId/platform. The tool is easy to understand at a high level, but not fully complete for autonomous invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description needed to compensate for undocumented parameters like deviceId and platform. It does not mention any parameter meanings; only the schema documents language and the platform enum. The description adds almost no parameter-level value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Get all text visible on screen using OCR.' This distinguishes it somewhat from siblings like get_semantic_hierarchy by specifying OCR as the extraction method, but it does not explicitly name or contrast sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'using OCR' implies this tool is for visual/text extraction from the screen rather than UI hierarchy or element lookup, but there is no explicit guidance on when to use it versus alternatives like get_semantic_hierarchy or find_element. No exclusions or alternative conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_semantic_hierarchyC
Get pruned, semantic UI hierarchy as JSON.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the output is JSON and that the hierarchy is 'pruned', but never explains what pruning removes or whether the operation requires a connected device. As a read-oriented 'get' operation this is likely safe, but nothing in the text confirms that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence is efficient and front-loaded with the action verb. However, the brevity borders on under-specification; it omits any behavioral or parameter detail while claiming to describe a meaningful operation, so the sentence does not fully earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two simple parameters, no output schema, and no annotations, the description is too thin. An agent is left without knowing what the 'pruned' JSON contains (element bounds, accessibility labels, etc.), when to prefer this over sibling tools, or what happens if the device is unavailable. The return format is stated only as 'JSON', which is vague.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds nothing about the two parameters. deviceId and platform are simple and partially self-evident from the enum, but an agent is not told where deviceId comes from (e.g., device_list) or that platform must match the device's OS. The description fails to compensate for the absent schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and a distinct resource ('pruned, semantic UI hierarchy'). The qualifiers 'pruned' and 'semantic' help set it apart from siblings like get_viewport and get_screen_text. However, it doesn't explicitly name any sibling it is not, so differentiation is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no guidance on when to use this tool versus alternatives such as get_viewport, find_element, analyze_layout_health, or get_screen_text. No context, exclusions, or mention of a device connection prerequisite is given. An agent would have to infer the appropriate call from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_viewportA
Capture screenshot of a device. Returns the image (resized to ~800px width for efficiency) and metadata with both resized and original dimensions. Use originalWidth/originalHeight for coordinate calculations when tapping. For Android, if logicalWidth/Height are provided, they represent the UI coordinate system which may differ from the physical screenshot pixels.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It transparently states that the image is resized to ~800px width for efficiency, returns both resized and original dimensions, and explains the Android logicalWidth/Height coordinate system discrepancy. This is valuable context beyond the tool name, though it doesn't mention side effects (likely read-only) or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no fluff. It front-loads the core action, then efficiently details the output format, resizing behavior, and the coordinate system caveat. Every sentence contributes directly to correct usage, and the structure is logical and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (2 simple parameters, no output schema), the description provides all essential information: what it does, what it returns (image and metadata), how the returned dimensions should be used, and a platform-specific nuance. An agent can confidently invoke this tool without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It mentions 'device' implicitly and the Android-specific note touches on the platform parameter, but it does not explicitly explain what deviceId or platform are used for. The parameters are simple and somewhat self-explanatory, but the description adds only partial contextual meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Capture screenshot of a device.' It also specifies the output (image and metadata), which distinguishes it from siblings like capture_diff or get_element_image. The purpose is unambiguous and immediately understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides implicit usage context by advising to use originalWidth/originalHeight for coordinate calculations when tapping, implying this tool is a precursor to tap actions. However, it does not explicitly mention when to use this tool over alternatives or any when-not-to-use cases. There is no comparison to sibling tools like capture_diff or get_element_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
install_depsC
Install dependencies using npx expo install.
| Name | Required | Description | Default |
|---|---|---|---|
| packages | Yes | ||
| projectPath | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reveals the underlying command but not the side effects. It does not disclose that installation typically modifies the project's manifest and lockfile, requires network access, or may fail when the project path is invalid. Since no annotations are provided, the description carries the full burden of behavioral disclosure and only restates that dependencies will be installed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition beyond the name. It is concise, though some of that brevity comes at the cost of missing behavioral and parameter context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a command-execution tool with two parameters, no annotations, and no output schema, the description is too thin. It does not provide prerequisites, project path semantics, side effects, or error behavior, leaving an agent without enough information to invoke it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema only provides types and required status. The description suggests that 'packages' refers to dependencies to install, but it does not clarify that these are package names for npx expo install, nor does it explain the meaning or default of 'projectPath'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Install dependencies') and the mechanism ('using npx expo install'), so an agent can identify what operation the tool performs. It does not explicitly differentiate from sibling tools, but none of the listed siblings appear to perform package installation, so this is not a practical concern.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives or what preconditions are needed. It does not mention that this is intended for Expo/React Native projects, whether it modifies package.json, or when agents should choose a different tool such as manage_bundler.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_app_lifecycleC
Launch, stop, install, or uninstall apps.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| target | Yes | Package ID/Bundle ID or file path | |
| verify | No | If true, capture and return a screenshot right after the action. | |
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It names mutating actions like install and uninstall but does not mention prerequisites, side effects, permission requirements, failure behavior, or what happens after an action completes. The verify screenshot option exists in the schema but is not surfaced in the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with no filler, which is good. However, it is so terse that it omits important operational context that a 5-parameter mutation tool needs, making it under-specified rather than optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with mutating operations, no annotations, and no output schema, the description should explain return values, verification behavior, target forms, and prerequisites. None of that is present, so an agent would need to inspect the schema and guess at important behavioral details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40%, with deviceId and platform undocumented. The description merely repeats the action terms and adds no guidance on when target should be a package/bundle ID versus a file path, so it fails to compensate for the schema's gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses concrete verbs—launch, stop, install, uninstall—and clearly identifies the resource as apps. It conveys the tool's core purpose well, but it does not differentiate it from overlapping siblings like open_deep_link or clear_app_data, so it falls just short of a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the action list: an agent can infer this tool is for basic app lifecycle operations. However, there is no explicit when-to-use guidance or mention of alternatives, which matters given siblings like open_deep_link also involve launching apps and clear_app_data affects app state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_bundlerB
Start, stop, or restart the Metro bundler. Platform logs (Android/iOS) auto-start by default. On Android, pass both deviceId and packageId to enable PID-based log filtering — this captures only your app's logs and eliminates all system/GMS noise.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| command | No | Optional custom command. Default: 'npx expo start'. To target a specific device use 'npx expo run:android --device <deviceId>' or 'npx expo run:ios --device <deviceId>'. | |
| deviceId | No | Device/emulator ID (from device_list). Pass with packageId to enable PID-based Android log filtering that eliminates system noise. | |
| logFilter | No | Optional raw log filter (adb logcat tags on Android, predicate on iOS) | |
| packageId | No | Package ID of the app (e.g. 'com.example.app'). Pass with deviceId on Android for precise per-app log filtering. | |
| projectPath | No | Optional path to project root | |
| showTerminal | No | Open the bundler in a new terminal window (default true) | |
| autoStartPlatformLogs | No | Auto-start platform log capture (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It adds useful context by noting that platform logs auto-start by default and that PID-based filtering removes system/GMS noise. However, it omits side effects such as what stopping or restarting does to a running session and whether starting blocks or returns immediately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: core action first, platform-log default second, and the advanced Android filtering tip last. Every sentence adds distinct value and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter control tool with no annotations and no output schema, the description is adequate but incomplete. It explains the core lifecycle and one advanced combination, but does not mention expected return behavior, failure modes, or how this tool relates to sibling log tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (88%), so the schema already documents the parameters. The description reinforces the key deviceId+packageId combination for Android, but largely repeats what the schema states, putting it at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Start, stop, or restart the Metro bundler.' This clearly identifies the tool's function. It does not explicitly differentiate from siblings like manage_platform_logs or get_bundler_logs, so it stops short of a top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose manage_bundler over alternatives such as manage_platform_logs or get_bundler_logs. The only actionable usage advice is parameter-level—the deviceId+packageId combination for Android—rather than tool-selection logic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_platform_logsA
Manually start/stop platform log capture (optional - auto-starts with bundler). On Android, pass both deviceId and packageId to enable PID-based filtering — this captures only your app's logs and eliminates all system/GMS noise. Without packageId, falls back to ReactNative/AndroidRuntime tag filtering which may include unrelated system warnings.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Start or stop log capture | |
| deviceId | No | Device/emulator ID (from device_list). Required for per-app PID filtering on Android. | |
| platform | Yes | Platform to capture logs from | |
| logFilter | No | Optional raw log filter (adb logcat tags on Android, predicate on iOS) | |
| packageId | No | Package ID to filter logs (e.g. 'com.example.app'). On Android, used with deviceId for precise PID-based filtering that eliminates system noise. | |
| projectPath | No | Optional path to project root | |
| showTerminal | No | Open the logs in a new terminal window (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It transparently explains the PID-based filtering behavior, the fallback to ReactNative/AndroidRuntime tag filtering, and the noise consequences. It does not mention return values or stop-tool edge cases, but the core behavioral branches are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no fluff. The most important information is front-loaded: what the tool does, when it is needed, and how to use the key Android parameters for optimal filtering.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter control tool with full schema coverage and no output schema, the description covers the necessary operational decisions: manual vs auto mode, Android filtering modes, and fallback behavior. It does not elaborate on iOS-specific behavior or stop semantics, but the critical guidance for correct invocation is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all 7 parameters, so the baseline is 3. The description adds value beyond the schema by explaining the relationship between deviceId and packageId and by describing what happens when packageId is omitted, which is not inferable from individual parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource pair ('start/stop platform log capture') and immediately clarifies that this is only a manual override because capture 'auto-starts with bundler'. This distinguishes it from bundler/log retrieval siblings and from the bare tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context on when manual invocation is needed by noting auto-start with the bundler, and it provides precise conditions for Android usage: pass both deviceId and packageId for PID-based filtering, or fall back to tag filtering. It does not explicitly name alternative log-reading tools, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_deep_linkC
Open a deep link or URL on the device.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| verify | No | If true, capture and return a screenshot right after opening the link. | |
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for disclosing behavior. It only states the action without mentioning prerequisites, side effects, or what happens after opening (e.g., app state changes, navigation results). This is insufficient for a tool that can alter device state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that conveys the core action efficiently. It is front-loaded and contains no redundant information, though its brevity may contribute to the lack of detail in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, no annotations, no output schema, and multiple sibling tools with overlapping capabilities, the description is far too minimal. It omits usage guidance, parameter semantics, and behavioral context, leaving the agent with insufficient information to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only 'verify' has a description). The tool description does not elaborate on what 'url', 'deviceId', or 'platform' mean or how they should be formatted. The agent must infer from parameter names alone, which is insufficient given the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (open) and the target (deep link or URL) on the device. It is specific enough to convey the core function, though it does not mention how this differs from other device interaction tools among the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as tap_on_element or device_type. The description gives no context on scenarios where deep links are appropriate or when other navigation tools should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_doctorC
Run npx expo-doctor.
| Name | Required | Description | Default |
|---|---|---|---|
| projectPath | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only names the command to run and does not mention side effects, network usage, output format, or whether the command modifies the project. This is a thin description for a tool that executes an external process.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only one short sentence and contains no wasted words, so it is concise in size. However, it is under-specified: it reads more like a command reminder than a tool definition. The brevity does not compensate for the missing context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a one-parameter tool with no annotations and no output schema, so the description needed to explain what running expo-doctor accomplishes and what projectPath is. It does neither, leaving an agent without enough context to reliably select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the only parameter, projectPath, is not described anywhere. The description does not explain how projectPath is used, whether it is optional, or what happens if it is omitted. With zero parameter description in the schema and zero mention in the description, the agent must guess the parameter's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Run npx expo-doctor' specifies a concrete action and the exact resource being invoked. It is clearly distinct from the sibling tools, several of which are device and layout operations. It loses a point because it never states what expo-doctor actually does or what outcome the agent should expect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus siblings such as run_maestro_flow, install_deps, or manage_bundler. There is no mention of project health checking, prerequisites, or when running expo-doctor is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_maestro_flowC
Run a complex Maestro flow via YAML.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| flowYaml | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavior on its own, but it only says it 'runs' a flow. It does not reveal whether the flow is executed on the device, whether it is long-running, what side effects it may have, or what happens on failure. The description conveys only the core action without behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler or redundancy, and the main action is front-loaded. However, it is concise to the point of being thin, so it earns a strong but not top score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema, annotations, or parameter documentation, the description is not complete enough. It omits what the YAML should contain, which device is targeted, and what the agent should expect after invocation. An agent would need additional knowledge to call this tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate, but it only hints that flowYaml is YAML-based. It does not explain that flowYaml likely contains the full Maestro YAML content, what deviceId refers to, or what format or constraints the parameters have. The parameter meanings are mostly left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Run a complex Maestro flow via YAML.' It identifies the mechanism (YAML) and clearly points to a flow-level operation rather than a single gesture. It does not explicitly differentiate itself from sibling tools like device_tap or device_swipe, but the wording makes the distinction reasonably inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as tap_on_element, device_type, or manage_bundler. It does not mention any prerequisites, exclusions, or conditions that would help an agent choose this over the many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_recordingC
Start screen recording on the device. Use an absolute path for localPath when stopping to ensure you can find the file.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, yet it only says that recording starts. It does not mention whether the recording must be explicitly stopped, whether starting a recording while one is already active fails, whether permissions are needed, or what side effects occur. The localPath note is forward-looking and not a behavior of this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences and is mostly concise. The first sentence is effective and front-loaded. The second sentence adds relevant workflow advice, though it is arguably misplaced in this tool's description since it refers to stopping rather than starting.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description leaves out key context: what happens when recording starts, whether a recording handle or identifier is returned, how long recording can run, and what state the device enters. The note about stopping is helpful but incomplete for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not explain deviceId or platform. The parameter names and the platform enum are somewhat self-explanatory, yet no additional meaning is added. The mention of localPath is actually confusing because localPath is not an input parameter of this tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Start screen recording on the device.' This clearly distinguishes the tool from its sibling stop_recording by the verb 'start.' However, the second sentence introduces a localPath concept that belongs to stopping, slightly muddying the focus on this tool's direct purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: call this when you want to begin capturing the device screen, then later stop with stop_recording. The guidance about an absolute path for localPath when stopping hints at the workflow, but no explicit alternatives, exclusions, or conditions are provided. The usage context is inferable but not clearly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_recordingA
Stop screen recording and save the file. Use an absolute path for localPath to ensure you can find the file.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | Yes | ||
| platform | Yes | ||
| localPath | Yes | Local destination path for the .mp4 file (e.g. C:\Users\...\recording.mp4) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of explaining behavior, and it does disclose the primary side effect: stopping the recording and writing the file to localPath. However, it omits preconditions (an active recording), overwrite behavior, and what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler; the primary action is front-loaded and the path caveat is a practical, necessary addition. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an unannotated tool with no output schema, the definition is too thin: it does not mention that a recording must already be active, what the call returns, or how deviceId and platform relate to the operation. An agent would have to guess several preconditions and confirmation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only localPath receives extra semantic guidance ('absolute path'), and that parameter is already described in the schema with a format and example. deviceId and platform are left to inference, and with schema description coverage at only 33%, the description does not compensate for the undocumented required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('stop screen recording') and states the outcome ('save the file'), so an agent immediately knows the core action. It is clearly distinguishable from the sibling start_recording even though that sibling is not mentioned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit 'use when' or 'use instead of' guidance; the intended pairing with start_recording is only implied by the tool name and sibling list. The absolute-path sentence is a parameter instruction, not tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stream_errorsA
Get recent error logs from Metro, Android, and iOS. To wait for a specific error, use wait_for_log(pattern: 'error|exception') in a background subagent instead of polling.
| Name | Required | Description | Default |
|---|---|---|---|
| deviceId | No | Restrict android/ios sources to this device's own log buffer. | |
| tailLength | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It indicates 'recent' and names the log sources, but it does not disclose whether this streams indefinitely, returns a snapshot, how errors are formatted, what happens without parameters, or any polling/retry characteristics. The wait_for_log hint implies polling is possible but does not explain the actual runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero filler. The core purpose is front-loaded, and the alternative-tool guidance is delivered immediately after in a compact second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with two optional parameters and no output schema, so the description is mostly adequate for basic invocation. However, it lacks return-format details and does not clarify tailLength semantics, which are important for an agent deciding how to call the tool correctly. The wait_for_log guidance adds useful context but does not fully compensate for the missing parameter and output behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: deviceId is described, but tailLength is an undocumented bare number. The description does not mention either parameter, so it adds no meaning beyond the schema. 'Recent' hints at tailing behavior but does not explain tailLength's meaning, default, or units.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Get'), a concrete resource ('recent error logs'), and the exact sources ('Metro, Android, and iOS'). It distinguishes this tool from sibling log tools like get_network_logs and get_bundler_logs by scope, and from wait_for_log by explicitly addressing the wait use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly says when to use this tool (when you need recent error logs) and explicitly redirects to wait_for_log(pattern: 'error|exception') for waiting on a specific error, discouraging polling. This is direct, actionable guidance with a named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tap_on_elementA
🥇 RECOMMENDED: Tap on a UI element by selector. Finds the element and taps its center automatically. Prefer over device_tap. If observing a log triggered by this tap, spawn the wait_for_log subagent BEFORE tapping. NOTE: text matching is exact — if an element renders with an emoji prefix (e.g. '🇫🇷 French A2'), passing 'French A2' will fail. Use get_semantic_hierarchy first to see the exact text, or use contentDescription/testId strategy instead.
| Name | Required | Description | Default |
|---|---|---|---|
| verify | No | If true, capture and return a screenshot right after the tap — saves a separate get_viewport round-trip when you're about to check the result anyway. | |
| deviceId | Yes | ||
| duration | No | Optional duration in ms. If > 0, performs a long press. | |
| platform | Yes | ||
| selector | Yes | Text, testId, or content description to find | |
| strategy | Yes | How to find the element: 'text' (visible text), 'testId' (accessibility ID), or 'contentDescription' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It does well by revealing that text matching is exact, that emoji prefixes can cause failures, that the tap targets the element center, that duration > 0 changes behavior to long press, and that verify captures a screenshot. A small gap is the lack of failure behavior or off-screen handling, but the disclosed caveats are highly relevant and actionable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: recommendation, behavior, sibling comparison, log-spawning order, exact-match caveat, and mitigation strategy. The key advice is front-loaded with 'RECOMMENDED' and the most important behavioral caveat appears before the trailing example. Slightly dense, but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, this description is remarkably complete. It tells the agent which sibling to prefer, when to use a different sibling, what precondition to check (exact visible text via get_semantic_hierarchy), what parameter combination to use as a fallback, and how verify optimizes the workflow. An agent has enough context to select and invoke this tool correctly without opening other tool definitions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, with deviceId and platform left to inference. The description adds meaning beyond the schema by explaining exact-match semantics for the selector, warning about emoji prefixes, and clarifying the verify option as a way to avoid a separate get_viewport round-trip. This genuinely supports parameter selection, though it does not add detail for the two unnamed parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Tap on a UI element by selector' with automatic center targeting. It also distinguishes itself from device_tap by explicitly recommending this tool over that sibling, so an agent can immediately tell what this tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Prefer over device_tap' and provides precise routing guidance: use get_semantic_hierarchy first when text matching may be affected by prefixes, or fall back to contentDescription/testId strategy. It also instructs the agent to spawn wait_for_log before tapping when observing a log triggered by the tap. This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_elementC
Wait for a UI element to appear (polls every 1s).
| Name | Required | Description | Default |
|---|---|---|---|
| timeout | No | Timeout in ms (default 20000) | |
| deviceId | Yes | ||
| platform | Yes | ||
| selector | Yes | ||
| strategy | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool polls every 1 second, but it does not describe behavior on timeout, failure, return values, or whether it waits for visibility vs. existence. This is minimal disclosure for a tool with no annotation safety net.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It is concise and directly states the core behavior. However, it is perhaps too terse given the tool's parameter complexity, so it earns a 4 rather than a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, no output schema, and no annotations, the description is notably incomplete. It omits return behavior, timeout failure semantics, and any context about how selectors or strategies are used. The schema provides some structural hints, but the description does not make the tool safely callable in varied scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20%, and the description adds no parameter-level meaning. It does not explain selector, strategy, platform, deviceId, or timeout semantics beyond the single schema description on timeout. The description fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action and resource: 'Wait for a UI element to appear'. It also adds a concrete detail about polling frequency. However, it does not explicitly distinguish itself from the sibling tool find_element, though the wait-vs-find distinction is somewhat implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives like find_element, or when waiting would be necessary. The polling detail implies a usage context, but there are no explicit conditions, exclusions, or comparisons to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_logA
Block until a log line matching a pattern appears, or timeout. Returns {matched, line, elapsed}.
ALWAYS use via background subagent — never call directly or it blocks the whole conversation:
Spawn background subagent: 'Call wait_for_log(pattern, timeout). Report result.'
Then perform the action (tap/navigate) in the main agent
Main agent continues; subagent notifies when pattern matched
Spawn subagent BEFORE the action that triggers the log, not after.
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | Which log buffer to watch (default 'all') | |
| pattern | Yes | Regex pattern to match against log lines (case-insensitive) | |
| timeout | No | Max wait time in ms (default 60000). Use at least 60000 to account for subagent startup overhead. | |
| deviceId | No | Restrict android/ios sources to this device's own log buffer. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly warns that calling directly blocks the whole conversation, explains the asynchronous subagent pattern, and describes the timeout behavior and return value. This is strong, candid transparency about a potentially dangerous blocking operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core behavior and return shape, then provides a compact numbered usage protocol. Every sentence earns its place; the repeated warning about spawning before the action is important enough to justify the slight redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the blocking behavior, the exact usage pattern, the ordering requirement, the return shape, and the timeout semantics. Combined with the fully described parameter schema, an agent has everything it needs to invoke this tool correctly in the intended async workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all four parameters, including defaults and usage notes like the 60000ms timeout recommendation for subagent startup. The main description adds operational context but does not add meaningful parameter-level meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it blocks until a matching log line appears or a timeout occurs, and explicitly lists the return shape. This distinguishes it from sibling log-retrieval tools like get_network_logs or get_bundler_logs, which fetch logs rather than wait on them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit, actionable instructions: always use via a background subagent, never call directly, spawn the subagent before triggering the action, and then perform the action in the main agent. It also includes a concrete invocation example and a warning about ordering, leaving no ambiguity about when and how to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
32 tool updates
v1.1.4- First observed
analyze_layout_health - First observed
capture_diff - First observed
clear_app_data - First observed
configure_ocr - First observed
device_list - First observed
device_pinch - First observed
device_press_key - First observed
device_rotate_gesture - First observed
device_swipe - First observed
device_tap - First observed
device_type - First observed
find_element - First observed
get_app_info - First observed
get_bundler_logs - First observed
get_element_image - First observed
get_network_logs - First observed
get_screen_text - First observed
get_semantic_hierarchy - First observed
get_viewport - First observed
install_deps - First observed
manage_app_lifecycle - First observed
manage_bundler - First observed
manage_platform_logs - First observed
open_deep_link - First observed
run_doctor - First observed
run_maestro_flow - First observed
start_recording - First observed
stop_recording - First observed
stream_errors - First observed
tap_on_element - First observed
wait_for_element - First observed
wait_for_log
TDQS
Scored across 32 tools
UI interaction tools like tap_on_element, device_tap, and find_element are differentiated by descriptions, but the logging cluster (manage_bundler, manage_platform_logs, get_bundler_logs, get_network_logs, stream_errors, wait_for_log) has substantial overlap in purpose and requires careful reading. get_semantic_hierarchy and get_screen_text also both retrieve visible content, though through different mechanisms.
Multiple conventions coexist: get_* for inspection, device_* for low-level actions, manage_* for lifecycle operations, and standalone verb phrases like tap_on_element and wait_for_log. The clusters are internally consistent, but the overall set mixes noun-first device_* names with verb-first action names, and device_list is a noun phrase rather than a verb_noun pattern.
32 tools is a heavy surface for a mobile testing server and exceeds the point where the toolset becomes hard to navigate. Several logging and UI inspection tools could reasonably be consolidated or split into sub-servers, even though most individual tools have a defined purpose.
The toolset covers core mobile testing workflows: device management, UI inspection and interaction, gestures, app lifecycle, logging, recording, deep links, and project tooling. Missing orientation control and a few native device actions are minor gaps that can be worked around with existing low-level gestures and commands.
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
MCP server for building and testing AI agents with multi-model experimentation and insights.
MCP-Native LLM Orchestration Agent
Mozark's MCP server for AI-powered app testing: device access, test automation, and QA insights.
Related MCP Servers
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables scalable mobile automation for iOS and Android through a platform-agnostic interface, allowing LLMs to interact with mobile applications via accessibility snapshots or screenshot-based inputs.1968,067 npm2Apache 2.0
- AlicenseNot gradedqualityDmaintenanceA MCP server that enables LLMs to control Android devices via ADB, supporting input, UI hierarchy, device management, and shell commands.16MIT
- AlicenseNot gradedqualityCmaintenanceMCP server that enables LLMs to control Android devices via ADB, providing tools for screen interaction and UI inspection.1MIT
- AlicenseAqualityCmaintenanceA Model Context Protocol server for ad-hoc UI testing of Android and iOS apps, enabling LLM agents to interact with mobile app UIs and react to observations.4013 npm3MIT