Skip to main content
Glama

mobilerun-mcp

An MCP server that controls phones: Android (physical, emulator, redroid; ARM or x86_64) over adb, iOS through ios-portal, and Mobilerun Cloud devices. Everything runs on the host, so nothing ARM-only has to run on the device.

94 core tools covering droidrun's mobilerun-core Device API, screen perception, gestures, apps, intents, and device control. Mobilerun agent actions and Mobilerun Cloud platform tools are detached into dedicated servers.

Architecture & Servers

  • mobilerun-mcp (Core Server, default): 94 tools focusing on mobilerun-core Device API, screen perception (read_screen, perceive_screen), input gestures, system intents, and app control.

  • mobilerun-agent-mcp (Agent Server): Detached server for mobilerun agent actions (get_state, click, type_secret), background agent tasks, and macro replay.

  • mobilerun-cloud-mcp (Cloud Server): Detached server for Mobilerun Cloud platform management (cloud devices, credentials, workflows).

  MCP client (Claude Code, Cursor, ...)
        |  stdio, or HTTP (--http)
  mobilerun-mcp ---- adb ------------------> Android device (Mobilerun Portal)
     |   '---------- mobilerun-core --------> iOS (ios-portal), Portal-HTTP-only Android,
     |                                        Mobilerun Cloud
     '-- host-side OCR (tesseract) and icon detector (OmniParser v2, onnxruntime)

Jump to: Quick start · Troubleshooting · Calling the tools · Tool reference · Configuration

Why this exists

Android MCP servers that run on the phone are convenient, but they depend on on-device native libraries built for ARM. On an x86_64 device (redroid, an emulator, an Android-x86 box) those libraries run under a translation layer that cannot execute some instructions, and the app crashes. This server keeps all the heavy lifting on the host: it reads the screen through the Mobilerun Portal accessibility service and drives gestures through the Portal and adb. Nothing ARM-only runs on the device.

It also works over any network adb works over (LAN, VPN, adb connect), not only the same Wi-Fi.

Related MCP server: Android MCP Server

Quick start

Requirements

Package

Needed for

Install

adb (platform-tools)

device access

download, apt install adb, brew install android-platform-tools

uv (or Python 3.11+)

environment and dependencies

install

mobilerun CLI

installing the Portal; run_task

uv tool install mobilerun

tesseract (optional)

OCR on screens with a sparse accessibility tree

install, apt install tesseract-ocr, brew install tesseract

Python dependencies (fastmcp, mobilerun-core[local], onnxruntime, numpy, pillow, ...) are installed with the project. The icon detector for perceive_screen(detail="full") (OmniParser v2, ~80 MB, AGPL-3.0) downloads on first use to ~/.cache/mobilerun-mcp.

You also need an Android device that adb can reach: a USB phone (enable USB debugging), an x86_64 emulator, or redroid:

docker run -itd --privileged -p 5555:5555 redroid/redroid:12.0.0-latest
adb connect localhost:5555

Install

adb devices                                   # note the serial

# Mobilerun Portal: the accessibility service the server reads the screen through
mobilerun setup -d <serial>
mobilerun ping -d <serial>                    # Portal is installed and accessible

# Server
git clone https://github.com/Hi-im-Connect/mobilerun-mcp.git && cd mobilerun-mcp
uv venv --python 3.13 .venv && uv pip install --python .venv/bin/python -e .

On Windows use .venv\Scripts\python.exe in place of .venv/bin/python.

Download the APK from Portal releases, then:

adb -s <serial> install -r <portal.apk>
# enable it: Settings > Accessibility > Mobilerun Portal, or headless (replaces other enabled services):
adb -s <serial> shell settings put secure enabled_accessibility_services com.mobilerun.portal/com.mobilerun.portal.service.MobilerunAccessibilityService
adb -s <serial> shell settings put secure accessibility_enabled 1

Register with your MCP client

Claude Code

claude mcp add --scope user mobilerun -e MOBILERUN_DEVICE=<serial> -- "$PWD/.venv/bin/python" -m mobilerun_mcp

Claude Desktop, Cursor, others: add to the client's MCP config and restart it.

{
  "mcpServers": {
    "mobilerun": {
      "command": "/path/to/mobilerun-mcp/.venv/bin/python",
      "args": ["-m", "mobilerun_mcp"],
      "env": { "MOBILERUN_DEVICE": "<serial>" }
    }
  }
}

Client

Config file

Claude Desktop (macOS)

~/Library/Application Support/Claude/claude_desktop_config.json

Claude Desktop (Windows)

%APPDATA%\Claude\claude_desktop_config.json

Cursor

~/.cursor/mcp.json

Client docs: Claude Code, Claude Desktop, Cursor.

MOBILERUN_DEVICE is optional when exactly one device is attached to adb. Every device tool also takes a device argument, so one server can drive several devices:

device / MOBILERUN_DEVICE

Target

emulator-5554, 192.168.1.20:5555, a USB serial

Android over adb (all tools)

http://host:8080 or android-http:<url>

Android through the Portal HTTP API only (MOBILERUN_ANDROID_PORTAL_TOKEN)

ios or ios:<url>

iOS via ios-portal (default http://127.0.0.1:6643)

cloud:<id> or a device UUID

Mobilerun Cloud device (MOBILERUN_CLOUD_API_KEY)

adb-only tools (dumpsys-based ones, shortcuts, files by path) return [unsupported] on the others.

HTTP instead of stdio: .venv/bin/python -m mobilerun_mcp --http serves http://127.0.0.1:4816/mcp.

Usage

Ask the agent in natural language ("open Settings and read the Android version"), or call tools directly. See Calling the tools and the Tool reference.

Troubleshooting

Symptom

Fix

[device_unreachable] no device selected

adb devices, then set MOBILERUN_DEVICE=<serial>

adb devices shows unauthorized

Accept the USB debugging prompt on the device; if it does not appear, adb kill-server and reconnect

adb devices empty or offline

Replug USB; for a container or remote device run adb connect <host>:<port>

adb: command not found

Install platform-tools, or set MOBILERUN_ADB_BIN

Mobilerun Portal is not enabled as an accessibility service

Enable it under Settings > Accessibility, or rerun mobilerun setup

mobilerun setup stalls after "Found Portal APK"

Play Protect is scanning the install (redroid, emulators with Google Play): adb -s <serial> shell settings put global verifier_verify_adb_installs 0 and ... package_verifier_enable 0, then rerun

Client lists no mobilerun tools

Restart the client and check the config path. Run .venv/bin/python -m mobilerun_mcp by hand: a FastMCP banner followed by waiting is healthy, a traceback is the cause

[stale_som_id]

An action changed the screen; call perceive_screen again

Slow first launch

Cold starts take 20 to 30 s on slow devices; launch_app waits for the app

web_search error

DuckDuckGo throttled the request; retry or set BRAVE_API_KEY

Sparse text on image-heavy screens

Install tesseract for OCR, or tap by coordinates from the screenshot

Issues: github.com/Hi-im-Connect/mobilerun-mcp/issues (include the error and adb devices output).

How it works

The agent works in a perceive, act, verify loop:

  1. perceive_screen returns a numbered list of everything tappable or readable (som_ids) and an annotated screenshot with the same numbers drawn on it.

  2. An action tool (tap, type_text, launch_app, ...) waits for the screen to settle, then returns a post_action_observation: foreground app, element count, keyboard state, the top labels on screen and whether the screen changed. That block is the verification step.

  3. som_ids describe one captured screen. After any action they are stale and the server refuses them (stale_som_id), so the agent can never tap something that has moved.

Details that matter in practice:

  • Elements without text. Icon-only buttons are numbered too, as long as the app exposes them to Android's accessibility service, which is the case for standard apps.

  • Cold starts. Launching an app waits for that app to reach the foreground. On a slow device a cold start can take 20 seconds; an app that is already on top returns immediately.

  • Gestures go through the Portal's accessibility gestures (fast, and accepted by system UI such as the notification shade), with adb input as the fallback.

  • Errors look like [code] message (hint: ...): device_unreachable, policy_blocked, stale_som_id, unknown_som_id, element_not_found, app_not_found, timeout, unsupported, invalid_argument, not_permitted, plan_incomplete.

Calling the tools

Every capability is an MCP tool: a name and a JSON arguments object. The client makes the call when you describe what you want, or you can invoke a tool by name. On the wire:

{"method": "tools/call", "params": {"name": "tap", "arguments": {"som_id": 2}}}

For every tool:

  • Optional arguments can be left out; null also means "not given".

  • device is an optional argument on every tool that touches a device (an adb serial). Leave it out to use MOBILERUN_DEVICE, or the only attached device.

  • State-changing tools (tap, swipe, type_text, launch_app, ...) wait for the screen to settle and return {"ok": true, "action": ..., "post_action_observation": {...}}.

  • Read tools return a JSON object; perceive_screen and the screenshot tools also return an image.

  • Errors come back as [code] message (hint: ...), for example [stale_som_id] ....

The post_action_observation block tells the agent what the screen looks like after the action:

Field

Meaning

foreground_app, package, activity

What is in front now

element_count

How many numbered elements are on screen

keyboard_visible

Whether the on-screen keyboard is up

top_labels

The first few labels on screen, in reading order

screen_changed

Whether the screen differs from before the action

loading_indicator_present

A spinner or progress bar is visible

settled, settle_ms

The screen stopped changing before the wait ended; how long that took

screen_changed_confidence

low when the screen never settled

seen_before

This screen was already seen N actions ago (going in circles?)

sensitive_foreground

A banking / payment / authenticator app is in front

hint

What to do next

A typical session

Search Contacts for "ali". Calls are written tool arguments; outputs are real.

1. Look at the screen.

perceive_screen {}

The reply is a JSON object with foreground_app, package, activity, keyboard_visible, screen_size, perception_tier, e ([x, y, name, flags] per som_id), mark_count, ocr_used and elements. Here e is [[56, 104, "Open navigation drawer"], [664, 104, "Search contacts"], [224, 104, "Contacts"], [360, 230, "A / Ali Omar", "l"], [632, 1096, "Create new contact"]] and elements contains:

  1 [button] "Open navigation drawer" @(56,104)
  2 [button] "Search contacts" @(664,104)
  3 [text] "Contacts" @(224,104)
  4 [button] "A / Ali Omar" @(360,230)
  5 [button] "Create new contact" @(632,1096)

The leading number is the som_id; @(x,y) is the tap point. An annotated screenshot with the same numbers is returned alongside the JSON.

2. Tap the search icon by its number.

tap {"som_id": 2}

The reply has "ok": true and keyboard_visible: true in post_action_observation. Any action invalidates the numbers: reusing som_id 2 now fails with stale_som_id until perceive_screen is called again.

3. Type.

type_text {"text": "ali"}
{
  "ok": true,
  "action": "type_text",
  "chars": 3,
  "post_action_observation": {
    "foreground_app": "Contacts",
    "package": "com.android.contacts",
    "activity": "PeopleActivity",
    "element_count": 5,
    "keyboard_visible": true,
    "top_labels": [
      "stop searching",
      "ali",
      "Clear search",
      "Ali Omar"
    ],
    "screen_changed": true,
    "loading_indicator_present": false,
    "settled": true,
    "settle_ms": 922,
    "sensitive_foreground": false,
    "hint": "Settled after the action. Judge the result from this observation ..."
  }
}

4. Verify.

verify_action {"expected": "Ali Omar"}
{
  "expected": "Ali Omar",
  "state": {"foreground_app": "Contacts", "element_count": 4, "keyboard_visible": true,
            "top_labels": ["stop searching", "ali", "Clear search", "Ali Omar"], "...": "..."},
  "verified": true,
  "evidence": "text=\"Ali Omar\"",
  "foreground": "com.android.contacts",
  "visible_text": ""
}

Tool reference

94 tools, plus adb when MOBILERUN_MCP_ENABLE_ADB=1. Every tool that acts on a device takes an optional device (adb serial, ios, cloud:<id>, or a Portal URL). [name=default] is optional. Generated by scripts/gen_reference.py.

Perception

See the screen. read_screen (text grid) or perceive_screen (annotated image) first; act by som_id.

Tool

What it does

Arguments

perceive_screen

LOOK at the screen: an annotated screenshot plus every element and its tap point.

[description] [detail] [include_image=true] [ocr=auto] [max_marks=150] [lang=eng]

read_screen

Read the screen now (waits for it to stop moving first): the screen drawn as a character grid, each element a box with its som_id and label, then a table of what can be acted on: som (tap by this), in (som_id of the smal

none

get_ui_tree

Compact accessibility tree (class, id, label, flags C/L/E/S/K/P, bounds).

[max_depth=8]

get_screenshot

Plain screenshot as an image.

none

screenshot

Plain screenshot.

[hide_overlay=false]

screenshot_path

Take a screenshot, save it as a PNG file and return the path.

none

perceive_screen {}
perceive_screen {"description": "search bar", "detail": "full"}
read_screen {}
get_ui_tree {}
get_screenshot {}
screenshot {}
screenshot_path {}

Gestures, typing and keys

Every action settles the screen and returns post_action_observation. Target with x/y, a som_id, or (mobilerun style) an index from get_state.

Tool

What it does

Arguments

tap

Tap at (x, y) or at the center of a numbered mark (som_id from perceive_screen / read_screen).

[x] [y] [som_id] [stealth=false]

double_tap

Double-tap at (x, y) or a mark.

[x] [y] [som_id]

long_press

Press and hold at (x, y), a mark (som_id) or a get_state element (index).

[x] [y] [som_id] [index] [duration_ms] [ms]

long_press_at

Long press at (x, y) (mobilerun agent action).

x y

swipe

Swipe from (x1, y1) to (x2, y2) over duration_ms / ms (default 300).

[x1] [y1] [x2] [y2] [duration_ms] [ms] [coordinate] [coordinate2] [duration]

scroll_down

Scroll the content down (reveal what is below): a centered swipe over half the screen (amount), or inside a scrollable mark (som_id).

[amount=0.5] [som_id]

scroll_up

Scroll the content up (reveal what is above).

[amount=0.5] [som_id]

scroll_left

Scroll the content left (reveal what is to the left).

[amount=0.5] [som_id]

scroll_right

Scroll the content right (reveal what is to the right).

[amount=0.5] [som_id]

scroll

Scroll the content in direction (up / down / left / right) by distance (fraction of the screen).

direction [distance=0.5] [ms=300] [verify=false]

scroll_to

Two modes.

[x1] [y1] [x2] [y2] [duration_ms=300] [text] [direction=down] [max_scrolls=8]

type_text

Type into the focused field (tap it first, or pass som_id).

text [clear=false] [submit=false] [som_id]

type

Type text (mobilerun).

text [index] [clear=false] [wpm] [stealth=false]

press_home

Press the Home button.

none

press_back

Press Back (also closes the keyboard without leaving the screen).

none

press_enter

Press Enter (submits search bars and forms).

none

open_recent_apps

Open the recent-apps overview.

none

key

Press a key by mobilerun-core name (back, home, menu, enter, delete, escape, tab, space, search, page_up, page_down, volume_up, volume_down, wakeup, media_play_pause, ...) or by Android keycode number.

name_or_code

tap {"som_id": 4}
tap {"x": 540, "y": 1200}
double_tap {}
long_press {"som_id": 4}
long_press {"index": 7, "ms": 800}
long_press_at {"x": 1, "y": 1}
swipe {"x1": 360, "y1": 1000, "x2": 360, "y2": 300}
swipe {"coordinate": [360, 1000], "coordinate2": [360, 300], "duration": 0.5}
scroll_down {}
scroll_up {}
scroll_left {}
scroll_right {}
scroll {"direction": "down"}
scroll_to {"text": "Battery"}
scroll_to {"x1": 360, "y1": 900, "x2": 360, "y2": 400}
type_text {"text": "hello", "som_id": 3, "submit": true}
type {"text": "hello", "index": 5, "clear": true}
press_home {}
press_back {}
press_enter {}
open_recent_apps {}
key {"name_or_code": "back"}

Tool

What it does

Arguments

launch_app

Open an app by name (fuzzy) or exact package_name.

[app_name] [package_name] [force=false] [package]

start_app

Start an app by id (Android package / iOS bundle id), optionally a specific activity.

[app_id] [activity] [package]

lookup_app

Search installed apps by name or package; returns ranked candidates with scores.

[app_name] [query] [limit=5]

list_apps

List installed apps (user apps only unless include_system_apps=true).

[include_system_apps=false] [include_protected_apps=false] [system=false]

list_app_deeplinks

Deep links into an app, best first.

[package_name] [app_name] [package]

resolve_deeplink

Which app would open this URI (or intent action such as android.settings.WIFI_SETTINGS)?

uri

open_deeplink

Jump straight to a screen via a URI, an app-shortcut://pkg/id from list_app_deeplinks, or an intent action.

uri [package_name] [app_name] [package]

launch_app {"app_name": "Clock"}
launch_app {"package_name": "com.android.settings", "force": true}
start_app {}
lookup_app {}
list_apps {}
list_app_deeplinks {}
resolve_deeplink {"uri": "https://example.com"}
open_deeplink {"uri": "android.settings.WIFI_SETTINGS"}
open_deeplink {"uri": "app-shortcut://com.android.settings/manifest-shortcut-wifi"}

System intents and contacts

Tool

What it does

Arguments

system_intent

One-call Android actions (action = the verb; verb= is accepted too).

[action] [verb] [hour] [minute] [seconds] [label] [phone_number] [body] [title] [start] [end] [location] [notes] [text] [subject] [destination] [mode=drive] [skip_ui=true]

resolve_contact

Find contacts by (partial) name and return their phone numbers.

name [limit=5]

system_intent {"action": "set_alarm", "hour": 7, "minute": 30, "label": "wake"}
system_intent {"action": "navigate", "destination": "Cairo Tower", "mode": "walk"}
resolve_contact {"name": "Ali"}

Notifications

Tool

What it does

Arguments

read_notifications

Current status-bar notifications, newest first, without touching the screen: key, app, title, text, action labels.

[package_name] [include_ongoing=false] [limit=20] [package]

dismiss_notification

Dismiss one notification (by key, or package/title) or every clearable one.

[key] [package] [title] [clear_all=false]

notification_action

Tap one of a notification's own buttons (reply, archive, stop...); reply_text fills an inline reply field and sends it.

action [key] [package] [title] [reply_text]

read_notifications {}
dismiss_notification {}
notification_action {"action": "list"}

Media and volume

Tool

What it does

Arguments

get_media_sessions

Active media sessions (app, playback state, title/artist) and the music volume.

[include_system=false]

media_control

Control playback in any app without touching the screen: play, pause, play_pause, next, previous, stop, rewind, fast_forward.

[command] [package_name] [action]

volume_up

Raise the music volume by steps.

[steps=1]

volume_down

Lower the music volume by steps.

[steps=1]

mute

Toggle mute on the media stream (muted=true/false forces a state).

[muted]

get_media_sessions {}
media_control {}
volume_up {}
volume_down {}
mute {}

Files

Tool

What it does

Arguments

find_files

Search the device's media index by name, newest first: images, videos, audio and documents (downloads included).

[query] [kind=any] [limit=10] [path] [max_depth=6]

open_file

Open a file in its default viewer.

[uri] [path]

find_files {}
open_file {}

Waiting and checking

Tool

What it does

Arguments

wait_for

LONG waits only (downloads, uploads, processing, status changes); gestures already settle.

[condition] [timeout_ms] [poll_interval_ms] [text] [package] [activity] [gone=false] [timeout] [interval]

watch_device_events

Collect device events for up to timeout_seconds (default 10, max 30), returning early once max_events (default 50) arrive: foreground app, keyboard, screen content, notifications posted/removed.

[timeout_seconds] [max_events=50] [duration] [interval=0.5] [kinds]

validate_action

Pre-check a planned action against the safety policy (and, for our action set, that its target exists) without doing it.

[gesture_type] [target] [action] [x] [y] [som_id] [text] [package] [app_name] [uri]

verify_action

Check an outcome against the live screen.

expected [kind=text] [timeout=3.0] [use_ocr=false]

wait_for {}
watch_device_events {}
validate_action {}
verify_action {"expected": "Settings is open"}

Plan, findings and research

Tool

What it does

Arguments

web_search

Search the web for how to do something in an app ('how to in android').

query [max_results] [topic=general] [limit=5]

set_plan

Start a plan checklist.

steps [goal] [deliverable] [target_count=0] [search_query]

mark_step

Update a plan step: pending / in_progress / done / skipped / failed.

index status [note]

record_finding

Record one item you found.

item quote

end_session

Mark the end of the task (the server keeps listening; the next call starts fresh).

[reason=agent-end] [outcome=success] [goal_type] [summary]

get_usage_guide

How to use this server well.

[topic]

web_search {"query": "wifi"}
set_plan {"steps": []}
mark_step {"index": 1, "status": "value"}
record_finding {"item": "Result 1", "quote": "exact text"}
end_session {}
get_usage_guide {}

mobilerun-core Device API

Same names and parameters as mobilerun_core.Device. Works on Android (adb or Portal HTTP), iOS and Mobilerun Cloud devices.

Tool

What it does

Arguments

ui

Raw UI snapshot (a11y_tree, phone_state, device_context, ...), as Device.ui().

[filter=true]

ui_json

The UI snapshot serialized as JSON text.

[filter=true] [indent]

ui_with_recovery

UI snapshot that retries past a dead or empty accessibility tree.

[filter=true]

capabilities

Backend, platform and the actions this device supports.

none

supports

Whether this device supports a Device action (e.g.

action

screen_size

[width, height] in pixels.

none

current_app_id

Package / bundle id of the foreground app.

none

time

The device clock.

none

find_nodes

Nodes matching every given filter (exact text/desc/resource_id/class_name, or *_contains substrings), including off-screen ones.

[text] [desc] [resource_id] [class_name] [text_contains] [desc_contains] [any_contains] [tree]

find_nodes_on_screen

Like find_nodes, limited to nodes inside the visible screen.

[text] [desc] [resource_id] [class_name] [text_contains] [desc_contains] [any_contains] [tree]

tap_text

Tap the first on-screen node whose text/description contains text.

text

tap_node

Tap the center of a node returned by find_nodes / find_nodes_on_screen.

node [stealth=true]

tap_and_wait

Tap a text (or node) and wait until the UI has been idle for idle seconds.

target [idle=2.0]

scroll_until

Scroll until a matching node is on screen; result is the node (or null).

[text] [text_contains] [any_contains] [resource_id] [direction=down] [max_swipes=10] [distance=0.35] [settle=0.5]

clear_input

Clear the focused text field.

none

assert_on

Fail unless app_id is in the foreground.

app_id

assert_text_visible

Fail unless text becomes visible on screen within timeout seconds.

text [timeout=5.0]

wait_for_app

Wait until app_id is in the foreground.

app_id [timeout=10.0] [poll=0.5]

wait_for_idle

Wait until the UI stops changing.

[timeout=5.0] [poll=0.5]

wait_for_screen_change

Wait until the UI differs from now.

[timeout=10.0] [poll=0.5]

wait_for_text

Wait until a node containing text exists (off-screen nodes count).

text [timeout=10.0] [poll=0.5]

wait_for_nodes

Poll find_nodes until something matches (or timeout, returning []).

[timeout=10.0] [poll=0.5] [text] [desc] [resource_id] [class_name] [text_contains] [desc_contains] [any_contains] [on_screen=false]

open_and_settle

Start an app and wait until it is in front and idle.

app_id [timeout=15.0] [idle=3.0]

stop_app

Force-stop an app; clear_data also wipes its data.

app_id [clear_data=false]

install_app

Install an APK (host path) on the device.

path [replace=false] [grant_permissions=true]

uninstall_app

Uninstall an app.

app_id

grant_permission

Grant a runtime permission (android.permission.*) to an app.

package permission

open_deep_link

Dispatch a deep link / intent (default action VIEW), optionally pinned to a package.

deep_link [package_name] [action]

execute_script

Run JavaScript in the foreground browser page and return its JSON result.

js

get_clipboard

The clipboard's text (Android needs the Mobilerun Keyboard as the active IME).

none

set_clipboard

Put text on the clipboard.

value

ui {}
ui_json {}
ui_with_recovery {}
capabilities {}
supports {"action": "list"}
screen_size {}
current_app_id {}
time {}
find_nodes {"text_contains": "Wi"}
find_nodes_on_screen {}
tap_text {"text": "Settings"}
tap_node {"node": {}}
tap_and_wait {"target": "Settings"}
scroll_until {}
clear_input {}
assert_on {"app_id": "com.android.settings"}
assert_text_visible {"text": "Settings"}
wait_for_app {"app_id": "com.android.settings"}
wait_for_idle {}
wait_for_screen_change {}
wait_for_text {"text": "Settings"}
wait_for_nodes {}
open_and_settle {"app_id": "com.android.settings"}
stop_app {"app_id": "com.android.settings"}
install_app {"path": "/sdcard/Download/a.apk"}
uninstall_app {"app_id": "com.android.settings"}
grant_permission {"package": "com.android.settings", "permission": "android.permission.CAMERA"}
open_deep_link {"deep_link": "https://example.com"}
execute_script {"js": "document.title"}
get_clipboard {}
set_clipboard {"value": "copied text"}

Devices and connection

Tool

What it does

Arguments

get_device_status

Battery, screen power, foreground app, size, storage, network addresses, volume.

none

list_devices

Devices you can control.

[scope=local] [state] [type] [name] [country] [page] [pageSize] [filters]

ping_device

Is the device reachable?

none

connect_device

(Re)connect adb and the Portal for a device; use after the network path came back.

none

disconnect_device

Disconnect a TCP/IP adb device (adb disconnect host:port) and drop its session.

none

setup_portal

Install and enable the Mobilerun Portal on the device (mobilerun setup); path installs a specific Portal APK.

[path]

doctor

Health check of adb, the Portal and the device (mobilerun doctor).

none

request_screen_capture_permission

Compatibility no-op: screenshots use the Portal / adb screencap, no prompt is needed.

none

echo

Returns text verbatim: a check that the MCP transport is alive (no device access).

[text] [message]

get_device_status {}
list_devices {}
ping_device {}
connect_device {}
disconnect_device {}
setup_portal {}
doctor {}
request_screen_capture_permission {}
echo {}

Compatibility

Tool

What it does

Arguments

press

Press home, back or enter (kept for old clients; prefer press_home/back/enter).

button

press {"button": "back"}

Raw adb

Only when MOBILERUN_MCP_ENABLE_ADB=1; refused while a safety policy is on.

Tool

What it does

Arguments

adb

Run an adb command against the device.

command

adb {"command": "shell dumpsys battery"}

Configuration

Variable

Default

Meaning

MOBILERUN_DEVICE

the only attached device

Default device (see the device table above)

MOBILERUN_MCP_POLICY

off

Safety policy: off, standard, strict

MOBILERUN_MCP_SCOPES

read,write

Set to read to expose only read-only tools

MOBILERUN_MCP_ENABLE_ADB

0

Set to 1 to expose the raw adb tool

BRAVE_API_KEY

unset

web_search uses Brave when set, DuckDuckGo otherwise

TAVILY_API_KEY

unset

web_search uses Tavily (synthesized answer) when set

MOBILERUN_CLOUD_API_KEY

unset

Mobilerun Cloud devices, tasks and the cloud platform tools

MOBILERUN_CREDENTIALS

config/credentials.yaml

Secrets file for type_secret (mobilerun format)

MOBILERUN_DETECTOR_MODEL

downloaded

Path to an OmniParser icon-detect .onnx

MOBILERUN_IOS_PORTAL_URL, MOBILERUN_IOS_PORTAL_TOKEN

http://127.0.0.1:6643

iOS portal for device="ios"

MOBILERUN_ANDROID_PORTAL_TOKEN

unset

Bearer token for Portal-HTTP-only Android targets

MOBILERUN_MCP_HTTP_HOST, MOBILERUN_MCP_HTTP_PORT

127.0.0.1, 4816

Address for --http

MOBILERUN_ADB_BIN, MOBILERUN_BIN

on PATH

Binary overrides

Safety policy

The policy is off by default, so an agent can sign in to accounts and use any app.

  • standard blocks banking, payment and wallet apps, authenticator apps and password managers, Luhn-valid card numbers, and fields asking for a card security code.

  • strict additionally refuses password and PIN fields and national-id numbers.

Blocked actions fail with [policy_blocked]. run_task and the raw adb tool are disabled while a policy is on, because they cannot be policed. Read mobilerun://policy for the active rules.

Limitations

  • Tested live on redroid 12 (Android 12, x86_64). Physical phones, other Android versions, iOS and cloud devices go through the same code paths but were not driven live here; the cloud tools are verified against mocked API responses.

  • Icon detection (detail="full") is a host-side guess: red boxes, not facts.

  • media_control with package_name goes to the active media session (adb cannot address one app's session); the reply says when that is a different app.

  • Notification actions and dismissal drive the notification shade, because adb cannot fire a PendingIntent. Ongoing notifications cannot be dismissed.

  • Local run_task runs the mobilerun CLI agent; outputSchema, apps, credentials, files and stealth apply to cloud tasks only. The agent's self-report can be wrong: verify on screen.

  • Launcher shortcuts come from dumpsys shortcut; Android elides the path of https shortcut URIs there, so those open the app without the exact page.

  • Volume commands succeed on redroid but have no audible effect.

Development

uv pip install --python .venv/bin/python -e ".[dev]"
.venv/bin/python -m pytest                          # unit tests, no device needed
MOBILERUN_DEVICE=<serial> .venv/bin/python -m pytest -m live     # drives a real device
.venv/bin/ruff check src tests && .venv/bin/ruff format --check src tests
  • Unit tests run against real output captured from a device (tests/fixtures; regenerate with scripts/capture_fixtures.py). The JavaScript snippets are syntax-checked with node when it is installed.

  • Live tests drive the device through an in-process MCP client and assert observable effects (foreground app, screen contents, notification state, page state), not just that a call returned. They post notifications, change the media volume and open apps, so use a scratch device.

src/mobilerun_mcp/
  adb.py portal.py          transports: adb wrapper, Portal HTTP client
  session.py observe.py     per-device state, settle-and-observe
  models.py marks.py        screen model, numbered marks, signatures
  parsers/                  pure parsers for accessibility state, dumpsys, intent filters, ...
  policy.py ledger.py       safety rules, plan ledger
  tools/                    one small module per tool group

Relationship to other projects

Independent; not affiliated with Mobilerun/droidrun or redroid.

  • mobilerun-core (Apache-2.0) is a dependency; its Device runs on this server's fast Android transport.

  • mobilerun (MIT): the agent's element indexing is adapted in src/mobilerun_mcp/agentui.py.

  • droidrun/mobilerun-mcp (Apache-2.0): the cloud tools in src/mobilerun_mcp/tools/cloud.py are a port of its tool layer.

  • Mobilerun Portal provides screen access on Android.

  • OmniParser v2 icon detector (AGPL-3.0), downloaded at runtime, not redistributed.

License

MIT. See LICENSE and NOTICE.

Available Tools

94 tools
assert_onAssert OnC

Fail unless app_id is in the foreground.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_idYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, but it does disclose the single most important trait: the tool raises/returns a failure rather than a boolean, which is non-obvious for an assertion helper. It is silent on timing (is it instantaneous?), retry behavior, timeout, and the actual error shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the failure condition and no filler. It is terse rather than wasted, though the terseness shades into under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no prose, but for a mutation-free assertion tool with no annotations and 0% parameter coverage the description should still cover synchronization semantics and the device parameter. As written it is too thin for an automation precondition check.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It conveys what app_id means (the package/identifier expected in the foreground) but says nothing about the second parameter, device, leaving half the parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific outcome: it fails unless the given app_id occupies the foreground. That is more precise than a restatement of the name 'assert_on', and an agent can infer it is a precondition check rather than a query. It does not, however, name the neighboring concepts it should not be confused with (current_app_id, wait_for_app, validate_action).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to prefer this over current_app_id (which reads the foreground app), wait_for_app (which waits for an app to appear), or validate_action. The agent must infer the usage context entirely from the word 'fail'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assert_text_visibleAssert Text VisibleB

Fail unless text becomes visible on screen within timeout seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
deviceNo
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load, and it does disclose the core behavior: the tool fails rather than merely returning, and it blocks for up to timeout seconds. It does not say how failure surfaces (exception vs. error result), whether it polls or waits once, or what happens on timeout expiry beyond 'fail'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the assertion condition and the timing constraint both land immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return-shape explanation is not required, but for a 3-parameter assertion tool with no annotations and 0% schema description coverage, leaving the device parameter and text-matching semantics unexplained is a real gap. It is minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate; it names 'text' and the timeout semantics ('within timeout seconds') but gives no matching rules for text (exact, substring, regex, case sensitivity) and never mentions the 'device' parameter or the 5-second default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (assert/fail) against a specific resource (text visible on screen) with the timeout condition, so the agent knows this is a blocking assertion, not a passive wait. It does not, however, name or contrast itself with close siblings like wait_for_text, assert_on, or verify_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this assertion versus wait_for_text, assert_on, or verify_action, nor any prerequisite (e.g., app must be foregrounded, screen state must be settled). The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capabilitiesCapabilitiesC

Backend, platform and the actions this device supports.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral-disclosure burden. It implies an informational query but does not state that it is read-only, whether it has side effects, or what permissions it requires. This is a significant gap for an agent deciding whether invoking it is safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and economical, so it is not bloated. However, it is a sentence fragment rather than a structured, front-loaded explanation, and its brevity comes at the cost of clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Still, with no annotations, 0% parameter coverage, and no usage guidance, the description is incomplete for helping an agent invoke the tool correctly in the context of many sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single optional 'device' parameter, and the description does not explain what the parameter accepts or how omitting it affects the result. The phrase 'this device' hints at device scoping but does not add real semantic meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the information categories ('Backend, platform and the actions this device supports'), which gives some sense of what the tool returns. However, it does not use a specific verb or resource action, and it does not distinguish this tool from siblings such as 'supports'. The purpose remains somewhat vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to use it, or which sibling tool is the appropriate alternative. The description provides no context for selecting it over related tools like 'supports' or 'get_device_status'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear_inputClear InputC

Clear the focused text field.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: it does not say whether the whole field is emptied or one character removed, whether it requires prior focus, what error occurs if nothing is focused, or whether the action is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is appropriate for a simple action tool. It is arguably too terse, but nothing is wasted or buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a tool with no annotations and an undocumented 'device' parameter, the description leaves the agent without enough information to invoke it confidently or predict its effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter ('device') at 0% description coverage, and the description says nothing about it — not its purpose, format, or what the null default implies. With a single undocumented parameter, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('clear') and resource ('the focused text field'), so an agent knows exactly what action is taken. It does not need sibling differentiation since no other sibling clears text input, but it also adds no scoping detail beyond the name/title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'focused text field' implies the precondition for use (a field must currently hold focus), which is useful context. However, there is no explicit when-to-use guidance, no mention of alternatives, and no statement of what happens when no field is focused.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connect_deviceConnect DeviceB

(Re)connect adb and the Portal for a device; use after the network path came back.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden of behavioral disclosure. It implies reconnection is idempotent and recovery-oriented, but does not explain side effects, prerequisites, failure behavior, or whether the operation disconnects before reconnecting.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that communicates the action and the trigger condition with no filler. Every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool with an output schema, the description is mostly adequate: it says what the tool does and when to use it. However, it leaves the optional device parameter's default behavior and the meaning of 'Portal' unexplained, so an agent may still be uncertain on some edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description only says 'for a device,' which adds minimal meaning beyond the parameter name. It does not explain what the device value should be, what null/default means, or what happens when the parameter is omitted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: (re)connecting adb and the Portal for a device, which distinguishes it from diagnostic siblings like ping_device or get_device_status. The term 'Portal' is somewhat domain-specific and unexplained, but the core verb and resource are clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage condition: use it after the network path has recovered. It does not mention alternative tools or explicit when-not-to-use cases, but the recovery context is enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

current_app_idCurrent App IdC

Package / bundle id of the foreground app.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not say the tool is a read-only query, what happens if no app is in the foreground (null? error?), or how the optional device parameter affects resolution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence fragment with no filler. It is efficient, though extremely terse given the tool's role.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, but the description leaves the optional device parameter and the no-foreground-app edge case entirely unaddressed. For a query tool with a nullable input, that gap matters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the sole `device` parameter, which defaults to null. The agent cannot tell whether omitting it targets the current device or something else.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and scope: the package/bundle id of the *foreground* app. An agent can distinguish this from list_apps or lookup_app. It lacks an explicit verb, but the noun phrase is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus siblings such as list_apps, get_device_status, or wait_for_app. Nothing tells the agent the context in which the foreground app id is useful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

disconnect_deviceDisconnect DeviceB

Disconnect a TCP/IP adb device (adb disconnect host:port) and drop its session.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the effect ('drop its session'), which is real behavioral information beyond the name, but says nothing about permissions required, behavior when the device is not connected, or idempotency of a repeated disconnect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the operation and its target front-loaded; no filler. It is arguably too lean rather than too verbose, which costs it a point on this dimension rather than earning one.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return values need not be explained. However, for a tool where the only parameter is optional with a null default, the description leaves the no-argument case (disconnect all vs. error) undefined, which is a material gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter is a bare anyOf(string, null) with default null. The description compensates by revealing the expected 'host:port' format, but it does not explain what omitting the optional device argument does.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (disconnect) and resource (TCP/IP adb device), and even names the underlying command it maps to (adb disconnect host:port). It does not explicitly name connect_device as the inverse sibling, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the 'TCP/IP adb device' qualifier tells the agent this is for network-attached devices, not USB ones, and the counterpart connect_device exists in the sibling list. No explicit when-to-use, prerequisites, or when-not guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dismiss_notificationDismiss NotificationC

Dismiss one notification (by key, or package/title) or every clearable one.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNo
titleNo
deviceNo
packageNo
clear_allNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It mentions 'clearable' but does not explain side effects, irreversibility, permissions, or how the tool handles the clear_all flag. It does not disclose whether dismissing is destructive or if it affects other system state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence, front-loaded with the main action. It avoids fluff and directly states the core functionality. However, it is so terse that it under-specifies important details, but that is a completeness issue rather than a structure issue.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters and no annotations, the description is incomplete. It does not explain the selection logic clearly, omits the device parameter, and does not describe return values or error conditions. The existence of an output schema does not compensate for the lack of parameter and usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It mentions key, package, title, and clear_all indirectly ('every clearable one'), but the phrasing 'by key, or package/title' is ambiguous about how selectors combine. The 'device' parameter is entirely unmentioned, and the logic for when clear_all applies is unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (dismiss) and the resource (notification), and distinguishes two modes: single (by key or package/title) or all clearable. It is specific and understandable, though it does not explicitly contrast with sibling tools like notification_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as read_notifications or notification_action. The description implies usage (dismissing notifications) but does not offer exclusions or context about when to prefer this over related tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

doctorDoctorC

Health check of adb, the Portal and the device (mobilerun doctor).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it does almost none of it. It doesn't say whether the check is read-only or side-effecting, whether it starts/restarts the Portal, how long it takes, or what failure modes look like — all of which matter for a multi-component diagnostic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with the purpose front-loaded; the parenthetical CLI hint adds a little orientation without bloat. Slightly terse given the tool's scope, but no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required. However, for a multi-component diagnostic with an undocumented device parameter and no annotations, the description leaves lingering questions about behavior and scoping that the structured fields do not answer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'device' parameter has 0% schema description coverage and a null default, and the description only refers to 'the device' in passing without explaining what supplying or omitting it does (e.g., check all devices vs. one). It adds essentially no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (health check) and enumerates the three things checked: adb, the Portal, and the device. It's a clear verb+resource, but it does not differentiate itself from diagnostic-looking siblings like get_device_status, ping_device, or capabilities, so an agent cannot tell which diagnostic to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. With siblings such as get_device_status, ping_device, and capabilities in the same list, the description never says when a full doctor run is preferable to a cheaper single-target check, nor does it state any preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

double_tapDouble TapC

Double-tap at (x, y) or a mark.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
deviceNo
som_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, but it only restates the tool name. It omits whether the double-tap is synchronous, how long to wait, whether it requires an active device, or what side effects occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and free of filler, but it is under-specified rather than appropriately sized for a four-parameter action tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, return values need not be explained. However, the description is not complete for a gesture tool with no annotations and undocumented parameters: usage alternatives, device semantics, and behavioral expectations are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for four parameters. The description hints that x, y, and a mark (likely som_id) are inputs, but it leaves device completely unexplained and gives no coordinate system, range, or format for the mark.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: double-tapping at coordinates or a mark. It is clearly distinguishable from unrelated siblings like volume_up or open_file, but it does not differentiate itself from tap, long_press, or tap_node, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use double_tap instead of tap, long_press, tap_node, or tap_text. No prerequisites, timing considerations, or exclusion conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

echoEchoA

Returns text verbatim: a check that the MCP transport is alive (no device access).

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
messageNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does disclose the key behavioral trait: output is verbatim input, with no device interaction or side effects. It is transparent about being a no-op diagnostic, though it says nothing about error behavior if the transport is down.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the behavior first and the purpose second, with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the tool is trivially simple. The description is nearly complete for its complexity; the only missing piece is guidance on the two input parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Two parameters ('text' and 'message') have 0% schema description coverage, and the description does not mention either one. It is unclear which parameter actually supplies the echoed value or how the two interact (defaults suggest both optional), so the description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and behavior ('Returns text verbatim') plus the operational purpose (transport liveness check). The parenthetical '(no device access)' distinguishes it from the many device-touching siblings like ping_device, read_screen, or get_device_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the tool as a liveness/transport check and rules out device access, which tells an agent when to reach for it over device tools. It stops short of naming an alternative or stating when not to use it, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

end_sessionEnd SessionA

Mark the end of the task (the server keeps listening; the next call starts fresh). reason: short summary of what was done. outcome: success (goal state verified) | partial | failure. goal_type: play_media | send_message | send_email | purchase | post | open_app | search | navigate | other. For send_message, send_email, purchase and post, success is refused unless you looked at the screen (read_screen / perceive_screen) after your last action. Success is also refused while fewer findings than the plan's target_count are recorded. failure is never refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
reasonNoagent-end
outcomeNosuccess
summaryNo
goal_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that the server keeps listening and the next call starts fresh, enumerates the accepted outcome values, and spells out the validation gates (screen read required after certain goal types, findings must meet the plan's target_count). These are exactly the behavioral traits an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The most important and surprising fact — that the server keeps listening — is front-loaded in the first clause, followed by parameter semantics. Content is dense and mostly earns its place, though the goal_type enumeration list and the refusal rules make it a fairly long block of prose rather than tightly chunked guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the refusal semantics an agent must anticipate. The remaining gap is that 'device' and the reason-versus-summary distinction are never explained, which matters for a tool that can be called against a specific target device.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for the three semantically rich parameters: reason (short summary), outcome (success/partial/failure with definitions), and goal_type (an explicit enum list absent from the schema). It does not explain the 'device' parameter or the difference between 'summary' and 'reason', leaving two of five parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource statement ('Mark the end of the task') and immediately clarifies the non-obvious scoping behavior ('the server keeps listening; the next call starts fresh'), which prevents an agent from assuming the session terminates. No sibling tool competes for this purpose, so differentiation is inherent. The definition is unambiguous about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit invocation rules — when success is refused (no screen read after send_message/send_email/purchase/post; fewer findings than the plan's target_count) and that failure is never refused. That is strong, actionable conditional guidance. It stops short of stating when not to call end_session at all (e.g., before completing the plan) or naming alternatives, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_scriptExecute ScriptC

Run JavaScript in the foreground browser page and return its JSON result.

ParametersJSON Schema
NameRequiredDescriptionDefault
jsYes
deviceNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that execution targets the foreground page and returns a JSON result, but says nothing about side effects, permissions, whether it can hang on arbitrary JS, or error behavior for a powerful script-execution primitive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the verb and target are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter script-execution tool with no annotations and no output schema, the description omits the meaning of `device`, safety/permission context, and how errors or non-JSON returns are handled. It is too thin for the tool's power.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and does not. The `device` parameter is entirely unexplained — nothing tells the agent what a device identifier means here or that a null default targets the active device. Only the `js` parameter is loosely implied by 'run JavaScript'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Run/Execute) and resource (JavaScript in the foreground browser page), which distinguishes it from UI-interaction siblings like tap or ui_json. However, it does not cross-reference any sibling tool, so selection between it and other script/UI tools is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to prefer running script versus using the dedicated UI tools (ui, tap_text, find_nodes), and no stated prerequisites or exclusions. The agent is given no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_filesFind FilesA

Search the device's media index by name, newest first: images, videos, audio and documents (downloads included). kind: image | video | audio | document | any. Returns content:// URIs for open_file (limit default 10, cap 25). With path, searches that shared-storage folder by file name instead (any file type).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoany
pathNo
limitNo
queryNo
deviceNo
max_depthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses useful behavior: result ordering (newest first), a default limit of 10 with a hard cap of 25, and that results are content:// URIs. It does not mention permission requirements, empty-result behavior, or cost/performance, so it is not complete, but it goes well beyond the bare schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense, front-loaded and free of filler, with the primary mode stated first and the path-based alternative last. The telegraphic 'kind: image | video | ...' fragment is compact but slightly terse for a prose description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required, and the tool is a moderately complex 6-parameter search. However, two parameters (device, max_depth) are undocumented in both schema and description, and the path-vs-index interaction could be clearer, so an agent lacks full information to invoke it precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does for most parameters: 'kind' enum values (image | video | audio | document | any), 'limit' default and cap, 'path' semantics, and 'query' implied by searching 'by name'. 'device' and 'max_depth' receive no explanation anywhere, leaving two of six parameters opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Search the device's media index by name') plus scope (images, videos, audio, documents) and ordering (newest first). It is clearly distinguishable from screen/UI siblings like find_nodes or perceive_screen, and it names the downstream consumer (open_file) of its results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: default search covers media types with downloads included, and supplying 'path' switches to a shared-storage folder search of any file type. It links output to open_file as the follow-up tool, but never states when NOT to use this versus alternatives like web_search or list_apps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_nodesFind NodesA

Nodes matching every given filter (exact text/desc/resource_id/class_name, or *_contains substrings), including off-screen ones.

ParametersJSON Schema
NameRequiredDescriptionDefault
descNo
textNo
treeNo
deviceNo
class_nameNo
resource_idNo
any_containsNo
desc_containsNo
text_containsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two meaningful behaviors: filters are ANDed ('matching every given filter') and off-screen nodes are included. It says nothing about read-only safety, result ordering, duplicates, or behavior when no filters are supplied (all nine parameters are optional).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the resource and matching rule, with the off-screen qualifier placed where it reads as a differentiator. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation here. However, for a nine-parameter tool with zero schema descriptions and no annotations, the description omits the 'tree' and 'device' inputs and gives no explicit routing guidance against find_nodes_on_screen or get_ui_tree.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does for most of the surface: it names text, desc, resource_id, class_name and the *_contains variants and explains exact-vs-substring semantics. It leaves two parameters, 'tree' and 'device', completely unexplained, which is a notable gap for an object-typed parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (find) and resource (nodes) plus the matching semantics: exact match on text/desc/resource_id/class_name or substring via *_contains. The clause 'including off-screen ones' implicitly distinguishes it from the sibling find_nodes_on_screen, though it never names that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the right tool when it wants nodes outside the current viewport, and that all filters are conjunctive. There is no explicit statement of when to prefer find_nodes over find_nodes_on_screen, get_ui_tree, or wait_for_nodes, nor any mention of required preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_nodes_on_screenFind Nodes On ScreenC

Like find_nodes, limited to nodes inside the visible screen.

ParametersJSON Schema
NameRequiredDescriptionDefault
descNo
textNo
treeNo
deviceNo
class_nameNo
resource_idNo
any_containsNo
desc_containsNo
text_containsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing beyond scope. It says nothing about what the operation returns, performance, truncation, or permission requirements for a 9-parameter query tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the scope constraint front-loaded and no filler. It is efficient, though it is arguably too terse to earn maximum credit for a 9-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but with no annotations and zero schema descriptions on 9 parameters, the description leaves the agent without the semantics needed to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the schema provides no meaning. The description adds only that results are screen-limited, giving no guidance on desc, text, tree, device, class_name, resource_id, or any of the *_contains filters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear operation (find nodes) and scopes it to the visible screen, which distinguishes it from the sibling find_nodes. However, it defines itself relative to find_nodes rather than describing what a 'node' is or what is returned, so an agent must already understand find_nodes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'limited to nodes inside the visible screen' implicitly tells when to prefer it over find_nodes (when off-screen nodes are irrelevant), but there is no explicit when/when-not guidance or mention of alternatives by name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_clipboardGet ClipboardC

The clipboard's text (Android needs the Mobilerun Keyboard as the active IME).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full behavioral burden. It usefully discloses the Android requirement for the Mobilerun Keyboard as active IME, but omits other behavioral details such as read-only nature, error conditions when clipboard is empty, or permission needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, front-loads the resource, and includes the Android prerequisite without wasted words. It is a sentence fragment, but structure is efficient for the small amount of content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, and the output schema exists, so return values need not be explained. However, the description omits any explanation of the optional 'device' parameter and provides no routing to the set_clipboard sibling, leaving gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'device' has 0% schema description coverage and the description does not mention it at all. No meaning is added beyond the raw schema type and default, so the description fails to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource returned ('the clipboard's text'), making the tool's purpose clear despite lacking an explicit verb. It adds an Android platform caveat but never contrasts with the sibling set_clipboard, so full sibling differentiation is absent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not say when to use this tool versus alternatives like set_clipboard or other clipboard-related operations. The Android IME note is a prerequisite, not usage guidance for selection between tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_device_statusGet Device StatusB

Battery, screen power, foreground app, size, storage, network addresses, volume.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it does convey what the report contains (a state snapshot across seven dimensions), which is useful. However, it does not disclose prerequisites (e.g., a connected device), the meaning or behavior of the null device default, failure modes, or confirm the read-only nature beyond what the name 'get' implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is eight words long with zero wasted tokens; every term carries information about the returned status. For a one-parameter read tool this is an appropriately minimal size, though the lack of a verb is a structural weakness already accounted for in purpose clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complexity is low: one optional parameter, an existing output schema, and a clear field list, so an agent can likely invoke it correctly with no arguments. Still, there are clear gaps — no usage routing against the many siblings and no explanation of the device parameter or null behavior — making it adequate but incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description was expected to compensate by explaining the device parameter, yet it is entirely silent on it. The schema's string/null type and default null weakly suggest an optional device selector, but the description adds no meaning about device identifiers or what null resolves to, so it misses the only parameter entirely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a fragment listing seven concrete status dimensions (battery, screen power, foreground app, size, storage, network addresses, volume), so it clearly conveys what the tool reports and avoids tautology. The verb must be inferred from the tool name since no verb is present, and the field list implicitly separates this state/health read from pixel- or UI-tree-focused siblings like get_screenshot, read_screen, and get_ui_tree, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives such as read_screen, get_ui_tree, ping_device, or volume_up. It names no alternatives, gives no exclusions, and never states that this is the read-only status query among the many action-oriented sibling tools; the intended usage is only implied by the tool name and the field list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_media_sessionsGet Media SessionsC

Active media sessions (app, playback state, title/artist) and the music volume.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
include_systemNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the behavioral disclosure burden. It lists the output contents but does not say whether the operation is read-only, whether device defaults to the active device, what include_system changes, or whether any side effects occur. Basic output information is present, but behavioral context is largely missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single terse sentence with no filler and the core resource is front-loaded. However, it is so abbreviated that it reads like a fragment rather than a complete sentence, and it omits parameter context. Still, the structure itself is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description leaves the two optional parameters unexplained and provides no usage context or behavioral notes. It is minimally acceptable for a simple getter, but for a tool that accepts a device selector and an include_system flag, the description is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention either parameter, device or include_system, at all. The parameter names offer limited hints, but the description adds no meaning about device targeting or what including system sessions would mean. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (media sessions and music volume) and the specific payload fields (app, playback state, title/artist), which clearly distinguishes it from sibling control tools like media_control or volume_up. It lacks an explicit verb like 'retrieves' or 'returns', but the tool name supplies the action, so the purpose is still clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus siblings such as media_control, volume_up, or get_device_status. The description implies it is a read action, but it does not state exclusions, prerequisites, or alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_screenshotGet ScreenshotC

Plain screenshot as an image.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only says the result is an image; it does not mention permissions, device applicability, side effects, failure modes, or whether the screenshot reflects the current screen. This is minimal disclosure for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with its main point, but it is under-specified rather than appropriately concise. The single sentence sacrifices essential usage and parameter information for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter, no output schema, and no annotations, the description should explain the device parameter and distinguish the tool from similar screenshot siblings. It does neither, leaving the agent without enough information to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one optional device parameter with 0% schema description coverage, and the description does not mention this parameter at all. The description completely fails to explain what device means, how it should be specified, or what the default behavior is.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool produces a screenshot as an image, giving the agent the resource and output form. However, it does not distinguish get_screenshot from sibling tools like screenshot, screenshot_path, or browser_screenshot, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives such as screenshot, perceive_screen, or read_screen. The word 'plain' weakly implies a raw image without analysis, but there is no explicit when-to-use or when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_ui_treeGet Ui TreeC

Compact accessibility tree (class, id, label, flags C/L/E/S/K/P, bounds).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
max_depthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It adds that the tree is 'compact' and lists output fields, but does not disclose how accessibility data is obtained, what happens when the tree is unavailable, whether permissions are needed, or how max_depth affects results. These are significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence with no filler and the key output concept is front-loaded. The field list is compact and informative, though the brevity comes at the cost of missing guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return structure, so explaining field contents is less critical. However, with zero annotations and zero parameter documentation, the description is too sparse to fully support correct invocation; it lacks device handling, depth semantics, and sibling differentiation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented device and max_depth parameters. It does not mention either parameter, leaving an agent unable to infer device selection or depth behavior from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (accessibility tree) and specifies the exact contents: class, id, label, flags, and bounds. It is not a tautology, but it does not differentiate from sibling tools like read_screen or perceive_screen.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives such as get_screenshot, read_screen, or perceive_screen. The description provides no context about suitable scenarios or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_usage_guideGet Usage GuideB

How to use this server well. Topics: overview, shortcuts, text_entry, failures, ledger, safety, stop, efficiency, full.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, but for a read-only documentation tool the risk surface is small. It does disclose the set of valid topic values, which is real behavioral information; however, it never states that the call is non-destructive, whether unknown topics error, or what the default (null) topic yields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the purpose followed by the topic list; no filler. The raw enumeration is dense but each token is the valid input value, so it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return formatting need not be described. Still missing for a zero-annotation, zero-coverage-parameter tool: behavior on omitted or invalid topic, and any hint of when this guidance is worth fetching.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single topic parameter has no schema description, no enum, and accepts any string. The description partially compensates by enumerating the valid topic names (overview, shortcuts, text_entry, failures, ledger, safety, stop, efficiency, full), but it does not say which topic is default or what happens for an unrecognized value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states what the tool returns ('How to use this server well') and enumerates the available topics, so an agent knows this is a self-documentation/help lookup rather than an action tool. It is distinguishable from siblings like capabilities or supports, though it doesn't explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'How to use this server well' — call it when guidance is needed — but there is no explicit when-to-use statement, no indication of when to call this versus capabilities/supports, and no note that it is safe to call speculatively.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grant_permissionGrant PermissionC

Grant a runtime permission (android.permission.*) to an app.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
packageYes
permissionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral disclosure burden. It says a grant occurs, but does not state whether the grant persists, whether it requires special device privileges, whether it can fail for undeclared permissions, or what side effects occur. It discloses only the basic mutation target.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no redundant clauses. It is appropriately terse for a simple grant action, though the brevity contributes to the missing guidance elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation tool with no annotations and no parameter descriptions, the description is incomplete. An output schema exists, so return values need not be explained, but the lack of device-parameter mention and prerequisite behavior leaves significant gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds useful meaning for the permission parameter by giving the android.permission.* format and implies the package parameter via 'to an app', but it completely omits the optional device parameter and any package format details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Grant) and resource (runtime permission) and scopes it to an app via the android.permission.* format. It clearly distinguishes itself from the related request_screen_capture_permission tool, though it does not explicitly name that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as request_screen_capture_permission or general app-management siblings. The description gives no prerequisites, no when-not conditions, and no indication of required state (e.g., app must be installed, permission must be declared).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_appInstall AppC

Install an APK (host path) on the device.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
deviceNo
replaceNo
grant_permissionsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It notes the APK path is a host path (useful), but says nothing about whether an existing install is replaced, that grant_permissions defaults to true, whether this is destructive, or what permissions/auth are needed for a mutating install operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, and the host-path clarification is placed inline where it matters. It is efficient, though the brevity contributes to the specification gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for a 4-parameter mutation tool with no annotations and 0% schema coverage, the description omits the behavioral and parameter details an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters. The description clarifies only that 'path' refers to a host-side APK; 'device', 'replace', and 'grant_permissions' (and its default of true) are left completely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Install) and resource (an APK on the device), which is clear enough to separate it from siblings like uninstall_app, launch_app, and start_app. It does not explicitly name or contrast those siblings, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use install_app versus launch_app/start_app, nor any prerequisite context such as whether the device must be connected or whether the app must be absent first. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

keyKeyC

Press a key by mobilerun-core name (back, home, menu, enter, delete, escape, tab, space, search, page_up, page_down, volume_up, volume_down, wakeup, media_play_pause, ...) or by Android keycode number.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
name_or_codeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, but it only enumerates accepted inputs. It says nothing about side effects (key events injected into the active app/window), whether a device must be specified, or error behavior for unrecognized names.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the core action first and the accepted value formats after. No filler, though the trailing ellipsis leaves the name list open-ended.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the primary input's format is explained. However, for a tool with 0% schema coverage and no annotations, the unexplained device parameter and absent routing guidance leave it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are two parameters. The description does add real meaning for name_or_code (enumerated mobilerun-core names and the Android keycode number form), which the schema alone would not convey. It says nothing about the 'device' parameter, so it only half-compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (press) and resource (a keyboard/device key) and enumerates the accepted key names plus the Android keycode alternative. It does not, however, differentiate itself from the many sibling press tools (press, press_home, press_back, press_enter, volume_up, volume_down).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this generic key tool versus the specialized siblings press_home, press_back, press_enter, volume_up or volume_down. No prerequisites or context are given; the agent must infer all routing decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

launch_appLaunch AppA

Open an app by name (fuzzy) or exact package_name. An ambiguous name returns ranked candidates instead of guessing. If the app is already in the foreground it is left as is (already_foreground=true) unless force=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNo
deviceNo
packageNo
app_nameNo
package_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does a good job: it discloses fuzzy matching, the disambiguation strategy (ranked candidates rather than a guess), and the idempotent foreground check with the force override. It omits failure modes (app not installed / no match) and any permission considerations, which are the notable remaining gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary action and lookup modes before the edge-case behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. For a launch/mutation-style tool with zero annotations, the description covers matching semantics, disambiguation, and the foreground/force interaction adequately; only the device parameter and error behavior are left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate, and it only partially does: it clarifies that app_name is fuzzy and package_name is exact, and explains force. It says nothing about device, and the schema also exposes an undocumented 'package' parameter that the description never mentions or reconciles with package_name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Open an app') and names the two accepted lookup forms (fuzzy app_name vs. exact package_name). It does not distinguish itself from the sibling start_app or lookup_app, so an agent still has to infer which launcher to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful conditional behavior (ambiguous names return ranked candidates instead of guessing; already-foreground is a no-op unless force=true), which implicitly tells the agent when the call is safe. However, it never says when to prefer this over start_app, open_deeplink, or open_and_settle, all of which appear to overlap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_appsList AppsA

List installed apps (user apps only unless include_system_apps=true). include_protected_apps is honoured on Mobilerun Cloud devices only (as in mobilerun-core).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
systemNo
include_system_appsNo
include_protected_appsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and does disclose the default filter (user apps only) and a platform-dependent limitation (include_protected_apps only honoured on Mobilerun Cloud). It omits other relevant behavior: whether a device must be connected, whether the result set is paginated, and error behavior on unreachable devices.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, both front-loaded with the most decision-relevant facts (default scope, then the platform caveat). No filler, though the protected-apps aside is parenthetical and slightly cryptic in its mobilerun-core reference.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. For a four-parameter read-only listing tool, the description covers the default filtering behavior and the one platform-specific constraint; the remaining gap around the `device` parameter is minor given the schema's anyOf/default null structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so all four parameters lack inline documentation, and the description compensates for only part of that: it explains include_system_apps and include_protected_apps semantics. The `device` parameter (format/selection) is never addressed, and the relationship between the `system` and `include_system_apps` booleans is left unreconciled despite looking redundant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('List installed apps') and immediately qualifies the default scope (user apps only). It is distinguishable from siblings like lookup_app, install_app, and list_app_deeplinks, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives the condition that changes default behavior (include_system_apps=true) and notes the cloud-only constraint for protected apps, which is useful operational guidance. It does not, however, say when an agent should prefer this over lookup_app or list_app_deeplinks, leaving sibling selection implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_devicesList DevicesA

Devices you can control. scope: local (adb devices) | cloud (Mobilerun Cloud, needs MOBILERUN_CLOUD_API_KEY) | all. Cloud filters: state (creating, assigned, ready, terminated, ...), type, name, country, page, pageSize (or a filters dict). Any listed id works as the device argument of every tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
pageNo
typeNo
scopeNolocal
stateNo
countryNo
filtersNo
pageSizeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, this description carries the full burden and does well: it discloses the auth prerequisite for cloud scope and, importantly, that any listed device id is valid as the device argument of every other tool — a non-obvious behavioral contract. It stops short of describing pagination behavior or result volume.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource, then scope semantics, then filters, then the cross-tool id note. It is dense and largely waste-free, though the pipe-delimited fragments read slightly telegraphically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-param tool with no annotations, the description covers scope, auth, filters, and the downstream use of returned ids; an output schema exists so return values need not be spelled out. Only pagination semantics and default behavior with no arguments are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: it names scope values, the cloud filter fields (state, type, name, country, page, pageSize, filters dict), and enumerates several state values. Only the precise behavior of the filters dict and defaults are left implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource ('Devices you can control') and, paired with the name list_devices, the operation is unambiguous. It does not need to differentiate from siblings since no other tool enumerates devices, but the fragmentary phrasing never states the verb outright, keeping it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent how to choose between the local, cloud, and all scopes, and flags that cloud requires MOBILERUN_CLOUD_API_KEY. It offers no exclusions or negative guidance (e.g. when not to call it), so it lands at a clear-context 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

long_pressLong PressB

Press and hold at (x, y), a mark (som_id) or a get_state element (index). Hold time: duration_ms or ms (default 1000).

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
msNo
indexNo
deviceNo
som_idNo
duration_msNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does surface the default hold time (1000 ms) and both duration parameter aliases. However, it says nothing about what a long press actually triggers (context menus, deletions, drag initiation), permission needs, or the device parameter's effect, which is a meaningful gap for an interaction tool with destructive potential.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and targeting options before the timing detail. Every clause carries information not available elsewhere in structured form.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. But for a 7-parameter, zero-coverage, annotation-free interaction tool, the description leaves 'device' undocumented and does not clarify mode precedence or the behavioral consequence of the hold, leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does for most parameters: x/y coordinates, som_id as 'a mark', index as 'a get_state element', and critically the ms/duration_ms alias and default of 1000 ms, none of which the schema conveys. It omits the 'device' parameter entirely, and does not state that the targeting modes are mutually exclusive alternatives.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Press and hold' with three concrete targeting modes (coordinates, som_id mark, get_state index). It is clear on its own, but it never distinguishes itself from the near-identical sibling 'long_press_at', which an agent must choose between.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement of which targeting mode takes precedence, and no routing to or away from alternatives such as 'press', 'tap', or the sibling 'long_press_at'. The agent must infer usage entirely from the gesture name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

long_press_atLong Press AtC

Long press at (x, y) (mobilerun agent action).

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it discloses nothing about press duration, whether a context menu or drag is triggered, permission requirements, or reversibility. The critical behavioral property of a 'long press' (how long, what it triggers) is absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the action. The trailing '(mobilerun agent action)' parenthetical is filler that adds no selection value but costs little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but the description still omits the undocumented 'device' parameter and any behavioral detail for a mutation-like gesture tool with zero annotation coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters. The description only echoes the required x/y coordinates and says nothing about the optional 'device' parameter, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (long press) and resource (a coordinate pair), which implicitly separates it from the node-based 'long_press' sibling. Sibling differentiation is implied by '(x, y)' rather than stated, but an agent can infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus 'long_press', 'tap', or 'double_tap'. No prerequisite or context conditions are given beyond the required coordinates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_appLookup AppC

Search installed apps by name or package; returns ranked candidates with scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
deviceNo
app_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden, and it delivers only that results are ranked with scores. It does not state read-only nature, permission requirements, whether an empty result is an error, or any rate/behavioral constraints. Ranking/scores is a genuine addition but far short of what an unannotated tool needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with the search target front-loaded and the return behavior appended. Nothing is wasted, though it is arguably too sparse for a four-parameter tool rather than elegantly brief.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But with zero annotation coverage and four entirely undocumented parameters, the description omits the parameter disambiguation an agent needs to call this correctly, leaving it materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for four parameters (limit, query, device, app_name). The description's 'by name or package' loosely gestures at query/app_name but never clarifies the difference between those two overlapping inputs, what device scopes, or what limit controls. It does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) plus resource (installed apps) and scope (by name or package), and discloses the return shape (ranked candidates with scores). It implies a distinction from list_apps through 'search' and 'ranked candidates', but never names that sibling or any other, so the differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no mention of alternatives such as list_apps (enumerate everything) or launch_app/start_app (act on a chosen app). The reader must infer from 'search' that this is the fuzzy-name resolution step before launching.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark_stepMark StepA

Update a plan step: pending | in_progress | done | skipped | failed. Put facts you read off the screen in note; the pixels are gone next turn.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
indexYes
deviceNo
statusYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that screen pixels are ephemeral ('the pixels are gone next turn'), which is a useful context for note-taking. However, it does not disclose other behavioral aspects such as side effects of updating a step, whether the change is reversible, or any permission requirements. It adds some value but remains thin for a state-changing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words. The purpose is front-loaded, and the note guidance is appended efficiently. Every sentence adds value, and the structure is clean and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 4 parameters, no schema descriptions, and no annotations. The description covers purpose and the ephemeral screen context, but it omits parameter meanings for index and device, and doesn't describe expected behavior beyond the status update. While an output schema exists (which might define return values), the input semantics are incomplete. For a tool of this complexity, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain parameter meanings. It covers 'status' by listing allowed values and explains 'note' as a place to put facts from the screen. However, it does not explain 'index' (likely the step number) or 'device' (which device to target), leaving two of four parameters ambiguous. The partial coverage is insufficient for a tool with no schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Update a plan step' with a specific list of valid status values (pending, in_progress, done, skipped, failed). This is a specific verb-resource pair that distinguishes it from sibling tools like set_plan (which likely creates a plan) and record_finding (which records findings). The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (when updating a plan step's status) but provides no explicit guidance on alternatives or when not to use it. It does offer a practical hint about using 'note' for facts read off the screen, which is a usage consideration, but it doesn't contrast with other plan-related tools like set_plan or record_finding. Clear context exists, but no exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

media_controlMedia ControlB

Control playback in any app without touching the screen: play, pause, play_pause, next, previous, stop, rewind, fast_forward. Goes to the active media session; package_name (from get_media_sessions) is checked against it and reported.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNo
deviceNo
commandNo
package_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a real behavioral trait — actions are dispatched to the *active* media session, and a supplied package_name is validated against that session and reported — but says nothing about failure behavior when no session is active or mismatched, or about prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary capability, and the compact action list is tightly packed. No filler, though the enumerated verbs would be more useful if tied to the correct parameter name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is not needed. However, with 4 undocumented parameters and an unresolved action/command ambiguity, the definition leaves an agent guessing about the exact call shape despite covering the high-level capability.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It usefully defines the accepted action vocabulary and explains the origin of package_name (from get_media_sessions), but it never disambiguates the ambiguous 'action' vs 'command' parameters, and leaves 'device' unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Control playback in any app') and enumerates the exact supported actions, so an agent immediately knows it is the media-session control tool. It is clearly distinct from the volume_* siblings, though it does not name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'control playback in any app without touching the screen', and it helpfully points at get_media_sessions as the source of package_name. It does not state when not to use it, nor how it relates to siblings like volume_up/mute.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

muteMuteB

Toggle mute on the media stream (muted=true/false forces a state). Muting remembers the previous level for unmute.

ParametersJSON Schema
NameRequiredDescriptionDefault
mutedNo
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses one genuinely useful behavioral trait—that the previous level is remembered for unmute—and clarifies toggle vs. forced state. It does not say which device/session is affected by default, whether it requires permissions, or what happens on an inactive stream.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and immediately followed by the forcing semantics and the unmute memory behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for an unannotated mutation-style control, the description omits device targeting, which stream/session is affected by default, and any permission or failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain both parameters. It does a good job on 'muted' (null = toggle, true/false = force), but the 'device' parameter is never mentioned, leaving half the input undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (toggle mute) and resource (the media stream), and clarifies that muted=true/false forces state rather than toggling. It does not, however, distinguish itself from siblings such as media_control, volume_up, or volume_down.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the resource named ('the media stream') and the toggle-vs-force distinction, but there is no explicit statement of when to prefer this over media_control or the volume tools, and no prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notification_actionNotification ActionB

Tap one of a notification's own buttons (reply, archive, stop...); reply_text fills an inline reply field and sends it. Best-effort: it drives the notification shade.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNo
titleNo
actionYes
deviceNo
packageNo
reply_textNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that the operation is best-effort and drives the notification shade, and that reply_text sends an inline reply. However, it omits important behavioral context such as prerequisites, failure modes, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, followed by a clarifying example and a brief caveat. Every sentence contributes meaningful information without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six parameters, zero schema descriptions, and no annotations, this description is incomplete. An agent cannot confidently determine how to identify the target notification using key/title/package, nor understand the limitations of the best-effort behavior beyond a vague hint.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains reply_text and gives examples for action. The key, title, device, and package parameters are left completely unexplained, which is a significant gap since they likely determine which notification is targeted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Tap one of a notification's own buttons' with concrete examples like reply, archive, and stop. This clearly differentiates it from generic tapping or notification dismissal tools, though it does not name a sibling tool explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: use when you need to click a notification's embedded action button, and reply_text is for inline replies. However, there is no explicit guidance on when not to use it or which alternative to choose, such as dismiss_notification or type_text.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_and_settleOpen And SettleB

Start an app and wait until it is in front and idle.

ParametersJSON Schema
NameRequiredDescriptionDefault
idleNo
app_idYes
deviceNo
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, but it does disclose real behavior: the call blocks until the app is foregrounded and idle. It omits failure modes (what happens on timeout, app not installed, permission needed) and does not explain what 'in front' or 'idle' concretely mean, which matters for a tool with side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the primary action verb comes first and the wait semantics follow. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a mutating four-parameter tool with zero annotation coverage and 0% schema description coverage, the definition is thin: it never explains the idle/timeout/device parameters, the failure behavior, or how it differs from sibling start/wait tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and none of the four parameters are documented anywhere. The description's mention of 'idle' loosely hints at the idle parameter but gives no units, thresholds, or meaning, and says nothing about idle/timeout interplay or the device selector.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start an app') plus the completion condition ('wait until it is in front and idle'), which is clearer than a bare name restatement. However, it does not distinguish itself from close siblings such as start_app, launch_app, and wait_for_idle, leaving the agent to guess which entry point is intended.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given, and no alternatives are named despite obvious overlap with start_app, launch_app, wait_for_app, and wait_for_idle. The agent must infer that this is a combined start-then-wait convenience tool from the description alone, which is exactly the kind of routing the description should make explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_fileOpen FileB

Open a file in its default viewer. uri: a content://media/... URI from find_files (never build one by hand); path: a file under shared storage.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriNo
pathNo
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It only names the viewer behavior; it does not disclose permission requirements, whether it launches an external app and disrupts current app state, or what side effects follow. That is thin for an action-launching tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the action and followed by per-parameter guidance. Every clause carries information; there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. However, the undocumented 'device' parameter and the absence of any side-effect or permission context leave gaps an agent would still need to resolve before calling confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It usefully documents uri (content://media/... format originating from find_files) and path ('a file under shared storage'), but the third parameter 'device' is undocumented, leaving part of the interface unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Open a file in its default viewer'), which clearly distinguishes it from lookup siblings like find_files. It does not explicitly name or contrast with alternative action tools (open_deeplink, system_intent), so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a strong sourcing rule for uri ('from find_files, never build one by hand'), which implies the intended workflow, but it never states when to prefer this tool over alternatives like open_deeplink or open_deep_link, nor any preconditions. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_recent_appsOpen Recent AppsC

Open the recent-apps overview.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It only states the action without disclosing side effects, prerequisites (e.g., device must be awake), or what happens if no recent apps exist. The optional 'device' parameter hints at multi-device support but the description doesn't explain behavior across devices.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no wasted words. It is front-loaded with the action. However, it is so brief that it sacrifices useful context, but for what it contains, it is concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema and one optional parameter, the description is minimal. It lacks context about device targeting, prerequisites, and expected outcomes. The output schema may cover return values, but the description doesn't help an agent understand when or how to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the 'device' parameter at all. The schema shows it's an optional string/null with a default of null, but the description adds no meaning about what device values are valid or how the parameter affects behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Open') and resource ('recent-apps overview'), which clearly identifies the tool's function. It distinguishes it from sibling tools like launch_app or open_deeplink, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like launch_app, press_home, or system_intent. The context is implied by the name and description, but there is no explicit when-to-use or when-not-to-use information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

perceive_screenPerceive ScreenB

LOOK at the screen: an annotated screenshot plus every element and its tap point.

e is an array whose INDEX is the som_id (first entry = som_id 1): [center_x, center_y, name, flags] (name/flags omitted when empty). elements is the same list as text. Box colours: BLUE tappable, GREEN text input (type_text), MAGENTA scrollable, AMBER toggle, GREY nothing declared, RED on-host vision (detector/OCR; a good guess, not a fact). Flags (only when true): e editable, c checked, o unchecked, d disabled, f focused, l long-pressable, ? low-confidence vision box, w scroll host (aim inside it). offscreen: text that exists but is not on screen (cannot be tapped). detail="full" adds the visual pass (OmniParser YOLOv8 icon detector + OCR) for icons the tree does not describe; perception_tier reports tree_only or full. description is logged only. ids go stale after any action.

ParametersJSON Schema
NameRequiredDescriptionDefault
ocrNoauto
langNoeng
detailNo
deviceNo
max_marksNo
descriptionNo
include_imageNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full behavioral burden and does so richly: som_id indexing, box colour semantics, flag meanings, offscreen elements being untappable, the cost/benefit of detail="full", perception_tier reporting, and that ids go stale after any action. It is silent on the screen-capture permission requirement (a sibling, request_screen_capture_permission, implies one) and on whether the visual pass is slow or expensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well front-loaded — the single-line summary comes first, then output anatomy, then the detail-mode caveat. Every sentence carries load-bearing information, though the colour/flag enumeration is terse to the point of requiring careful re-reading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Return-value interpretation is thoroughly covered, which matters given there is no output schema, but the description is thin on the input side for a 7-parameter tool and omits any permission or prerequisite context. An agent can call it and read the result, but cannot reason about half its knobs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must compensate, and it only covers two: detail="full" and the fact that description is logged only. ocr, lang, device, max_marks, and include_image are left completely undocumented, including ocr's "auto" value and lang's "eng" default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource — 'LOOK at the screen: an annotated screenshot plus every element and its tap point' — which is far more concrete than a tautology. It does not explicitly distinguish itself from closely related siblings such as get_ui_tree, read_screen, ui_json, or screenshot, so an agent must infer the boundary from the detailed output anatomy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: 'ids go stale after any action' hints that this tool should be re-run before acting, and detail="full" is framed as an escalation for undescribed icons. There is no explicit when-to-use guidance and no named alternative (e.g. get_ui_tree for a cheap text-only pass, get_screenshot for pixels only).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ping_devicePing DeviceC

Is the device reachable? For adb devices, reports the Portal transport (http or content_provider).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that the tool reports Portal transport for adb devices, but it does not state whether the ping is read-only, whether it has side effects, how failures are represented, or any timeout/auth behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loaded as a direct question, with no wasted wording. It is appropriately sized for a one-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but the description is incomplete for routing among many device-related siblings. It omits usage conditions and parameter semantics, leaving an agent with insufficient context to choose it confidently over alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single optional 'device' parameter, and the description never explains what the parameter means, how to supply it, or what happens when it is omitted. The description does not compensate for the schema's lack of parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose: testing whether a device is reachable and reporting the Portal transport for adb devices. It identifies the resource and output dimension, but it does not distinguish this tool from siblings like get_device_status, list_devices, or connect_device.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives such as get_device_status or doctor. The phrase 'For adb devices' hints at a context, but it does not state prerequisites, exclusions, or routing criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pressPressA

Press home, back or enter (kept for old clients; prefer press_home/back/enter).

ParametersJSON Schema
NameRequiredDescriptionDefault
buttonYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It states the action (press) and the targets (home/back/enter) but does not disclose any side effects, permissions, or behavior for invalid inputs. For a simple action this is adequate but minimal; no contradiction with annotations (none provided).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with two clauses: the action and the usage note. It is front-loaded with the core function and wastes no words. Perfectly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple legacy action, the description covers the core purpose, the button values, and usage preference. The 'device' parameter is not explained, but it is optional and likely obvious from context. With an output schema present, return values are not needed. Overall adequate, with a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning for the 'button' parameter by specifying the allowed values (home, back, enter), which the schema does not enumerate. However, it does not explain the 'device' parameter, leaving it to inference. Since schema coverage is 0%, the description partially compensates but not fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool presses home, back, or enter – a specific verb and resource. It also names the preferred alternatives (press_home/back/enter), distinguishing it from siblings. This is explicit and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'kept for old clients; prefer press_home/back/enter', giving clear when-to-use and when-not-to-use guidance. It tells the agent to use this only for legacy compatibility and prefer the dedicated tools otherwise.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_backPress BackB

Press Back (also closes the keyboard without leaving the screen).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It usefully discloses that with the keyboard open, Back closes the keyboard instead of leaving the screen, which is a non-obvious trait. It does not describe the ordinary back-navigation effect, device selection behavior, or error handling, but the disclosed nuance is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and the useful side effect is placed in a parenthetical, but the opening phrase largely restates the tool name and title. There is no padding or unnecessary detail, though the wording could be more informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, one-parameter action, the description is minimally viable: the action and one important keyboard behavior are clear, and an output schema presumably explains return values. The gap is the undocumented device parameter, which prevents the description from being fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage and the description never mentions the optional device parameter. An agent cannot tell whether device selects a target device, what null means, or why it is optional, so the description adds no semantic value for the only parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete action, pressing the Back control, and the parenthetical adds a distinguishing effect (closing the keyboard without leaving the screen), separating it from siblings like press_home and press_enter. It does not explicitly contrast with all navigation siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The keyboard-dismissal note implies a useful scenario: if the keyboard is open, pressing Back will close it without navigating away. However, it does not name alternatives or give explicit when-to-use and when-not-to-use guidance, so usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_enterPress EnterA

Press Enter (submits search bars and forms).

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It discloses the core action (pressing Enter) but does not mention potential side effects (e.g., navigation, requiring a focused element) or how the device parameter affects behavior. It is minimally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words. The core action and purpose are front-loaded, making it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool, the description covers the basic function, but it omits explanation of the device parameter and does not address potential prerequisites (e.g., a focused input). The output schema exists, so return format is not required, but the description is incomplete for full correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter (device) with 0% description coverage, and the description does not explain it at all. An agent is left guessing whether device selects a specific device or has other meaning. The description fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (press Enter) and its purpose (submits search bars and forms). It is specific and distinguishes from generic 'press' by naming the exact key and its typical use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use the tool (when you need to submit a search bar or form), but does not explicitly compare to alternatives like tapping a submit button or using other key press tools. No exclusions or conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_homePress HomeC

Press the Home button.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the action itself and does not mention whether a device must be specified, what side effects occur, or what the tool returns after pressing Home.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and action-first, with no filler words. However, it is terse to the point of omitting parameter and context details, making it concise but under-specified rather than well-structured and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema, the core action is stated, but the missing device parameter semantics, absence of usage guidance, and lack of behavioral notes leave an agent guessing about invocation context. The description is minimally adequate for a human but not complete for an agent navigating many UI-action siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one optional device parameter with no descriptions and 0% schema description coverage. The tool description never mentions the parameter, so it adds no meaning beyond the schema's name, type, and default value, and it fails to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a clear imperative sentence naming both the action ('press') and the target ('Home button'), so an agent can understand the tool's function. It does not explicitly differentiate itself from siblings like press_back or press_enter, but the specific Home target is unambiguous enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance about when to use press_home versus alternatives such as press, press_back, or open_recent_apps. The only implied context is 'when you need to press Home,' with no exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_notificationsRead NotificationsA

Current status-bar notifications, newest first, without touching the screen: key, app, title, text, action labels. Ongoing ones (music, navigation, downloads) only with include_ongoing=true; package_name filters to one app; limit default 20, cap 30. With the safety policy on, banking and authenticator notifications are withheld.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
deviceNo
packageNo
package_nameNo
include_ongoingNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and discloses ordering ('newest first'), the non-screen-touching read behavior, the safety-policy withholding of sensitive notifications, and the ongoing-notification condition. It stops short of stating required permissions or whether any state is modified beyond 'without touching the screen'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core read operation and return fields, then uses semicolons to add conditions compactly. Every sentence fragment carries useful information with no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value details are optional, and the description still names the returned fields. It covers ordering, filtering, limits, and safety withholding, but two optional input parameters (device and package) remain undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It usefully documents include_ongoing, package_name, and limit (default 20, cap 30), but leaves device and package unexplained, and does not clarify how package differs from package_name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: reading current status-bar notifications. It adds scope and output fields (key, app, title, text, action labels) and distinguishes the read-only nature from siblings like dismiss_notification and notification_action by saying it does so 'without touching the screen'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage conditions: ongoing notifications require include_ongoing=true, package_name filters to one app, and banking/authenticator notifications are withheld under the safety policy. It does not explicitly name alternative tools or when not to use this one, which keeps it from a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_screenRead ScreenA

Read the screen now (waits for it to stop moving first): the screen drawn as a character grid, each element a box with its som_id and label, then a table of what can be acted on: som (tap by this), in (som_id of the smallest box containing it), flg, label (only when it did not fit on the grid). Flags: * tappable, e text input (type_text, not tap), S scrollable, c toggle ON, o toggle OFF, l long-pressable, d disabled, - nothing declared (usually still tappable). N+k = som_id N plus k more elements with exactly those bounds; tap N. Header IDLE/BUSY: BUSY means it was still moving when the wait expired. Use perceive_screen instead for how something looks, when this ends with ESCALATE, or when what you need is missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses the blocking wait ('waits for it to stop moving first'), the BUSY/IDLE header semantics and what BUSY implies, the 'N+k' aggregation notation and the instruction to tap N, and the ESCALATE outcome. These are behavioral facts an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core action and its blocking behavior come first, then the output contract, then the flag legend, then the routing note. The flag glossary and N+k note are long but each entry encodes a distinct actionable signal, so little is wasted; only mild compression is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description defines the output format (character grid, element boxes with som_id/label, action table), decodes every flag, explains the header state, and gives the escalation/fallback route. An output schema exists, yet the description still adds the interpretive layer the schema cannot supply, making the definition complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one optional parameter exists (device, default null) and the description says nothing about it. However, with a single optional device selector and a rich, fully-specified output contract, the omission is minor; the description compensates for the 0% schema coverage by exhaustively defining the return structure the caller must interpret.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the screen now') and immediately scopes it with the wait-for-still behavior. It also names the sibling it is not (perceive_screen) and the reason to prefer that sibling, so an agent can route correctly without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the alternative and the three conditions selecting it: 'Use perceive_screen instead for how something looks, when this ends with ESCALATE, or when what you need is missing.' This is textbook when-to-use/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_findingRecord FindingB

Record one item you found. quote must be copied exactly from the CURRENT screen.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemYes
quoteYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden, and it does add one useful behavioral constraint: the quote must be copied exactly from the CURRENT screen, which tells the agent how to source that argument. However, it does not disclose what recording does (e.g., session logging vs. side effects), how the exact-match requirement is enforced, or any device-related behavior — partial coverage at best.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler. The core purpose is front-loaded in the first sentence and the critical quote-copying constraint follows immediately; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, but the tool has three parameters at 0% schema coverage, no annotations, and no explanation of what constitutes an 'item', when recording is warranted, or what 'device' controls. An agent would be guessing on half the input contract, making this incomplete for a tool with this little structured support.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all three parameters, but it only adds meaning to 'quote' (exact copy from current screen). 'item' is glossed merely as 'the item you found' with no definition of what qualifies, and 'device' is never mentioned at all, leaving its purpose entirely to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('record') and resource ('item you found'), and this clearly distinguishes it from the sibling set, which is entirely UI interaction, screenshot, and app management tools — no other sibling records findings. It falls short of 5 because 'item' is never defined, leaving the scope of what counts as a finding ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives — no exclusions, no named alternatives, no conditions for skipping. 'Record one item you found' merely restates the purpose with an implied trigger; the only hint of a prerequisite (a 'CURRENT screen' must exist) is buried in a parameter constraint rather than framed as usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_screen_capture_permissionRequest Screen Capture PermissionA

Compatibility no-op: screenshots use the Portal / adb screencap, no prompt is needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It fully discloses that the tool performs no actual permission request, has no prompt, and is only a compatibility shim. This gives an agent complete knowledge of side effects and expected behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence delivers the essential message: the tool is a no-op, why it exists, and what the real mechanism is. There is zero wasted wording, and the most important information ('no-op') appears first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a compatibility no-op with no annotations and a single optional parameter, the description is complete enough: it explains the tool's purpose, its lack of side effects, and the reason it exists. The output schema exists, so return-value details are not the description's responsibility.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one optional 'device' parameter with 0% description coverage, and the description does not mention it at all. Even though the tool is a no-op, the description fails to explain whether the parameter is ignored, validated, or affects behavior. The description should compensate for the bare schema but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is a 'Compatibility no-op' and explains why (screenshots use Portal / adb screencap, no prompt needed). This precisely differentiates it from the many screenshot-related siblings, which actually capture or read the screen, by identifying this as a stub.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: because screenshots are handled elsewhere and no prompt is needed, calling this tool is unnecessary. It stops short of explicitly naming an alternative tool or an exact condition for using this one, but 'compatibility no-op' strongly implies it should only be invoked when some external contract requires it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_contactResolve ContactA

Find contacts by (partial) name and return their phone numbers.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
limitNo
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It does disclose partial-name matching and that phone numbers are returned, which is core behavior. But it leaves important traits unstated: whether it reads local device contacts, whether permissions are required, how duplicates or unmatched names are handled, and how limit/device affect results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one concise sentence with no filler. The main action is front-loaded, and every phrase adds relevant information: 'contacts', 'partial name', and 'return their phone numbers'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple lookup tool with an output schema, the description is minimally adequate: it states what to pass and what comes back. However, with no annotations and no parameter descriptions, the missing semantics of 'limit' and 'device', plus no mention of no-match behavior, make it incomplete for robust autonomous use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the bare input schema. It adds meaning only to the 'name' parameter via 'partial'. It does not explain 'limit' or 'device', leaving their semantics to be guessed from their names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Find contacts'), a specific resource ('contacts'), a matching rule ('partial name'), and the return value ('phone numbers'). It also clearly distinguishes resolve_contact from sibling tools like resolve_deeplink because it names the exact subject matter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The use case is implied: use it when you need contact phone numbers from a partial or full name. However, the description gives no explicit when-to-use/when-not-to-use guidance, no mention of alternatives, and no prerequisites such as required permissions or device context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotScreenshotC

Plain screenshot. hide_overlay hides the Portal's element overlay first.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
hide_overlayNo

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses only that hide_overlay hides the Portal's overlay first; it does not state permissions, return format, side effects, or whether the call is safe/read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, which is appropriate, but the first sentence is largely tautological and does not earn its place. The second sentence is useful but the whole remains under-specified for a two-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter description coverage, the description is incomplete. It does not explain the return value, distinguish from screenshot siblings, or document the device parameter, all of which an agent needs to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It partially clarifies hide_overlay but says nothing about the device parameter, leaving half the inputs undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Plain screenshot" essentially restates the tool name and title without a distinct verb or resource scope. It does not differentiate this tool from close siblings such as get_screenshot, screenshot_path, read_screen, or perceive_screen, so an agent cannot tell when this one is the right choice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no exclusions, and no named alternatives. It mentions hide_overlay but does not explain when to set it versus using another screenshot-related sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshot_pathScreenshot PathC

Take a screenshot, save it as a PNG file and return the path.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the core action but omits critical details such as where the PNG is saved (local vs. device), whether it overwrites existing files, permission requirements, or behavior on failure. The device parameter is not mentioned, so its effect on the screenshot is unknown. This is a significant gap for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the primary action and output. However, it is concise to the point of omitting important details (parameter semantics, side effects), so it is not fully optimal. It earns a 4 for structure but loses a point for under-specification that could be addressed without verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has one optional parameter (device) and many closely related sibling tools, the description is insufficiently complete. It does not explain the device parameter, differentiate from siblings, or provide usage context. While an output schema exists (so return format is covered), the tool's overall behavior and applicability are left underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one parameter 'device' with 0% description coverage, and the tool description does not mention it at all. Since the description fails to explain the purpose or effect of the parameter, the agent has no semantic context for it. The description adds zero value beyond the raw schema, and with no schema descriptions, the parameter is effectively undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action (take a screenshot), the output format (PNG file), and the return value (path). This distinguishes it from siblings like 'screenshot' (which may not save to a file) or 'get_screenshot' (which might retrieve an existing one). The verb+resource+output structure is explicit and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'screenshot', 'get_screenshot', or 'perceive_screen'. The description does not mention any conditions, exclusions, or comparisons to other tools, leaving the agent to infer usage context on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screen_sizeScreen SizeD

[width, height] in pixels.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It adds no information about side effects, permissions, whether the value is current or static, or any other behavioral trait beyond a vague output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, which is good for conciseness. However, its structure is a bare phrase rather than a clear tool description, so it does not fully earn its place as meaningful guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and can document the return values, the description still fails to explain what the tool does or how to invoke it. For a callable tool with an undocumented parameter, this is inadequate context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one optional parameter ('device') with 0% description coverage. The description does not mention this parameter or explain what it controls, leaving its semantics entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description '[width, height] in pixels.' merely restates the tool name and title as an output format. It never states the action (e.g., retrieving the current screen dimensions), so it reads as a tautology rather than a clear verb+resource definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to use it, or which sibling tools might be alternatives. The description offers no usage context at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrollScrollC

Scroll the content in direction (up | down | left | right) by distance (fraction of the screen). verify=true reports whether the screen actually moved (mobilerun-core).

ParametersJSON Schema
NameRequiredDescriptionDefault
msNo
deviceNo
verifyNo
distanceNo
directionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that verify=true reports whether the screen actually moved, which is not in the schema, but it omits other relevant behavior such as whether the action is destructive, what permissions or app context are needed, and what happens on unsupported screens.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single compact sentence plus a parenthetical note, with the core action and important parameters front-loaded. The '(mobilerun-core)' tag is minor noise but the description is otherwise efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but the tool still has five parameters with zero schema descriptions and no annotations. The description leaves ms, device, and the choice among sibling scroll tools unexplained, which is inadequate for a tool with this much structured surface area.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains direction, distance, and verify. The ms and device parameters remain completely undocumented in both the schema and the description, leaving an agent without semantics for two parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scroll the content'), and enumerates the parameterized directions and distance unit. It does not explicitly name or contrast with the dedicated siblings scroll_up, scroll_down, scroll_left, scroll_right, scroll_to, or scroll_until, so the agent must infer when the generic form is preferred.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does but gives no when-to-use guidance, prerequisites, or comparison to alternative scrolling tools. It never says when to prefer this general tool over the specialized scroll_down/scroll_up/scroll_left/scroll_right or scroll_to.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_downScroll DownB

Scroll the content down (reveal what is below): a centered swipe over half the screen (amount), or inside a scrollable mark (som_id).

ParametersJSON Schema
NameRequiredDescriptionDefault
amountNo
deviceNo
som_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does add real mechanism detail beyond the name — the gesture is a centered swipe over half the screen — which is genuinely useful. However it never states what happens at the end of content, whether the scroll is animated/instant, or any side effects on UI state, so it is only partially transparent for a UI-mutating action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the action and then qualifies it. Very little waste, though the parenthetical placement of '(amount)' and '(som_id)' makes the parameter mapping slightly choppy to read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained. Still, with no annotations, 0% schema description coverage, and an undocumented 'device' parameter, the definition leaves gaps an agent must guess at — notably target-device handling and end-of-content behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it partly does: it maps 'amount' to a fractional screen swipe and 'som_id' to a scrollable mark target. The third parameter (device) is never mentioned, and the amount's numeric scale/units are only vaguely implied by 'half the screen'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Scroll the content down') with an operational gloss ('reveal what is below') and names the two execution modes (centered swipe via amount, in-place scroll via som_id). The direction word 'down' implicitly separates it from scroll_up/left/right siblings, though it never names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description distinguishes the two calling modes (amount vs som_id), which is useful selection guidance internal to the tool, but it gives no when-to-use guidance relative to the many scrolling siblings (swipe, scroll, scroll_up, scroll_to, scroll_until, scroll_until). Nothing says when this is preferable to raw swipe or to scroll_to.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_leftScroll LeftC

Scroll the content left (reveal what is to the left).

ParametersJSON Schema
NameRequiredDescriptionDefault
amountNo
deviceNo
som_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It mentions the effect ('reveal what is to the left') but does not explain how the 'amount' parameter influences scrolling, what device or som_id are for, or any side effects or limitations. This is minimal behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no filler. It is appropriately concise and front-loads the action. However, its brevity borders on under-specification, which is more a completeness issue than a conciseness issue. The structure is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three parameters, no annotations, and no parameter explanations, the description is incomplete. It does not cover parameter semantics, usage context, or behavioral details beyond the basic action. The presence of an output schema mitigates the need to explain return values, but the lack of parameter guidance is a significant gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the description provides no information about the three parameters (amount, device, som_id). The description does not compensate for the lack of schema documentation, leaving the agent without any understanding of what these parameters mean or how to set them correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('scroll') and resource ('content') with a direction ('left'), and adds 'reveal what is to the left' to clarify the effect. It is distinguishable from other scroll tools like scroll_right and scroll_up, though it does not explicitly name them. This is clear and purposeful, not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like scroll_to, scroll_right, or swipe. No context is provided for choosing this over a generic swipe or scroll. The description only states the action without any situational direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_rightScroll RightB

Scroll the content right (reveal what is to the right).

ParametersJSON Schema
NameRequiredDescriptionDefault
amountNo
deviceNo
som_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral disclosure burden. It does state the intended observable effect—scroll right and reveal content to the right—but it does not disclose how the amount parameter affects the scroll, whether the scroll is smooth, or how device/som_id scoping works.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. The parenthetical earns its place by clarifying the direction and purpose, and the entire definition is appropriately sized for such a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the tool is simple and an output schema exists, the definition is incomplete because it leaves all parameter semantics unexplained and offers no guidance for choosing this tool over related scroll/swipe tools. An agent would not know what amount=0.5 means or whether device/som_id are needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no meaning for any of the three parameters (amount, device, som_id). The default 0.5 for amount is unexplained, and device/som_id are completely undocumented, leaving the agent to guess their semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Scroll' with a clear object 'the content' and direction 'right', and the parenthetical 'reveal what is to the right' reinforces the intended outcome. This clearly distinguishes it from siblings like scroll_left, scroll_down, and scroll_up.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when content is clipped or hidden on the right side, but it does not explicitly state when to use this tool versus alternatives like scroll_to, swipe, or the other directional scroll tools. It gives context but no exclusions or routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_toScroll ToA

Two modes. With x1, y1, x2, y2: drag the content from one point to the other (precise scroll). With text: scroll in direction until an element containing text is visible and return its mark.

ParametersJSON Schema
NameRequiredDescriptionDefault
x1No
x2No
y1No
y2No
textNo
deviceNo
directionNodown
duration_msNo
max_scrollsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose meaningful behavior: the coordinate-drag mode, the scroll-until-visible mode, and that it returns the element's 'mark'. It omits direction defaults, the roles of duration_ms/max_scrolls/device, and failure behavior, leaving notable gaps for a mutation-style interaction tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with 'Two modes' and no wasted words. Each clause maps directly to a distinct invocation path.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The existence of an output schema reduces the need to explain return values, and the description usefully flags the returned mark. However, for a nine-parameter tool with zero schema coverage and no annotations, the description leaves several parameters and failure/timing behavior unexplained, so it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It gives meaning to x1/y1/x2/y2, text, and direction, but says nothing about device, duration_ms, or max_scrolls, leaving a third of the nine parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource (scroll) and enumerates two distinct operating modes with their triggering parameters, which lets an agent tell it apart from a plain scroll. It does not, however, differentiate itself from near-neighbors like scroll_until or swipe, so sibling disambiguation is incomplete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implicitly tells you when each mode applies (coordinates vs. text), which is useful routing guidance. But it never names alternatives such as scroll_until, scroll_down, or swipe, nor states when-not to use this tool, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_untilScroll UntilC

Scroll until a matching node is on screen; result is the node (or null).

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
deviceNo
settleNo
distanceNo
directionNodown
max_swipesNo
resource_idNo
any_containsNo
text_containsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden and largely fails it. It reveals only that the call ends with a node or null; it does not say that it performs repeated swipe gestures, whether it alters screen state, what happens on timeout, or side effects of scrolling the UI.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the core behavior front-loaded and no filler. It is efficient, though its brevity is more under-specification than true economy for a 9-parameter looping tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return format need not be re-explained, but the tool is a multi-swipe loop with 9 undocumented parameters and no annotations. For that complexity the description leaves far too much undefined for an agent to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Nine parameters with 0% schema description coverage, and the description names none of them. 'Matching node' vaguely implies the match criteria exist but gives no clue about the relationship between text, text_contains, any_contains, and resource_id, nor the meaning of settle, distance, direction, or max_swipes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: scrolling until a matching node appears, and even names the return value (node or null). It does not differentiate itself from close siblings like scroll_to, find_nodes, or wait_for_nodes, which an agent must pick between, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus the many alternatives in the sibling list (scroll_to, scroll_down, wait_for_text, find_nodes). The only hint is the tool name itself implying a loop-until-found semantics, which the agent must infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_upScroll UpC

Scroll the content up (reveal what is above).

ParametersJSON Schema
NameRequiredDescriptionDefault
amountNo
deviceNo
som_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does explain the intended effect, but it is silent on how the amount parameter influences the scroll, how device and som_id determine the target context, and what observable side effects or limits exist. For a UI action with three parameters, this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler, and the parenthetical adds meaningful clarification rather than redundancy. However, it is so terse that it leaves all parameter-level information unaddressed, so it is not fully appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists, the definition is not complete enough for correct invocation: parameter semantics are undocumented and there is no guidance for choosing among the many sibling navigation tools. Scroll direction alone is insufficient for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions amount, device, or som_id. An agent cannot determine whether amount is a fraction of the viewport, a pixel distance, or a scroll unit, nor what device and som_id refer to. The description adds no value over the raw parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('scroll'), a resource ('content'), and a direction ('up'), with the parenthetical clarifying that this reveals content above the current viewport. This clearly separates it from scroll_down, scroll_left, and scroll_right by direction, though it does not explicitly contrast it with scroll_to or swipe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as swipe, scroll_to, or other scroll directions. The phrase 'reveal what is above' weakly implies a use case, but no explicit conditions, exclusions, or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_clipboardSet ClipboardC

Put text on the clipboard.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and delivers almost nothing: no mention of whether the clipboard is local or device-scoped, overwrite semantics, or any permission requirement. A mutation tool with zero annotation coverage needs far more than 'Put text on the clipboard.'

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the action. It is efficient, though its brevity here reflects under-specification rather than disciplined concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for a 2-parameter mutation tool with no annotations and an undocumented 'device' parameter, the description leaves critical invocation details missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and does not. The 'device' parameter in particular is completely opaque — whether it selects a remote device, defaults to local, or accepts an ID is left entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Put') and resource ('text on the clipboard'), so the core action is unambiguous. It does not distinguish itself from the sibling get_clipboard, which an agent could easily confuse with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives offered. The only implicit signal is the existence of get_clipboard as the read counterpart, which the agent must infer on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_planSet PlanC

Start a plan checklist. target_count > 0 means 'N items must be recorded' before end_session(success) is allowed. With 3+ steps and a search_query, the first web search rides along in the reply.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNo
stepsYes
deviceNo
deliverableNo
search_queryNo
target_countNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It reveals two key behaviors: the target_count constraint on ending sessions and the side-effect of an accompanying web search. However, it does not mention whether this tool mutates state, requires authentication, or has side effects beyond the stated ones, leaving gaps in the behavioral picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, using three short sentences that front-load the core purpose. It avoids unnecessary verbosity, though the phrase 'rides along' is informal and could be clearer. The structure is efficient, but it lacks a logical separation between the primary purpose and the behavioral constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six parameters and no annotations, the description is incomplete. It explains a few specific behaviors but does not clarify how the plan integrates with related tools like mark_step, record_finding, or end_session, nor does it describe the output or expected response. An agent would struggle to understand the full lifecycle and prerequisites for using this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the missing parameter documentation. It explicitly explains target_count and search_query, and implicitly refers to steps, but leaves goal, device, and deliverable completely unexplained. This partial coverage is insufficient for a tool with six parameters, especially since the schema provides no descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the primary action ('Start a plan checklist'), giving a specific verb and resource. It does not explicitly distinguish from siblings like 'run_task' or 'mark_step', but the added behavioral details imply a distinct role in orchestrating a task plan. The term 'plan checklist' is somewhat ambiguous but still understandable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It mentions conditions like target_count and search_query but does not explain when a user should invoke this tool instead of other plan-related tools. There is no mention of prerequisites, exclusions, or typical use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

setup_portalSetup PortalC

Install and enable the Mobilerun Portal on the device (mobilerun setup); path installs a specific Portal APK.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It says the tool installs and enables software but omits side effects, whether the device must be connected/authorized, whether an existing Portal is replaced, and whether a restart or elevation is required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with the primary action front-loaded before the parameter caveat. No filler, though it could be marginally clearer with a short prerequisite clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, but for a no-annotation mutation tool with an undocumented device parameter and no stated prerequisites, the definition is too thin to call the tool safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain both parameters. It clarifies that path installs a specific Portal APK, but the device parameter is left entirely undefined in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('install and enable') and resource ('Mobilerun Portal'), and explains the path parameter's effect. It does not differentiate itself from the nearby install_app sibling, which also installs software on the device.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Conveys context implicitly through the '(mobilerun setup)' CLI hint and the two modes (default install vs path-specific APK), but never states when to prefer this over install_app or what prerequisites (connected device, permissions) must hold.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_appStart AppC

Start an app by id (Android package / iOS bundle id), optionally a specific activity.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_idNo
deviceNo
packageNo
activityNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does not state whether the call blocks until the app is foregrounded, what happens if the app is not installed, whether it requires a connected device, or how the device parameter interacts with the operation. Only the minimal 'start an app' semantics are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the core verb and resource lead. It is efficient but borders on under-specification given the tool's four parameters and unclear sibling overlap.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained. However, for a 4-parameter tool with zero schema coverage, no annotations, and a near-duplicate sibling in launch_app, the description omits the device/package semantics and any disambiguation, leaving real gaps an agent must fill by guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description must compensate. It explains app_id and activity, but leaves device and package entirely undocumented. Worse, the coexistence of app_id and package is never reconciled, leaving ambiguous which one to pass and whether one is an alias for the other.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (start) and resource (app), and clarifies the id format as Android package / iOS bundle id. However, it does not differentiate itself from the sibling launch_app, which appears to serve a nearly identical purpose, leaving the agent unable to choose between them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use start_app versus launch_app, open_and_settle, or open_deeplink. The optional activity parameter is mentioned but no condition for supplying it is given. The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stop_appStop AppB

Force-stop an app; clear_data also wipes its data.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_idYes
deviceNo
clear_dataNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the destructive side effect that 'clear_data also wipes its data', which is the key risk an agent must know. However, it omits that a force-stop kills the running process and loses unsaved state, permission requirements, and any failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the primary action front-loaded and the destructive caveat appended. No wasted words, though the extreme brevity leaves behavioral gaps rather than padding them.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. For a mutation tool with no annotations and undocumented parameters, the description covers the most important destructive consequence but leaves the device targeting and preconditions unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across three parameters, so the description must compensate. It explains the semantic effect of clear_data (wipes app data), which is more than the schema's bare boolean, but says nothing about app_id format or how the device parameter selects a target.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Force-stop an app'), which cleanly separates it from launch_app/start_app and other app-lifecycle siblings. It does not explicitly name alternatives, but the verb itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus start_app, uninstall_app, or press_home, and no prerequisites (e.g., app must be running/installed). Usage is only implied by the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

supportsSupportsC

Whether this device supports a Device action (e.g. execute_script, get_clipboard).

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It only states the query purpose and does not disclose whether the call is read-only, has side effects, requires permissions, or how errors are handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the core purpose and includes examples without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the simple two-parameter structure is mostly covered by the schema. However, the lack of annotations and the completely absent usage context leave the description only minimally adequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate. It gives useful examples for the required 'action' parameter but completely omits the optional 'device' parameter, leaving its semantics (e.g., defaulting to the current device) unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific check: whether the device supports a Device action, with examples like execute_script and get_clipboard. This clearly identifies the verb and resource, but it does not distinguish the tool from the sibling 'capabilities', which likely also reports device capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives like 'capabilities' or 'validate_action'. The examples clarify the 'action' parameter but do not explain the intended workflow, such as checking support before attempting an action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

swipeSwipeB

Swipe from (x1, y1) to (x2, y2) over duration_ms / ms (default 300). mobilerun agent form: coordinate=[x, y], coordinate2=[x, y], duration in seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
msNo
x1No
x2No
y1No
y2No
deviceNo
durationNo
coordinateNo
coordinate2No
duration_msNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It does disclose a concrete default (300 ms for the duration) and the unit semantics (seconds vs. ms), which is genuine added value, but it says nothing about return behavior, device targeting, or required permissions/state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the coordinate semantics and the duration default. Slightly dense phrasing ('duration_ms / ms') but no wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers most of the 10 parameters. Gaps remain for the device parameter and for when to choose this gesture over scroll siblings, which for a 10-parameter tool is a noticeable omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: it maps x1/y1/x2/y2 to the drag path, documents the duration_ms/ms default of 300, and explains the coordinate=[x,y]/coordinate2=[x,y] alias form plus the seconds unit for duration. Only the device parameter is left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific gesture (swipe) and spells out its start/end coordinates, so an agent immediately knows what the tool does. It does not, however, distinguish itself from gesture siblings such as scroll, scroll_down, or tap, which perform overlapping drag-like actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer swipe over scroll/scroll_down/scroll_until or tap-based interaction. Usage can only be inferred from the tool name, and no preconditions (e.g. screen must be settled) are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

system_intentSystem IntentA

One-call Android actions (action = the verb; verb= is accepted too). Verbs: set_alarm(hour, minute, label), set_timer(seconds, label), dial(phone_number), compose_sms(phone_number, body), add_calendar_event(title, start, end, location, notes; ISO datetimes), share_text(text, subject), navigate( destination, mode drive|walk|bike|transit). dial/compose_sms only prefill; the user sends.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
bodyNo
hourNo
modeNodrive
textNo
verbNo
labelNo
notesNo
startNo
titleNo
actionNo
deviceNo
minuteNo
secondsNo
skip_uiNo
subjectNo
locationNo
destinationNo
phone_numberNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, and it does disclose one important behavior: dial/compose_sms only prefill, the user sends. However it says nothing about what the other verbs actually do (execute silently via skip_ui=true?), permission needs, whether the device param targets a remote device, or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core idea ('One-call Android actions') and then packs the verb catalog densely into a compact form; the alias note and prefill caveat are placed where they matter. The verb list is long but every entry carries unique routing information, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the per-verb parameter mapping is thorough. Gaps remain for a 19-parameter tool: device targeting, skip_ui semantics, and whether verbs succeed silently or require user confirmation are all unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: it maps which parameters belong to which verb, gives types (hour/minute integers, seconds), ISO datetime format for start/end, the mode value set drive|walk|bike|transit, and clarifies action vs the accepted verb= alias. Only device and skip_ui are left unexplained out of 19 params.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific capability — dispatching one-call Android actions — and enumerates the supported verbs with their argument signatures, so an agent knows exactly what kinds of intents it can execute. It is distinguishable from low-level siblings like tap/press/launch_app by being the high-level intent layer, though it never explicitly says that.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implicit guidance comes from the verb list — use this when the user wants an alarm, timer, call, SMS, event, share or navigation rather than UI manipulation. But there is no explicit when-to-use vs alternatives (e.g. launch_app, web_search, ui) and no prerequisites or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tapTapC

Tap at (x, y) or at the center of a numbered mark (som_id from perceive_screen / read_screen). stealth=true uses mobilerun-core's humanized tap.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
deviceNo
som_idNo
stealthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses one behavioral modifier (stealth uses mobilerun-core's humanized tap), but says nothing about what a tap does as a side effect, whether it waits for the UI to settle, or whether it can fail/be retried. That is thin for an action tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the primary coordinate/mark behavior front-loaded and the stealth caveat trailing. Zero filler, though the parenthetical makes it slightly dense to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. However, with 5 parameters, no annotations, no defaults documented, and an unexplained device param and x/y vs som_id precedence, the definition is only minimally complete for an action tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 params, so the description must compensate. It adds real meaning for som_id (source tools) and stealth (humanized tap), and loosely covers x/y, but the device parameter and the x/y-with-som_id precedence rule are left entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: tap at (x, y) or at a numbered mark via som_id, and names the sibling tools that produce som_id (perceive_screen / read_screen). It implicitly distinguishes itself from double_tap, long_press, and tap_node by being the plain single tap, though it never says so explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells the agent where som_id comes from and that stealth=true switches to a humanized tap, but gives no guidance on when to prefer this over the many siblings (double_tap, long_press, tap_text, tap_node, tap_and_wait) or what prerequisites exist. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tap_and_waitTap And WaitB

Tap a text (or node) and wait until the UI has been idle for idle seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
idleNo
deviceNo
targetYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden; it does disclose the key trait that the call blocks until the UI is idle for `idle` seconds, which is genuinely useful. It omits other relevant traits such as whether the tap target must already be on screen, device routing, and what happens if no idle state is reached.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler that names the action, the target types, and the wait condition in order. Brief, though the brevity trades off against the missing detail noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the core wait semantics are stated. But for a mutating tap with zero annotations and 0% schema description coverage, the definition is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate; it explains `idle` ('idle seconds') and hints at `target` accepting a text string or a node object. The `device` parameter and the accepted target formats are undocumented, leaving part of the gap unfilled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Tap a text (or node)') plus the distinctive wait behavior, so an agent can distinguish it from plain `tap` by the settle-after-action semantics. It does not, however, explicitly differentiate itself from close siblings like `tap_text`, `tap_node`, `wait_for_idle`, or `open_and_settle`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given despite a crowded sibling set of tap and wait tools (`tap`, `tap_text`, `tap_node`, `wait_for_idle`, `open_and_settle`). The agent must infer that this is the combined tap-then-settle variant purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tap_nodeTap NodeB

Tap the center of a node returned by find_nodes / find_nodes_on_screen.

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeYes
deviceNo
stealthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It reveals only that the tap targets the node's center; it says nothing about whether the tap waits for the UI to settle, what happens if the node is stale/offscreen, what the stealth flag does, or what the tool returns on success or failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action and the input source front-loaded. No filler, no repetition of the title, nothing that could be cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the core purpose is conveyed. But for an unannotated, 0%-documented mutation tool, the definition omits device targeting, the meaning of stealth, and failure behavior (e.g., node not found or not tappable), leaving real gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters. The description usefully explains where the required 'node' object comes from, but it says nothing about 'device' or the 'stealth' boolean (default true), which is the most opaque parameter and would be the prime candidate for explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (tap) plus the precise target (center of a node) and it names the sibling tools that produce the node object (find_nodes / find_nodes_on_screen). This lets an agent place it in the find-then-act workflow, though it does not explicitly contrast with the plain tap or tap_text siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'a node returned by find_nodes / find_nodes_on_screen', which tells the agent the expected input provenance. However it never states when to prefer this over tap (coordinates), tap_text (text match), or long_press, nor any precondition such as requiring a prior find call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tap_textTap TextC

Tap the first on-screen node whose text/description contains text.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that matching is substring-based and that only the first match is tapped, but says nothing about failure behavior (no match, multiple matches, off-screen nodes), whether it waits for the node, or where on the node the tap lands. For a mutation tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the matching semantics front-loaded and zero filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the description covers the core selection rule. However, with no annotations and an undocumented 'device' parameter, an agent lacks the failure-mode and targeting details needed to call this reliably in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does clarify that 'text' is a contains-match against text or description, but the 'device' parameter is left completely unexplained in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (tap) and resource (on-screen node) plus the matching rule: text/description contains the given string. An agent can distinguish it from coordinate-based 'tap' and node-id-based 'tap_node'. It stops short of naming those siblings, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over 'tap', 'tap_node', 'find_nodes', or 'wait_for_text', nor any note about what to do when the text is not present. The agent must infer usage entirely from the name and one sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

timeTimeD

The device clock.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It says nothing about whether this is a read-only or mutating operation, whether permissions are needed, or what the result represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short and front-loaded, but the single sentence is under-specified rather than concise. It does not earn its place because it conveys almost no actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a tool with an undocumented optional parameter and no annotations, the description is too incomplete to guide correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one parameter ('device') with 0% description coverage, and the description does not mention it at all. An agent cannot tell what the parameter does or whether it is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'The device clock' names a resource but no action, so it does not clearly state whether the tool reads, sets, or otherwise uses the device clock. It is essentially a restatement of the tool name/title with a minor qualifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance at all on when to use this tool, when not to use it, or which sibling tools might be alternatives. The description provides no context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

typeTypeC

Type text (mobilerun). index taps that get_state element first; stealth=true types key by key like a person (mobilerun-core), at wpm words per minute.

ParametersJSON Schema
NameRequiredDescriptionDefault
wpmNo
textYes
clearNo
indexNo
deviceNo
stealthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it only partially delivers: it explains the stealth/keystroke and index/tap-then-type behaviors. It says nothing about what 'clear' does to existing field content, how 'device' selects a target, or any permission/latency characteristics of this input-mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is tight and front-loaded into a single sentence with no filler, which is good. But the dense parenthetical jargon and clipped clauses ('get_state element', 'mobilerun-core') trade clarity for brevity, making it harder to parse than a cleaner phrasing would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But for a 6-parameter mutation tool with zero annotation coverage and 0% schema description coverage, the agent is left without enough to call it confidently, and half the parameters remain undocumented anywhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, so the description must compensate and does so only partly. It clarifies index, stealth, and wpm, but leaves text, clear, and device completely unexplained, and even the covered ones lack format detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb+resource ('Type text') and tags the backend (mobilerun), so the basic action is understood. However, it does not distinguish itself from close siblings such as type_text, tap_text, or key, and the '(mobilerun)' qualifier is opaque jargon rather than a real differentiator.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implicitly tells the agent when the 'index' and 'stealth' behaviors matter ('index taps that get_state element first', 'stealth=true types key by key like a person'), which is usable context. But there is no explicit statement of when to prefer this tool over the sibling type_text, and no prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

type_textType TextA

Type into the focused field (tap it first, or pass som_id). submit presses Enter after.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
clearNo
deviceNo
som_idNo
submitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It mentions the submit behavior and the need for focus, but fails to disclose what the clear parameter does, whether text is appended or replaces existing content, or any side effects. This is a significant gap for a tool with multiple behavioral options.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no waste. It front-loads the action and the submit behavior, and the prerequisite is clearly stated. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 5 parameters and an output schema, but the description only addresses 2 of them (som_id and submit). It omits clear and device entirely, and doesn't mention potential edge cases like handling special characters or long text. With zero schema descriptions and no annotations, this is incomplete for safe and correct use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains som_id (as an alternative to tap) and submit (presses Enter), but leaves clear and device completely unexplained. The agent cannot know what clear does or how to use device without additional context. This is insufficient given the zero coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Type into the focused field' with a specific resource, and differentiates from siblings like tap and press_enter by explaining the field focus mechanism (tap or som_id) and the submit behavior. It is unambiguous and concise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a prerequisite (tap first or pass som_id) and explains the submit flag's effect. It doesn't explicitly state when to avoid using this tool or compare to alternatives, but the context for when to use it (typing text) is implied. This is sufficient for most agents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

uiUiC

Raw UI snapshot (a11y_tree, phone_state, device_context, ...), as Device.ui().

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
filterNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it says almost nothing: it doesn't state that this is a non-destructive read, whether the call can be slow, or how 'filter' changes the returned snapshot. The cryptic 'as Device.ui()' reference is a Python API pointer, not behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no wasted filler, which is good. But the brevity comes at the cost of under-specification rather than genuine economy, and the trailing '...' and 'as Device.ui()' reference add little for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be re-explained. Still, with zero annotation coverage, 0% parameter documentation, and five-plus competing snapshot siblings, the description leaves the agent unable to decide when or how to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description explains neither parameter. The 'filter' boolean defaulting to true is completely opaque — the agent cannot know whether it strips nodes, prunes the tree, or limits the payload — and 'device' (nullable string selector) is likewise unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource (a raw UI snapshot) and enumerates some of its contents (a11y_tree, phone_state, device_context), which is more than a tautology of the name 'ui'. However, it gives no verb (fetch/return) and fails to distinguish itself from close siblings like ui_json, ui_with_recovery, or get_ui_tree, so an agent cannot tell which snapshot tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives, despite several near-identical siblings (ui_json, ui_with_recovery, get_ui_tree, perceive_screen, read_screen). The agent gets no signal about why it would choose this over those.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_jsonUi JsonC

The UI snapshot serialized as JSON text.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
filterNo
indentNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden, yet it only says the result is a JSON-serialized snapshot. It does not state that it is a read-only operation, how large the output may be, whether `device` defaults to the active device, or what `filter` removes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no filler, so it is concise. However, the brevity comes at the cost of substance — there is no front-loaded verb or usage context to anchor it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but the definition is still incomplete: parameters are undocumented, there is no usage guidance against sibling tools, and no behavioral traits are disclosed despite the absence of annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Three parameters (`device`, `filter`, `indent`) have 0% schema description coverage, and the description mentions none of them. 'Serialized as JSON text' loosely gestures at formatting but explains neither the boolean `filter` nor the `indent` width, so an agent cannot infer their effect.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a noun phrase ('The UI snapshot serialized as JSON text') with no verb, so it only passively restates the title 'Ui Json' with a little extra detail. It gives no way to tell this apart from close siblings such as `ui`, `ui_with_recovery`, or `get_ui_tree`, which also deal with UI state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative guidance. With several overlapping UI-retrieval siblings present, the agent is left to guess which of `ui_json`, `ui`, `ui_with_recovery`, or `get_ui_tree` to invoke.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui_with_recoveryUi With RecoveryC

UI snapshot that retries past a dead or empty accessibility tree.

ParametersJSON Schema
NameRequiredDescriptionDefault
deviceNo
filterNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose one genuine trait beyond structured data — automatic retry on a dead/empty tree — but says nothing about permission needs, timeouts, how many retries, or failure modes. That is a real contribution but far from complete coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded clause with no filler; the recovery behavior is stated immediately. It is efficient, though almost to the point of under-specification rather than exemplary concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but with zero annotation coverage and 0% parameter description coverage the definition leaves meaningful gaps: what the two parameters do, and how the retry/recovery behaves on repeated failure. For a 2-param, no-annotation tool this is too thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both 'device' and 'filter'. It adds no meaning at all for either parameter — an agent cannot tell what 'device' accepts or what 'filter=true' actually filters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete verb and resource ('UI snapshot') and adds the distinctive behavioral trait 'retries past a dead or empty accessibility tree,' which differentiates it from the plain 'ui', 'ui_json', and 'get_ui_tree' siblings. It never names those siblings explicitly, so the differentiation must be inferred from the recovery clause.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'retries past a dead or empty accessibility tree' implies the use case (when a normal snapshot might fail) but gives no explicit when-to-use, when-not-to-use, or sibling comparison against ui/ui_json. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

uninstall_appUninstall AppD

Uninstall an app.

ParametersJSON Schema
NameRequiredDescriptionDefault
app_idYes
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but discloses nothing beyond the basic action. It does not mention whether uninstallation is destructive, whether user data is removed, what permissions are required, or how the optional device parameter affects behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is concise but severely under-specified for a two-parameter destructive operation. Brevity here reflects missing information rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists to cover return values, the description is incomplete for a destructive app-removal tool with no annotations and undocumented parameters. It omits usage context, safety implications, and parameter details that an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description adds no meaning for either parameter. It does not explain the required app_id or the optional device parameter, leaving both undocumented in the schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Uninstall an app.' restates the tool name and title without adding specificity. It does not distinguish this from siblings like install_app, stop_app, or launch_app, leaving the agent with only the verb to infer purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as stop_app or install_app. The description provides no context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_actionValidate ActionB

Pre-check a planned action against the safety policy (and, for our action set, that its target exists) without doing it. gesture_type is the action (tap, type_text, launch_app, open_deeplink, ...); target is the app name/package, deep-link URI or text you plan to use. Returns allowed=false with a category for blocked apps and text.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
uriNo
textNo
actionNo
deviceNo
som_idNo
targetNo
packageNo
app_nameNo
gesture_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose key traits: it is non-mutating ('without doing it'), it may check target existence, and blocked inputs return allowed=false with a category. It omits whether policy evaluation needs network/device state or how partial/ambiguous actions are treated, so it is strong but not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then parameter hints in a compact three-clause sentence. No filler, though the parenthetical 'and, for our action set, that its target exists' is slightly convoluted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description does cover the central semantics. However, for an 11-parameter, zero-coverage, unannotated tool, the majority of inputs remain ambiguous, leaving the definition only adequately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 11 parameters, so the description must compensate and only partially does: it explains gesture_type and target (app name/package, deep-link URI, or text), but leaves x, y, uri, text, action, device, som_id, package, and app_name entirely undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Pre-check a planned action against the safety policy... without doing it.' The dry-run nature is explicit, so an agent can distinguish it from the execution siblings (tap, launch_app, type_text) even though it never names verify_action directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'pre-check ... without doing it' framing implies it should be called before performing an action, but there is no explicit when-to-use/when-not statement and no reference to the obvious alternative verify_action. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_actionVerify ActionC

Check an outcome against the live screen. kind: text (visible), gone (not visible), app (foreground package or name), activity, changed (the last action changed the screen).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNotext
deviceNo
timeoutNo
use_ocrNo
expectedYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the transparency burden. It usefully discloses behavioral semantics such as 'gone (not visible)', 'app (foreground package or name)', and 'changed (the last action changed the screen)'. However, it does not explain matching behavior, side effects, permission needs, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loads the core purpose before giving the kind list. It contains no filler, though the single run-on sentence is dense and could be structured more clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the tool context is underspecified: no annotations, no guidance for key parameters, no alternative routing, and only minimal behavioral detail. An agent cannot fully determine correct usage, especially around OCR, device targeting, and timeout behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all five parameters. It only adds real meaning to 'kind' and partially to 'expected' by implication. 'device', 'timeout', and 'use_ocr' receive no explanatory treatment, leaving the agent to guess their semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: verify an outcome against the live screen. It enumerates supported verification kinds (text, gone, app, activity, changed), making its purpose concrete. However, it does not distinguish itself from the similarly named sibling 'validate_action'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus sibling alternatives like validate_action, wait_for, or read_screen. It implies usage through the kind list, but does not state exclusions, prerequisites, or recommended scenarios beyond the terse kind definitions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

volume_downVolume DownC

Lower the music volume by steps.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsNo
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose potential side effects (e.g., volume range limits, whether it affects media sessions), device targeting behavior, or what happens if steps exceeds available volume.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely brief (one sentence) and front-loads the key action. The backticks around 'steps' are inconsistent but minor. It is concise but lacks critical details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema but no annotations, and the description must cover behavioral context. It omits device semantics, possible errors, and effects on other volumes, making it incomplete for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description only explains 'steps' implicitly through the verb, but 'device' is entirely unexplained. There is no guidance on how device affects the operation or the units/range for steps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (lower), the resource (music volume), and the controlled parameter (steps). It distinguishes from volume_up and mute by indicating it decreases volume, though it doesn't explicitly mention alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for lowering music volume but provides no explicit context on when to use this tool versus volume_up or mute, and no mention of device selection or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

volume_upVolume UpC

Raise the music volume by steps.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsNo
deviceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states 'Raise the music volume' and does not mention volume limits, device selection behavior, whether it is a media vs system volume change, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short, front-loaded sentence with no filler or redundancy. It is concise, though somewhat under-specified relative to the tool's parameter set.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is too terse to be complete for a tool with no annotations, two parameters with no schema descriptions, and no usage guidance. It omits device selection semantics, volume boundaries, and any interaction with sibling volume controls, leaving an agent with important gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal meaning for the 'steps' parameter by indicating it is the increment amount, but it offers no details on range, units, or behavior when the parameter is omitted. The 'device' parameter is entirely unexplained, and schema description coverage is 0%, so the description does not compensate for that gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Raise') and the target resource ('the music volume'), with the increment amount parameter 'steps'. It is distinguishable from sibling tools like volume_down and mute through the direction of the action, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives such as volume_down, mute, or media_control. There are no conditions, exclusions, or context cues beyond the implicit meaning of the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_forWait ForA

LONG waits only (downloads, uploads, processing, status changes); gestures already settle. With text / package / activity: wait until that is on screen (gone=true: until it disappears). Otherwise wait until loading finishes (no progress bar or "Loading" text) and the screen is still; condition is echoed back, you judge the returned state. Timeouts: timeout_ms (default 5000, max 30000) / poll_interval_ms (default 500, min 100), or timeout / interval in seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
goneNo
textNo
deviceNo
packageNo
timeoutNo
activityNo
intervalNo
conditionNo
timeout_msNo
poll_interval_msNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does well: it discloses that the condition is echoed back and that the agent must judge the returned state, and it documents timeout defaults/limits (5000ms default, 30000ms max) and polling behavior (500ms default, 100ms min). It omits failure semantics—what the return looks like when a timeout expires.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the scoping rule and then the condition semantics, then timeouts—a sensible order. The middle sentence is long and packs several clauses, but each sentence contributes information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, and it covers the condition model, mutual exclusion of text/package/activity, and timing controls adequately for a 10-param tool. The main remaining gap is the undocumented device parameter and timeout-failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 10 parameters, so the description must compensate, and it explains the semantics of text/package/activity (wait until on screen), gone (invert to until-disappears), timeout_ms/poll_interval_ms defaults and bounds, and the timeout/interval seconds aliases. It leaves device and condition's accepted formats underspecified, so it is strong but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: wait until a condition holds, with concrete condition categories (downloads, uploads, processing, status changes). It distinguishes itself from gesture tools by declaring gestures already settle, but does not explicitly differentiate from near-siblings like wait_for_text, wait_for_idle, or wait_for_screen_change.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use constraint ('LONG waits only') and a when-not-to-use rule ('gestures already settle'), plus branch conditions for text/package/activity vs. generic loading. It stops short of naming the sibling tools that cover the excluded quick-settle cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_appWait For AppC

Wait until app_id is in the foreground.

ParametersJSON Schema
NameRequiredDescriptionDefault
pollNo
app_idYes
deviceNo
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and provides only the wait condition. It omits what happens on timeout (raise, return false, hang?), whether the call blocks, and how 'foreground' is determined — all critical for a blocking wait tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. Its brevity is a virtue structurally, though it borders on under-specification rather than optimal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, but for a 4-parameter blocking wait with no annotations the description should cover timeout behavior and polling. As written it is too thin for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters. The description implies app_id is the target but gives no format or matching semantics, and says nothing about poll (default 0.5) or timeout (default 10), leaving half the parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Wait') and resource ('app_id ... in the foreground'), which distinguishes it from siblings like wait_for_text, wait_for_idle, and current_app_id. It is clear what the tool does, though it doesn't explicitly contrast itself with the other wait_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus wait_for, wait_for_idle, wait_for_screen_change, or current_app_id, and no mention of prerequisites or the timeout failure mode. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_idleWait For IdleC

Wait until the UI stops changing.

ParametersJSON Schema
NameRequiredDescriptionDefault
pollNo
deviceNo
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a blocking wait until the UI is stable but says nothing about polling behavior, timeout handling, what happens if the timeout is reached, or whether device selection affects the wait. The presence of an output schema covers return values but not runtime behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence and is front-loaded, but it is under-specified rather than appropriately concise. For a tool with multiple parameters and no annotations, the brevity leaves critical usage and behavioral details unstated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with no annotations and three undocumented parameters, the description is not complete enough for an agent to invoke the tool confidently. It omits parameter meanings, timeout behavior, and differentiation from sibling wait tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are 3 parameters with 0% schema description coverage, and the description does not mention any of them. The meaning and effect of poll, device, and timeout must be inferred entirely from their names and default values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Wait until the UI stops changing.' An agent can understand the core action. However, it does not distinguish this tool from siblings like wait_for_screen_change, wait_for_nodes, or wait_for, which appear to address similar waiting needs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives. It does not mention prerequisites, when idle waiting is appropriate, or how it differs from other wait-related sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_nodesWait For NodesC

Poll find_nodes until something matches (or timeout, returning []).

ParametersJSON Schema
NameRequiredDescriptionDefault
descNo
pollNo
textNo
deviceNo
timeoutNo
on_screenNo
class_nameNo
resource_idNo
any_containsNo
desc_containsNo
text_containsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the timeout branch and its return value ([]), which tells the agent this call can fail non-exceptionally after blocking. However, it omits poll-interval semantics, whether it blocks the caller, and which underlying find_nodes parameters are forwarded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the core mechanism first and the fallback behavior in parentheses. Nothing is wasted, though the extreme brevity leaves meaning unstated rather than trimmed of redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter, zero-coverage blocking tool with children semantics inherited from find_nodes, one sentence is not enough. The presence of an output schema excuses it from describing return shape, but the parameter model, polling behavior, and relationship to the near-identical wait-style siblings remain unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 11 parameters (desc, text, device, on_screen, timeout, poll, class_name, resource_id, and the *_contains variants), yet the description explains none of them and only implies they are forwarded to find_nodes. The agent gets no help distinguishing text vs text_contains, on_screen, or the units/defaults of poll and timeout.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific mechanism (poll find_nodes until a match) and names the base tool it wraps, so the agent knows this is a blocking/polling variant of find_nodes rather than a one-shot query. It stops short of differentiating itself from other wait-style siblings such as wait_for_text or wait_for_screen_change, which also match on conditions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to choose this over find_nodes, find_nodes_on_screen, wait_for_text, or wait_for_screen_change. The 'until something matches' phrasing implies a wait-for-appearance use case, but the agent must infer the selection criteria from sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_screen_changeWait For Screen ChangeC

Wait until the UI differs from now.

ParametersJSON Schema
NameRequiredDescriptionDefault
pollNo
deviceNo
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses almost nothing: what happens on timeout (error, null, or stale return), whether it captures a baseline snapshot first, and how polling interacts with the wait are all unstated. 'Differs from now' is the only behavioral hint, implying a comparison against a captured initial state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no wasted words, but its brevity is under-specification rather than efficiency given the tool's 3-parameter surface and many siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be spelled out, but the absence of annotations and any parameter or timeout-behavior explanation leaves the definition incomplete for a tool whose semantics hinge on polling and timeout.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters. The description never mentions poll, timeout, or device, so nothing compensates for the undocumented schema — an agent cannot know what units 'poll'/'timeout' use or what 'device' selects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'wait' plus the resource 'screen change' make the core action understandable, and 'differs from now' clarifies that it blocks on a UI delta. However, it never distinguishes itself from the many sibling wait tools (wait_for, wait_for_idle, wait_for_text, wait_for_nodes), so an agent cannot tell which waiter to pick from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no conditions, and no named alternative. Given the crowded sibling set of wait_* and perceive/read_screen tools, the description leaves the routing decision entirely to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_textWait For TextC

Wait until a node containing text exists (off-screen nodes count).

ParametersJSON Schema
NameRequiredDescriptionDefault
pollNo
textYes
deviceNo
timeoutNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose one real trait not in the schema (off-screen nodes count), but omits what happens on timeout (throw vs. return false), polling behavior, and whether a non-visible/off-screen match is treated as success for the caller.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the parenthetical scoping note is the only elaboration and it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but with 0% parameter coverage, no annotations, and overlapping siblings, the description is too thin for a timing/wait tool where timeout and failure semantics matter most.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description only implicitly references 'text'. It says nothing about poll, timeout, or device, leaving three of four parameters undocumented in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: wait until a node containing text exists, and adds the scope qualifier that off-screen nodes count. However it does not distinguish itself from close siblings like wait_for_nodes or assert_text_visible, so the agent gets a clear purpose but no routing signal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative is named despite several overlapping siblings (wait_for, wait_for_nodes, wait_for_idle, assert_text_visible). The agent must infer when this is preferable to a general wait or an assertion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_device_eventsWatch Device EventsA

Collect device events for up to timeout_seconds (default 10, max 30), returning early once max_events (default 50) arrive: foreground app, keyboard, screen content, notifications posted/removed. kinds filters: foreground, keyboard, screen, notifications.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindsNo
deviceNo
durationNo
intervalNo
max_eventsNo
timeout_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it specifies the default and maximum timeout, the default max_events, and the early-return condition. It also lists the exact event types and available filters, giving the agent a clear picture of runtime behavior, though it omits any mention of permissions or resource consumption.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tightly packed sentences with no filler, front-loading the core collection behavior and constraints before listing filters. Every phrase adds useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the core streaming behavior is covered. However, with no annotations and 0% schema description coverage, the omission of device, duration, and interval leaves the definition incomplete for a six-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains three of six parameters (timeout_seconds, max_events, kinds). It leaves device, duration, and interval completely unexplained, missing key semantics for half the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Collect') and resource ('device events') and enumerates the event categories (foreground app, keyboard, screen content, notifications). It is clear what the tool does, but it does not explicitly differentiate from sibling tools that also deal with events or notifications.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by detailing the streaming behavior and filterable kinds, but it never states when to choose this tool over alternatives like read_notifications or wait_for_screen_change. There are no explicit when-to-use or when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 94 tool updatesv0.1.0
    • First observedassert_on
    • First observedassert_text_visible
    • First observedcapabilities
    • First observedclear_input
    • First observedconnect_device
    • First observedcurrent_app_id
    • First observeddisconnect_device
    • First observeddismiss_notification
    • First observeddoctor
    • First observeddouble_tap
    • First observedecho
    • First observedend_session
    • First observedexecute_script
    • First observedfind_files
    • First observedfind_nodes
    • First observedfind_nodes_on_screen
    • First observedget_clipboard
    • First observedget_device_status
    • First observedget_media_sessions
    • First observedget_screenshot
    • First observedget_ui_tree
    • First observedget_usage_guide
    • First observedgrant_permission
    • First observedinstall_app
    • First observedkey
    • First observedlaunch_app
    • First observedlist_app_deeplinks
    • First observedlist_apps
    • First observedlist_devices
    • First observedlong_press
    • First observedlong_press_at
    • First observedlookup_app
    • First observedmark_step
    • First observedmedia_control
    • First observedmute
    • First observednotification_action
    • First observedopen_and_settle
    • First observedopen_deep_link
    • First observedopen_deeplink
    • First observedopen_file
    • First observedopen_recent_apps
    • First observedperceive_screen
    • First observedping_device
    • First observedpress
    • First observedpress_back
    • First observedpress_enter
    • First observedpress_home
    • First observedread_notifications
    • First observedread_screen
    • First observedrecord_finding
    • First observedrequest_screen_capture_permission
    • First observedresolve_contact
    • First observedresolve_deeplink
    • First observedscreen_size
    • First observedscreenshot
    • First observedscreenshot_path
    • First observedscroll
    • First observedscroll_down
    • First observedscroll_left
    • First observedscroll_right
    • First observedscroll_to
    • First observedscroll_until
    • First observedscroll_up
    • First observedset_clipboard
    • First observedset_plan
    • First observedsetup_portal
    • First observedstart_app
    • First observedstop_app
    • First observedsupports
    • First observedswipe
    • First observedsystem_intent
    • First observedtap
    • First observedtap_and_wait
    • First observedtap_node
    • First observedtap_text
    • First observedtime
    • First observedtype
    • First observedtype_text
    • First observedui
    • First observedui_json
    • First observedui_with_recovery
    • First observeduninstall_app
    • First observedvalidate_action
    • First observedverify_action
    • First observedvolume_down
    • First observedvolume_up
    • First observedwait_for
    • First observedwait_for_app
    • First observedwait_for_idle
    • First observedwait_for_nodes
    • First observedwait_for_screen_change
    • First observedwait_for_text
    • First observedwatch_device_events
    • First observedweb_search

TDQS

C2.6/5.0

Scored across 94 tools

Disambiguation2/5

Many tools overlap heavily: ui/ui_json/ui_with_recovery/get_ui_tree/perceive_screen/read_screen all return screen state, screenshot/screenshot_path/get_screenshot are near-duplicates, and tap/tap_text/tap_node/tap_and_wait plus press/press_home/press_back/press_enter/key blur action boundaries. The descriptions do hint at distinctions (raw vs annotated, coordinate vs text), but the sheer number of near-synonyms with legacy aliases ('kept for old clients', 'prefer ...') makes misselection likely.

Naming Consistency4/5

Almost everything follows a snake_case verb_noun convention (get_screenshot, launch_app, list_devices, press_home). Deviations are minor: bare nouns (ui, capabilities, time, echo), a single-word verb 'press' vs compound 'press_home', and a few legacy synonyms. Overall predictable and readable.

Tool Count1/5

94 tools is far beyond what a device-control surface needs, and much of the count comes from redundant aliases and overlapping variants (multiple screenshot tools, multiple screen-read tools, multiple tap/wait/key variants). This is an extreme mismatch that burdens selection.

Completeness5/5

Coverage is exhaustive: input gestures, UI perception, app lifecycle (install/uninstall/stop/grant), files, notifications, media, clipboard, device management, deep links, safety validation, and a task-planning/verification loop. Nearly every lifecycle operation for the domain is present with no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Enables AI agents to interact with Android devices through UI manipulation, screen capture, touch gestures, text input, and app management via ADB. Provides comprehensive mobile automation capabilities including element detection, navigation, and application control for Android device testing and interaction.
    9
    4
    -
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to fully control Android devices through over 30 tools for app management, UI automation, and vision-based analysis via ADB. It supports multi-device management, action recording, and smart execution strategies ranging from UI hierarchy parsing to coordinate-based interaction.
    37
    120 npm
    1
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to control, inspect, and automate Android devices, Waydroid containers, and AVD emulators over ADB. Provides tools for screenshots, UI-hierarchy text-based tapping, gestures, key presses, text input, app management, and raw shell commands.
    12
    MIT