Skip to main content
Glama
CT-mao

uitree-local

by CT-mao

uitree-local — 本地 Uitree Portal 工具链

为精简为本地-only 的 Uitree Portal APK(v0.7.24)配套的 MCP server、CLI agent 与 skills。全部流量走本机(HTTP 8080 / WebSocket 8081),无任何云端依赖。

  • MCP server 基于 mcp 2.x SDK,实现 2026-07-28 最新协议: 无状态(无 handshake / Session-Id)、server/discover 服务发现、JSON Schema 2020-12。

  • 27 个工具覆盖:UI 树读取、截图、点按/滑动/按键、输入、剪贴板、App 管理、文件、 APK 安装、防休眠。

  • agent/ 提供自然语言任务循环(OpenAI 兼容 LLM,可选 vision),与 MCP 共用同一套工具。

官方 api.uitree.ai/v1/mcp 是托管闭源服务(dr_sk_ key 计费)。 本项目是自建替代:MCP server + skills + agent 全部在本地跑,直连 Portal APK 的本地 API。

目录结构

mcp-server/
├── pyproject.toml            # mcp>=2.0.0, websocket-client, httpx; pytest dev
├── src/uitree_mcp/
│   ├── portal_client.py      # Portal HTTP/WS 客户端(token 解析 + adb forward 自动)
│   ├── tools.py              # 27 个工具定义(MCP server 与 agent 共用)
│   ├── server.py             # MCP server(stdio / streamable-http)
│   └── agent.py              # CLI agent(LLM 循环 + vision 截图)
├── skills/uitree/SKILL.md # opencode/anthropic 风格 skill
├── agent/                    # agent 说明与示例
└── tests/                    # 17 个单测(假 Portal + 假 LLM,无需真机)

Related MCP server: waydroid-mcp

安装 & 运行

前置:adb devices 可见设备(如 emulator-5554);Portal APK 已安装并启用无障碍服务; token 自动解析(显式 > UITREE_TOKEN > adb content provider),端口自动 forward。

cd mcp-server
uv sync

# 1) MCP server(stdio —— 供 Claude Desktop / opencode 等客户端接入)
uv run uitree-mcp
# 客户端配置示例(claude mcp add / opencode mcp add):
#   uitree-local: {"type": "stdio", "command": "uv", "args": ["run", "--project", "/path/to/mcp-server", "uitree-mcp"]}

# 2) MCP server(streamable-http,2026-07-28 无状态)
uv run uitree-mcp --transport streamable-http --port 8083
# 客户端连接: http://127.0.0.1:8083/mcp

# 3) CLI agent(自然语言任务;LLM 用 OpenAI 兼容端点,Ollama 也可)
UITREE_LLM_BASE=https://api.openai.com/v1 UITREE_LLM_MODEL=gpt-4o \
UITREE_LLM_API_KEY=sk-... \
uv run uitree-agent "打开设置,滑动到存储,截图汇报占用情况"

# Ollama 本地:
UITREE_LLM_BASE=http://localhost:11434/v1 UITREE_LLM_MODEL=llama3.2-vision \
UITREE_LLM_VISION=1 uv run uitree-agent "看下当前屏幕有什么"

验证

uv run pytest tests/            # 17 个单测(无需设备)
curl -s http://127.0.0.1:8080/ping -H "Authorization: Bearer $KEY"   # → pong
uv run uitree-mcp --transport streamable-http --port 8083 &       # 手工联调

关键设计

  • 单一工具源:tools.py 的 ToolDef 同时驱动 MCP server(server.py 把 schema 逐字映射为 wire inputSchema)与 agent(进程内直调,零 MCP 往返)。

  • token 安全:token 不出本机;uitree-mcp 也可用 --token 显式传入,避免 adb shell。

  • WS 专有能力:install_apk、set_keep_awake 走 WebSocket JSON-RPC({id, method, params})。

  • 协议最新:mcp 2.x 为 2026-07-28 规范重写的 SDK(MCPServer 取代 FastMCP), 自动处理 server/discover、响应缓存、无状态 HTTP。

Available Tools

30 tools
delete_fileC

Deletes a file on the device.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states the mutation but omits whether deletion is permanent/irreversible, what happens if the path does not exist, error behavior, and whether special permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words, which is structurally clean. Its brevity, however, reflects under-specification rather than deliberate concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no annotations, no output schema, and an undocumented parameter, the description is far too thin. It should at minimum cover path conventions, irreversibility, and error handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the single required 'path' parameter is documented nowhere. The description does not say whether the path is absolute or relative, what format is expected, or which directory roots are valid, leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (deletes) and resource (a file on the device), which is unambiguous. However, it offers no differentiation from the nearby file siblings (read_file, write_file, list_files) and no scoping detail such as whether it targets app sandbox or device-wide paths.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, no prerequisites, and no mention of alternatives or sibling tools. The agent must infer everything about usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_clipboardB

Reads the current clipboard content.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. 'Reads' implies a safe, non-destructive operation, but the description does not state what happens with an empty clipboard, whether it requires the app to be focused, or what format the content is returned in.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero filler. Every word earns its place for a tool of this simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should ideally describe the return value (text? image? MIME type?) and any failure conditions. For a simple zero-parameter read, the one-liner is minimally adequate but leaves the return contract unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty with 100% coverage, so there is nothing for the description to clarify. Baseline 4 applies since no parameter explanation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Reads') and resource ('current clipboard content'), so an agent immediately knows what it returns. It does not explicitly name the contrast with the sibling set_clipboard, but the read/write distinction is self-evident from the verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to, or how it relates to siblings like set_clipboard or keyboard_input. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_element_treeB

Returns only the accessibility element tree of the current screen as JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoReturn only the simplified tree (default false).

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only operation by saying 'Returns,' but does not state permissions, side effects, limitations, or how the filter setting changes behavior beyond what the schema already says.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It communicates the return type and scope immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter getter with full schema coverage, the description is adequate but thin. It does not explain the returned JSON shape, the effect of the filter parameter, or how this differs from sibling screen-inspection tools, and there is no output schema to compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single filter parameter is already documented structurally. The description does not add any meaning about the filter beyond the schema, which matches the baseline score for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: it returns the accessibility element tree of the current screen as JSON. The word 'only' hints that this is narrower than a full screen-state capture, but no sibling tool is named explicitly, so it falls short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool returns but not when to use it instead of alternatives like get_screen_state or get_screenshot. There is no explicit when-to-use, when-not-to-use, or prerequisite guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_keep_awake_statusA

Returns whether screen keep-awake is currently active.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral burden. 'Returns' implies a read-only query, but it does not explicitly state that it has no side effects, required permissions, or return type.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded with the core purpose and no wasted words. Appropriate for a zero-param status getter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the description conveys the status query, but because there is no output schema and no annotations, the lack of explicit return type/format leaves a small gap. Still adequate for a simple boolean getter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has zero parameters with 100% description coverage, so the schema provides all necessary structural information. Per the rubric, zero params baseline is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Returns') and resource ('whether screen keep-awake is currently active'), which is distinct from sibling set_keep_awake and get_screen_state. An agent can identify it as a status query without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or alternatives named. The getter nature makes usage implied, but the description does not state when to call it (e.g., before set_keep_awake) or that no alternatives exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_phone_stateA

Returns phone state: current app, screen on/off, battery, connectivity.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Returns' correctly signals a read-only query, and the enumerated fields tell the agent what is disclosed, but there is no mention of permissions, cost, rate limits, or whether reads are cached/live.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the return payload with zero filler. Nothing can be trimmed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter getter with no output schema, the description enumerates the meaningful return fields, which is exactly what the agent needs to decide and consume the call. It falls just short of complete because it never addresses the overlap with get_screen_state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to document; the baseline for a 0-param tool applies. The description correctly implies no inputs are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Returns') and resource ('phone state') and enumerates the returned fields (app, screen, battery, connectivity). However, it does not differentiate itself from the sibling get_screen_state, which overlaps on the screen on/off signal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no alternative named, and no mention of the relationship with get_screen_state or other state-query siblings. The usage context is only implicit from the tool's purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_screenshotA

Takes a screenshot of the device screen and returns it as an image. Use this for visual verification of UI before/after actions.

ParametersJSON Schema
NameRequiredDescriptionDefault
hide_overlayNoHide the Portal overlay in the screenshot (default true).

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the return type (image) and that it captures the device screen, but does not state read-only safety, permissions, failure modes, or image formatting details. It adds basic behavioral context but leaves several traits implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the action and return value, followed by the intended use. No redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one optional parameter, no nested objects) and the schema fully documents its only parameter. The description explains the core action, output type, and use case, but does not elaborate on image format or read-only behavior; given the low complexity and schema coverage, it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the single optional parameter hide_overlay is fully described in the schema, including its default and effect. The description adds no parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('takes') and resource ('screenshot of the device screen') and clarifies the output ('returns it as an image'). It implies differentiation from get_screen_state through visual verification, but it does not explicitly name the sibling, so it is clear but not fully sibling-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs to 'Use this for visual verification of UI before/after actions,' giving clear usage context. It does not state when not to use it or name alternatives like get_screen_state, so no exclusion guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_screen_stateA

Returns the current UI accessibility tree (text, class names, bounds) plus phone state and screen dimensions. Call this first to understand what is on screen, then use tap / swipe / keyboard_input with coordinates taken from the returned bounds.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoReturn only the simplified tree (default false).

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return payload and implies a read-only observation, but says nothing about permissions, rate limits, or whether the tree can be stale or large. It is enough to call the tool safely but thin on behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler. The primary purpose leads and the workflow directive follows, which is the right front-loading for an orientation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description does the necessary work of naming returned fields and describing how the result is consumed downstream. It is nearly complete for a zero-required-param read tool, though it omits the filter toggle that affects the payload.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single filter parameter is fully documented in the schema, so the baseline is 3. The description never mentions the filter option or what 'simplified tree' means, adding no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns the current UI accessibility tree') and enumerates the returned content (text, class names, bounds, phone state, screen dimensions). It does not explicitly differentiate itself from overlapping siblings like get_element_tree, get_phone_state, or get_screenshot, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear ordering directive ('Call this first to understand what is on screen') and names the downstream tools that consume its output (tap / swipe / keyboard_input via returned bounds). Strong context, but no exclusion criteria for when to prefer get_screenshot, get_element_tree, or get_phone_state instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_versionA

Returns the installed Portal app version.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the burden; 'Returns' correctly signals a non-mutating read. However it does not state the return shape (e.g., a version string) or whether it reads from the running app vs. an installed package, which matters when 'installed' could be ambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivial zero-arg getter with no output schema, the description is nearly complete. The one residual gap is that, absent an output schema, it could say what the return value looks like, but the complexity is low enough that this barely matters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline 4 case; there is nothing for the description to disambiguate. The empty schema is fully consistent with the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Returns') and a precise resource ('the installed Portal app version'). It is unambiguous against siblings like list_installed_apps or get_phone_state, since none of them return the app's own version.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance is given. For a zero-parameter read-only getter the usage context is largely self-evident, so the omission is minor, but the description leaves it to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_apkA

Installs one or more APKs from URLs on the device (WebSocket required). Automatically accepts the install confirmation via the Portal.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
hide_overlayNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a genuinely non-obvious side effect: the install confirmation is automatically accepted via the Portal, meaning the user will not be prompted. It also flags the WebSocket dependency. It still omits failure modes, permissions, and whether the call blocks until install completes, so it is good but not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences with the core action front-loaded and qualifications appended. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no annotations, no output schema, and 0% schema description coverage, the description covers the action and one important side effect but leaves hide_overlay, install outcome/error behavior, and timing semantics undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 2 parameters, so the description must compensate. It explains the gist of 'urls' ('one or more APKs from URLs') but gives no format expectations, and it says nothing at all about the hide_overlay parameter or its default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (installs), resource (APKs from URLs), and target (on the device). No sibling tool performs installs, so it is trivially distinguishable from launch_app, list_installed_apps, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(WebSocket required)' gives a prerequisite, which is useful. However, there is no explicit when-to-use framing, no statement about when it would fail or be inappropriate, and no routing to alternatives (none exist, but the agent is left to infer this).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is_overlay_visibleA

Returns whether the Portal overlay is currently visible.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, but the phrasing 'Returns whether...' clearly signals a non-mutating read query, which is the key behavioral trait. It does not clarify what counts as 'visible' (e.g., occlusion, transparency) or the freshness of the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter boolean query with no output schema and no annotations, the description is nearly sufficient — it states the return meaning. A brief note on what 'visible' means or on related tools would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema baseline is 4. There are no parameter semantics for the description to add or omit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (returns) and resource (whether the Portal overlay is visible), so the agent knows exactly what question it answers. It does not distinguish itself from siblings like get_screen_state or get_element_tree, which could also reveal overlay state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as get_screen_state, nor any preconditions. The agent must infer usage entirely from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

keyboard_clearC

Clears the currently focused text field.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNo复查屏幕树确认变更生效 (default true).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full disclosure burden. It does not say whether existing content is destroyed irreversibly, what happens if no field is focused, or that the verify parameters trigger screen-tree polling, which is meaningful behavior for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no waste, but for a mutation tool with three parameters it is arguably too thin to be considered well-structured rather than merely brief.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and three verification parameters, the description should explain mutation semantics, failure behavior, and the verification flow. It covers none of these, leaving significant gaps for an agent deciding whether and how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the three verify-related parameters are already documented in the schema (including defaults). The description adds nothing about them, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: clearing the currently focused text field. It is clearly distinct from the text-entry siblings like keyboard_input and keyboard_key, though it never explicitly names them as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus keyboard_input or keyboard_key, and no stated prerequisite such as how the target field becomes focused. The phrase 'currently focused' hints at a precondition but leaves the agent to infer it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

keyboard_inputA

Types text into the currently focused field using the Portal keyboard bridge. Set clear=false to append instead of replacing the field content.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to type.
clearNoClear the field first (default true).
verifyNo复查屏幕树确认变更生效 (default true).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the meaningful mutation behavior (existing field content is replaced by default, appended with clear=false) and the bridge mechanism, but it is silent on the verification/polling behavior implied by verify, verify_poll_ms and verify_timeout_ms, and on what happens when no field is focused or verification fails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler, front-loading the action and target before the one actionable option. Nothing needs trimming and nothing important is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no annotations and no output schema, the description covers the core action and the destructive/append distinction but leaves the entire verification sub-behavior (verify, polling interval, timeout) unexplained and says nothing about error or no-focus outcomes. It is adequate but has clear gaps an agent would want filled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters, including defaults. The description only restates the `clear` semantics ('append instead of replacing'), which the schema covers as 'Clear the field first', and adds nothing about the verify parameters. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Types text into the currently focused field') and names the underlying mechanism (Portal keyboard bridge). It is clearly distinguishable in practice from keyboard_key and keyboard_clear, though it never explicitly differentiates itself from those siblings by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the tool operates on the 'currently focused field', which hints that focus must be established first (e.g. via tap/tap_element), and the clear flag tells you how to append. There is no explicit when-to-use statement, no guidance on choosing this over keyboard_key or set_clipboard, and no stated preconditions or failure conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

keyboard_keyB

Sends a key press. Use named keys: back, home, app_switch, menu, enter, del, tab, space, escape, search, volume_up, volume_down, arrow_up/down/left/right, or a raw Android keycode integer.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesNamed key or raw keycode (e.g. 'enter', '66').
verifyNo复查屏幕树确认变更生效 (default true).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. Beyond 'sends a key press' it discloses nothing about where the key goes (focused element?), what the verify behavior does, or any side effects or permissions, leaving the three verify parameters entirely unexplained at the description level.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action followed immediately by the accepted value domain. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations and no output schema, the description adequately covers the required key argument but says nothing about the verify mechanism that governs the other three parameters, so an agent must read the schema to understand default post-press verification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description still adds real value by listing the complete set of accepted named keys, which the terse schema example ('enter', '66') does not. It adds nothing about verify/verify_poll_ms/verify_timeout_ms, but those are schema-documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Sends a key press') and enumerates the accepted named-key domain, which concretely defines the tool's scope. It doesn't explicitly distinguish itself from the sibling keyboard_input or tap, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the key enumeration but never states when to choose this over keyboard_input (text entry) or tap (coordinate press). No exclusions, prerequisites, or alternative routing are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

launch_appB

Launches an app by package name (and optional activity). Use list_installed_apps to find package names.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNo复查屏幕树确认变更生效 (default true).
packageYesPackage name, e.g. com.android.settings.
activityNoOptional activity class name.
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).
stop_before_launchNoForce-stop the app before launching (default false).

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden but only says what gets launched. It does not disclose that verify polling, timeouts, and stop_before_launch exist as behaviors, nor what happens when the package is missing or what the call returns. Those traits live only in schema field descriptions, leaving the narrative thin for a side-effecting action tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, and the core action is front-loaded ahead of the helper reference. Nothing extraneous is included.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the primary parameters but leaves out the behavioral story that matters for a 6-parameter, annotation-free action tool with verification polling and stop_before_launch. It is adequate to invoke the tool, but not complete enough to predict its side effects or failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameter meaning is already fully documented, and the description only restates package/activity at a high level. Baseline 3 is appropriate since the description adds no format, defaulting, or interaction detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Launches an app by package name') with the optional activity qualifier, so an agent immediately knows the action. It names list_installed_apps as a helper for obtaining the package, but does not distinguish itself from action siblings that also open apps, such as open_deep_link.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'Use list_installed_apps to find package names' gives a concrete precondition for the required parameter, which is real guidance. However, it offers no when-not guidance and never contrasts this tool with open_deep_link, install_apk, or stop_app, so selection among overlapping siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_filesB

Lists files in a directory on the device.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute directory path.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states only that files are listed and does not disclose whether the operation is read-only, how hidden or special files are handled, sort order, error behavior, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single front-loaded sentence with no redundant or filler content. It communicates the core action efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read operation, the description is minimally viable. However, with no annotations and no output schema, it could do more to clarify return format, listing behavior, or edge cases that an agent might need when invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single 'path' parameter already documents that it expects an 'Absolute directory path.' The description adds no further parameter meaning or format details beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Lists files in a directory on the device.' This clearly distinguishes it from siblings such as read_file, write_file, and delete_file, but it does not explicitly differentiate itself from other listing tools or explain its scope beyond a single directory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance. The description implies the tool is for listing files, but it does not mention alternatives like read_file or list_installed_apps, nor does it specify prerequisites or appropriate contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_installed_appsA

Lists installed packages. Returns a large list; grep it for the app you want.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that the return is a large list with no pagination or filtering, which tells the agent to expect volume and post-process. However, it says nothing about ordering, format, or cost of repeated calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste; the purpose is front-loaded and the operational hint follows immediately. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and no parameters, the description gives enough to call the tool correctly: what it returns and that it will be large. It could be marginally more complete by naming the downstream tool for launching a found package.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There are no arguments whose meaning needs explaining, and the description correctly implies no filtering capability exists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: listing installed packages. It is clearly distinct from sibling list tools like list_files or get_element_tree. It only lacks an explicit contrast with those siblings to reach a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The advice 'grep it for the app you want' implies the intended workflow of finding an app before launching it, but it never names launch_app or states when this tool should be preferred over, say, directly calling launch_app with a known package. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

long_pressA

Long-presses at (x, y) for duration_ms (default 600). Opens context menus, selects text, or grabs an icon before dragging. For element targets with humanized random points use long_press_element.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
verifyNo复查屏幕树确认变更生效 (default true).
duration_msNoHold duration in ms (default 600).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral-disclosure burden. It states the default hold duration and common effects, but omits verification behavior, polling/timeout semantics, and any risk or mutating side-effect profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded core action, then effects, then sibling alternative. Every sentence earns its place with no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Core invocation and sibling routing are covered, but with no output schema and no annotations the description leaves verification parameters and result/return behavior unexplained. The schema partially compensates for optional parameter details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%. The description names x, y and repeats the duration_ms default, but does not add meaning for verify, verify_poll_ms, or verify_timeout_ms beyond what the schema descriptions already provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific verb 'long-presses' with coordinate and duration inputs, describes intended outcomes (context menus, text selection, drag grab), and names the sibling for element targets. An agent can distinguish it from long_press_element and other gesture tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear use cases (context menus, text selection, dragging) and explicitly routes element-target use to long_press_element. It does not cover prerequisites or when coordinate long press should be preferred over other siblings like swipe or tap beyond the element alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

long_press_elementA

Long-presses an element identified by bounds (random point inside the bounds, edge margin excluded). For e.g. icon long-press menus.

ParametersJSON Schema
NameRequiredDescriptionDefault
boundsYes[left, top, right, bottom],来自 get_screen_state 节点的 boundsInScreen。
marginNoExclude this many pixels from each edge (default 4).
verifyNo复查屏幕树确认变更生效 (default true).
duration_msNoHold duration in ms (default 600).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does add non-obvious behavioral detail: a random point inside the bounds is used, with edge margin excluded, and hints at the typical effect (context menus). However, it omits permissions, side effects, failure behavior, and verification semantics, so the disclosure is only partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and an immediately useful example. There is no filler, and the key implementation detail is placed early.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter UI action tool with full schema coverage but no annotations or output schema, the description covers core purpose and selection strategy. It does not mention the verification behavior available in the schema, which is a notable omission when no annotations are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameter meanings are already documented in the schema. The description adds only point-selection and edge-exclusion context, which slightly reinforces bounds/margin semantics but does not go beyond the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (long-presses) and target (element identified by bounds), and clarifies how the press point is chosen. It does not explicitly name or differentiate from the coordinate-based long_press sibling, so it is clear but not maximally distinguishable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an example use case ('icon long-press menus') but no explicit when-to-use versus alternatives or when-not-to-use guidance. Usage is implied rather than fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pingB

Checks whether the Portal device API is reachable and responding.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully implies the call is a non-mutating probe with no side effects, which is real behavioral signal, but says nothing about timeouts, failure modes, authentication requirements, or what a negative result means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every word earns its place for a zero-parameter health-check tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivially simple ping tool this is nearly sufficient, but with no output schema and no annotations, the agent still does not know the return shape (boolean? status object? latency?) or what counts as 'responding'. A brief note on the response would close the only real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case per the rubric. The description correctly implies no inputs are needed, and there is nothing further to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('checks') and resource ('Portal device API') plus the scope of the check (reachable and responding). It is unambiguous on its own, though it does not contrast itself with any sibling tool — none of the 30+ siblings overlap with a health check, so the omission is minor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance: it does not say whether this is a pre-flight connectivity check before other calls, a retry probe, or a diagnostic. No alternatives or exclusions are named. Usage is only faintly implied by the word 'checks'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_global_actionB

Performs a system-level action. Valid actions: back, home, recents, notifications, quick_settings, power_dialog, toggle_split_screen, lock_screen, take_screenshot, headphone_button, dock_and_hold. Uses Android GLOBAL_ACTION_* semantics, requires the Portal accessibility service to be enabled.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesOne of: back, home, recents, notifications, quick_settings, power_dialog, toggle_split_screen, lock_screen, take_screenshot, headphone_button, dock_and_hold.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully discloses the prerequisite ('requires the Portal accessibility service to be enabled') and the underlying GLOBAL_ACTION_* semantics, but says nothing about side effects of impactful actions (lock_screen, take_screenshot, power_dialog) or the verify/rollback behavior implied by the schema's verify fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then the valid action list, then the semantics/requirement note. Every sentence has a role, though the action enumeration duplicates the schema's own list, so it is slightly redundant rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema and no annotations, the description covers the prerequisite and action semantics adequately, but omits what happens on failure (e.g., whether an unsupported action errors, whether lock_screen/take_screenshot have special permission needs), which matters for these system-level actions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds the mapping to Android GLOBAL_ACTION_* constants, which gives real meaning to the action strings beyond the schema's bare enumeration, but adds nothing about the verify parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Performs a system-level action') and enumerates the exact valid actions, which lets an agent distinguish it from siblings like tap, swipe, and keyboard_key. It stops short of explicitly saying it is the navigation/global-action tool rather than the input-simulation tool, but the action list conveys that.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the action names (home, back, recents are clearly navigation actions), but there is no explicit guidance on when to prefer this over tap/swipe or keyboard_key for navigating the device's UI. No exclusions or prerequisites for choosing this tool are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_fileA

Reads a file from the device and returns its content base64-encoded as data URI. Decode it if you need the raw bytes.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute file path.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the important non-obvious behavior: the content comes back base64-encoded as a data URI. However it says nothing about failure modes (missing file, unreadable path), permissions, or size limits on large files.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, and the return-format caveat is front-loaded immediately after the core action so the agent knows what it will get before deciding to call it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must cover the return value, and it does via the base64/data-URI explanation. For a one-parameter read tool this is nearly complete; only error behavior for a bad path is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'path' param is documented as an absolute file path, so the schema does the work. The description adds no path-format or constraint detail beyond it, which is the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Reads) plus resource (a file from the device) and even the return encoding, which cleanly separates it from write_file, delete_file and list_files among the siblings. It does not name a sibling explicitly, so it stops just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is 'Decode it if you need the raw bytes,' which is a hint about consuming the output rather than when to choose this tool over list_files or get_clipboard. No exclusions or prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_clipboardC

Writes text to the device clipboard.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
verifyNo复查屏幕树确认变更生效 (default true).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden, and it discloses almost nothing beyond the mutation itself. It does not mention that verification polling runs by default (verify=true, 300ms, 3000ms timeout), whether existing clipboard contents are overwritten, or any permission/platform constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste, but the brevity is under-specification rather than disciplined conciseness for a 4-parameter tool with verification logic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and three behavioral parameters that materially change tool behavior (screen-tree verification with polling), the description is too thin to let an agent invoke this tool correctly or predict its effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, below the high-coverage threshold, and the description contributes no parameter meaning at all. The verify/verify_poll_ms/verify_timeout_ms semantics are only in the schema, and the description never explains that verification behavior is configurable or what polling accomplishes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Writes text to the device clipboard'), which is enough to distinguish it from the get_clipboard sibling by direction (write vs. read). It stops short of naming that sibling or describing scope, so it is clear but not maximally differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no prerequisites, and no mention that get_clipboard is the counterpart. The agent must infer the use case entirely from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_keep_awakeB

Prevents the screen from sleeping while the connection is active (WebSocket required).

ParametersJSON Schema
NameRequiredDescriptionDefault
enabledYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but does add real behavioral context: the effect is transient and scoped to the active connection, and a WebSocket is mandatory. It still omits what happens on disconnect, whether the device retains any state, or what permissions/errors are involved.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the core effect comes first and the prerequisite is appended compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple state-setting tool with no output schema, the description covers the effect and one prerequisite. It remains incomplete regarding the boolean's on/off semantics and the relationship to get_keep_awake_status, which an agent would need to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter with 0% schema description coverage, so the description must explain it. It never mentions the 'enabled' boolean nor clarifies that passing false releases the keep-awake lock, so the schema parameter remains semantically opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Prevents the screen from sleeping') and is clearly distinguishable from the sibling get_keep_awake_status. However, it only describes the enabling direction, while the required boolean parameter implies the tool can also turn keep-awake off, leaving that half of the purpose unstated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It discloses a precondition ('WebSocket required'), which is genuine usage context. But it never says when to call this versus get_keep_awake_status, nor when keep-awake should be set or unset, so the agent must infer the workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_overlay_offsetC

Adjusts the Portal floating-offset (only meaningful when the overlay is enabled).

ParametersJSON Schema
NameRequiredDescriptionDefault
offsetYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not say what units the offset uses, whether negative values are allowed, whether the change persists across sessions, or what happens if the overlay is disabled—all material for a mutating setter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with the important precondition attached rather than buried. Nothing is wasted, though the brevity is partly the cause of the semantic gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and an undocumented required parameter, the description is too thin. Units, valid range, and interaction with overlay visibility are all missing details the agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%: the single required 'offset' integer has no description, type range, or units. The description only restates it as a 'floating-offset', adding no meaning such as pixels, direction convention, or valid bounds, so the agent cannot choose a value confidently.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (adjusts) and resource (Portal floating-offset), and adds the precondition that it only matters when the overlay is enabled. No sibling tool overlaps with this function, so differentiation is not the issue; it stops short of five only because 'offset' itself is never explained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical gives a real condition ('only meaningful when the overlay is enabled'), which implies the agent should verify overlay state first. It never points at the obvious companion tool (is_overlay_visible) or states when not to call it, so usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stop_appC

Force-stops an app by package name.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNo复查屏幕树确认变更生效 (default true).
packageYes
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. 'Force-stops' hints at a mutating, potentially destructive operation, but the description says nothing about required permissions, side effects such as process termination or data clearing, or how the built-in verification works.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for a tool whose purpose is straightforward.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and multiple verification parameters, the description is too thin. It omits usage context, side effects, and verification behavior that an agent would need to invoke the tool safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the baseline is 3. The description mentions the required 'package' parameter ('by package name'), but adds no syntax or format details beyond what the schema already conveys, and it ignores the verify-related parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Force-stops an app') and scopes it by package name. It is clear what the tool does, but it does not differentiate itself from siblings like launch_app or get_screen_state, so it falls short of the explicit sibling-routing seen in top-scoring definitions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention when a forced stop is appropriate, when a graceful stop would be preferred, or what prerequisites exist before calling it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

swipeA

Swipes (or drags) from one point to another over the given duration. A slow swipe (>400ms) reads as a drag in most apps.

ParametersJSON Schema
NameRequiredDescriptionDefault
end_xYes
end_yYes
verifyNo复查屏幕树确认变更生效 (default true).
start_xYes
start_yYes
duration_msNoDuration in milliseconds (default 300).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses one genuine behavioral trait (slow swipes register as drags), which is valuable, but says nothing about the verification side effect (verify=true default), the coordinate space, or whether the gesture is a mutation with side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and followed by the actionable duration heuristic. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter gesture tool with no annotations and no output schema, the description covers the core motion and one tuning heuristic but omits explanation of the verify/verify_poll_ms/verify_timeout_ms side-effect machinery and coordinate expectations, leaving meaningful behavioral gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (duration_ms and the three verify params have descriptions; the four coordinate params do not). The description adds a meaningful threshold for duration_ms (>400ms = drag) but offers nothing for the coordinates or the verify/poll/timeout behavior, which is the real gap for an 8-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('swipes/drags from one point to another') with clear scope, and the parenthetical 'or drags' clarifies the gesture family. It implicitly distinguishes itself from tap/long_press siblings by naming the two-point motion, though it does not name a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note 'a slow swipe (>400ms) reads as a drag' gives useful context for choosing a duration, but there is no explicit when-to-use guidance relative to alternatives like tap, long_press, or tap_element, and no prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tapA

Taps at the given screen coordinates (pixels, 0,0 = top-left). Read coordinates from get_screen_state bounds. Prefer tap_element when you have a node's bounds: it picks the point randomly inside the element.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesX coordinate in pixels.
yYesY coordinate in pixels.
verifyNo复查屏幕树确认变更生效 (default true).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It usefully contrasts the exact-point tap with tap_element's randomized interior point, but says nothing about whether the tap blocks, whether it verifies the result, or what happens on a miss. The schema does document the verify/verify_poll_ms/verify_timeout_ms behavior, which partially compensates, but the description itself adds limited behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, no filler, with the core action and coordinate convention front-loaded and the sibling routing at the end. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter action tool with no output schema and no annotations, the description covers the essentials: what it does, coordinate convention, where to get coordinates, and when to use the sibling instead. Only the verification/wait behavior is left implicit and handled by the schema rather than the prose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description still adds value beyond the schema by clarifying that coordinates are pixels with origin at top-left, which the 'X coordinate in pixels' schema text does not establish, and by pointing to get_screen_state as the source of valid values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Taps at the given screen coordinates') and pins down the coordinate semantics ('pixels, 0,0 = top-left'). It also distinguishes itself from the sibling tap_element, so the agent can separate the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says where to get input ('Read coordinates from get_screen_state bounds') and names the alternative with the condition that selects it ('Prefer tap_element when you have a node's bounds'). Both when-to-use and when-to-prefer-another are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tap_elementA

Taps an element identified by its accessibility bounds. The tap point is chosen randomly inside the bounds after excluding margin pixels on each edge (center fallback for small elements), so repeated taps never land on identical pixels. Get bounds from get_screen_state.

ParametersJSON Schema
NameRequiredDescriptionDefault
boundsYes[left, top, right, bottom],来自 get_screen_state 节点的 boundsInScreen。
marginNoExclude this many pixels from each edge (default 4).
verifyNo复查屏幕树确认变更生效 (default true).
verify_poll_msNo轮询树的时间间隔毫秒 (default 300).
verify_timeout_msNo等待变更生效的超时毫秒 (default 3000).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful non-obvious behavior: the tap point is randomized within the bounds, margin is excluded on each edge, and small elements fall back to center. It does not cover failure behavior (e.g., element not found) or what the tool returns, which leaves some gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and immediately following with the mechanics and the bounds source. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-param mutation tool with no annotations and no output schema, the description covers the core mechanic and bounds source well. It stops short of describing the verify/poll behavior's purpose or return values, which the schema only partially compensates for.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all five parameters including margin, verify, and the timeout/poll defaults. The description's margin explanation ('excluding margin pixels on each edge') largely restates the schema rather than adding new meaning, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Taps') and resource ('element identified by its accessibility bounds'), which cleanly separates it from the coordinate-based `tap` and the `long_press`/`long_press_element` siblings. An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one concrete prerequisite ('Get bounds from get_screen_state'), which is genuinely useful routing guidance. However, it never states when to prefer this over `tap` (coordinate-based) or `long_press_element`, and offers no exclusions or failure conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_fileB

Writes base64-encoded bytes to a file on the device (max ~4MB decoded).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute file path.
data_b64YesBase64-encoded content.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two real constraints: the base64 input format and a ~4MB decoded size limit. It says nothing about overwrite vs. append behavior, required permissions, where the file lands, or what happens on failure — significant gaps for a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, and the encoding plus size constraint are front-loaded. Efficient, though it is arguably too terse for a write operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Parameters are fully covered by the schema and there is no output schema to explain, so the core is adequate. However, for a mutation tool with zero annotations, the missing overwrite/permission/error semantics leave the definition only minimally sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented in the schema (absolute path, base64 content). The description's encoding mention reinforces data_b64 but adds no format or boundary detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('writes ... bytes to a file on the device') plus the exact encoding contract (base64-encoded). It is distinguishable from read_file, delete_file and list_files by the verb, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and the encoding requirement, but there is no explicit when-to-use guidance, no prerequisites, and no statement about how this differs from any alternative. An agent can infer the purpose but gets no routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 30 tool updatesv0.1.0
    • First observeddelete_file
    • First observedget_clipboard
    • First observedget_element_tree
    • First observedget_keep_awake_status
    • First observedget_phone_state
    • First observedget_screen_state
    • First observedget_screenshot
    • First observedget_version
    • First observedinstall_apk
    • First observedis_overlay_visible
    • First observedkeyboard_clear
    • First observedkeyboard_input
    • First observedkeyboard_key
    • First observedlaunch_app
    • First observedlist_files
    • First observedlist_installed_apps
    • First observedlong_press
    • First observedlong_press_element
    • First observedopen_deep_link
    • First observedping
    • First observedpress_global_action
    • First observedread_file
    • First observedset_clipboard
    • First observedset_keep_awake
    • First observedset_overlay_offset
    • First observedstop_app
    • First observedswipe
    • First observedtap
    • First observedtap_element
    • First observedwrite_file

TDQS

B3.2/5.0

Scored across 30 tools

Disambiguation3/5

Several tools overlap: get_screen_state already includes element tree and phone state, making get_element_tree and get_phone_state potentially confusing; keyboard_key and press_global_action both handle back/home; tap vs tap_element and long_press vs long_press_element split the same action by target type. Descriptions clarify intended use, but an agent still faces multiple plausible choices.

Naming Consistency4/5

Names are consistently snake_case and mostly follow verb_noun or noun_verb patterns (get_screen_state, launch_app, tap_element). Minor deviations exist: keyboard_clear and keyboard_key put the noun first, and bare verbs tap/swipe/ping lack a noun. Overall readable and predictable.

Tool Count2/5

30 tools is high for a device-automation MCP and exceeds the typical 3-15 well-scoped range; several getters are redundant because get_screen_state already bundles element tree and phone state. This inflates surface area and increases selection cost.

Completeness4/5

The surface covers core device control well: screen inspection, taps/swipes, keyboard, global actions, apps, files, clipboard, overlay, and keep-awake. Minor gaps remain (e.g., pinch/zoom or multi-touch gestures, wait-for-idle), but agents can work around them with existing swipe/tap primitives.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables MCP-compatible agents to control an Android device over the network via ADB, providing tools for shell commands, screen capture, UI inspection, file operations, and input simulation.
    16 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables agents to interact with Android via text-based UI trees instead of screenshots, supporting taps, swipes, input, macro recording, and device control through an MCP server and CLI.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables an AI agent to operate a spare Android phone through MCP: capture screenshots, inspect UI controls, tap/swipe, input Chinese text, launch apps, and optionally run rooted shell commands or transfer files. It also supports a live web console for remote viewing and control.
    18
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables controlling an Android device from an MCP client via ADB, including wireless debugging discovery, UI-tree inspection, and performing taps, swipes, typing, app actions, and shell commands without relying on screen coordinates.
    1 npm
    MIT