Skip to main content
Glama

mobile-agent-harness

Android device automation for AI agents. MCP server, CLI, and a hot-reloadable plugin runtime on top of a self-healing uiautomator2 bridge. No device-side app required.

中文文档:README.zh-CN.md

This is the base runtime only. Plugins live in the community catalog: awesome-mobile-agent-harness-plugin. App business knowledge for AI lives in app-lore, which this repo's knowledge/ loader reads directly.

Install

git clone https://github.com/Aelindra/mobile-agent-harness
cd mobile-agent-harness
pip install -e .          # add ".[ocr]" for vision.ocr / ui2.*

Requires Python 3.9+ and adb on PATH with USB debugging enabled.

Related MCP server: adb-mcp-server

Quick start

  1. Point the bridge at your device:

export AGENT_SERIAL_DEFAULT=192.168.1.23:5555   # or "emulator-5555"
  1. List the tool surface (works without a device) and run the offline tests:

python -m core.cli tools
python tests/test_runtime.py
  1. Call a tool:

python -m core.cli call ui.snapshot
python -m core.cli call ui.click --args '{"selector":"text@^Network & internet$","pkg":"com.android.settings"}'
  1. Attach to any MCP client:

{
  "mcpServers": {
    "phone": {
      "command": "python",
      "args": ["-m", "core.cli", "mcp"],
      "cwd": "/path/to/mobile-agent-harness",
      "env": { "AGENT_SERIAL_DEFAULT": "192.168.1.23:5555" }
    }
  }
}

All tools return a {"ok": bool, "error"?, ...} envelope. ok=false is a business result (e.g. error_type=selector_not_found) — adapt instead of retrying.

Tools

Group

Tools

Observation

ui.snapshot ui.dump ui.find ui2.state ui2.screen_evidence ui2.timeline ui2.find_templates vision.screenshot vision.ocr vision.diff

Action

ui.click ui.click_handle ui.text_handle ui.hold_read ui.set_text ui.scroll ui.back ui.home input.tap input.swipe input.key

App

app.launch app.current app.probe app.list app.stop app.wait_idle

Shell & data

shell.run state.prefs state.db file.push file.pull net.http

Events

events.tail events.wait events.capture events.record_start events.record_stop

Knowledge & tasks

knowledge.list knowledge.guide ui2.check_states ui2.orient task.begin task.note task.end template.add sys.capabilities

ui.snapshot returns a compact accessibility-tree listing with element handles (e0, e1, ...); ui.click_handle acts on a handle with drift detection. ui2.state fuses the accessibility tree with OCR for canvas-drawn UIs. ui2.screen_evidence collects measurable screen features (overlay level, motion, color, layout density) without interpreting them. app.probe returns deterministic side signals (package, version, orientation); ui2.timeline captures N frames of lightweight evidence in time order; ui2.orient probes context, routes to matching knowledge packs, and evaluates their state discriminators — returning matched states or ranked hypotheses with the missing evidence listed.

Selectors

Selectors resolve against the live accessibility tree. Coordinates are not used.

# ui.click selector examples
"id@d98 && pkg@com.example.app"       # resource-id, short form completed with pkg@
"text@^Settings$ && pkg@com.example.app"
"class@android.widget.EditText && [clickable=true]"
"desc@^Search$"                        # content-desc regex

Harnesses and plugins

A harness file declares per-app flows as selector steps with pre/postconditions; distill drafts one by exploring an app, lint checks selector quality, and failed postconditions mark tools stale for re-distillation. See harnesses/com.android.settings/harness.json5 for the built-in example.

Plugins drop into plugins/ and register tools or capability providers at load time; registrations are reversible and hot-reloaded on file change. See the Plugin development guide in README.zh-CN.md and knowledge/README.md for the knowledge-pack format.

For business context on apps where model priors are unreliable, point agents at app-lore — a community catalog of structural app surveys in the open Agent Skills format, split per feature module for complex apps. Mobile scenarios in particular are private-domain and underrepresented in training data, so agents benefit from loading the relevant survey before operating an unfamiliar app. The local knowledge/ directory remains this repo's runtime layer (packs, icon templates, state discriminators).

Configuration

Variable

Description

Example

AGENT_SERIAL_DEFAULT

Device serial used when --serial/ANDROID_SERIAL absent

192.168.1.23:5555

AGENT_SERIAL_USB

USB serial fallback

0123456789abcdef

MAH_FRAME_STREAM

1 enables the H.264 frame stream for observation tools

1

MAH_VISION_MODE

Vision enhancement: off / client / endpoint

client

MAH_VLM_URL / MAH_VLM_MODEL

External vision endpoint (OpenAI-compatible)

http://127.0.0.1:11434/v1/chat/completions

How it compares

  • vs mobile-mcp / agent-device: they are generic device toolsets; this project adds the layer above transport — per-app knowledge as files (harnesses), a hot-reloadable plugin runtime with capability seams, and an explicit adb/root privilege model.

  • vs droidrun: droidrun requires a device-side accessibility app; this project is zero-install over adb.

  • Root support is declarative: if the device already has su, the shell seam gains a root provider gated by a command allowlist. This project does not provide root.

Known limitations

Messaging, banking, and e-commerce apps run active anti-automation risk control; the adb tier is unreliable on them. Target use is development, testing, and your own device workflows.

Contributing

Plugins are contributed to awesome-mobile-agent-harness-plugin; app business knowledge to app-lore — see their contributing guides.

License

MIT.

Available Tools

60 tools
app.currentB

当前前台应用

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: it does not confirm this is a read-only query, nor does it describe what happens if no app is in the foreground or whether the call is cheap/fast. Only the implicit read semantics of 'current' hint at the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single five-character phrase with zero waste and the key qualifier ('current') front-loaded. It is terse rather than bloated, though it borders on under-specification for a tool whose entire purpose is its return value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameters, the description is the only place the return value could be documented, and it does not say whether the result is a package name, an activity name, a display label, or null when nothing is focused. That omission matters for a query-only tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate on the input side.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase '当前前台应用' (current foreground app) names a specific resource and, by implication, a query verb, and it is clearly distinct from siblings like app.list, app.launch, and app.stop. It is a noun phrase rather than an explicit verb, but an agent can tell what it retrieves without opening anything else.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives (e.g. app.list for all apps), and no stated preconditions. The agent must infer from the name alone that this is the tool for checking which app is currently on screen.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app.launchB

启动应用并等待前台空闲

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It discloses one important behavior — that it blocks until the foreground is idle — which is useful. But it does not state what happens if the package does not exist, whether an existing instance is reused, or whether it fails/returns an error.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no wasted words. The core action and the blocking behavior are both stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-param tool with no annotations or output schema, the description is minimal but acceptable. It is missing failure modes and whether it waits for a specific activity, which would help an agent recover from launch failures.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter 'pkg' has no description. The description does not add meaning beyond the parameter name, but with only one obvious param (package identifier) the baseline is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: '启动应用' launches an app. It adds the scope '并等待前台空闲' (waits for foreground idle), distinguishing it from app.wait_idle or app.current. Sibling differentiation is implicit rather than explicit, keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus app.list, app.current, or app.wait_idle. The description implies launch-then-wait, but does not state exclusions, prerequisites, or that app.wait_idle can be called separately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app.listB

列出已安装应用包名(third_party=true 仅用户应用)

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNo包名子串过滤(大小写不敏感)
third_partyNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the third_party semantics, but says nothing about ordering, result limits, or the return format beyond 'package names', leaving behavior under-specified for a list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste. It delivers purpose plus the key parameter constraint efficiently, though it is arguably terse to the point of missing usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-param read tool with no output schema, the description covers purpose and the third_party switch adequately. It stops short of noting what the listing returns (full package list vs paginated) or its ordering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%; the filter param is documented in the schema while third_party is not. The description compensates by explaining 'third_party=true 仅用户应用' (only user apps), adding real meaning the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: '列出已安装应用包名' (list installed app package names). This is clearly distinct from siblings like app.launch, app.current, and app.stop, though it doesn't name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this vs alternatives. The parenthetical explains third_party behavior but not the scenario in which listing apps is the right call over app.current or knowledge.list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app.stopD

停止应用

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing. It does not say whether the stop is graceful or forced, whether background processes and cached state are cleared, whether it requires the app to be currently running, or what error occurs if the package is not installed or already stopped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At four characters there is no wasted padding, but brevity here is under-specification rather than conciseness — the single sentence carries no actionable detail. It also mixes languages with an otherwise English tool surface.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and an undocumented required parameter, the definition is entirely inadequate. An agent cannot determine the argument format, the effect on running state, or the failure modes before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter 'pkg' is undocumented anywhere. The description does not explain that the argument is a package identifier (e.g. com.android.settings), whether it accepts a display name, or whether multiple packages can be passed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description '停止应用' is a direct restatement of the tool name app.stop, adding no information an agent could not already infer from the identifier. It names a verb and resource but provides zero differentiation from siblings such as app.launch, app.current, or app.list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus app.launch (its inverse) or app.list/app.current, nor any precondition such as 'the target app must be running'. Usage is only implied by the tool name itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app.wait_idleC

等待应用前台且稳定

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgYes
timeoutNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that the wait covers both foreground state and stability, which is useful, but says nothing about polling frequency, what 'stable' means, what happens on timeout (error, return value, or silent exit), or whether the call blocks the harness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the core action front-loaded and zero padding, which is structurally clean. However, at this length the brevity tips into under-specification rather than disciplined conciseness, leaving the agent without enough to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and two undocumented parameters, the description is too thin for a blocking wait primitive. It omits timeout semantics, failure behavior, and any distinction from the many sibling wait tools, so an agent cannot call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with two parameters, so the description must compensate and does not: 'pkg' is never defined (package name? activity?) and 'timeout' is never mentioned, including whether it is in seconds or milliseconds or what occurs when it elapses. The word 'app' only loosely gestures at 'pkg'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (wait) with two conditions (foreground + stable) on a named resource (the app), which is more informative than a tautology. It does not, however, differentiate itself from sibling wait tools such as ui.wait_for, events.wait, or ui2.wait_change, so an agent must infer which wait is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and names no alternative. There is nothing telling the agent to prefer this over ui.wait_for or events.wait when the goal is app-level readiness rather than element-level or event-level waits.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

com.android.settings__launchA

[harness:com.android.settings] 打开设置并等待其稳定到前台

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the behavioral burden. It usefully discloses that the tool waits for Settings to stabilize in the foreground, which is more than a bare name. However, it omits side effects such as interrupting the current screen, device-state requirements, and whether the call blocks until idle.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the action and the wait condition with no filler. The harness prefix is metadata, but the substance is maximally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter launch tool with no output schema or annotations, the description covers the essential action and the stabilization wait. It could still be strengthened with a note about prerequisites or how it relates to sibling launch tools, but it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 per the rubric. There are no parameter semantics to explain and the empty schema introduces no ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (open/打开) and resource (com.android.settings), plus the post-condition of waiting for the app to stabilize in the foreground. It is clear what the tool does, but it does not distinguish itself from sibling launchers like app.launch or the settings search tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call this when you need the Settings app foregrounded. There is no explicit when-not guidance, no mention of prerequisites such as device unlocked, and no routing to com.android.settings__search or app.launch as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

device.wake_unlockA

亮屏 + 按需解除非安全锁屏 + 清理残留弹窗(agent 自身 setup,非人工干预)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that it wakes the screen, only unlocks when the lock is non-secure, and cleans leftover dialogs, plus that this is agent-side setup rather than human interaction. It omits failure behavior (e.g., what happens with a secure lockscreen), side effects, and whether it is idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the primary action front-loaded and the scope note in parentheses. Nothing is wasted, though the phrasing is terse enough that subtleties (ordering of the three steps, conditional unlock) require careful reading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, annotation-free tool with no output schema, the description identifies what the routine does but leaves gaps around preconditions (secure vs non-secure lockscreen), return signals, and error states that an agent orchestrating this setup step would want to know.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is no parameter semantics to document. The description correctly describes a no-argument behavior without inventing parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete multi-step action (turn on screen, dismiss non-secure lockscreen, clear residual popups) with a clear verb-plus-resource pattern. It distinguishes itself from the surrounding ui.* / input.* siblings by being a self-setup routine rather than an interactive automation step. It stops short of full 5 because the naming convention (dotted namespace) is not explained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(agent 自身 setup,非人工干预)' implies this is a preparatory step the agent runs for itself, which gives implicit when-to-use context. However, it never states when it should be called relative to alternatives (e.g., before app.launch or input.* sequences) nor any conditions that would make it unnecessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

events.captureC

执行动作同时捕获瞬时 Toast/弹层(Monitor 接线;action=click 可带 selector;expect 为正则过滤)

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo
actionNoclick
expectNotext/desc 命中正则,缺省收全部瞬态
timeoutNo
selectorNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints at 'Monitor 接线' (Monitor wiring) but never explains what is returned, how captured popups are surfaced, what happens on timeout, or any side effects of the click action. This is a significant gap for an action-performing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the core purpose front-loaded, followed by parenthetical parameter hints. No wasted words, though the packed parentheticals slightly reduce readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 5 parameters at 20% coverage, the description is too thin. It omits the meaning of pkg and timeout, the return shape, and any behavioral guarantees, leaving an agent under-informed for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'expect' is documented), so the description must compensate. It usefully clarifies that action=click can carry a selector and that expect acts as a regex filter, covering three of five params. However, 'pkg' and 'timeout' receive no explanation in either the schema or the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: execute an action while capturing transient Toast/popup layers. This distinguishes it reasonably from siblings like events.snapshot, events.tail, and events.wait, which deal with ongoing event streams rather than transient UI popups. It is clear but does not explicitly name which sibling it replaces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no alternatives are named despite many closely related siblings (events.wait, events.tail, ui.wait_for). Usage must be inferred from the parenthetical hints about action=click and the expect regex.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

events.record_startB

后台连拍+状态流开始(record_stop 收结果)

ParametersJSON Schema
NameRequiredDescriptionDefault
intervalNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose two meaningful traits: capture runs in the background and it is a start/stop pair requiring record_stop to finish. It is silent on whether the recording persists, resource cost, concurrency limits, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler; the start action and its terminating sibling are the first things read. It is terse to the point of under-specification rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and an undocumented parameter leave the agent without enough to call this correctly beyond the start/stop pairing. Return-value retrieval, interval semantics, and lifecycle behavior are all missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'interval' parameter has 0% schema description coverage and no default documented in the schema body, yet the description never mentions it or its unit (seconds?). With a coverage gap this large the description needed to compensate and does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific action (start) on a specific resource (background burst capture + state stream), and explicitly marks the terminal counterpart 'record_stop 收结果'. This lets an agent separate it from siblings like events.capture/snapshot/tail, though what the stream actually records is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical 'record_stop 收结果' implies this must be paired with events.record_stop to retrieve output, which is usable routing context. However there is no explicit when-to-use vs events.capture or events.tail, and no indication of prerequisites or duration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

events.record_stopA

停止后台记录,返回 toasts/帧数/目录

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the return payload (toasts, frame counts, directory), which is useful behavioral context, but it omits what happens if no recording is active, whether the artifacts are flushed/persisted, and whether the call is idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact clause with no filler, and the action is front-loaded ahead of the return-value detail. It is terse almost to a fault, but there is no wasted wording to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the brief mention of returned artifacts partially compensates, and there are no parameters to explain. However, for a stateful session-terminating tool with no annotations, the description is thin on edge cases and side effects, leaving the agent to guess at failure and no-op behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there are no parameter semantics to clarify or obscure. The agent needs no argument-level guidance to invoke it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource (停止后台记录 / stop background recording) and even names the return payload, so the action is unambiguous. It does not explicitly differentiate itself from sibling tools like events.record_start, events.snapshot, or events.capture, relying on the naming pairing to convey that role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent must infer this terminates a session previously opened with events.record_start. There is no explicit when-to-use statement, no mention of prerequisites (an active recording), and no comparison against alternative ways to stop or collect events.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

events.snapshotB

单次前台状态快照(含游离窗口节点=toast/悬浮窗候选)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It merely states that a snapshot is taken and mentions included floating nodes, but does not disclose whether the operation is read-only, whether it blocks, any side effects, or return characteristics beyond the purpose itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. The parenthetical clarification is compact and directly follows the main purpose statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and many sibling tools, the description is too thin to be complete. It names the tool's scope but omits usage guidance, behavioral traits, and any sense of what the snapshot returns or how it differs from alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so per the rubric the baseline is 4. The description cannot add parameter meaning beyond what the empty schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: '单次前台状态快照' (single foreground state snapshot). The parenthetical scope note '含游离窗口节点=toast/悬浮窗候选' clarifies the content included. However, it does not explicitly distinguish this from sibling tools like ui.snapshot, events.capture, or ui2.state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit when-to-use guidance or alternatives. It only implies single-call usage through the word '单次', but does not say when to choose this over events.capture, ui.snapshot, or other state-observation tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

events.tailB

事件面:最近的设备事件(page_change/app_start/app_crash/visual_change)——'刚才发生了什么'的场景感知,推送非轮询

ParametersJSON Schema
NameRequiredDescriptionDefault
nNo
typeNo过滤事件类型,缺省全部

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the event taxonomy and that delivery is push-based rather than polled. It does not state return shape, retention window, ordering, or how n interacts with the stream, leaving meaningful behavioral gaps for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence: it names the resource first, then enumerates types, then gives the use-case and delivery model. Dense but no filler. The em-dash structure packs a lot without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-param, all-optional read tool with no output schema, the description is adequate: it names the event types and the push delivery model. But with no annotations and no output schema it should say more about what an event record looks like and how many are returned beyond the default n=20.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'type' is documented but 'n' has none. The description's parenthetical list (page_change/app_start/app_crash/visual_change) supplies concrete filter values the schema's 'type' field lacks, which is genuine added meaning. It says nothing about the 'n' limit or ordering, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific verb+resource: '最近的设备事件' (recent device events) and enumerates the concrete event kinds (page_change/app_start/app_crash/visual_change). This is far more informative than a bare 'tail'. However, it does not explicitly contrast itself with close siblings like events.snapshot, events.capture, or events.wait, so differentiation relies on inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase "刚才发生了什么"的场景感知 (situational awareness of 'what just happened') and '推送非轮询' (push, not polling) imply when it is useful, and the push-vs-poll hint gestures at behavior. But there is no explicit when-to-use/when-not and no named alternative among events.snapshot/events.capture/events.wait, so guidance stays implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

events.waitA

等一个匹配事件(推送,非 sleep 轮询):type=page_change/app_start/app_crash/visual_change;可限定 pkg。事件到后仍应用帧/快照验证防误报

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo
typeYes
timeoutNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses that event delivery is push-based and that frame/snapshot verification is still needed to guard against false positives, but it omits timeout behavior and match/return semantics for this blocking wait.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with semicolon-separated clauses, front-loading the main action and event types. Every clause adds information; no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% schema coverage, the description covers the core event types, pkg filtering, and false-positive verification. It is still incomplete because timeout semantics and what a successful or timed-out wait returns are not described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates by listing valid type values and noting that pkg narrows the wait. The timeout parameter (default 15) is not explained, so a significant gap remains.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-resource ('等一个匹配事件') and enumerates event types (page_change/app_start/app_crash/visual_change), with the push-vs-polling distinction. It is clear but does not distinguish itself from sibling wait tools such as ui.wait_for, app.wait_idle, or ui2.wait_change, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '推送,非 sleep 轮询' and the supported event types imply an event-driven waiting context, and pkg is offered as a scope filter. However, it gives no explicit when-to-use vs alternatives or when-not-to-use guidance, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file.pullA

手机→PC 传文件:base64 读出解码(绕开 adb pull 的不可靠/权限问题),MD5 校验

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNo源在 /data/data 等受限目录时 true(su 读)
localYesPC 侧目标路径
max_mbNo超过则拒绝(大文件请 shell.run 分卷)
remoteYes手机侧源文件绝对路径

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the base64 read/decode method and MD5 verification, plus that it bypasses adb pull's permission problems. It does not say whether root/su is required (that lives in the schema), what happens on checksum failure, or how errors surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence: direction first, mechanism second, verification last. Every clause earns its place and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, so the description should do more. It covers the transfer mechanism and integrity check adequately for a 4-param tool, but says nothing about the return value (e.g., whether the decoded file path or MD5 is reported) or failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (root, local, max_mb, remote) are already documented in the schema, including the root/su behavior and the shell.run fallback. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with direction ('手机→PC 传文件') and the mechanism (base64 读出解码 + MD5 校验), which implicitly separates it from the sibling file.push (PC→phone). It does not, however, explicitly name file.push as its counterpart, so sibling differentiation relies on the reader inferring direction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '绕开 adb pull 的不可靠/权限问题' implies the use case (when adb pull is unreliable or blocked), which is useful. But the description itself names no alternative tool, and the fallback for large files (shell.run 分卷) appears only in the schema, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file.pushB

PC→手机传文件:base64 分块管道(WiFi adb push 二进制不可靠的替代),MD5 校验

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNosu 写入;仅 /data/local/tmp、/sdcard 免 allow_risk,其余路径需 allow_risk=true
chunkNobase64 分块字节(上限受单条 shell 命令 128KB 限制)
localYesPC 侧源文件路径
remoteYes手机侧目标绝对路径

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full behavioral burden. It does disclose meaningful mechanism (chunked base64 transport, MD5 verification, reliability rationale), but omits overwrite/reversibility behavior, permission implications, and failure/response format. The su/allow_risk constraints live only in the schema, not the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense, front-loaded sentence with no wasted words — the mechanism and verification claim are packed efficiently. It borders on cryptic (treats the reader as already familiar with adb push pitfalls), but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the transport mechanism but leaves out what an agent needs to call it safely: overwrite behavior, error/verification reporting, and permission prerequisites (only hinted at in the schema). Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (root, chunk, local, remote) are already documented in the schema, including the allow_risk constraint and 128KB shell limit. The description adds no parameter-level syntax or format detail beyond that, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with direction (PC→手机传文件) and names the implementation mechanism (base64 分块管道) plus MD5 verification. The direction implicitly distinguishes it from the file.pull sibling, but the sibling is never named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a conditional rationale — this is the fallback when WiFi adb push of binaries is unreliable — which is implied usage guidance. But there is no explicit when-to-use framing, no reference to file.pull or shell.run as alternatives, and no stated exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

harness.check_versionsA

读取设备上各 harness 目标 app 的 versionName,与 versionRange(>=/<=/* /通配) 比对;不匹配 → stale 需重蒸馏

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden. It usefully discloses the comparison logic and the downstream meaning of a mismatch ('stale, needs re-distillation'), which is real behavioral context. However it does not state read-only nature, return shape, or whether wildcard semantics differ per target.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the action (read versionName) and then the comparison and its consequence. Every clause earns its place, though the internal jargon ('重蒸馏') assumes domain knowledge.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, no-output-schema tool, the description explains the logic but not the actual return value (is it a list of stale targets, a boolean, a per-app report?). Some essential detail is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline rule parameter semantics start at 4. The description references versionRange concepts but there are no arguments to clarify, so no further credit is available.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: reads versionName of each harness target app and compares against a versionRange. The comparison logic and the notion of a 'stale' result are concrete. It does not explicitly distinguish itself from sibling harness.list or harness.run, keeping it just short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when the tool is useful (detecting targets whose installed version falls outside the required range and thus need re-distillation), but it never states when to call it versus harness.list or harness.run, nor any prerequisites. Usage is inferable rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

harness.listB

列出已加载的 harness 与工具

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It conveys only that results are limited to already-loaded harnesses, but says nothing about read-only safety, return shape, ordering, or whether the 'tools' portion overlaps with other listing tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no filler. It is appropriately sized for a zero-parameter list operation, though it is thin enough that it borders on under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description is the only source of information, and it does not explain what a returned entry looks like or what 'harness' means in this system. For a simple enumeration tool this is minimally adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema declares zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('列出'/list) and resource ('已加载的 harness 与工具'), which is clearly distinct from harness.run (execute a harness) and harness.check_versions (version check). It does not, however, explicitly contrast itself with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus harness.run, harness.check_versions, or knowledge.list. The agent must infer that this is a discovery/enumeration call from the word 'list' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

harness.runC

运行某 harness 工具(steps 由语义选择器组成)

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgYes
toolYes
paramsNo
dry_runNo
allow_untrustedNo

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses almost nothing: no side effects, no permission or trust requirements, no indication that the run may be destructive. The only behavioral hint is that steps are composed of semantic selectors, which is thin for an execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short clause with no filler, so it is not bloated. But the brevity is under-specification rather than economy; the one sentence it contains does not front-load enough information to be useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, a nested object, no annotations, and no output schema, a single clause is far from complete. Critical details — what the run returns, what dry_run previews, what trust boundary allow_untrusted crosses — are entirely absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, including a nested object ('params') and two security-relevant booleans ('dry_run', 'allow_untrusted'). The description adds no meaning for any parameter — not even which package/tool identifiers are expected or what allow_untrusted unlocks — so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('运行') and resource ('harness 工具'), which is clearer than the bare name, and the parenthetical hints that steps are semantic selectors. However, it never explains what a harness tool actually is or how it differs functionally from siblings like harness.list or harness.check_versions, leaving the agent to infer the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus harness.list or harness.check_versions, no prerequisites (e.g. must the harness be listed/version-checked first?), and no exclusions. The agent gets a verb but no selection context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

input.clipboardB

剪贴板读写(set=写入设备剪贴板;get 读回,需 u2 支持)

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
actionYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose one behavioral dependency (get requires u2 support), which is valuable. It does not state that set overwrites the existing clipboard contents, nor what get returns, leaving key mutation semantics implied rather than stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact parenthetical line, front-loaded with the resource and split into the two modes. Efficient, though the parenthetical packing makes it slightly dense and it omits the text parameter entirely.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema and no annotations, the description covers the action enum and one dependency but omits the second parameter and any statement of what get returns. It is minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema documents neither parameter. The description compensates for action by defining set vs get, but adds nothing about text (format, length, what it supplies for which action), so half the parameter surface is undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (剪贴板/clipboard) and the two verbs (read/write), and maps each enum value to its operation, so an agent can immediately tell this is the device clipboard accessor. It is clear without reference to any sibling, though no sibling performs this exact function so differentiation is trivial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It distinguishes the two modes (set=write, get=read) and adds a precondition that get requires u2 support, which is genuine usage information. However, it gives no guidance on when an agent should prefer this over, e.g., ui.set_text or file.push, nor any exclusion or prerequisite for set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

input.keyC

按键注入(back/home/menu/volume_up/volume_down/power/enter... 或键码整数)

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral disclosure. It mentions possible keys but omits crucial context: whether the key injection requires an active unlock state, whether it can turn the device on/off, timing constraints, or what happens if the key is invalid. For a general input tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the action ('按键注入') and provides the acceptable values. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and incomplete parameter documentation, the description is not sufficient for an agent to use the tool safely and correctly. It lacks details on device state prerequisites, supported key names beyond examples, integer representation, and error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'key' is just typed as string. The description partially compensates by listing example values (back, home, menu, volume_up, volume_down, power, enter) and allowing integer keycodes, but it does not explain the format of integer keycodes or how to combine them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool injects key presses, and lists supported key names (back/home/menu/volume_up/volume_down/power/enter) and the integer keycode option. It does not distinguish from siblings like ui.back or ui.home, which cover a subset of these keys, so it loses a point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus dedicated siblings such as ui.back, ui.home, app.launch, or ui.click. The description simply lists valid values.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

input.swipeC

裸坐标滑动(通用层原语)

ParametersJSON Schema
NameRequiredDescriptionDefault
x1Yes
x2Yes
y1Yes
y2Yes
duration_msNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. The parenthetical '(generic-layer primitive)' faintly signals that this bypasses UI-semantic layers, but nothing is said about whether the gesture is synchronous, how duration affects fling vs drag behavior, or what happens on a failed swipe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded phrase with zero waste, which is structurally clean. But the extreme brevity here reflects under-specification rather than disciplined conciseness — there is simply not enough content for a 5-parameter gesture tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 undocumented parameters, no annotations, and no output schema, a four-character description is wholly inadequate. An agent cannot determine coordinate conventions, duration semantics, or any behavioral contract from it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate and largely does not. 'Raw coordinates' implies x1/y1/x2/y2 are screen pixel start/end points, which adds slight meaning beyond the bare 'integer' schema, but duration_ms (default 300) and coordinate units/origin are entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'raw-coordinate swipe (generic-layer primitive)' names a specific verb+resource and hints at the low-level nature of the operation. However, it offers no differentiation from the many sibling gesture tools such as ui.scroll, ui.click, or input.tap, so an agent cannot tell from the description alone why it would choose this over ui.scroll.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all. Nothing tells the agent when a raw-coordinate swipe is preferable to ui.scroll or the higher-level gesture tools, nor whether a device/app context or prior wake-unlock is required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

input.tapB

裸坐标点击(通用层原语:不解析选择器,直接注入;长按调大 duration_ms)

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
duration_msNo>120 视为长按

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully discloses the injection mechanism ('直接注入', no selector parsing) and the long-press threshold hint, but says nothing about coordinate frame (absolute vs normalized), required device state, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with a parenthetical; the core operation is front-loaded and there is no filler. Slightly dense but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple primitive with no output schema and no annotations, this covers the essential what and the long-press mechanism, but leaves notable gaps: coordinate reference frame, device-state prerequisites, and error behavior. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only duration_ms documented, as '>120 视为长按'). The description reinforces duration_ms semantics and implies x/y are raw coordinates, but adds no detail on coordinate system or units to compensate for the undocumented x/y parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (raw coordinate tap) and adds scope ('通用层原语:不解析选择器,直接注入'), which meaningfully distinguishes it from selector-based siblings like ui.click/ui.long_click. It stops short of naming the alternative explicitly, but the contrast is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage context by calling itself a raw-coordinate primitive that bypasses selector parsing, and notes that long-press is achieved by raising duration_ms. However, it never states when to prefer this over ui.click/ui.long_click, nor any exclusions or preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

knowledge.guideA

读取知识包的 guide(上下文面:业务流程/使用时机/注意事项)。用某 app 前先看对应包的 guide

ParametersJSON Schema
NameRequiredDescriptionDefault
packYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. '读取' implies a read-only operation and the content surface is described, but it says nothing about required permissions, where pack identifiers come from, or return behavior. It adds useful context but not full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the core action and the when-to-use guidance front-loaded. Every clause earns its place; no redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with no annotations and no output schema, the description adequately covers purpose and timing. It is incomplete on how an agent obtains a valid 'pack' value and what form the guide content takes, which are the main remaining gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter ('pack') with 0% schema description coverage, so the description must compensate. It only says to use the pack corresponding to the app; it does not explain the format, source, or valid values of a pack identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (读取) and resource (知识包的 guide), and even enumerates the guide's content areas (业务流程/使用时机/注意事项). It is clear what the tool returns, but it does not explicitly differentiate itself from the sibling knowledge.list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit usage condition: '用某 app 前先看对应包的 guide' (read the corresponding pack's guide before using an app). That is clear when-to-use guidance, though it does not name alternatives or when-not-to-use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

knowledge.listB

列出已加载知识包(guide=上下文面,templates/states=运行时面)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. '列出' (list) implies a read-only enumeration, and the parenthetical adds useful structural context about how packages are categorized, but there is no statement about cost, side effects, or output volume. Adequate but thin for a zero-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the core action front-loaded and the classification detail packed into a parenthetical. No filler and no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and no output schema, the description should ideally sketch the return shape (e.g. a list of package names with their categories). It gestures at this via the guide/templates/states taxonomy but never says what a returned entry looks like, so an agent cannot fully anticipate the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate at the parameter level, and the schema is trivially complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: list knowledge packages that are currently loaded. The parenthetical clarifies the two package categories (guide = context face, templates/states = runtime face), which helps an agent understand what will come back. It does not, however, distinguish itself from the sibling knowledge.guide, leaving the boundary between the two ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as knowledge.guide, harness.list, or sys.capabilities. The description only says what the tool returns, not under what circumstances an agent should reach for it. Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

net.httpC

发 HTTP 请求(proxy 参数可接 mitmproxy 等审计代理;verify 默认关闭以适配自签代理证书)

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
bodyNo请求体(文本)
proxyNo置空字符串则直连http://127.0.0.1:8080
methodNoGET
headersNo
timeoutNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that verify is off by default and why (self-signed proxy certificates), which is useful behavioral context. However, it omits critical details such as authentication requirements, rate limits, error handling, or side effects. For a network tool with no annotations, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that efficiently conveys the core purpose and two important configuration notes. It is concise and avoids unnecessary words. However, the inclusion of 'verify' (not a parameter) could confuse and slightly detracts from structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and only 33% schema coverage, the description should provide more behavioral context. It mentions proxy and verify but omits authentication, error handling, rate limits, and return format. It also references a non-existent parameter (verify), which is misleading. The description is not complete enough for an agent to use the tool reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 33%, so the description must compensate. It adds meaning for the proxy parameter (can accept mitmproxy, empty string for direct connection) and for verify (default off for self-signed certs), but verify is not even a parameter in the schema—this is a discrepancy. The description does not explain other parameters like method, headers, timeout, or body. It partially fills gaps but leaves much undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action: sending an HTTP request. It distinguishes this from siblings like shell.run or knowledge.list by being a direct network operation. However, it does not explicitly differentiate from other potential HTTP-like tools (none exist in siblings), so it's clear but not maximally distinctive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions that the proxy parameter can accept mitmproxy and that verify is disabled by default to accommodate self-signed proxy certificates, but it does not provide explicit when-to-use or when-not-to-use guidance. There is no mention of alternatives or conditions for selecting this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shell.runA

执行 shell(root=true 走 su,设备需已 root);风险命令需 allow_risk。script=多行脚本模式:base64 推到手机 /data/local/tmp 执行(规避嵌套引号/管道符转义地狱),与 cmd 二选一

ParametersJSON Schema
NameRequiredDescriptionDefault
cmdNo
rootNo
scriptNo多行 shell 脚本(推荐:可含管道/引号)
allow_riskNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses real behavior beyond the schema: root mode routes through su and presupposes a rooted device, risky commands are gated behind allow_risk, and script mode base64-pushes to /data/local/tmp to dodge quote/pipe escaping. It does not explain what allow_risk actually bypasses or how errors/output are surfaced, leaving a gap for a tool with arbitrary-command execution power.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A compact, front-loaded passage: the core action leads, with mode and precondition details packed into parentheticals. Dense but every clause carries information, so it earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations and no output schema, so the description is the sole carrier of behavioral context. It covers preconditions and mode mechanics well but says nothing about return values, exit status, or failure handling, which matters for an arbitrary shell-execution tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (only script is documented in-schema), so the description must compensate and largely does: it explains root's su semantics, allow_risk's risk gating, and script's multiline/base64 mechanism. cmd is only referenced obliquely via '二选一', so one of four parameters stays thin, but the compensation is substantial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (执行 shell) and immediately qualifies it with the two operational modes (root via su, script mode). Clear and specific. There is no sibling shell tool in the list, so sibling differentiation is largely unnecessary here; the description does its job without it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives conditional guidance that is really parameter-level: root=true requires an already-rooted device, risky commands require allow_risk, and script/cmd are mutually exclusive (二选一). It never states tool-level when-to-use versus alternatives, but no sibling tool competes with it, so the omission is minor.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

state.dbC

root 拉取应用 sqlite 库到 PC 临时目录并执行只读 SQL

ParametersJSON Schema
NameRequiredDescriptionDefault
dbYes
pkgYes
sqlNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose two meaningful traits: root access is required and the SQL is read-only, plus the db is copied to a PC temp directory. However it omits temp-dir location, cleanup/lifetime of the copy, permission failures, and error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the action front-loaded and no filler. It is terse to the point of under-specification, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With three undocumented params, no annotations, and no output schema, the description leaves the agent guessing about what pkg/db mean and what the tool returns. For a tool that copies a database and executes arbitrary SQL, this is substantially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three params, and the description only vaguely gestures at db and sql without defining pkg, db format, or sql constraints. It does not compensate for the total schema documentation gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific actions (pull the app's sqlite db to a PC temp directory, execute read-only SQL) with a clear verb+resource. Distinguishable from state.prefs, though it never names or contrasts with that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this versus state.prefs or other state/data tools, and no prerequisites beyond the implied root requirement. Usage must be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

state.prefsC

root 读应用 shared_prefs XML 并解析

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgYes
fileYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses one useful trait, that root access is required, but omits read-only vs mutating semantics, permission failures, and what parsing yields or how errors surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that is front-loaded with the action and resource, with no wasted words. It is only held back from a 5 by being so terse that it verges on under-specification rather than crisp completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and 0% parameter coverage, the description is the only source of context and is inadequate for a two-required-parameter tool. An agent cannot confidently supply pkg/file formats or predict the return shape from this text alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both required parameters (pkg, file) have 0% schema description coverage, and the description only vaguely implies pkg is the app package and file is the shared_prefs XML. It adds almost no usable syntax, format, or example detail to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'root 读应用 shared_prefs XML 并解析' (reads an app's shared_prefs XML and parses it). This is far more concrete than a tautology, though it does not differentiate itself from the similarly named sibling state.db or explain the relationship between reading and parsing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, when-not-to-use, or alternative routing. The only implicit signal is the leading 'root', hinting at an elevated-privilege context, but no conditions or alternatives (e.g. state.db) are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sys.capabilitiesB

能力发现:各缝当前可用 provider 及 privilege/stealth 元数据;插件热载入口

ParametersJSON Schema
NameRequiredDescriptionDefault
reload_pluginsNotrue 时先热载插件目录再报告

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the key trait that the tool can hot-reload the plugin directory (a state-changing side effect) and that it reports provider/privilege/stealth metadata, implying a read-only discovery default. However, it omits the shape of the returned capability report and any cost/permission considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single densely packed line with the discovery purpose front-loaded ahead of the plugin-reload secondary role. Nothing is wasted, though the compressed phrasing borders on cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should help convey what a 'capability report' contains; it gestures at providers and privilege/stealth metadata but not structure, count, or format. Adequate for identifying the tool's role but thin for an agent relying on it to plan subsequent calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is fully documented in the schema, so the baseline is 3. The description adds only a loose pointer by calling the tool the 'plugin hot-load entry,' which corroborates but does not extend the schema's reload_plugins semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific function: capability discovery of currently available providers per seam plus privilege/stealth metadata, and identifies itself as the plugin hot-load entry point. The verb+resource (discovery of providers) is concrete enough that an agent knows this is an introspection tool, though it doesn't explicitly contrast itself with any sibling (none of which are close in function).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement of when the reload path should be preferred, and no mention of prerequisites or alternatives. The only implicit signal is that it's the plugin hot-load entry point, which an agent must infer rather than being told.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task.beginA

开始一个业务任务:声明目标(如'领所有商品的券')。返回 task_id,后续 task.note/task.end 回传它

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYes业务目标一句话

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It does disclose the key behavior an agent needs: this call returns a task_id that must be threaded through task.note/task.end, which is genuine lifecycle context. However, it says nothing about side effects, idempotency, concurrency (can two tasks run at once?), or failure modes for a tool that clearly opens a stateful session.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: purpose, example, and the return/follow-up contract in essentially two clauses. No filler, though the example is embedded after the fact rather than being a separate line an agent could skip.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, fully-documented schema with no output schema, the description covers the essentials: what it opens, what it returns (task_id), and how the id is reused downstream. Only the stateful side effects of beginning a task are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description earns a bump by supplying a concrete example goal ('领所有商品的券') that clarifies the expected abstraction level beyond the schema's terse '业务目标一句话'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (开始/begin) and resource (业务任务/business task), and distinguishes itself from its siblings by explicitly naming task.note and task.end as the follow-up consumers of the returned task_id. An agent can identify this as the task-lifecycle entry point without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the workflow ordering clear: declare the goal here, then pass the returned task_id to task.note/task.end. The example goal ('领所有商品的券') shows the required granularity of intent. It stops short of stating when NOT to use it (e.g. if a task is already open), so it is not a full when/when-not treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task.endB

结束任务:实际结果与目标对照(成功/失败/原因)

ParametersJSON Schema
NameRequiredDescriptionDefault
okYes
outcomeNo实际结果一句话
task_idNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that an outcome comparison is recorded, but says nothing about side effects, whether an active task is required, what happens if no task exists, or whether the end state is reversible — significant gaps for a state-mutating lifecycle terminator.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact line with the core action front-loaded and no filler. It is efficient, though the terse parenthetical is slightly compressed rather than fully readable prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must stand on its own for a 3-parameter mutation tool. It covers the essential intent but omits the role of task_id and the behavior when invoked outside an active task, leaving the definition minimally viable rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only 'outcome' is documented in the schema). The description partially compensates by explaining that 'ok' encodes success/failure and that a reason is captured, but the task_id parameter is left entirely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (end task) and clarifies the semantics as recording the actual result against the goal with success/failure/reason. It is understandable in isolation, but it does not differentiate itself from sibling lifecycle tools task.begin or task.note, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the 'end task' framing and the goal-vs-result comparison, which hints at when it should be called. However, there is no explicit guidance on prerequisites (must a task be active?) or when to use this versus task.note for recording intermediate observations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task.noteB

记录当前想法/判定/预期(业务化历史的核心):比如'我判断这个按钮是搜索入口,点击预期出现搜索框'。可能错,先声明后验证

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes当前想法/判断
aboutNo相关句柄/元素,如 e12
expectNo预期会发生什么
task_idNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It conveys that notes may be wrong and are provisional, but says nothing about persistence, whether notes are append-only, permissions, or what a call returns. For a write-style memory tool this is a clear gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A short two-clause definition with the purpose front-loaded and a compact illustrative example; nothing is wasted. Slightly less tight than a maximally economical definition, but well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 4-parameter note tool with no output schema, the description covers purpose and the provisional nature of notes adequately. It leaves behavioral details (persistence, return value) unaddressed, which matters more given zero annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already documents text, about, and expect. The description's example ('this button is a search entry, click expects a search box') illustrates how text/expect relate, adding modest value, but no syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: recording current thoughts/judgments/expectations, with a concrete example that clarifies the intended granularity. It doesn't explicitly distinguish itself from siblings like knowledge.list or events.capture, but the 'declared hypothesis' framing makes its role reasonably identifiable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase '先声明后验证' (declare before verifying) implies a usage pattern – jot down a hypothesis before checking it – but there is no explicit when/when-not guidance and no named alternatives. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

template.addA

声明图标模板:从截图(缺省最近一次 vision.screenshot)按 box 裁剪图标,写入 templates// 并更新 manifest。这是'图标→语义'声明层的一次性成本,之后 ui2.find_templates 永久可匹配(无需 hover)

ParametersJSON Schema
NameRequiredDescriptionDefault
boxYes[l,t,r,b] 截图上的裁剪框
descNo用途说明(声明层核心字段)
fromNo截图路径,缺省最近一次 vision.screenshot
gameYes
kindNo
nameYes模板名(ascii slug,如 flash)
labelNo显示名(如 闪现)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description must carry the behavioral load, and it does disclose meaningful traits: it crops from the last vision.screenshot by default, writes files into templates/<game>/, mutates a manifest, and the result persists for future matching. It omits permission/auth needs, overwrite behavior for an existing template name, and failure modes when the box is invalid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action is front-loaded in the first clause and the entire definition is two tight sentences with no filler. The parenthetical and the second sentence about permanence both earn their place, though the content is dense enough that it borders on packing two ideas into one sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description still tells the agent what is produced (a crop written under templates/<game>/, manifest updated) and why it matters, which is enough to invoke it correctly. It could say more about the returned identity of the created template and what happens on a name collision, but nothing essential is missing for a 7-parameter file-writing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 71%, so the structured fields already do most of the work; the description's contributions (box as a crop region, from defaulting to the last screenshot, game as the target directory) largely restate schema descriptions. It adds no meaning for the undocumented parameters such as kind, and offers nothing on coordinate space or box validation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: declaring an icon template by cropping from a screenshot by box, writing to templates/<game>/ and updating the manifest. It also names the downstream sibling (ui2.find_templates), so an agent can tell this declaration tool apart from the matching tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Frames the tool as a one-time 'icon→semantic' declaration cost whose payoff is that ui2.find_templates can then match permanently without hover, which is a clear when-to-use story tied to an alternative. It stops short of an explicit when-not-to-use (e.g. one-off matching without declaring, which other siblings such as ui2.scene_match cover), so it is good but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui2.check_statesB

按知识包的状态判别式求值当前屏幕:先采证据(同 ui2.screen_evidence),再对声明状态逐个判别式求值——全满足=ok,缺证据单列(不出具结论)。这是'证据与判读分离'的声明面

ParametersJSON Schema
NameRequiredDescriptionDefault
gameNo
packNo只求值该包;缺省求值全部已加载包

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses evidence collection behavior (same as ui2.screen_evidence), the one-by-one predicate evaluation, and result semantics — all satisfied = ok, missing evidence is listed separately and no conclusion is issued. It does not cover side effects or permission needs, but for a check/evaluation tool this is a strong disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the core purpose before detailing the process. It is efficient, though the closing conceptual phrase adds jargon rather than operational detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must carry the full burden. It explains evaluation and partial-result semantics well, but leaves the game parameter completely unexplained and does not clarify what packs are loaded or how the caller should supply context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: pack has a schema description but game has none, and the tool description never mentions either parameter. It adds no meaning beyond the schema, so it fails to compensate for the undocumented game parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: evaluate the current screen against knowledge-pack state predicates. It also signals its relationship to ui2.screen_evidence for evidence collection, but it does not explicitly distinguish itself from other siblings like ui2.state or ui2.scene_match.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the process but never states when to use this tool versus ui2.screen_evidence, ui2.state, or ui2.scene_match. There are no exclusions or alternative-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui2.diffC

界面变更检测:对比两次融合状态的文本元素增删

ParametersJSON Schema
NameRequiredDescriptionDefault
gapNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does not say whether the operation is read-only, where the 'two fused states' come from, whether prior state calls are required, or what the delta output contains. It only restates the comparison in prose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the operation and its scope with no filler. It is efficient, though the density comes partly at the cost of the missing detail noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, an undocumented parameter, and no explanation of how the two states are supplied leaves an agent unable to call this correctly. For a diff tool that must consume prior state, this is a significant shortfall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'gap' has 0% schema description coverage and is never mentioned in the description, so its meaning (tolerance? frame gap between captures?) is completely undocumented. With one undocumented parameter, the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (对比/diff) and a specific resource (两次融合状态的文本元素增删), so an agent knows it produces a text-element delta between two fused UI states. It does not distinguish itself from the adjacent siblings that also compare screens (vision.diff) or watch for change (ui2.wait_change), so it falls short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this rather than ui2.state, ui2.wait_change, ui2.check_states, or vision.diff, and no prerequisites (e.g. whether two ui2.state calls must be made first). The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui2.find_templatesA

图标模板匹配:在当前屏幕上找声明图集里的图标(无需 hover,图标在屏即命中),返回位置+置信度+manifest 里声明的语义。templates 缺省匹配图集全部;图集用 template.add 从截图裁剪声明

ParametersJSON Schema
NameRequiredDescriptionDefault
gameNo图集名(templates/<game>/),缺省扫全部
max_perNo
templatesNo模板名列表,缺省全部
thresholdNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the burden. It discloses return shape (position + confidence + manifest semantic) and the no-hover matching behavior, but says nothing about read-only nature, thresholds' effect, failure when nothing matches, or rate/limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and then the key constraints/defaults in two dense sentences. Every clause adds information, though the run-on structure could be tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description helpfully names the return values and the template.add dependency, making the tool callable end-to-end. The main gap is the unexplained threshold/max_per parameters for a 4-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50% (max_per and threshold undocumented), yet the description only restates that templates defaults to the whole collection and links it to template.add. The two tuning knobs (max_per=6, threshold=0.78) get no explanation, so the coverage gap is not compensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: template-based icon matching on the current screen, returning position/confidence/semantic. It implies its niche (no hover required, icon-on-screen = hit) versus element-based siblings like ui.find, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives useful context (no hover needed, hits on presence) and a prerequisite (collections are declared via template.add), but offers no explicit when-to-use-vs-alternatives guidance against ui.find or ui2.scene_match.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui2.scene_matchA

场景/花纹零样本分类(小模型层):给候选标签短语(如 '设置页面'/'聊天列表页'),CLIP 类模型计算图文相似度返回各候选得分——加新场景=加一个标签,无需训练。定位:色彩/布局等像素统计概括不了的花纹与场景语义。需可选依赖:pip install 'mobile-agent-harness[scene]';模型经 MAH_SCENE_MODEL 配置(中文标签用 OFA-Sys/chinese-clip-vit-base-p16,英文标签默认 openai/clip-vit-base-patch32)

ParametersJSON Schema
NameRequiredDescriptionDefault
boxNo[l,t,r,b] 只对裁剪区域分类
topNo
fromNo截图路径,缺省最近一次 vision.screenshot
candidatesYes候选场景/花纹标签短语(封闭集合声明)

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses an optional dependency (pip install 'mobile-agent-harness[scene]'), the model configuration mechanism (MAH_SCENE_MODEL) and defaults per label language. It does not mention load latency, GPU needs, or failure behavior when the dependency is absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then positioning, then dependency/model config – a logical order. It is dense and long, but each clause (dependency install, model selection) carries operative information for actually invoking the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey returns; 'returns scores for each candidate' is brief but sufficient, and the dependency/config details round out what an agent needs to call it. Minor gap: it does not describe score format/range (e.g. softmax probabilities vs raw similarity).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the baseline is 3. The description adds meaning to 'candidates' (closed-set label phrases) and reinforces the zero-shot angle, but contributes nothing for box, top, or from beyond what the schema already documents, and 'top' remains undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: zero-shot scene/pattern classification that returns per-candidate image-text similarity scores, with concrete label examples ('设置页面'/'聊天列表页'). It positions itself against pixel-statistics tools ('色彩/布局等像素统计概括不了的花纹与场景语义') but does not name specific siblings like vision.vlm or vision.ask.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the positioning sentence – use it for pattern/scene semantics that color/layout statistics cannot capture – and through 'no training needed when adding labels'. However, there is no explicit when-to-use vs alternatives guidance, and no named sibling (e.g. vision.vlm) to route the agent against.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui2.screen_evidenceB

屏幕证据采集(只出证据不下结论):overlay=灰屏/暗化覆盖(HSV)、motion=两帧运动量、color=色相分布、filter=vignette/flash、layout=UI 热区网格、hud=动作栏图标在位性(图集 kind=hud 模板)。防幻觉纪律:缺证据时列备选假设,不得凭单一特征断言状态

ParametersJSON Schema
NameRequiredDescriptionDefault
gameNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares that the tool only produces evidence and never asserts state, and it discloses the concrete signals computed per category. It omits output format, cost, and permission requirements, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then lists evidence modes compactly, then closes with the discipline rule. Dense but every clause earns its place; the colon-delimited enumeration is terse rather than bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex six-mode tool with no output schema and no annotations, the evidence-category explanation is solid, but the undocumented 'game' parameter and the absent return-shape description leave real gaps an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'game' has 0% schema description coverage and is never mentioned in the description, so an agent gets no meaning for it from either source. The description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (屏幕证据采集 / screen evidence collection) and enumerates the six evidence categories it computes (overlay, motion, color, filter, layout, hud). The clause 只出证据不下结论 ('only emits evidence, does not conclude') meaningfully separates it from conclusion-oriented siblings like ui2.state and ui2.check_states, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The anti-hallucination rule (缺证据时列备选假设) guides how to interpret results, but there is no explicit when-to-use / when-not-to-use or named alternative among the many ui2.* and vision.* siblings. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui2.stateB

结构化 UI 状态 v2:a11y×OCR 融合——a11y 提供零误报可交互性,OCR 补 canvas/自绘文本(src=ocr 的元素为启发式推断)。元素带句柄,ui.click_handle 直接操作。canvas/自绘界面必用

ParametersJSON Schema
NameRequiredDescriptionDefault
max_elementsNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavioral traits: a11y gives zero-false-positive interactability, OCR-sourced elements (src=ocr) are heuristic inferences, and elements carry handles usable by ui.click_handle. It stops short of covering cost, size limits, or return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense, front-loaded sentence with no filler; the fusion mechanism, handle capability, and usage directive are all packed efficiently. It is information-dense but still readable and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe returns, and it partially does (handles, src markers, interactability signal). It still omits the shape of the returned state and the effect of max_elements, leaving meaningful gaps for a state-dump tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter max_elements (default 80) has 0% schema description coverage and is not mentioned anywhere in the description, so an agent cannot learn its meaning or effect. With 1 undocumented parameter the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource ('structured UI state v2') and explains the mechanism (a11y×OCR fusion), which lets an agent understand what the output represents. It implies distinction from ui.dump/ui.snapshot via the fusion/OCR angle but never names or contrasts those siblings explicitly, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one clear directive condition ('must use for canvas/self-drawn interfaces') and hints that OCR fills gaps a11y misses. However it does not state when NOT to use it versus the many nearby siblings (ui.dump, ui.snapshot, ui.get_text, vision.ocr), leaving routing largely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui2.wait_changeB

等画面变化(点了按钮后确认生效),变化时返回新界面的可交互元素

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
baseline_textNo当前页面的标志文本(消失=变化)

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that a change is detected by the disappearance of baseline_text and that the new interactive elements are returned. It does not describe timeout behavior, polling cadence, or what is returned when nothing changes within the timeout.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that packs purpose, trigger context, and return value with no wasted words. Slightly dense but appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and no output schema, the description covers purpose, trigger, and return content, but leaves failure/timeout behavior and the undocumented timeout parameter unaddressed. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: baseline_text is documented ('disappearance = change') but timeout has no description beyond its default. The description adds no parameter meaning at all, so it fails to compensate for the undocumented timeout parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: it waits for a screen change and returns the new interface's interactive elements. The parenthetical use case ('after clicking a button to confirm the effect') makes the intent concrete. It does not, however, distinguish itself from the sibling ui.wait_for or events.wait.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete usage scenario ('confirm a button click took effect'), which implies when to reach for it. But it names no alternatives (ui.wait_for, events.wait) and states no when-not conditions, so the agent must infer the boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.backC

系统返回键

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing beyond the name's own meaning. There is no mention of what happens when no back stack exists, whether it is a device-level or app-level key, or any side effects/dependencies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short phrase with zero waste and the essential idea front-loaded. It is concise to the point of under-specification, but the dimension rewards economy and there is no filler to cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivial zero-parameter tool with no output schema, the description is minimally sufficient to call it correctly. It lacks any failure-mode or interaction context (e.g., behavior at the root screen), leaving clear gaps for an automation agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case. There is nothing further a description could clarify about arguments for a no-arg navigation action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"系统返回键" names a single, recognizable action (triggering the system back key), so the agent can infer the tool's purpose. But it is a bare noun phrase with no verb, no scope, and no differentiation from near-identical siblings like ui.home or input.key, which would also perform device-level navigation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this versus ui.home, app.stop, or ui.click, and no conditions or prerequisites. Nothing is misleading, but the agent must infer the entire usage context from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.clickC

节点级点击(选择器解析,非裸坐标)

ParametersJSON Schema
NameRequiredDescriptionDefault
nthNo
pkgNo
timeoutNouia 选择器服务端等待秒数
selectorYes
long_clickNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only notes that resolution is selector-based. It does not disclose what happens on timeout, whether the click waits for the element, or how it interacts with the sibling ui.long_click given the long_click parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single terse parenthetical with no filler, which is efficient. However its brevity stems from under-specification rather than disciplined conciseness, leaving most required context unstated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter interaction tool with no annotations and no output schema, the description is too thin. It omits selector syntax, behavior on failure, and the meaning of long_click/nth/pkg, none of which are recoverable from structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%; selector, nth, pkg, and long_click are undocumented in the schema AND absent from the description. The description adds no parameter meaning beyond the timeout field the schema already covers, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) and resource (node/selector), and distinguishes its mechanism from coordinate-based tapping via '非裸坐标'. It clearly separates itself from input.tap-style tools, though it does not name the specific sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical implies you should use this when you have a selector to resolve rather than raw coordinates, which implicitly routes against input.tap. However no explicit when/when-not conditions or named alternatives are given, so usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.click_handleA

按快照句柄点击/长按(a11y 元素先重定位防漂移,纯视觉元素用快照中心)

ParametersJSON Schema
NameRequiredDescriptionDefault
handleYes
long_clickNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses a non-obvious mechanism – a11y elements are re-located first to prevent drift, visual elements use the snapshot center – but says nothing about stale/invalid handle behavior, failure semantics, or side effects, so it is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded clause with the key action first and the qualifying mechanism in a compact parenthetical. No sentence is wasted and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and no output schema, the description covers purpose, parameters and the core drift-avoidance behavior. It omits what happens on failure, whether the action waits for idle, and how it relates to ui.wait_for/app.wait_idle, leaving the agent with gaps for edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does map both parameters semantically: handle is specified as a snapshot handle (快照句柄) and long_click is covered by 长按. It stops short of explaining where the handle comes from or its format, so not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (点击/长按) and the resource it operates on (快照句柄 / snapshot handle), which is enough to tell it apart from ui.find, ui.dump and ui.snapshot. It does not explicitly distinguish itself from close siblings like ui.click, ui.long_click or ui.text_handle, so it falls one notch short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the phrase 按快照句柄 tells the agent to use it when it holds a snapshot handle, and the a11y-vs-visual clause hints at when each branch applies. There is no explicit when-to-use/when-not guidance versus ui.click, ui.long_click or ui.text_handle, and no prerequisites (e.g. needing a fresh snapshot).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.dumpC

全量节点树 JSON(实时桥数据):class/id/text/desc/bounds/checked/clickable 等语义

ParametersJSON Schema
NameRequiredDescriptionDefault
save_toNo可选,完整 JSON 落盘路径
max_nodesNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It mentions that data is real-time, but does not disclose read-only status, performance impact, output size limits, or what happens when max_nodes is exceeded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that efficiently lists the returned fields. It is concise with no wasted words, though the dense field list could be slightly better structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description helpfully enumerates some returned fields, but it omits return structure, truncation behavior from max_nodes, and save_to semantics. It is minimally viable but incomplete for a dump tool with two parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: save_to is documented in the schema, but max_nodes has no description. The tool description mentions neither parameter and adds no meaning beyond the schema, failing to compensate for the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific resource (full node tree JSON with real-time bridge data) and lists the semantic fields returned. It distinguishes itself from search-oriented siblings like ui.find or ui.get_text by implying a complete tree dump, but it does not explicitly name an alternative to route against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as ui.snapshot or events.snapshot. The description only describes the output, leaving the agent to infer the use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.findC

语义查找节点(GKD 风格选择器),返回节点句柄摘要列表

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo
timeoutNo
selectorYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It discloses only that results come back as a handle-summary list; nothing is said about blocking/timeout behaviour, whether a single match or all matches are returned, what happens on no match, or whether pkg scopes the search.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the core purpose, so there is no waste. However it reads as under-specification rather than deliberate brevity at this parameter and behaviour complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema and three undocumented parameters, the description should do far more. It omits return shape details, selector syntax expectations, and how the handles it returns are consumed by siblings like ui.click_handle or ui.hold_read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters. The description loosely clarifies that `selector` is a GKD-style selector rather than CSS/XPath, but `pkg` and `timeout` (default 0) are left entirely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (semantic node lookup) and even names the selector dialect (GKD-style), plus what it returns (a list of node handle summaries). It does not explicitly distinguish itself from near neighbors like ui.dump, ui.snapshot or ui.get_text, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is given. The agent cannot tell from the text whether to reach for ui.find, ui.snapshot or ui.dump to locate a node.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.get_textC

读取选择器命中节点的文本(只读不点击)

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo
limitNo最多返回节点数
selectorYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that the tool is read-only and does not click, but omits return format, selector-failure behavior, and any limits beyond the schema's brief limit description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with no wasted words, and the core purpose plus read-only constraint are front-loaded. Its brevity leaves gaps, but as a conciseness measure it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and low schema coverage mean the description should carry more. It omits selector semantics, pkg/limit meaning, and return details, leaving it incomplete for a 3-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: just 'limit' is documented in the schema. The description refers to the selector concept but adds no format or meaning for selector or pkg, and does not compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb '读取' and resource '选择器命中节点的文本', and adds the qualifier '只读不点击'. It distinguishes itself from click/write siblings such as ui.click and ui.set_text, but does not explicitly name alternatives like ui.find or ui.snapshot, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use, when-not-to-use, or alternative tools. The parenthetical '只读不点击' hints at a read-only use case, but does not guide tool selection or provide prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.hold_readA

长按读悬浮提示(hover/tooltip):后台线程长按目标,按住期间读屏对比,返回新出现的提示文本。适合悬浮说明、图标长按菜单等按住才显示的提示。ocr=false 只比 a11y 树(快,~1s,普通 app 用);ocr=true 加 OCR 对比(~6s,canvas/自绘界面用)

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
ocrNo
pkgNo
handleNo快照句柄(与 selector/x,y 三选一)
hold_msNo
selectorNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden and does so well: it discloses the background-thread hold mechanism, the screen-comparison approach, what is returned (newly appeared tooltip text), and per-mode latency (~1s a11y vs ~6s OCR). It omits permission/auth needs and hold_ms defaults, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly-packed sentences that front-load the core purpose, then the suitable scenarios, then the ocr tradeoff. No filler, though the density makes it slightly harder to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, no-annotation, no-output-schema tool, the description covers the return value and the ocr decision well but leaves most targeting parameters (x, y, selector, pkg, handle) and hold_ms unexplained. It is adequate but has clear gaps an agent must infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 14% (only handle is described), so the description must compensate. It explains the ocr boolean meaningfully with timing and use-case, but leaves x, y, pkg, selector, hold_ms, and the x/y-vs-selector-vs-handle relationship undocumented, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource: long-press a target on a background thread, read the screen during the hold, and return the newly-appeared tooltip text. It clearly distinguishes this from siblings like ui.long_click (which just presses) and ui.get_text (which reads existing text), so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use it (hover explanations, icon long-press menus, tips that only appear while held) and gives clear ocr=false vs ocr=true selection guidance with timing tradeoffs. It does not, however, name alternative sibling tools or state when-not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.homeB

回到桌面

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether the current app is killed or backgrounded, whether state is preserved, whether the gesture works from any screen, or whether it can fail. For a UI-navigation action this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four characters is front-loaded and waste-free, but it is under-specified rather than efficiently concise, and the description is in Chinese while all sibling tool names and identifiers are English, which adds a small parsing cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument, no-output UI action the description is arguably sufficient to invoke it correctly, but it omits any note about failure conditions or interaction with the navigation stack, which is the only real risk with this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is nothing for the description to disambiguate beyond the action itself. Schema coverage is 100% trivially.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and destination (return to the home/desktop screen), so an agent can tell what it does without opening the schema. It does not, however, distinguish itself from the nearest sibling ui.back, which also navigates away from the current view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives such as ui.back or app.launch, even though those are the obvious competing ways to change screen state. The agent must infer the selection criteria entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.long_clickC

节点级长按

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo
selectorYes

TDQS

C2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: no press duration, no expected side effects, no permission or timing requirements, and no indication of what a successful press does. '节点级长按' is a restatement of the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short but by under-specification, not by efficient compression. A four-character phrase cannot be considered front-loaded structure when essential invocation context is absent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter interaction tool with no annotations and no output schema, the description leaves everything the agent needs (selector syntax, pkg usage, press semantics) unspecified. It is inadequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter is documented. The phrase 'node-level' loosely implies that selector targets a UI node, but the format of selector and the role of pkg are entirely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description '节点级长按' (node-level long press) states a specific verb (long press) and scope (node-level), which separates it from coordinate-based siblings like input.tap. However, it gives no differentiation from ui.click or ui.hold_read and relies on the agent knowing what 'node-level' means in this framework.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus ui.click, ui.hold_read, or input.tap/input.swipe. The agent must infer that a long press is needed and that this variant targets a node rather than coordinates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.scrollC

滚动(down=查看下方内容,手指上滑)

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo
containerNoscrollable 容器选择器,缺省自动找
directionNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It says nothing about what happens on scroll, whether the pkg argument changes behavior, what occurs at the edge of scrollable content, or what the tool returns. The only behavioral hint is the direction-to-gesture mapping.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely terse — a single clause. Front-loads the verb, which is good, but it is arguably under-specified rather than efficiently concise given three parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and only 33% schema description coverage. The description does not explain the pkg parameter, the container auto-find fallback behavior, or scroll edge behavior, leaving the agent with clear gaps for a 3-parameter UI tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so compensation is expected. The parenthetical usefully explains that direction=down means viewing content below via an upward finger swipe, which is real added meaning beyond the bare enum. However, pkg remains entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (scroll) and clarifies the direction semantics in a parenthetical. An agent can tell it apart from ui.click or input.swipe, though the description never explicitly contrasts with input.swipe, which could also produce scrolling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use ui.scroll versus scrolling sibling tools like input.swipe, nor any preconditions. The direction hint implies usage but does not state it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.set_textB

精确输入到指定组件(优先 ACTION_SET_TEXT;无 selector 时输入到聚焦框)

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo
textYes
clearNo
selectorNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the preferred underlying action and the no-selector fallback, which is genuine behavioral context. But for a mutation tool it stays silent on critical traits: whether existing content is overwritten (clear defaults to true), permissions, and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the mechanism and fallback front-loaded; no filler. It is arguably a touch too terse for a four-parameter mutation tool, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and zero schema description coverage across four parameters, the description only addresses element targeting. The destructive default of clear and the role of pkg remain undocumented, leaving an agent unable to predict the tool's effect on existing input content.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for four parameters. The description explains selector semantics and the fallback when it is absent, which is real added value, but says nothing about clear (default true, i.e. whether existing text is wiped), pkg, or the text payload format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: precisely input text into a designated component. It further names the underlying mechanism (ACTION_SET_TEXT) and the fallback target (the focused field), which distinguishes it from sibling input tools. No explicit sibling naming, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description tells the agent how targeting resolves (selector present vs. absent), which implies when this tool applies. However, it never states when to prefer ui.set_text over input.key or input.clipboard, nor any preconditions such as the element being visible.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.snapshotA

读屏默认首选:a11y 语义裁剪快照,元素编号 e0/e1/...(token 约为 ui.dump 的 1/10)。k 类型 btn/input/toggle/scroll/icon/text;i=可交互 e=可输入 s=可滚动 v=开关态。用 ui.click_handle/ui.text_handle 按句柄操作

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo缺省取当前前台应用
max_elementsNo
include_boundsNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden. It discloses the output shape (element numbering e0/e1/...), the per-element flag legend (k type, i/e/s/v states) and the token-cost characteristic, which is meaningful behavioral context. It does not state pagination/truncation behavior (max_elements) or that it is read-only, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the decisive fact ('读屏默认首选') and then densely packed with the comparison and output legend; every clause earns its place. The telegraphic shorthand is information-dense rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description does the work of explaining the return format (element handles plus a flag legend), which is the key missing piece for an agent. Input-side detail (max_elements/include_bounds) is not covered, keeping it short of 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% and the description adds no input-parameter meaning: the k/i/e/s/v legend concerns output fields, not pkg, max_elements or include_bounds. Two of three parameters are undocumented in both schema and description, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('a11y 语义裁剪快照' - an accessibility-semantic-trimmed snapshot) and marks itself as the '读屏默认首选' (default first choice for screen reading). It also explicitly distinguishes itself from the sibling ui.dump by token cost (~1/10), so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Declares itself the default for screen reading and gives the comparative condition vs ui.dump (~1/10 tokens), implying when to prefer this over the heavier dump. It also routes the agent onward to ui.click_handle/ui.text_handle for acting on handles. No explicit 'when-not' exclusion is stated, so 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.text_handleB

按句柄向输入元素写文本(a11y 元素走 set_text 原生通道链;纯视觉元素先点中心聚焦)

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
clearNo
handleYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the internal focusing/tapping behavior for pure visual elements, which is non-obvious. However, it omits whether it clears existing text by default (only inferable from the `clear` parameter), what happens on failure, and whether it waits for the keyboard.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact parenthetical sentence, front-loaded with the primary action and followed by the two routing cases. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, 0% parameter coverage, and an overlapping sibling, this description is too thin. It should state the `clear` default effect, failure behavior, and how it differs from `ui.set_text`.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters, so the description must compensate. It clarifies the `handle` and `text` semantics implicitly, but says nothing about the `clear` parameter's default behavior or meaning, leaving a required piece undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('write text to an element by handle') and even distinguishes two internal routing mechanisms (a11y elements vs visual elements). It lacks explicit differentiation from the sibling `ui.set_text`, which appears to overlap heavily.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use it (when you have a handle and want to input text), but it does not state when to use this rather than the sibling `ui.set_text` or `ui.find`. The two-path behavior (set_text native channel vs center-tap focus) is described but not framed as guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ui.wait_forA

轮询等待选择器出现/消失(等弹窗、加载完成、页面跳转),超时返回 ok=false

ParametersJSON Schema
NameRequiredDescriptionDefault
pkgNo包名(id 短名补全用)
goneNotrue=等待其消失
timeoutNo最长等待秒数
selectorYesGKD 选择器

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses failure semantics ('超时返回 ok=false'), which is real value, but says nothing about polling interval, blocking behavior, or whether it throws versus returns. Adequate but incomplete for an annotation-free tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence front-loads purpose, adds concrete use cases, and closes with failure behavior. Zero waste; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-annotation, no-output-schema tool with four well-documented params, the description covers purpose, typical scenarios, and timeout outcome. It could do slightly more (e.g., polling cadence or selector format), but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters (pkg, gone, timeout, selector) with defaults and meaning. The description only mirrors the appear/disappear concept already captured by the 'gone' parameter, adding no syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource ('轮询等待选择器出现/消失' – poll-wait for selector appear/disappear), which is more precise than the bare name. It doesn't explicitly contrast with close siblings like app.wait_idle, events.wait, or ui2.wait_change, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical scenarios ('等弹窗、加载完成、页面跳转' – popups, loading complete, page navigation) give concrete when-to-use context. However, no alternatives or when-not conditions are given, so an agent must infer the choice between this and app.wait_idle/ui2.wait_change.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vision.askA

通用识图增强入口(增强层,非主链路)。MAH_VISION_MODE 决定策略:client=截图以 MCP image 内容块返回,由具备视觉能力的客户端模型自己判读(零额外部署);endpoint=转发问题到外部部署的视觉端点(MAH_VLM_URL/MAH_VLM_MODEL,同 vision.vlm);off=关闭(默认,仅用 ui.snapshot/ui2.*/图集等确定性证据层)

ParametersJSON Schema
NameRequiredDescriptionDefault
boxNo[l,t,r,b] 只取该区域
fromNo截图路径,缺省最近一次 vision.screenshot
promptNo问题(endpoint 模式用)

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It usefully discloses mode-dependent behavior, the default (off), and that client mode returns an MCP image content block judged by the client model. It does not cover permissions, rate limits, or failure behavior for the endpoint path, leaving meaningful gaps for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, but the definition is a single dense block overloaded with parentheticals and environment-variable names. It is information-rich yet hard to scan, with the mode mechanics and the off-state caveat competing for attention.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an enhancement tool with no output schema and no annotations, the description covers the mode strategies, the default state, and the return format in client mode. What an agent needs to decide whether to invoke it is largely present, though endpoint-path behavior remains thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents box, from, and prompt (including that prompt is used in endpoint mode). The description adds no syntax or semantics beyond what the schema provides, so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (a universal image-recognition enhancement entry) and explicitly frames it as an enhancement layer rather than the main path. It distinguishes itself from the deterministic evidence siblings (ui.snapshot/ui2.*) and notes endpoint mode is 'same as vision.vlm', giving partial sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains when each strategy applies via MAH_VISION_MODE (client/endpoint/off) and states the default is off with reliance on deterministic evidence layers instead. It hints at the vision.vlm alternative but does not give an explicit 'prefer X over this tool when…' rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vision.diffB

界面变化检测:当前截图 vs 基线的像素差异(变化占比+差异框);ocr=true 附文本增删

ParametersJSON Schema
NameRequiredDescriptionDefault
ocrNotrue 时对两图各跑 OCR,返回文本集合差异
currentNo对比图路径;缺省当场截图
baselineNo基线图路径;缺省用最近一次 vision.screenshot
thresholdNo像素差灰度阈值(0-255),低于视为噪声

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the output shape (change ratio, diff boxes) and that ocr=true adds text add/delete, which is genuine behavioral context. However it omits anything about baseline selection behavior, tolerances being noise-based, or the read-only/analysis nature of the operation beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded clause states the operation, the output, and the optional ocr behavior with zero filler. The semicolon structure reads cleanly and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description covers the essentials an agent needs: what is compared, what is returned, and the ocr extension. The main omission is disambiguation from the similar ui2.diff sibling, which leaves a routing gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (ocr, current, baseline, threshold) are already documented in the schema. The description's mention of ocr=true only restates the schema's ocr semantics, adding no new syntax or default info, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: pixel-difference change detection between current screenshot and baseline, plus what is returned (change ratio + diff boxes). An agent understands the operation immediately, though the description never differentiates itself from the sibling ui2.diff, which sounds like an overlapping change-detection tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied (visual regression / change detection). There is no explicit when-to-use, no when-not-to-use, and no guidance on choosing this over siblings such as ui2.diff, ui2.wait_change, or vision.screenshot. The agent must infer the scenario by itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vision.ocrB

OCR(懒加载 rapidocr_onnxruntime;未安装则明确报错);path=本地图片离线 OCR,缺省截屏

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo本地图片路径(离线模式,无需设备连接)
regionNo

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses lazy-loading of rapidocr_onnxruntime, an explicit error when the dependency is missing, offline/no-device-connection operation, and the screenshot fallback. It omits performance/rate characteristics, but the operational behaviors that matter for invocation are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded with the core purpose, using semicolons to pack dependency, error, and default behaviors into one line. No wasted text, though the heavy parenthetical phrasing slightly compresses readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param tool with no output schema, the description covers purpose, dependency loading, error behavior, offline mode, and default input. It is incomplete on the region parameter and says nothing about the return format (recognized text structure), leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The path parameter is clarified (offline local image, defaults to screenshot when omitted), adding value beyond the schema. But the region parameter (array of integers, e.g. a crop box) is undocumented in both schema and description, leaving half the parameters without meaning at 50% coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource (OCR on a local image), and clarifies the default behavior when path is omitted (screenshot). It does not explicitly differentiate itself from vision-adjacent siblings like vision.vlm or vision.ask, which also extract information from images, so it stops short of full sibling disambiguation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use it to OCR a local image, or a screenshot by default. The description notes offline mode requires no device connection, which hints at when this is preferable. However, there is no explicit when/when-not guidance or named alternative (e.g. use vision.vlm for semantic understanding instead of raw text).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vision.screenshotB

截屏存盘(视觉兜底/证据);路径会记住,作为 vision.diff 缺省基线

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It does disclose real behavior beyond the schema: it writes a file to disk and the path is remembered as persistent state feeding vision.diff's default baseline. It omits what is returned, overwrite behavior, and whether the path is optional or auto-generated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact clause with the action front-loaded and the cross-tool consequence appended. No filler, though it is terse enough to feel underspecified rather than optimally structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 1-param tool with no output schema, the description covers the core action and the vision.diff linkage. It still leaves gaps on return value and the optional path default, which matter for calling it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'path' param has 0% schema description coverage. The description adds meaning by explaining that the path is remembered and reused as vision.diff's default baseline, but it doesn't state the format, whether it is optional, or what happens when omitted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (screenshot saved to disk) and adds a distinguishing role: it becomes the default baseline for vision.diff. This differentiates it from vision.ocr/vlm/ask, though it doesn't explicitly contrast against ui2.screen_evidence or the other capture tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'视觉兜底/证据' (visual fallback/evidence) implies when to reach for it – when structured UI queries fail or proof is needed – and it names its downstream relationship to vision.diff. However there is no explicit when-not guidance or named alternative to prefer first, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vision.vlmA

本地视觉小模型问图(crop-and-ask):缺省整屏,传 box 只裁剪该区域提问(推荐:小图快且准)。需本地 OpenAI 兼容视觉端点:设 MAH_VLM_URL 与 MAH_VLM_MODEL(如 ollama:MAH_VLM_URL=http://127.0.0.1:11434/v1/chat/completions,MAH_VLM_MODEL=qwen2.5vl:3b)。定位:图集外未知图标的语义兜底层

ParametersJSON Schema
NameRequiredDescriptionDefault
boxNo[l,t,r,b] 只裁剪该区域
fromNo截图路径,缺省最近一次 vision.screenshot
promptNo问题;缺省=描述该界面元素及用途
timeoutNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a real prerequisite beyond structured data: a local OpenAI-compatible vision endpoint configured via MAH_VLM_URL/MAH_VLM_MODEL, with a concrete ollama example. That is valuable. But it omits return format, latency/failure behavior, and what happens if the endpoint is unconfigured, so a 3 fits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core crop-and-ask behavior, then setup requirements, then positioning. The env-var example is verbose but genuinely useful for invocation. Overall appropriately sized with little wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must cover a lot, and it does well on setup and defaults. The main gap is the return-value shape (what the model returns: free text? structured?), which an agent calling a Q&A tool would benefit from knowing. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%. The description adds meaning beyond the schema: box semantics are reinforced ('only crop that region', 'recommended: small images fast and accurate') and the default-scope behavior (full screen when box absent) is explicit, which the schema only hints at. The from/prompt/timeout params are largely covered by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific capability: a local VLM 'crop-and-ask' that takes a screenshot (default full screen) or a cropped box and asks a vision model about it. It also positions itself as the 'semantic fallback layer for unknown icons outside the gallery', which helps distinguish it from template/icon-matching siblings. It does not explicitly name the nearest siblings (vision.ask, vision.ocr), so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives useful context: default is full screen, pass a box to crop that region, and cropping is recommended for speed/accuracy. The 'positioning' line implies it is a fallback when gallery/icon matching fails. However it never explicitly says when to choose this over vision.ask, vision.ocr, or ui2.scene_match, so usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 60 tool updatesv0.5.0
    • First observedapp.current
    • First observedapp.launch
    • First observedapp.list
    • First observedapp.stop
    • First observedapp.wait_idle
    • First observedcom.android.settings__launch
    • First observedcom.android.settings__search
    • First observeddevice.wake_unlock
    • First observedevents.capture
    • First observedevents.record_start
    • First observedevents.record_stop
    • First observedevents.snapshot
    • First observedevents.tail
    • First observedevents.wait
    • First observedfile.pull
    • First observedfile.push
    • First observedharness.check_versions
    • First observedharness.list
    • First observedharness.run
    • First observedinput.clipboard
    • First observedinput.key
    • First observedinput.swipe
    • First observedinput.tap
    • First observedknowledge.guide
    • First observedknowledge.list
    • First observednet.http
    • First observedshell.run
    • First observedstate.db
    • First observedstate.prefs
    • First observedsys.capabilities
    • First observedtask.begin
    • First observedtask.end
    • First observedtask.note
    • First observedtemplate.add
    • First observedui.back
    • First observedui.click
    • First observedui.click_handle
    • First observedui.dump
    • First observedui.find
    • First observedui.get_text
    • First observedui.hold_read
    • First observedui.home
    • First observedui.long_click
    • First observedui.scroll
    • First observedui.set_text
    • First observedui.snapshot
    • First observedui.text_handle
    • First observedui.wait_for
    • First observedui2.check_states
    • First observedui2.diff
    • First observedui2.find_templates
    • First observedui2.scene_match
    • First observedui2.screen_evidence
    • First observedui2.state
    • First observedui2.wait_change
    • First observedvision.ask
    • First observedvision.diff
    • First observedvision.ocr
    • First observedvision.screenshot
    • First observedvision.vlm

TDQS

C2.5/5.0

Scored across 60 tools

Disambiguation2/5

Many tools overlap heavily across layers: ui.click vs ui.click_handle vs input.tap; ui.set_text vs ui.text_handle; ui.dump vs ui.snapshot vs ui2.state; vision.ocr vs vision.vlm vs vision.ask. Descriptions try to explain preferred/default vs fallback vs primitive, but an agent still faces multiple plausible choices for the same action.

Naming Consistency3/5

Most names use a readable namespace.verb_noun pattern such as ui.get_text, app.launch, vision.screenshot, but the set mixes ui vs ui2, dot-delimited namespaces, and harness-specific double-underscore names like com.android.settings__launch. The convention is largely understandable but not uniform.

Tool Count1/5

With 60 tools, the server is far beyond a well-scoped set and likely overwhelms tool selection. The layering of ui, ui2, input, vision, events, harness, and fallback tools suggests internal architecture leaking directly into the exposed surface.

Completeness4/5

The surface covers app lifecycle, UI interaction, vision/OCR, files, shell, state, events, tasks, templates, harnesses, and knowledge packages, which is very comprehensive for phone automation. Some device-level domains such as telephony, notifications, sensors, or media remain absent, but core agent workflows are well represented.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    A comprehensive MCP server that enables AI agents to interact with Android devices through Android Debug Bridge (ADB), offering 198 tools for device control, app management, diagnostics, and more.
    100
    104 npm
    19
    Apache 2.0
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables MCP-compatible agents to control an Android device over the network via ADB, providing tools for shell commands, screen capture, UI inspection, file operations, and input simulation.
    16 npm
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    A powerful MCP server that provides comprehensive Android device automation capabilities through ADB, enabling AI agents to interact with Android devices for testing, automation, and device control tasks.
    1
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents and test runners to control Android devices via local ADB, including listing devices, tapping, swiping, typing, sending system keys, launching apps, and dumping UI hierarchy. Communicates over MCP stdio without exposing network listeners.
    -