ui2.diff
Compare two captured UI states to detect added or removed text elements, helping identify interface changes during Android automation.
Instructions
界面变更检测:对比两次融合状态的文本元素增删
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| gap | No |
Compare two captured UI states to detect added or removed text elements, helping identify interface changes during Android automation.
界面变更检测:对比两次融合状态的文本元素增删
| Name | Required | Description | Default |
|---|---|---|---|
| gap | No |
Changes observed during successful MCP inspections.
v0.5.0Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does not say whether the operation is read-only, where the 'two fused states' come from, whether prior state calls are required, or what the delta output contains. It only restates the comparison in prose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the operation and its scope with no filler. It is efficient, though the density comes partly at the cost of the missing detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, an undocumented parameter, and no explanation of how the two states are supplied leaves an agent unable to call this correctly. For a diff tool that must consume prior state, this is a significant shortfall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'gap' has 0% schema description coverage and is never mentioned in the description, so its meaning (tolerance? frame gap between captures?) is completely undocumented. With one undocumented parameter, the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (对比/diff) and a specific resource (两次融合状态的文本元素增删), so an agent knows it produces a text-element delta between two fused UI states. It does not distinguish itself from the adjacent siblings that also compare screens (vision.diff) or watch for change (ui2.wait_change), so it falls short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this rather than ui2.state, ui2.wait_change, ui2.check_states, or vision.diff, and no prerequisites (e.g. whether two ui2.state calls must be made first). The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.