Skip to main content
Glama

ui.snapshot

Capture a compact accessibility snapshot of the current Android screen, labeling interactive elements with handles for UI inspection and click/text operations.

Instructions

读屏默认首选:a11y 语义裁剪快照,元素编号 e0/e1/...(token 约为 ui.dump 的 1/10)。k 类型 btn/input/toggle/scroll/icon/text;i=可交互 e=可输入 s=可滚动 v=开关态。用 ui.click_handle/ui.text_handle 按句柄操作

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pkgNo缺省取当前前台应用
max_elementsNo
include_boundsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden. It discloses the output shape (element numbering e0/e1/...), the per-element flag legend (k type, i/e/s/v states) and the token-cost characteristic, which is meaningful behavioral context. It does not state pagination/truncation behavior (max_elements) or that it is read-only, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the decisive fact ('读屏默认首选') and then densely packed with the comparison and output legend; every clause earns its place. The telegraphic shorthand is information-dense rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description does the work of explaining the return format (element handles plus a flag legend), which is the key missing piece for an agent. Input-side detail (max_elements/include_bounds) is not covered, keeping it short of 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% and the description adds no input-parameter meaning: the k/i/e/s/v legend concerns output fields, not pkg, max_elements or include_bounds. Two of three parameters are undocumented in both schema and description, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('a11y 语义裁剪快照' - an accessibility-semantic-trimmed snapshot) and marks itself as the '读屏默认首选' (default first choice for screen reading). It also explicitly distinguishes itself from the sibling ui.dump by token cost (~1/10), so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Declares itself the default for screen reading and gives the comparative condition vs ui.dump (~1/10 tokens), implying when to prefer this over the heavier dump. It also routes the agent onward to ui.click_handle/ui.text_handle for acting on handles. No explicit 'when-not' exclusion is stated, so 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.