Skip to main content
Glama

ui2.state

Capture a structured Android UI state combining accessibility and OCR to identify interactive elements and canvas text. Returns element handles for direct actions, including on custom-drawn screens.

Instructions

结构化 UI 状态 v2:a11y×OCR 融合——a11y 提供零误报可交互性,OCR 补 canvas/自绘文本(src=ocr 的元素为启发式推断)。元素带句柄,ui.click_handle 直接操作。canvas/自绘界面必用

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
max_elementsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.0

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavioral traits: a11y gives zero-false-positive interactability, OCR-sourced elements (src=ocr) are heuristic inferences, and elements carry handles usable by ui.click_handle. It stops short of covering cost, size limits, or return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense, front-loaded sentence with no filler; the fusion mechanism, handle capability, and usage directive are all packed efficiently. It is information-dense but still readable and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe returns, and it partially does (handles, src markers, interactability signal). It still omits the shape of the returned state and the effect of max_elements, leaving meaningful gaps for a state-dump tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter max_elements (default 80) has 0% schema description coverage and is not mentioned anywhere in the description, so an agent cannot learn its meaning or effect. With 1 undocumented parameter the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource ('structured UI state v2') and explains the mechanism (a11y×OCR fusion), which lets an agent understand what the output represents. It implies distinction from ui.dump/ui.snapshot via the fusion/OCR angle but never names or contrasts those siblings explicitly, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one clear directive condition ('must use for canvas/self-drawn interfaces') and hints that OCR fills gaps a11y misses. However it does not state when NOT to use it versus the many nearby siblings (ui.dump, ui.snapshot, ui.get_text, vision.ocr), leaving routing largely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.