Skip to main content
Glama

ui_tree

List visible UI widgets with IDs, roles, names, and center coordinates to locate GUI controls for automation. Filter by app or window, or restrict to the active window.

Instructions

List visible widgets as [id] role "name" @(x,y) with center coordinates in screenshot space. Filter by app/window name substring, or active_window_only=True. Ids are valid until the next ui_tree call. Apps only appear if accessibility is enabled (Chrome/Electron need --force-renderer-accessibility).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
appNo
windowNo
max_depthNo
only_interactiveNo
active_window_onlyNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the id lifecycle ('Ids are valid until the next ui_tree call'), the coordinate space, and the accessibility prerequisite that silently hides apps. It does not state that this is a non-mutating read, but the list semantics make that reasonably clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the output format, then filters, then the id lifecycle and the accessibility caveat. Four dense sentences, no filler, and the most decision-relevant facts come first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no further explanation, and the id-invalidation warning is valuable. However, for a 5-parameter tool with zero schema descriptions and zero annotations, the unexplained only_interactive default and max_depth leave real gaps an agent would hit when widgets it expects are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate. It explains app/window substring filtering and active_window_only, but says nothing about max_depth or only_interactive, whose default of true materially changes results (non-interactive widgets are silently omitted).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (list visible widgets) and even pins down the exact output shape `[id] role "name" @(x,y)` in screenshot space. It is clearly distinguishable in substance from siblings like screenshot or list_windows, though it never names an alternative to route between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Offers concrete usage mechanics (filter by app/window substring, or active_window_only=True) and a prerequisite (accessibility must be enabled, with the Chrome/Electron flag), but gives no explicit when-to-use-vs-screenshot/list_windows guidance or when-not-to-use condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.