Skip to main content
Glama

Ui Find Element

ui_find_element
Read-only

GUI automation — control a native app's interface. Finds an element (button, field, menu…) in an app's accessibility tree by role and/or label. Scope with app_bundle_id or window_id. Returns an opaque element_ref (usable by ui_click / ui_get_element this session) plus role, label, bounds, focused, and enabled only when the app publishes AXEnabled (otherwise enabled_unknown: true, which does NOT mean disabled). found=false when the app is reachable but no element matches; app_not_found is an explicit error. Requires Accessibility permission.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
roleNoAX role, e.g. AXButton, AXMenuItem, AXTextField.
indexNoWhich match to return if several (default 0).
labelNoAX title/description to match.
matchNoDefault contains.
window_idNoAlternatively scope by a window_id from list_windows.
app_bundle_idNoScope the search to this app.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the read-only and non-destructive annotations, the description explains session-limited element_ref lifetime, the enabled vs enabled_unknown distinction, the found=false result when no match occurs, the app_not_found error, and the Accessibility permission requirement. This is rich behavioral disclosure that materially helps an agent anticipate outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries operational value: purpose, scoping, return value, edge cases, and permissions. The front-loaded 'GUI automation' context orients the agent immediately, and the length is appropriate for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the burden of explaining return semantics, error cases, and required permissions. It covers the key operational concerns an agent needs to invoke the tool correctly and interpret results, including subtle cases like enabled_unknown and app_not_found.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents role, label, index, match, window_id, and app_bundle_id. The description adds a useful framing of role and/or label and scoping, but does not go beyond the schema in explaining individual parameter semantics. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the operation: finding an element in a native app's accessibility tree by role and/or label, with optional scoping by app_bundle_id or window_id. It also specifies the return artifact and how it relates to sibling tools like ui_click and ui_get_element, making the tool's identity unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: locating a UI element for later interaction and scoping the search to a specific app or window. It does not explicitly enumerate alternatives like ui_read_tree or ui_wait_for_element, so it lacks exclusions, but the use case is well implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources