Skip to main content
Glama
conorluddy

XC-MCP: XCode CLI wrapper

by conorluddy

Describe Accessibility Tree

idb-ui-describe
Read-onlyIdempotent

Query iOS accessibility tree to discover tappable elements, text fields, and their coordinates, enabling targeted UI automation without screenshots.

Instructions

idb-ui-describe

šŸ” Query UI accessibility tree - discover tappable elements and text fields for precise automation

What it does

Queries iOS accessibility tree to discover UI elements, their properties (type, label, enabled state), coordinates (frame, centerX, centerY), and accessibility identifiers. Returns full tree with progressive disclosure (summary + cache ID for full data), element-at-point queries for tap validation, and data quality assessment (rich/moderate/minimal) to guide automation strategy. Automatically parses NDJSON output to extract all elements (not just first), includes AXFrame coordinate parsing for precise tapping, and caches large outputs to prevent token overflow.

Progressive Filtering: Supports 4 filter levels for element discovery - start conservative with moderate filtering (default), escalate to permissive/none if minimal data found.

iOS Compatibility: Recognizes iOS-specific accessibility fields (role, role_description, AXLabel, AXFrame) in addition to standard fields.

Why you'd use it

  • Discover all tappable elements from accessibility tree - buttons, cells, links identified by JSON element objects

  • Get precise tap coordinates (centerX, centerY) for elements without needing screenshots

  • Assess data quality before choosing automation approach - rich data enables precise targeting, minimal data requires screenshots

  • Validate tap coordinates by querying elements at specific points before execution

  • Progressive disclosure prevents token overflow on complex UIs - get summary first, full tree on demand

  • Progressive filter escalation - start with moderate filtering, escalate to permissive/none if minimal data found

Parameters

Required

  • operation (string): "all" | "point"

Point operation parameters

  • x (number, required for point operation): X coordinate to query

  • y (number, required for point operation): Y coordinate to query

Optional

  • udid (string): Target identifier - auto-detects if omitted

  • screenContext (string): Screen name for context (e.g., "LoginScreen")

  • purposeDescription (string): Query purpose (e.g., "Find tappable button")

  • filterLevel (string): "strict" | "moderate" | "permissive" | "none" (default: "moderate")

    • strict: Only obvious interactive elements via type field (original behavior)

    • moderate: Include iOS roles (role, role_description) - DEFAULT, fixes iOS button detection

    • permissive: Any element with role/type/label information

    • none: Return everything (debugging)

Returns

For "all": UI tree summary with element counts (total, tappable, text fields), data quality assessment (rich/moderate/minimal), top 20 interactive elements preview with centerX/centerY coordinates, uiTreeId for full tree retrieval, current filter level, and guidance on automation strategy including suggestions to escalate filter level if minimal data found.

For "point": Element details at coordinates including type, label, value, identifier, frame coordinates (x, y, centerX, centerY), enabled state, and tappability.

Examples

Query full UI tree with default moderate filtering

const result = await idbUiDescribeTool({
  operation: 'all',
  screenContext: 'LoginScreen',
  purposeDescription: 'Find email and password fields'
});
// Result includes elements with centerX, centerY for direct tapping

Progressive filter escalation pattern

// 1. Start with default (moderate)
let result = await idbUiDescribeTool({ operation: 'all' });

// 2. If minimal data, try permissive
if (result.summary.dataQuality === 'minimal') {
  result = await idbUiDescribeTool({
    operation: 'all',
    filterLevel: 'permissive'
  });
}

// 3. If still minimal, try none (return everything)
if (result.summary.dataQuality === 'minimal') {
  result = await idbUiDescribeTool({
    operation: 'all',
    filterLevel: 'none'
  });
}

// 4. If STILL minimal, fall back to screenshots
if (result.summary.dataQuality === 'minimal') {
  // Use screenshot-based approach
}

Validate element at tap coordinates

const element = await idbUiDescribeTool({
  operation: 'point',
  x: 200,
  y: 400
});
// Element includes frame coordinates if available
  • idb-ui-tap: Tap discovered elements using centerX/centerY coordinates

  • screenshot: Capture screenshot for visual element identification

  • idb-ui-find-element: Semantic element search by label/identifier

  • accessibility-quality-check: Quick assessment before choosing approach

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
udidNo
operationYes
screenContextNo
purposeDescriptionNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.1.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already set readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable context beyond these: it mentions caching large outputs to prevent token overflow, progressive disclosure, NDJSON parsing, and data quality assessment. It also explains the filtering behavior and when to escalate, which is not present in annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with clear sections (What it does, Why you'd use it, Parameters, Returns, Examples, Related Tools) and effectively front-loaded. Every sentence contributes necessary information for a tool of this complexity. The examples are concise and illustrate key usage patterns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description must explain return values, and it does thoroughly for both operations. It covers all parameters, return structure, progressive filtering, and even includes code examples. It also relates to sibling tools, completing the decision-making context. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has zero descriptions and parameter coverage is 0%, so the description carries the full burden. It meticulously explains each parameter: operation enum, x/y coordinates, udid auto-detection, screenContext, purposeDescription, and filterLevel (even though filterLevel is missing from the schema, the description documents it thoroughly with four levels and defaults). It adds meaning far beyond what the bare schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb ('query UI accessibility tree') and resource (iOS UI elements), and explicitly distinguishes itself from siblings like idb-ui-tap and idb-ui-find-element. It clearly defines what it does: discover tappable elements, text fields, and coordinates. No ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Why you'd use it' section provides concrete scenarios, and the 'Related Tools' section names alternatives with their purposes. It gives explicit fallback guidance (e.g., 'if minimal data, fall back to screenshots') and a progressive filter escalation pattern, leaving no doubt about when to use this tool versus others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.