Skip to main content
Glama

native-describe-screen

Read a running app's raw native accessibility tree and point-space geometry to debug low-level UI data behind the public describe screen tool.

Instructions

Read the running app's native accessibility screen description via injected native devtools.

Returns a flat list of accessibility leaf elements with:

  • raw frame and tapPoint in points, on the portrait axes of the screen

  • normalizedFrame and normalizedTapPoint on the same axes. gesture-tap uses these axes, also when the UI is landscape

  • top-level screenFrame metadata: the screen bounds on these axes

  • traits and optional labels/identifiers

This is a low-level native inspection tool. The normalized fields are intended to help with backend migration work, but the public describe contract is still separate.

Use when you are evaluating or debugging the lower-level native data behind the public describe tool, or when you need its raw point-space geometry rather than describe's normalized contract.

If status is restart_required: follow the message (usually restart-app), then retry. If status is service_stale: the app is already injected, so restarting it cannot help — restart the tool-server (argent server stop && argent server start --detach) and retry. If the same status comes back after that restart, stop restarting: follow the message, which names the terminal fallback. If status is connect_pending: the app is injected and still connecting — do not restart it, wait a few seconds and retry. If status is init_failed: the simulator's native-devtools environment could not be initialised — follow the message (re-boot the simulator) rather than retrying this tool. A not-connected or not-running app comes back as one of those statuses rather than a failure. Failures are separate: an Apple system app is rejected outright (terminal — never retry it), and the screen query itself can error or time out.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
udidYesSimulator UDID
bundleIdYesBundle ID of the app
skipClassesNoExact UIView class names whose entire subtree should be pruned (e.g. ["UIImageView"] to drop image-heavy branches)
skipClassPrefixesNoClass name prefixes to prune entire subtrees. For SwiftUI apps use ["_TtGC7SwiftUI"] to drop mangled SwiftUI generic type subtrees while keeping UIKit bridges.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.15.0

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully covers behavioral expectations: it lists output fields, then enumerates status codes (`restart_required`, `service_stale`, `connect_pending`, `init_failed`) with exact follow-up actions, and explicitly states that non-connected apps return statuses rather than failures and that Apple system apps are terminal. This goes beyond a simple read-only hint and gives an agent actionable operational semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with a purpose sentence, an output bullet list, a low-level/describe contrast, a use-when line, and a status-handling paragraph. It is long, but every sentence provides distinct value — the status block, while dense, covers the only operational edge cases an agent might hit. No word is wasted, though it could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-level inspection tool with no annotations and no output schema, the description is self-sufficient: it defines the returned element list, coordinate-axis normalization, status handling, and terminal failure conditions. The only omitted topic is parameter behavior, but the schema covers that completely. An agent has all necessary context to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes all four parameters (udid, bundleId, skipClasses, skipClassPrefixes) with 100% coverage, so the description does not need to add parameter details. The description adds no extra parameter semantics, which is acceptable given the schema's completeness; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Read the running app's native accessibility screen description via injected native devtools.' It further details a flat list of accessibility leaf elements with raw/normalized frames, and explicitly labels itself as a low-level tool separate from the 'public describe contract,' distinguishing it from the sibling `describe` tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides an explicit use-case sentence: 'Use when you are evaluating or debugging the lower-level native data behind the public describe tool, or when you need its raw point-space geometry rather than describe's normalized contract.' It also gives situational retry guidance for statuses, but it does not enumerate when to choose `native-full-hierarchy` or `native-find-views`, so it stays at 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.