Skip to main content
Glama

Observe the screen

observe
Read-onlyIdempotent

Inspect desktop app and browser GUIs via accessibility tree to list interactive elements with refs and pixel boxes, capture screenshots or Set-of-Mark views, and expand containers for automation.

Instructions

Observe the GUI: interactive elements with refs (e7) and [x,y,w,h] boxes in image pixels. mode auto returns the tree, or a Set-of-Mark screenshot when the tree is thin; screenshot/som attach an image. root_ref expands a container; scope='screen' lists windows for cross-app work. Start every step with this.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
appNoApp to observe; default: the last observed window or the frontmost app
modeNoauto: tree, or Set-of-Mark screenshot when the tree is thinauto
pageNoSet-of-Mark page (80 marks each)
scopeNoscreen only for cross-app work (deny-listed apps masked)window
root_refNoExpand this container: list all its descendants
window_idNoWindow from apps(action='windows')
max_elementsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/non-destructive, so the bar is lower. The description adds real behavior: auto mode falls back to a Set-of-Mark screenshot when the tree is thin, screenshot/som attach an image, and scope='screen' masks deny-listed apps. It does not mention pagination limits or output size caps, which would have pushed this higher.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences with no filler, and the core purpose plus the 'start every step with this' directive are front-loaded. The middle clauses about mode and refs are packed tightly enough to be slightly hard to scan, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the essential workflow positioning plus mode/scope/root_ref semantics for a 7-parameter tool. It leaves minor gaps around default app resolution and result size limits, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 86%, so the baseline is 3, but the description adds coordinate-frame semantics ([x,y,w,h] in image pixels) and ref format (e7) that the schema does not state. Mode, scope, and root_ref behavior are largely restated from the schema, so the gain is incremental rather than comprehensive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (observe) and resource (the GUI), then describes exactly what is returned: interactive elements with refs like e7 and [x,y,w,h] boxes in image pixels. This clearly separates it from action siblings such as click, type_text, and scroll, which operate on refs that observe produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit workflow directive ('Start every step with this') and conditions for specific modes (scope='screen' for cross-app work, root_ref to expand a container). It does not mention the locate sibling or explain when to prefer locate over observe, so it falls short of full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.