Skip to main content
Glama

glass

CI glass MCP server

Give your coding agent hands for the app it's building — launch it, drive it, and verify the result, without burning a screenshot on every step.

A Rust MCP server that gives an AI coding agent a closed build → see → interact → debug loop over external native GUI applications.

glass lets an agent launch a GUI app, capture what is on screen, inject mouse and keyboard input, read the app's logs, and detect visual changes — so a coding agent can build and debug UI applications independently instead of asking the user "does this look right?".

glass drives apps as an external black box, so it works with any native GUI app regardless of toolkit or language. It has two Linux backends (X11 and Wayland), a Windows backend, an Android backend (an AVD emulator, driven over adb from any host), an iOS backend (native apps in the Simulator over xcrun simctl, with input and the accessibility tree via idb_companion, including a two-finger pinch), and a macOS backend, behind a platform-agnostic core.

See it

An agent debugging a GTK app under glass

An agent building a GUI app runs it under glass, reproduces a bug from the accessibility tree (no screenshots), fixes the code, and re-verifies — the loop it otherwise can't close on its own. Try it yourself.

Related MCP server: Playwright MCP

Try it in 60 seconds

  1. Download glass for your platform from the Releases page and connect it to your agent — glass works with any MCP host (see which are verified).

  2. Get the example app — clone this repo, or download examples/tasks-demo/tasks_demo.py (on Linux it needs sudo apt install python3-gi gir1.2-gtk-4.0).

  3. Paste this to your agent:

    Use glass to run examples/tasks-demo/tasks_demo.py with accessibility on. There's a bug: clicking Add doesn't add the typed task. Reproduce it by driving the UI and checking the accessibility tree (don't just screenshot), then find and fix the bug in the code and verify a task actually appears.

Your agent launches the app, reproduces the bug from the accessibility tree, fixes the one-line wiring bug, and confirms the task appears — the whole build → see → interact → debug loop, start to finish. (glass-mcp doctor checks your environment if anything's off.)

The loop in practice

An agent can fill in a form, click Save, and check that the app reports success. When the app exposes an accessibility tree, the agent works with named controls and verifies changes from text:

glass_start { "run": ["python3", "app.py"] }
glass_do { "actions": [
  { "action": "set_value",
    "target": { "query": "account", "role": "TextField", "states": ["enabled"] },
    "text": "hello" },
  { "action": "click_element",
    "target": { "query": "Save", "role": "Button", "states": ["enabled"] },
    "mode": "auto" },
  { "action": "wait_for_element", "name": "Saved" }
] }
glass_logs

Use glass_find_elements to inspect candidates when the intended target is not unique or known well enough to act on directly. Use glass_a11y_snapshot to explore the app's controls. See the tool reference for click modes and platform behavior.

For a canvas or custom-rendered app with no accessibility tree, drive it by pixels instead — glass_screenshot, glass_click {x,y}, and glass_diff, which returns changed_pct + a bbox as text, so routine checks between screenshots cost no vision tokens. Why the loop is shaped this way: the build → see → interact → debug loop.

To review a run after the app stops, opt into session evidence recording with --trace-dir. Glass retains tool inputs and requested results in bounded storage, then glass-mcp trace inspect and trace export validate the trace and package its evidence in a ZIP. Recording is off by default; retained app content can contain sensitive data.

Install at a glance

Download the latest build for your platform from the Releases page, then set up your host:

Every asset is listed in docs/reference/platforms.md. Prefer to compile, or on an architecture with no published asset? See docs/how-to/build-from-source.md — it is a single cargo build.

Then connect glass to your agent and run glass-mcp doctor to check the environment. New here? Follow the tutorial for a guaranteed first success.

The full toolbox remains the default. For fewer agent tool definitions, start with glass-mcp --tool-profile lean and use glass_do for actions. Inspect either inventory without starting a session using glass-mcp tools --json; see tool profiles.

Drive it well — the glass-drive skill

glass needs no app integration and no skill to run, but an agent drives it far more reliably with the open glass-drive Agent Skill — it stops the agent spending its first turns rediscovering the verify-cheaply-then-look loop. Installing it is the single highest-leverage thing you can add when pointing an agent at glass.

Platform support

✓ supported · ◑ partial · – not supported · 🚧 planned.

Capability

Linux (X11 + Wayland)

Windows

Android (AVD)

iOS (Simulator)

macOS

Capture · input · windows · clipboard · logs

✓

✓

✓

✓

✓

Accessibility (semantic addressing)

✓ AT-SPI

✓ UI Automation

✓ UIAutomator

✓ idb

✓ AX

Containment / sandboxing

✓ bubblewrap

✓ Sandboxie

✓ the emulator VM

✓ the Simulator

✓ Seatbelt

Display isolation (app off your desktop)

✓ headless Xvfb / sway

◑ virtual display · VM tier

✓ headless emulator

✓ headless simctl boot

🚧

Full matrix, per-capability detail, and system requirements: docs/reference/platforms.md. Transport is MCP over stdio (default) or network HTTP.

Documentation

The full docs — tutorial, how-to guides, reference, and explanations — are under docs/. See CHANGELOG.md for release notes, and Stability and versioning for what a 1.0 release guarantees.

Contributing? CONTRIBUTING.md has the gates a PR has to pass, and Verify a change covers platform code your own host cannot run.

License

glass is open core, licensed Apache-2.0 — see LICENSE-APACHE.

Available Tools

31 tools
glass_a11y_marksA
Read-only

Capture a screenshot with numbered accessibility boxes and a text legend. Use the IDs with a click_element action; they share the latest snapshot. Boxes follow toolkit geometry and can drift; inspect action results. Errors without a tree; use glass_screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so no safety contradiction. The description adds valuable behavioral context beyond annotations: boxes 'can drift' and the user should 'inspect action results', plus the error condition 'without a tree'. This informs the agent about reliability and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences: the core action, the intended usage of results, and a caveat plus alternative. Every sentence earns its place, and the most critical information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool with no output schema, the description tells the agent what the tool produces, how to consume the result, what can go wrong, and which sibling to fall back to. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is trivially 100% and the baseline is 4. The description adds no parameter syntax, but none is needed; it instead clarifies that the produced IDs are output artifacts, not inputs, which is relevant for post-processing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Capture a screenshot with numbered accessibility boxes and a text legend.' This clearly distinguishes it from siblings like glass_screenshot (plain screenshot) and glass_a11y_snapshot (accessibility tree), and it adds the distinctive detail that the boxes carry numbers usable with click_element.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use and when-not-to-use guidance: 'Use the IDs with a click_element action' names the downstream interaction pattern, and 'Errors without a tree; use glass_screenshot' explicitly tells the agent to fall back to another tool when the accessibility tree is unavailable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_a11y_snapshotA
Read-only

Read current semantic state: IDs, roles, name/description, bounded editable value, bounds and states. Values can be absent, redacted or truncated; use wait_for_element for exact verification. IDs refresh each snapshot. A childless Document notice calls for one fresh read, then pixels; an unpublished-content placeholder requires pixels. Errors without an accessibility tree; use glass_screenshot. Batch known actions and checks with glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_nodesNoNode cap; omit for server default, 0 for unlimited. Changing the cap renumbers IDs; re-read them.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only convey readOnlyHint=true; the description adds substantial behavioral context: values can be absent/redacted/truncated, IDs renumber on each snapshot, a childless Document notice requires one fresh read followed by pixels, and errors without an accessibility tree should fall back to glass_screenshot. These are exactly the failure modes and re-read behaviors an agent cannot infer from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five dense sentences, each carrying distinct information, with the core purpose front-loaded in the first clause. The collapsed workflow quirk sentence ('A childless Document notice... an unpublished-content placeholder') is terse to the point of requiring careful parsing, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully covers return contents (IDs, roles, names, values, bounds, states), value fidelity caveats, ID refresh behavior, error conditions and their fallback, and batching guidance. For a read-only snapshot tool whose safety profile is already covered by annotations, nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: the schema for max_nodes already documents the node cap, server default, 0-for-unlimited behavior, and the ID-renumbering consequence of changing the cap. The description itself adds no parameter detail, but with the schema carrying the full burden, the baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb ('Read') and a specific resource (the current semantic state), then enumerates exactly what is returned: IDs, roles, name/description, bounded editable value, bounds, and states. This makes it immediately distinguishable from siblings like glass_screenshot (pixels), glass_wait_for_element (exact verification), and glass_a11y_marks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description names three explicit alternatives with the conditions that select them: wait_for_element when exact verification is needed (values may be absent/redacted/truncated), glass_screenshot when there is no accessibility tree or content is image-based (childless Document notice, unpublished-content placeholder), and glass_do for batching known actions and checks. This is textbook when-to-use-vs-alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_baseline_saveA

Save the whole current window as a named pixel baseline for glass_diff or glass_wait_for_region. Settle first if animating. Replaces an existing name silently. Baselines survive glass_stop and expire when the server exits.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesASCII letters/digits/-/_ only. Replaces an existing baseline silently; used by glass_diff and glass_wait_for_region.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are thin (only safety hints), so the description carries the behavioral burden — and it delivers: silent replacement of existing names, persistence across glass_stop, and expiry at server exit are all explicitly disclosed. 'Settle first if animating' also pre-announces a likely failure mode for unstable windows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four terse sentences, each carrying distinct information: the core action, the animation prerequisite, replacement behavior, and lifecycle semantics. Primary purpose is front-loaded with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter save operation with no output schema, the description covers action, prerequisite, side effects, and lifecycle. The only unaddressed gap is the success/return signal, which is minor for a tool this simple.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — the schema already documents the ASCII character restriction, silent replacement, and downstream consumers for the name parameter. The description repeats the consumers rather than adding new parameter-specific meaning, matching the baseline baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb + resource ('Save the whole current window as a named pixel baseline') and names the downstream consumers (glass_diff, glass_wait_for_region), differentiating it from capture-only siblings like glass_screenshot. The 'whole current window' scope makes the action precise and immediately actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context by naming the two tools that consume baselines and gives a concrete prerequisite ('Settle first if animating') that ties into sibling tools like glass_wait_stable. It doesn't explicitly state when-not-to-use or name an alternative, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_capabilitiesA
Read-only

Report live backend operation status and the exposed tools it affects, plus tool_profile. Status: supported; degraded (reduced fidelity); requires_setup; unsupported. Notes explain limits/remedies. No session required; backend defaults to the active/default backend. Profile membership does not imply backend support.

ParametersJSON Schema
NameRequiredDescriptionDefault
backendNox11, wayland, windows, macos, android or ios. Default active/default backend. Valid but unbuilt backends report available:false.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavioral context beyond the readOnlyHint/openWorldHint annotations: no session required, default backend behavior, status semantics, and the caveat that profile membership does not imply backend support. It does not contradict annotations and gives an agent realistic expectations about tool behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, then efficiently adds status categories, notes, session requirements, and caveats. Every sentence contributes distinct information with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with one optional parameter and no output schema, the description explains what will be reported (status per tool, tool_profile, notes) and key call semantics like default backend and session requirements. It is complete enough for correct invocation, though it could be slightly more explicit about the response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage for the single backend parameter, including valid values and default behavior, so the description does not need to compensate. It adds the profile-membership caveat, but much of the backend defaulting information repeats the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and identifies the resource: live backend operation status, affected exposed tools, and tool_profile. It also enumerates the status categories, which clarifies scope. However, it does not differentiate from sibling tools such as glass_doctor, so it falls short of full sibling distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used to check backend capability/status and notes that no session is required and the default backend is used. It does not explicitly say when to prefer this over sibling diagnostic tools or provide when-not conditions, leaving usage selection partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_clickA

Click a window-relative point, optionally with a button, click count and held modifiers. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesClick x in window-relative pixels.
yYesClick y in window-relative pixels.
countNoClick count (default 1, range 1 through 10); 2 double-clicks.
buttonNo"left" (default), "right", or "middle".
modifiersNoHeld modifiers, e.g. ["ctrl", "shift"].

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey readOnlyHint=false and destructiveHint=false. The description adds the useful window-relative coordinate context, but it does not disclose further behavioral details beyond what a click operation implies, and the optional button/count/modifiers mostly restate schema properties.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler: the primary action is front-loaded, optional parameters are summarized in one clause, and the alternative tool is routed in the second sentence. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple click tool, the description plus full parameter schema covers the coordinate space, optional arguments, and multi-step routing via glass_do. No output schema exists, but a return-value explanation is not essential for a click operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter already described in the input schema. The description's mention of button, click count, and modifiers adds no new meaning beyond those schema descriptions, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Click') and a clear resource ('a window-relative point'), and it lists the optional inputs. It also distinguishes itself from glass_do, which is the main sibling an agent might confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states a routing rule: 'For 2+ known steps, use glass_do.' This tells the agent when to choose an alternative, leaving the single-click case clearly as the context for this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_click_elementA

Click exactly one id or target. A target resolves fresh and uniquely within timeout_ms (10 seconds by default). Mode auto prefers native then pointer fallback; native requires native action; pointer waits for stability and reports actionability, including unproven checks. ID targets remain immediate. Native actions may bypass occlusion or focus text editors. method/native_fallback/actuated_id disclose the path; popover actions restore the prior window. Glass never retries after possible dispatch. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoLatest snapshot ID, exclusive with target; re-read after UI changes. Native action may focus text editors; fallback clicks the center. Popover actions restore the prior window.
modeNo
returnNonone (default): no observation; settle: text-only visual stability; snapshot: settle and refresh/fold the accessibility tree.
targetNo
max_nodesNo
timeout_msNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations only set readOnlyHint=false, destructiveHint=false, leaving the description to carry the full behavioral burden. The description does this well: it discloses that native actions may bypass occlusion or focus text editors, that pointer waits for stability and reports actionability, that method/native_fallback/actuated_id disclose the path, that popover actions restore the prior window, and that Glass never retries after possible dispatch. These are rich behavioral traits beyond what annotations offer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then provides a dense set of behavioral notes. Every sentence contributes new information without redundancy, though the density may require careful parsing. It is appropriately sized for a tool with complex behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, no output schema, and minimal annotations, the description covers many critical aspects: target resolution, timeout, modes, retry behavior, and fallback mechanics. It does not detail max_nodes or return behavior, but those are partially covered by schema descriptions. The description is complete enough for an agent to call the tool correctly in most scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, so the description must compensate. It explains timeout_ms (10 seconds default) and mode (auto/native/pointer semantics), which adds value. However, it does not explain max_nodes or the exact use of target/id beyond the exclusivity mention, leaving some parameters under-specified. The description adds meaningful but incomplete parameter context, warranting a score above baseline but not high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Click exactly one id or target,' which clearly states the verb (click), the resource (one element identified by id or target), and the constraint of exactly one. It also names a sibling alternative, glass_do, for multi-step actions, helping distinguish the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance for when to use this tool vs glass_do ('For 2+ known steps, use glass_do') and explains the mode selection logic (auto/native/pointer). It does not mention the simpler glass_click sibling, but the core usage context is clear enough for an agent to decide when to invoke this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_clipboard_getA
Read-only

Read the app's clipboard as text ("" if empty). Also the cheap text-extraction path: glass_do ctrl+a then ctrl+c, then read here (beats OCR for selectable text). Returns Unsupported where the backend can't provide clipboard access.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint=true, so the description adds value beyond that by disclosing edge cases: "" if empty and "Unsupported where the backend can't provide clipboard access." It also explains the ctrl+a/c Ctrl+c workflow as an implementation detail. No contradiction with readOnlyHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary purpose, then the secondary use case. Every clause contributes information: return value, empty-handling, workflow hint, OCR comparison, and failure mode. No waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only tool with no output schema, the description covers the essential behavioral contract: what it returns (text), empty case, and the Unsupported state. It also ties into the broader tool ecosystem with the glass_do reference, making it complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and schema coverage is trivially 100%. The description cannot add parameter-level meaning, but it does not need to. The baseline for 0-param tools is 4, and the description doesn't introduce any ambiguity about inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ("Read") and resource ("the app's clipboard"), and specifies the return type (text) and empty-string behavior. It also differentiates from siblings like glass_clipboard_set by framing itself as the read counterpart and adding a specific text-extraction use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a concrete usage scenario: "the cheap text-extraction path" using glass_do ctrl+a then ctrl+c, and explicitly compares to OCR ("beats OCR for selectable text"). It lacks explicit exclusion criteria or when-not-to-use guidance, hence not a 5, but the context is clear enough for an agent to select correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_clipboard_setA

Write text to the app's clipboard so it can paste it. Returns Unsupported where the backend can't provide clipboard access.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text to write to the clipboard.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a write operation (readOnlyHint=false, destructiveHint=false). The description adds a specific error condition ('Returns Unsupported where the backend can't provide clipboard access'), which goes beyond the annotations. This is useful context about a failure mode that an agent would not otherwise know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The primary action is in the first sentence, and the error condition in the second. Every word contributes to the tool's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter write operation with no output schema, the description covers the core behavior and the only notable edge case (Unsupported backend). It does not mention return values on success, but that is implied for a write operation and not essential. The absence of output schema means the description isn't required to explain return format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'text' parameter, with a clear description ('The text to write to the clipboard.'). The tool description does not add any additional meaning or nuance beyond the schema's own documentation, so it meets the baseline but adds no extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Write') and resource ('text to the app's clipboard'), with an explicit purpose ('so it can paste it'). It is easily distinguishable from its sibling 'glass_clipboard_get' since it names the opposite action (write vs. get).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied by the tool's name and description (write text to clipboard), but no explicit guidance is given about when to use this versus the sibling 'glass_clipboard_get' or any exclusions. The description does not mention alternatives, so it relies on the agent to infer from the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_diffA
Read-only

Compare current pixels with a named baseline; returns change stats and bbox. A single comparison does not establish transition completion or quiescence. Use glass_wait_for_region to wait; include_image:true returns the changed crop only when pixels differ.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo`"perceptual"` (default) or `"exact"`.
nameYesSaved baseline name; unsaved names error.
ignoreNoWindow-relative excluded rects intersected with region. changed_pct uses remaining pixels; ignored_pixels reports the excluded count.
regionNoWindow-relative comparison area; omit for whole window. Returned bbox is region-relative.
max_widthNoMaximum returned image width; shrinks after crop, preserving native comparison pixels.
thresholdNoPerceptual sensitivity for `mode="perceptual"`, 0..1 (default 0.1; smaller = stricter).
toleranceNoPer-channel tolerance for `mode="exact"` (default 0).
max_heightNoMaximum returned image height; omit both limits for native output.
include_imageNoAlso return the current frame cropped to the changed region (default false). No image is returned when nothing changed.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, and the description adds behavioral context beyond that: it notes that include_image:true returns the changed crop only when pixels differ, and that a single comparison is insufficient for quiescence. It doesn't detail return stats format, but the read-only nature is covered by annotations and the description adds meaningful caveats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no waste. The core function is front-loaded, the critical caveat about quiescence follows immediately, and the include_image behavior is stated precisely. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only comparison tool with 100% schema coverage and no output schema, the description covers the essential behavioral caveats: single comparison is not quiescence, use wait_for_region for waiting, and image return behavior. It could mention what 'change stats' includes, but the schema and read-only annotation carry much of the burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 9 parameters. The description adds value by explaining the include_image behavior (crop only when pixels differ) and the mode semantics, but most parameter meaning is already in the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Compare current pixels with a named baseline') and resource ('named baseline'), and distinguishes itself from siblings by naming glass_wait_for_region as the tool to use for waiting. It clearly identifies the tool's core function: returning change stats and bbox.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says a single comparison does not establish transition completion or quiescence, and directs the agent to use glass_wait_for_region to wait. This provides clear when-to-use and when-not-to-use guidance, differentiating it from the waiting tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_doA

Batch 2+ known actions or waits here; observe before choosing dependent steps. Run one action or fixed static ordered actions: click, move, drag, scroll, type, key, settle, click_element, set_value, wait_for_element, scroll_to_element. Maximum 64 actions and 65536 compact argument bytes. timeout_ms defaults to 30000; range 1 through 120000; one deadline includes terminal observations. Fail-fast on action errors, deadline or unmatched batched predicates; standalone predicates remain soft. Success without then returns ordered steps with result and content_blocks; ok:true means all actions completed. Failures and then retain completed, failed, unexecuted and terminal_steps; inspect before recovery. Preflight invalid_sequence has no step outcomes. Terminal settle/diff/screenshot runs after actions; terminal failure retains completed action outcomes. No branching, bindings, loops, retries or generated steps. Semantic targets resolve fresh; targeted type requires confirmed focus; set_value requires confirmation. Never replay completed or possibly dispatched mutations; observe after uncertainty.

ParametersJSON Schema
NameRequiredDescriptionDefault
thenNoTerminal observation in order: settle, diff, screenshot.
actionsYesNon-empty ordered actions; failure stops remaining steps. Completed mutations may already have landed.
timeout_msNoOverall budget in ms (default 30000, range 1..120000), shared by all actions and terminal observations.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the annotations, which only declare readOnlyHint=false, destructiveHint=false, openWorldHint=false. It discloses fail-fast behavior, the shared deadline semantics, default timeout and range, what 'ok:true means', the structure of success vs. failure returns, terminal settle/diff/screenshot behavior, preflight invalid_sequence behavior, semantic target freshness, and a critical warning: 'Never replay completed or possibly dispatched mutations; observe after uncertainty.' This is precisely the kind of behavioral context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and packs a lot of information, but every sentence contributes a distinct fact: limits, defaults, failure modes, terminal observation, and restrictions. It is front-loaded with the core purpose and the accepted action list. It is not padded, though it is a long wall of text for a tool definition; a slightly more structured layout (e.g., headings) could improve scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex batching tool with no output schema, the description covers all the essentials an agent needs to call it correctly: the action vocabulary, size limits, timeout behavior, success/failure return structures, terminal observation semantics, and critical caveats about mutations and semantic targets. It even explains what happens on preflight invalid_sequence. Nothing critical is missing given the schema already documents individual action arguments.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema covers 100% of parameters with descriptions, the tool description adds operational meaning not present in the schema: 'Maximum 64 actions and 65536 compact argument bytes,' the default timeout (30000ms) and range, the single shared deadline for all actions and observations, and the return contract ('ordered steps with result and content_blocks'). It also clarifies the semantics of `then` ordering (settle → diff → screenshot) and how failures retain `completed, failed, unexecuted` steps. This is substantial additional meaning, not mere repetition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: 'Batch 2+ known actions or waits here; observe before choosing dependent steps.' It distinguishes itself from the 30 sibling tools by positioning glass_do as the batching/sequencing tool, and enumerates the exact action types it accepts (click, move, drag, scroll, type, key, settle, etc.), leaving no ambiguity about what it does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use it: when you have multiple known actions or waits that should run as an ordered sequence, with observation between batches. It also states what it cannot do ('No branching, bindings, loops, retries or generated steps'), which helps an agent avoid misusing it. It doesn't explicitly say 'use this instead of calling standalone sibling tools individually,' but the context of batching makes that evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_doctorA

Diagnose setup or launch failures. Returns report text, structured sections/checks with optional remedy/remedy_action, and overall (the verdict to branch on). Check status is ok, warn, fail or skip; non-default backend failures only warn in overall. deep also starts and tears down the default display.

ParametersJSON Schema
NameRequiredDescriptionDefault
deepNoStart and tear down the default display to prove it starts (default false).

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavior beyond the annotations: returns structured report sections and an overall verdict, reports statuses such as ok/warn/fail/skip, and warns that non-default backend failures only affect the overall verdict as warnings. It also explicitly notes that deep starts and tears down the default display, which is a useful side-effect disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and information-dense: three sentences cover purpose, return shape, statuses, backend-failure nuance, and the deep flag's side effect. It is appropriately sized for the tool's behavior, though the output-structure explanation could be slightly more organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one parameter and no output schema, the description covers the important return fields, overall verdict, status values, and deep behavior. It also gives enough context for branching on 'overall.' It does not explain the exact values of 'overall' or when 'deep' should be used, but these are minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes the single 'deep' parameter with 100% coverage. The description adds only the phrase 'deep also starts and tears down the default display,' which largely mirrors the schema's own description, so the parameter semantics do not get much additional value from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Diagnose setup or launch failures' and explains the return format. It is clear, but it does not explicitly distinguish this tool from diagnostic-looking siblings such as glass_logs or glass_capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: when diagnosing setup or launch failures. It does not mention alternatives or exclusion criteria, but it does explain the optional 'deep' mode and its effect, giving the agent enough context to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_dragA

Drag one pointer from (x1,y1) to (x2,y2), holding the button and modifiers throughout. Motion spans duration_ms. Either endpoint outside the window is refused. Use glass_gesture for multi-touch. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
x1YesPress x in window-relative pixels.
x2YesRelease x in window-relative pixels.
y1YesPress y in window-relative pixels.
y2YesRelease y in window-relative pixels.
buttonNoButton held for the drag: "left" (default), "right", or "middle".
modifiersNoHeld modifiers, e.g. ["ctrl", "shift"].
duration_msNoMotion duration in ms (default 200). Faster motion gives the app fewer sampled frames.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the description must carry the behavioral burden. It does well by disclosing the holding of the button/modifiers, the duration control, and the endpoint refusal condition. It also notes in the parameter description that faster motion gives the app fewer sampled frames, which is valuable. It does not disclose side effects (e.g., whether this triggers navigation or state changes) but for a drag tool the key behaviors are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct - three sentences that cover the core action, a key constraint, and routing to alternatives. It is front-loaded with the primary action and immediately provides differentiation. Every sentence earns its place, with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and the complexity is moderate, the description covers the core action, constraints, and sibling routing. It does not specify the return value or error behavior, but that is often implicit. The guidance on duration effects and endpoint refusal is sufficient for an agent to call it correctly. Missing a brief note on the default button and modifier behavior is negligible because the schema covers those defaults.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all parameters clearly. The description adds minimal extra semantic value: it explains the holding behavior and the effect of duration on sampling, but the parameter details are already comprehensive. Thus a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (drag one pointer from (x1,y1) to (x2,y2)) and the resource (a pointer in a window), with specific details about holding the button and modifiers. It explicitly differentiates from siblings by naming glass_gesture for multi-touch and glass_do for 2+ known steps, which are the primary alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: for multi-touch use glass_gesture, for 2+ known steps use glass_do. It also mentions the constraint that either endpoint outside the window is refused, which sets expectations for when the tool cannot be used. This is thorough and specific.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_find_elementsA
Read-only

Find ranked accessibility candidates from one fresh read when the target is approximate, duplicated, or not yet identified. Query matches name, description and non-secure value; within must match one semantic scope. max_results defaults to 10, capped at 20; timeout_ms optionally waits. Returns actionable IDs, compact context and explicit truncation in an untrusted match array. Total text is capped at 8 KiB.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNoNormalized target role.
queryNoApproximate case-insensitive semantic text. Optional when role or states are supplied.
statesNoTarget state predicates combined with AND.
withinNoOptional unique semantic scope resolved in the same fresh tree.
max_nodesNoExisting accessibility walk limit semantics; 0 removes the node-count limit.
timeout_msNoOptional wait for at least one match; default 0 performs one fresh read.
max_resultsNoMaximum ranked matches before the byte budget; default 10, range 1 through 20.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/openWorld annotations, the description discloses matching fields, default and cap for max_results, timeout behavior, the return shape ('actionable IDs, compact context and explicit truncation'), the 'untrusted' nature of matches, and the 8 KiB total-text cap. This is rich behavioral context that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three dense sentences with no filler: purpose and trigger come first, then matching semantics and parameter limits, then return behavior. Every sentence earns its place, and key operational constraints are packed efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description does a good job of describing what the caller gets back: ranked IDs, compact context, explicit truncation, and a bounded result set. It could be clearer about what 'ranked' means and the exact structure of the returned matches, but it is sufficient for selecting and invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the input schema already documents query matching, within's semantic scope, max_results range, and timeout semantics. The description largely restates these facts ('max_results defaults to 10, capped at 20') and adds only minor gloss such as 'within must match one semantic scope.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('accessibility candidates') with a concrete verb ('Find') and ties it to a clear triggering condition: 'when the target is approximate, duplicated, or not yet identified.' It does not explicitly name or contrast sibling tools such as glass_wait_for_element or glass_a11y_snapshot, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear use context: this is the one-shot ranked search for ambiguous or unknown targets, and it clarifies the single-read timing model. It does not state exclusions or explicitly point to alternatives, but the 'one fresh read' phrase distinguishes it from wait/snapshot-style tools without naming them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_gestureA

Perform a multi-touch gesture: 2–10 pointers, each a straight from→to segment in window-relative px, all down together at t=0 and up at duration_ms. Pinch = two pointers toward/apart; rotate = two on an arc; two-finger swipe = two parallel segments; a from==to pointer is held. Multi-touch isn't available on every backend — it returns a clear Unsupported error where the active backend can't do it.

ParametersJSON Schema
NameRequiredDescriptionDefault
pointersYes2–10 simultaneous pointers; each a straight from→to segment. Pinch = two pointers moving toward/apart; rotate = two on an arc; two-finger swipe = two parallel segments.
duration_msNoSpan the gesture over this many ms (all pointers down at 0, up at duration). Default 250.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=false), the description reveals important behavioral traits: all pointers go down simultaneously at t=0 and up at duration_ms, a from==to pointer is held, and multi-touch may return an Unsupported error depending on backend. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense with no wasted words. It front-loads the core purpose ('multi-touch gesture'), packs constraints into the first sentence, and uses the second sentence for examples and a backend caveat. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and only two parameters, the description covers what an agent needs: the gesture model, pointer count range, timing behavior, held-pointer semantics, and the backend limitation. Together with the fully described schema, this is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the pointers array and duration_ms. The description adds value by describing gesture patterns (pinch, rotate, swipe) and clarifying that all pointers are synchronized, but most parameter meaning is already present in the schema, so the additional contribution is moderate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Perform') and resource ('multi-touch gesture') with precise scope: 2–10 pointers, each a straight from→to segment in window-relative px, all down together at t=0 and up at duration_ms. This clearly differentiates it from single-pointer siblings like glass_drag, glass_click, and glass_scroll.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete use cases: pinch, rotate, two-finger swipe, and held pointers. It clearly implies this tool is for multi-touch gestures rather than simple single-pointer interactions, though it does not explicitly name when-not-to-use it or point to specific sibling tools for single-pointer alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_keyA

Press a key chord (for example ctrl+s or Return), releasing modifiers afterwards. Unknown keys/modifiers fail before input. Use type for literal text; it cannot express shortcuts. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
chordYesModifiers plus one key: ctrl+s, Return, alt+F4. Modifiers: ctrl/shift/alt/super (cmd/win/meta aliases). Key: named key, F1-F12 or printable ASCII; case-insensitive names.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds behavioral detail beyond annotations: 'releasing modifiers afterwards' and 'Unknown keys/modifiers fail before input'. These are not present in the minimal annotations and help the agent predict side effects and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler. The main action is front-loaded, followed by key behavioral notes and sibling routing. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter tool, the description adequately covers purpose, usage, and behavior. It does not mention return values, but the absence of an output schema and the tool's nature make this less critical. Minor gap, not a major omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, fully describing the 'chord' parameter with modifier list and examples. The tool description adds no new semantic information about the parameter beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Press a key chord' with examples. It distinguishes itself from glass_type ('Use type for literal text') and glass_do ('For 2+ known steps'), clearly differentiating among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use alternatives: use glass_type for literal text, use glass_do for sequences. It also clarifies that this tool handles shortcuts, which is a clear usage condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_list_windowsA
Read-only

List app windows: id, title, class, geometry and active state. Re-list after windows open/close; do not cache IDs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotation Contradiction: the description instructs agents to re-list after windows open/close, implying the output changes with external world events, but the annotation openWorldHint=false explicitly declares the tool's output is not affected by world changes. This is a direct contradiction with the description's stated behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence names the action and output fields; the second gives a critical caching caveat. Information density is high and the key behavioral note is placed at the end without burying the purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with no output schema, the description adequately enumerates the returned fields and warns about ID stability. However, the contradiction with openWorldHint leaves the expectations about world dynamics ambiguous, preventing a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so the schema already fully describes the input. The description need not add parameter detail; the baseline of 4 applies because there is nothing missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'List app windows' – a specific verb and resource – and enumerates the returned fields (id, title, class, geometry, active state). This clearly differentiates it from siblings like glass_select_window (which selects) and glass_a11y_snapshot (which captures accessibility data) without needing to open schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides actionable guidance: 'Re-list after windows open/close; do not cache IDs.' This tells the agent when to invoke the tool again and warns against reusing stale window IDs. It does not explicitly name alternatives or exclusion cases, but the context for correct usage is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_logsA
Read-only

Read buffered stdout/stderr now to verify completed actions. Filter stream/contains; cursor resumes reading. Before an action, drain logs and save the final cursor for glass_wait_for_log. Old lines age out. App text is untrusted.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoResume from a prior returned cursor; omit for oldest buffered line.
streamNo"stdout", "stderr", or "both" (default).
containsNoCase-sensitive substring filter applied before the line cap.
max_linesNoLine cap (default 200). Returned cursor resumes at the first unread line.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation readOnlyHint=true already covers the non-destructive nature, but the description adds valuable operational details: buffered lines age out, cursor-based resumption, filtering, and that app text is untrusted. These behaviors are not available from the annotations alone and materially affect how an agent should use the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences cover purpose, filtering, cursor semantics, workflow integration, retention behavior, and a trust warning. The most important information is front-loaded and every sentence contributes new value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with four optional parameters and no output schema, the description provides sufficient context: how to start reading, resume with a cursor, filter, and prepare for waiting. It does not fully spell out the return payload shape, but the cursor-resumption language and pointer to glass_wait_for_log make the main workflow clear without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the structured definitions already explain cursor, stream, contains, and max_lines. The description reinforces the cursor/filter concepts and adds workflow context around cursors, but it does not materially enrich parameter meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Read'), resource ('buffered stdout/stderr'), and immediate purpose ('verify completed actions'), making the tool's role unmistakable. The phrase 'now' and the reference to glass_wait_for_log help distinguish it from the waiting-oriented sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete workflow guidance: drain logs before an action, save the final cursor, and use that cursor with glass_wait_for_log. It clearly frames when to read logs ('now to verify completed actions'), though it does not explicitly enumerate when not to use this tool or name direct alternatives beyond glass_wait_for_log.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_moveA

Move the pointer to a window-relative point. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesDestination x in window-relative pixels.
yYesDestination y in window-relative pixels.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states the core behavioral trait: moving the pointer without clicking or dragging. It also clarifies the window-relative coordinate semantics, which is useful context even though the schema already mentions it. With readOnlyHint=false and destructiveHint=false, the mutation semantics are consistent and adequately disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The core action comes first and the alternative-tool routing is appended as a single conditional clause. Every word contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter pointer-move tool, the description plus full schema coverage is sufficient. No output schema is needed for a side-effect operation, and the sibling guidance covers the main ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with x and y each described as 'Destination ... in window-relative pixels'. The description adds no new parameter-level detail beyond the schema, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Move') and resource ('the pointer') plus a precise coordinate space ('window-relative point'). It clearly distinguishes this from sibling tools like glass_click or glass_drag by framing it as a pure pointer move.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly names the alternative glass_do and gives the condition for choosing it ('For 2+ known steps'). This gives the agent actionable routing guidance beyond what the schema or tool name provides.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_screenshotA
Read-only

Capture current visual evidence as lossless WebP, not semantic state or transition completion. Off-display captures are clipped: returned dimensions disclose the actual size; fully off-screen surfaces error. Optional window_id observes without switching the active window.

ParametersJSON Schema
NameRequiredDescriptionDefault
regionNoOptional window-relative sub-rectangle to capture; omit for the whole window.
max_widthNoMaximum returned image width; shrinks after crop, preserving native comparison pixels.
window_idNoObserve this current glass_list_windows ID without selecting it; omit for active window.
max_heightNoMaximum returned image height; omit both limits for native output.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description adds significant behavioral context beyond that: output format (lossless WebP), clipping behavior for off-display captures, error conditions for fully off-screen surfaces, and the non-intrusive nature of window_id. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero waste. The core purpose is front-loaded, followed by two high-value behavioral notes. Every sentence earns its place, and the structure is highly scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a screenshot tool with no output schema. It covers output format, clipping and error behavior, and window selection semantics. An agent has everything needed to invoke it correctly and interpret results. No gaps that would cause mis-calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter (region, max_width, window_id, max_height) is already documented. The description adds minimal extra parameter-specific meaning—the window_id observation behavior is already in the schema. It does clarify overall output behavior (dimensions disclosure) but that is not parameter semantics. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Capture current visual evidence as lossless WebP') and clearly distinguishes it from 'semantic state or transition completion', which sets it apart from sibling tools like glass_a11y_snapshot or glass_wait_for_region. This is a precise verb+resource statement that leaves no ambiguity about the tool's function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it ('when you need visual evidence') and explicitly contrasts it with semantic state checks. It also provides guidance on window_id usage ('observes without switching the active window'). However, it does not name specific alternative tools or explicit exclusion conditions, so it falls short of a perfect 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_scrollA

Scroll the container under (x,y) by horizontal/vertical wheel notches, optionally holding modifiers. Notches are not pixels. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesWindow-relative anchor x; selects the container under the pointer.
yYesWindow-relative anchor y.
dxNoHorizontal wheel notches, -100 through 100 (positive right, negative left); typically 1-5, not pixels.
dyNoVertical wheel notches, -100 through 100 (positive down, negative up); typically 1-5. App determines distance/zoom.
modifiersNoHeld modifiers, e.g. ["ctrl", "shift"].

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, indicating a mutating but non-destructive operation. The description adds that scrolling uses wheel notches rather than pixels, which clarifies the effect. It does not describe potential side effects or error conditions, but given the annotations cover the safety profile, the added context is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first states the purpose and key parameters, the second clarifies units and provides the alternative. No wasted words, and the critical scoping information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-scroll action with no output schema, the description covers the purpose, the alternative for multi-step sequences, and the unit semantics. Nothing essential is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all parameters including dx/dy as notches. The description reinforces the 'not pixels' distinction, which adds a critical clarification beyond the schema's 'wheel notches' wording. This adds value without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Scroll the container under (x,y)') with a specific verb and resource, and distinguishes itself from the sibling glass_do by explicitly noting when to use that alternative. It also clarifies the unit of measurement (notches vs pixels), which is essential for correct usage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit guidance on when to use an alternative: 'For 2+ known steps, use glass_do.' This directly informs the agent of the boundary between this tool and a related sibling. It also clarifies the notches unit, which is a usage constraint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_scroll_to_elementA

Scroll until a semantic element is actually on-screen; returns matched, elapsed_ms, element and scrolled details. Sweeps the chosen/inferred direction, then reverses. No match after both ends or timeout returns matched:false; no tree errors. Batched use fails the sequence on an unmatched predicate. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoWindow-relative anchor x; supply both x/y for a container. Default target row/column, or window center if absent.
yNoScroll anchor y (window-relative). See `x`.
nameNoSubstring of the target element's accessible name (selector).
roleNoElement role filter, e.g. "ListItem", "Button", "Document" (selector).
stepNoWheel notches per step (default 3); large steps can skip realized rows.
directionNoup/down/left/right. Default infers off-screen direction, or down then up if absent. Reverses at the first end.
timeout_msNoTimeout in ms (default 20000). Standalone: matched:false; batched: fails the sequence.
descriptionNoAccessible-description substring; can select unnamed controls.
value_containsNoValue substring; requires name, description and/or role.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses sweep-then-reverse behavior, return fields, timeout and end-of-scroll failure modes, the absence of tree errors, and batched sequence failure semantics. The annotations are sparse, so this description carries the behavioral burden effectively.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Roughly sixty words, front-loaded with the core action, followed by return values, failure modes, batched behavior, and a routing note. Every sentence contributes information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers action, return fields, direction handling, timeout behavior, failure semantics, and sibling routing despite having no output schema. Minor gaps remain around what 'on-screen' means exactly and the shape of the returned element, but the context is strong for a complex automation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the description need not re-explain parameters. It adds useful context about direction sweeping and timeout behavior, but no per-parameter semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Scroll until a semantic element is actually on-screen'. It clearly distinguishes the tool from glass_do by naming the alternative, and the unique role is evident from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to use the tool and explicitly routes the user to glass_do for '2+ known steps'. It could more fully contrast with sibling scroll/find tools, but the main alternative is covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_select_windowA
Idempotent

Select a window by current ID. Subsequent actions and observations target it; input coordinates are relative to it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesCurrent ID from glass_list_windows; re-list after window changes.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses a meaningful behavioral trait: selection is stateful and influences all later actions and observations, including coordinate interpretation. This is useful context that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences deliver the core purpose and the key behavioral consequence with no filler. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter selection tool with rich annotations and full schema coverage, the description covers what selection does, how it affects subsequent operations, and how coordinates are interpreted. No critical operational detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully documents the single 'id' parameter, including that it comes from glass_list_windows and should be re-listed after window changes. With 100% schema coverage, the description does not need to add parameter details, and its 'current ID' mention is consistent but not additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Select a window'), a specific resource ('window'), and the key qualifier ('by current ID'). It also explains the consequence of selection, which distinguishes this from listing or inspecting windows among the sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when the tool matters: before subsequent actions and observations, and when input coordinates should be interpreted relative to the selected window. It does not explicitly name alternatives or exclusions, but the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_set_valueA

Set exactly one editable id or target with backend confirmation. A target resolves fresh and uniquely (10 seconds by default); IDs remain immediate. Writes directly or focuses/clears/types as supported, then reads back. Errors distinguish mismatch from unconfirmed writes; post-write uncertainty is terminal. Do not write again on uncertainty: observe where input landed before recovery. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoLatest snapshot ID, exclusive with target.
textYesText, slider/spin number, toggle boolean (true/false/on/off/1/0), or case-insensitive combo option label. Toggle writes are idempotent; combos open and choose the option.
returnNonone (default): no observation; settle: text-only visual stability; snapshot: settle and refresh/fold the accessibility tree.
targetNo
max_nodesNo
timeout_msNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the minimal false hints, it discloses backend confirmation, fresh target resolution, immediate ID behavior, supported write modes, read-back behavior, and terminal post-write uncertainty. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence carries operational weight, from the core contract to recovery guidance and the sibling reference. It is dense but not bloated, with the most important constraints front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It covers the tool's core behavior, error taxonomy, and recovery policy, making the mutation tool much safer to invoke. The main gap is the absence of exact output/error shape, which matters more given there is no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers id, text, and return, but leaves target, max_nodes, and timeout_ms largely undocumented. The description adds value by explaining target resolution freshness/uniqueness and ID immediacy, though it does not clarify max_nodes or timeout_ms.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Set...') with a bounded scope: exactly one editable id or target. It also distinguishes itself from the multi-step sibling by explicitly naming glass_do as the alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'For 2+ known steps, use glass_do,' giving a clear routing rule. It also provides when-not-to guidance by warning not to write again on uncertainty and to observe first before recovery.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_startA

Build, launch and locate an app; returns window geometry. Accessibility is enabled by default. window_hint can select among windows or locate a handoff to another process. See parameters for backend and containment choices.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory for build and app; defaults to the server directory.
envNoExtra {KEY: VALUE} environment for build and app. Android applies it only to the host build, not the app.
runYesDesktop: [executable, args...]. iOS: [.app-or-bundle-id, args...]. Android: [apk?, package/.Activity] in either order, e.g. ["/absolute/path/app.apk", "com.example.app/.MainActivity"].
a11yNoEnable the private accessibility bus (default true). False skips it for canvas-only apps. Linux only; other backends read accessibility ambiently.
buildNoOptional shell command to run (in `cwd`) before launching.
backendNox11/wayland (Linux), windows (Windows), macos (macOS), android (any host), ios (macOS). Default: GLASS_BACKEND, else host default (x11 on Linux).
sandboxNodefault: filesystem/process containment, network on; strict: also no network; off: uncontained. Default GLASS_SANDBOX or default. GLASS_SANDBOX_FLOOR raises omitted levels and refuses explicit lower levels.
timeout_msNoWindow-publication timeout in ms (default 10000); does not bound build.
window_hintNoSelect a window by title/class, including process handoffs. Omit for the first window owned by the process or a followable descendant.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that accessibility is enabled by default and that window_hint can locate handoffs—not obvious from the schema alone. It also notes that backend and containment choices are in the parameters, which adds behavioral context. The annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false) are not contradicted; the tool performs actions but is not destructive. The description adds useful nuance beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences and front-loaded with the core action (build, launch, locate) and return value (window geometry). It avoids unnecessary fluff. The only minor deduction is that the last sentence is somewhat vague ('See parameters for backend and containment choices') rather than giving specific guidance, but it's still efficient overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 9 parameters, a nested object, and no output schema, the description does not detail return format, error cases, or platform specifics beyond its brief mention. However, the schema carries most of the parameter documentation, and the description covers the high-level workflow. It could be more complete by explaining what 'returns window geometry' implies in practice (e.g., coordinates, size), but given the schema's richness, this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds meaning by summarizing how window_hint works (select among windows or locate a handoff) and pointing to backend and containment choices. It does not enumerate each parameter, but it gives a high-level semantic framework that helps agents understand the tool's logic without reading every schema detail. Given the schema is already thorough, this adds value without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Build, launch and locate an app; returns window geometry.' It uses specific verbs and a clear resource (an app), and mentions returning window geometry. It also distinguishes itself from siblings like glass_select_window and glass_window by focusing on the full build/launch/locate workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context on when to use this tool (to build, launch, and locate an app) and hints at alternatives via window_hint for selecting windows or handoffs. However, it does not explicitly state when not to use it or name specific sibling tools as alternatives, such as glass_select_window for pure window selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_stopA
DestructiveIdempotent

End the session: request app close, then terminate if necessary. Read logs first; logs and element IDs are discarded. Baselines survive until server exit. No resume; glass_start creates a fresh session. Errors without an active session.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and idempotentHint=true, but the description adds valuable behavioral context: it requests app close first, then terminates if necessary, and clarifies that logs and element IDs are discarded while baselines survive. It also notes that errors occur without an active session. This goes beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action ('End the session'), followed by essential behavioral details. Every sentence adds value, and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description covers the key aspects: what it does, side effects, prerequisites, and error conditions. It could potentially mention what the return value looks like, but the absence of an output schema makes that less critical. The description is complete enough for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description doesn't need to explain parameter semantics. The baseline for 0 params is 4, and the description appropriately focuses on behavior rather than parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: ending the session by requesting app close and terminating if necessary. It distinguishes itself from siblings like glass_start by explicitly noting that glass_start creates a fresh session and there is no resume.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool: at the end of a session. It also gives important context about prerequisites (read logs first) and what happens to data (logs and element IDs are discarded, baselines survive until server exit). This helps the agent decide when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_typeA

Type text once. Without target, focused-window typing sends keystrokes to current focus. A target resolves one fresh unique element; focus_mode auto, native or pointer confirms focus then types once. Unconfirmed focus never types or tries another path. Newlines do not press Return: use a key action. No value confirmation; verify the resulting field. Never replay after uncertain dispatch. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesSynthetic keystrokes, not paste. When `target` is omitted: current keyboard focus. When `target` is supplied: resolves and focuses the target, confirms focus, then types. Newlines do not press Return.
returnNonone (default): no observation; settle: text-only visual stability; snapshot: settle and read the current accessibility tree, including app-visible text, with secure values redacted.
targetNo
max_nodesNo
focus_modeNo
timeout_msNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations providing only readOnly/destructive hints, the description discloses key behavioral traits: single dispatch only, unconfirmed focus never types, fresh-element target resolution, no value confirmation, and no replay after uncertain dispatch. This gives the agent a clear safety model for a write-like operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action ('Type text once') and then delivers a dense set of caveats in compact, purposeful sentences. Every sentence either clarifies behavior, prevents a misuse, or routes to an alternative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write-like tool with no output schema, the description covers the most important behavioral constraints and alternatives. It does not explain the return/snapshot semantics or the undocumented timeout/max_nodes parameters, though the schema partially covers return and the parameter names are suggestive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 33%, so the description must compensate. It adds meaningful semantics for text, target, and focus_mode, including the 'one fresh unique element' behavior and the three focus modes. However, max_nodes and timeout_ms receive no added explanation in either schema or description, leaving a partial gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

'Type text once' is a specific verb+resource statement that immediately establishes the tool's core behavior. It also distinguishes itself from siblings by explicitly routing multi-step work to glass_do and newline handling to a key action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use context: type once into current focus or a resolved target, and never replay after uncertain dispatch. It also names an alternative ('use glass_do' for 2+ steps) and a fallback for newlines ('use a key action'), which is strong usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_wait_for_elementA
Read-only

Wait for semantic transition completion, not pixels/stability. Select by name, description and/or role; value is exact, value_contains is a substring. Supports appears/disappears and state conditions. Returns matched, elapsed_ms and the element; timeout is matched:false. Waits for initial tree publication, but errors if no tree appeared. Batched use fails the sequence on an unmatched predicate. For 2+ known steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSubstring of the element's accessible name (selector).
roleNoElement role filter, e.g. "Button", "ProgressBar", "Document" (selector).
valueNoExact case-sensitive accessible value, requiring another selector and excluding `value_contains`.
conditionNoDefault appears. appears|disappears|enabled|disabled|checked|unchecked|selected|unselected|expanded|collapsed|focused|visible|hidden. checked/unchecked require a real checkable state.
timeout_msNoTimeout in ms (default 10000). Standalone: matched:false; batched: fails the sequence.
descriptionNoAccessible-description substring; can select unnamed controls.
interval_msNoPoll interval (default 200ms — an a11y snapshot per tick).
value_containsNoValue substring; requires name, description and/or role; exclusive with value.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, it discloses timeout behavior (matched:false), batched-sequence failure, error when no accessibility tree appears, and the exact return fields (matched, elapsed_ms, and the element). These are meaningful behavioral details not present in the annotations, and there is no contradiction with readOnlyHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, and each sentence covers a distinct concern: scope, selectors, matching semantics, conditions, return values, edge cases, and the sibling alternative. There is no filler, and the brief redundancy with schema timeout descriptions is acceptable given the absence of an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, it explicitly states what is returned and how timeout is represented. Together with the fully documented input schema and read-only annotations, an agent has enough information to invoke the tool correctly, understand failure modes, and distinguish it from related tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all 8 parameters at 100% coverage, so the baseline is 3. The description adds useful selection semantics by stating that name, description, and/or role can be used as selectors and by clarifying exact-value versus substring matching for value and value_contains, which goes slightly beyond the schema's individual property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb and resource: wait for an element's semantic transition, explicitly contrasted with pixel/stability waiting. It enumerates selection fields and supported conditions, and its scope is clearly distinct from sibling tools like glass_find_elements, glass_wait_stable, and glass_do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit exclusion ('not pixels/stability') and a direct alternative for multi-step flows ('For 2+ known steps, use glass_do'). However, it does not explicitly compare against related sibling wait tools such as glass_wait_stable, glass_wait_for_log, or glass_wait_for_region, leaving some routing implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_wait_for_logA
Read-only

After an action, use cursor:0 for buffered logs; for repeated events drain glass_logs before acting and pass its final cursor. Omit cursor for future lines only. Match a substring; timeout: matched:false. Resume from returned cursor.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoCursor after draining logs before the action; 0 includes retained history. Omit for future lines only.
streamNo"stdout", "stderr", or "both" (default).
containsYesSubstring to wait for (required, non-empty).
timeout_msNoGive up after this long (default 10000ms); returns `{matched:false}`.
interval_msNoPoll interval (default 100ms).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description surfaces important behavioral details beyond the readOnlyHint: cursor semantics, the need to drain glass_logs in repeated-event scenarios, substring matching behavior, timeout results (`matched:false`), and resuming from a returned cursor. These are non-obvious traits that materially affect invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences deliver dense operational guidance with no filler. The most important usage decision (cursor:0 for buffered logs) is front-loaded, and each sentence covers a distinct necessary behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the key waiting, cursor, timeout, and resume behaviors, and the schema fills parameter details. With no output schema, it could more explicitly define the success return shape, but the mention of `matched:false` and 'returned cursor' gives enough guidance for correct use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all five parameters, so the description is not the primary source of parameter meaning. However, the description adds workflow context around cursor: 'for repeated events drain glass_logs before acting and pass its final cursor' and 'Resume from returned cursor,' extending the schema's definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description implies the tool waits for a log line matching a substring, and it clearly associates cursor behavior with log positions. It references glass_logs, which helps distinguish it from the related log-reading sibling, though it never states 'wait for a log' as an explicit verb+resource declaration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete, scenario-based usage instructions: use cursor:0 for buffered logs, drain glass_logs before repeated actions and pass its final cursor, and omit cursor to wait only for future lines. This is explicit operational guidance that lets an agent decide exactly how to invoke the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_wait_for_regionA
Read-only

Wait for pixel transition completion: changes from the initial frame or matches a saved baseline (required for matches). Returns matched, changed_pct, bbox and elapsed_ms as text; include_image opts into an image on match. Timeout returns matched:false. Does not prove semantics or subsequent quiescence; use a settle action for the latter.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo"perceptual" (default) or "exact".
untilNo"changes" (default; diverge from reference) or "matches" (converge to baseline).
ignoreNoWindow-relative excluded rects intersected with region. Off-area rects clamp/drop silently; check ignored_pixels. changed_pct uses remaining pixels.
regionNoWindow-relative sub-rectangle to watch; omit for the whole window.
baselineNoSaved baseline name to compare against; omit to use the frame at call start.
max_widthNoMaximum returned image width; shrinks after crop, preserving native comparison pixels.
thresholdNoPerceptual sensitivity (default 0.1; smaller = stricter).
toleranceNoExact per-channel tolerance (default 0).
window_idNoObserve this current glass_list_windows ID without selecting it; omit for active window.
max_heightNoMaximum returned image height; omit both limits for native output.
timeout_msNoGive up after this long (default 10000ms); returns `{matched:false}`.
interval_msNoPoll interval (default 100ms).
include_imageNoOn match, also return the watched region as an image (default false).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the agent knows this is a safe read operation. The description adds valuable behavioral context beyond that: it explains the return values (matched, changed_pct, bbox, elapsed_ms), the timeout behavior (returns matched:false), the include_image opt-in, and the limitation that it does not prove semantics or quiescence. This is rich, honest behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core purpose is in the first sentence, followed by return values, opt-in behavior, timeout behavior, and a clear limitation with an alternative. Every sentence earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter tool with no output schema, the description covers the key behavioral aspects: what it returns, timeout semantics, the baseline requirement, and the limitation. It doesn't enumerate every parameter, but the schema already does that. The only minor gap is that it doesn't explain the difference between 'perceptual' and 'exact' modes, but the schema covers that. Overall, complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 13 parameters. The description adds some semantic context (e.g., 'changes from the initial frame or matches a saved baseline', 'include_image opts into an image on match', 'Timeout returns matched:false'), but it does not systematically explain each parameter's meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('wait for'), a precise resource ('pixel transition completion'), and the two modes ('changes' vs 'matches'). It also names the sibling it is not ('use a settle action for the latter'), which distinguishes it from glass_wait_stable. This is a clear, specific definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: wait for pixel transition completion, and when not: 'Does not prove semantics or subsequent quiescence; use a settle action for the latter.' It also explains the 'matches' mode requires a saved baseline. This gives the agent clear routing guidance relative to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_wait_stableA
Read-only

Wait for visual quiescence and return the last frame, not that an expected semantic state or pixel design was reached. Use include_image:false for text-only metadata. A timeout returns settled:false. window_id observes without selecting that window. For 2+ known active-window steps, use glass_do.

ParametersJSON Schema
NameRequiredDescriptionDefault
ignoreNoWindow-relative rects excluded from comparison and saw_motion, intersected with stability_region. Off-area rects clamp/drop silently; check ignored_pixels for misplaced masks.
regionNoOptional window-relative sub-rectangle for the returned frame.
max_widthNoMaximum returned image width; shrinks after crop, preserving native comparison pixels.
toleranceNoPer-channel difference allowed (0-255, default 0).
window_idNoObserve this current glass_list_windows ID without selecting it; omit for active window.
max_heightNoMaximum returned image height; omit both limits for native output.
timeout_msNoGive up after this long (default 5000ms); returns `{settled:false}` rather than erroring.
interval_msNoHow long to wait between capture ticks (default 100ms).
include_imageNoReturn image (default true). False returns settled/saw_motion/observed_ms/ignored_pixels/dimensions as text; region then has no effect.
settle_framesNoConsecutive unchanged frames required (default 3).
stability_regionNoWindow-relative area watched for settling, independent of returned-image region.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true and openWorldHint=false already in annotations, the description adds important behavioral nuance: it returns the last frame rather than confirming a semantic state, a timeout returns settled:false, and window_id observes without selecting. These go beyond what annotations and schema already state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five short sentences, front-loaded with core purpose, and each sentence contributes either a key distinction, a parameter usage tip, or alternative routing. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 11-parameter tool with no output schema, the description covers core purpose, timeout behavior, window_id semantics, and routing to glass_do. It does not enumerate all return metadata fields, but those are documented in the schema's include_image parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and each parameter already has a thorough description. The description re-states some parameter behaviors (include_image, window_id, timeout) but adds no new meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('wait') and resource ('visual quiescence'), clarifies it returns the last frame, and explicitly contrasts with waiting for semantic state or pixel design, distinguishing it from siblings like glass_wait_for_element and glass_wait_for_region.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit alternative condition ('For 2+ known active-window steps, use glass_do') and advises on include_image:false for text-only metadata. It doesn't enumerate all sibling comparisons, but provides clear context for when the tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

glass_windowB

Focus/resize/move the window or read its geometry. op: focus|resize|move|geometry.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoScreen-relative left edge; required only for move.
yNoScreen-relative top edge; required only for move.
opYesOne of: "focus", "resize", "move", "geometry".
widthNoWidth in pixels; required only for resize.
heightNoHeight in pixels; required only for resize.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations mark the tool as not read-only, and the description reinforces that by listing mutating operations (focus, resize, move) alongside read-only geometry. However, it doesn't disclose side effects, window targeting assumptions, or what the geometry operation returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with the key verb and resource front-loaded, followed by a compact op enumeration. The second sentence is slightly redundant but serves as a useful quick reference.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is workable for a four-operation dispatcher with a well-covered schema, but it omits important context such as which window is affected and how geometry results are returned. Since there is no output schema, this leaves some ambiguity for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter already described including conditional requirements like 'required only for move' and 'required only for resize'. The description's op list adds little beyond what the schema already provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (window) and the specific operations it supports (focus, resize, move, geometry) with a compact enumeration. It is understandable and distinct from many siblings, though it doesn't explicitly contrast with tools like glass_move or glass_select_window.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to choose this tool versus sibling tools such as glass_move or glass_select_window. The operation list implies its intended use, but no exclusions or alternative routing are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 25 tool updatesv1.8.0
    • Changedglass_a11y_snapshot1 field changed
      • changedInput schema / properties / max_nodes / description
        Previous value: -"Maximum number of elements to include. Omit for the default cap (protects the token\nbudget). Pass a larger number to raise it, or `0` for the full tree (no limit). A\nsnapshot renumbers ids, so re-read them after changing this."New value: +"Node cap; omit for server default, 0 for unlimited. Changing the cap renumbers IDs; re-read them."
    • Changedglass_baseline_save1 field changed
      • changedInput schema / properties / name / description
        Previous value: -"Name to file this baseline under, reused by `glass_diff` and\n`glass_wait_for_region`. ASCII letters, digits, `-` and `_` only; saving over\nan existing name replaces it without warning."New value: +"ASCII letters/digits/-/_ only. Replaces an existing baseline silently; used by glass_diff and glass_wait_for_region."
    • Changedglass_capabilities1 field changed
      • changedInput schema / properties / backend / description
        Previous value: -"Which backend to report: `x11`, `wayland`, `windows`, `macos`, `android`, or\n`ios`. Omit for the active/default backend (`GLASS_BACKEND`, else the host\ndefault). A valid name for a backend not built into this binary reports\n`available: false`."New value: +"x11, wayland, windows, macos, android or ios. Default active/default backend. Valid but unbuilt backends report available:false."
    • Changedglass_click4 fields changed
      • changedInput schema / properties / count / description
        Previous value: -"Consecutive clicks at this point (default 1, valid range 1 through 10); pass 2 for a\ndouble-click."New value: +"Click count (default 1, range 1 through 10); 2 double-clicks."
      • changedInput schema / properties / modifiers / description
        Previous value: -"Modifier keys to hold during the action, e.g. [\"ctrl\"] or [\"ctrl\",\"shift\"] for multi/range-select."New value: +"Held modifiers, e.g. [\"ctrl\", \"shift\"]."
      • changedInput schema / properties / x / description
        Previous value: -"Click point x, window-relative — 0 is the window's left edge, not the screen's."New value: +"Click x in window-relative pixels."
      • changedInput schema / properties / y / description
        Previous value: -"Click point y, window-relative — 0 is the window's top edge, not the screen's."New value: +"Click y in window-relative pixels."
    • Changedglass_click_element9 fields changed
      • addedInput schema / $defs
        Added value: +{
        +  "ActionModeArg": {
        +    "enum": [
        +      "auto",
        +      "native",
        +      "pointer"
        +    ],
        +    "type": "string"
        +  },
        +  "ActionScopeArgs": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "query": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "role": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "states": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": [
        +          "array",
        +          "null"
        +        ]
        +      }
        +    },
        +    "type": "object"
        +  },
        +  "ActionTargetArgs": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "query": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "role": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "states": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": [
        +          "array",
        +          "null"
        +        ]
        +      },
        +      "within": {
        +        "anyOf": [
        +          {
        +            "$ref": "#/$defs/ActionScopeArgs"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ]
        +      }
        +    },
        +    "type": "object"
        +  }
        +}
      • changedInput schema / properties / id / description
        Previous value: -"Element `#id` from the latest `glass_a11y_snapshot`.\nRe-snapshot after UI changes.\nPopover-owned targets route to their window and restore the prior active window.\nThe role-appropriate native accessibility operation handles occluded or off-screen targets\nand separate labels through their enclosing control.\nUnavailable native operations fall back to a pointer click at the target center.\nText editors may receive focus without activation.\n`method:\"native-action\"` labels any native path.\n`native_fallback` explains pointer fallback.\n`actuated_id` identifies a substituted enclosing control."New value: +"Latest snapshot ID, exclusive with target; re-read after UI changes. Native action may focus text editors; fallback clicks the center. Popover actions restore the prior window."
      • changedInput schema / properties / id / type
        Previous value: -"integer"New value: +[
        +  "integer",
        +  "null"
        +]
      • addedInput schema / properties / max_nodes
        Added value: +{
        +  "format": "uint32",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / mode
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionModeArg"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • changedInput schema / properties / return / description
        Previous value: -"Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default)."New value: +"none (default): no observation; settle: text-only visual stability; snapshot: settle and refresh/fold the accessibility tree."
      • addedInput schema / properties / target
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionTargetArgs"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • addedInput schema / properties / timeout_ms
        Added value: +{
        +  "format": "uint64",
        +  "maximum": 120000,
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • removedInput schema / required
        Removed value: -[
        -  "id"
        -]
    • Changedglass_diff9 fields changed
      • changedInput schema / $defs / RegionArgs / properties / height / description
        Previous value: -"Height in pixels, extending down from `y`."New value: +"Height in pixels."
      • changedInput schema / $defs / RegionArgs / properties / width / description
        Previous value: -"Width in pixels, extending right from `x`."New value: +"Width in pixels."
      • changedInput schema / $defs / RegionArgs / properties / x / description
        Previous value: -"Left edge in pixels, window-relative — 0 is the window's left edge, not the screen's."New value: +"Left edge in window-relative pixels."
      • changedInput schema / $defs / RegionArgs / properties / y / description
        Previous value: -"Top edge in pixels, window-relative — 0 is the window's top edge, not the screen's."New value: +"Top edge in window-relative pixels."
      • changedInput schema / properties / ignore / description
        Previous value: -"Window-relative rectangles to exclude from the comparison. Use for\nperpetually animating content — a blinking text caret, a clock, a\nspinner — which otherwise keeps `changed_pct` permanently non-zero.\n`changed_pct` is measured over the pixels that remain; the excluded\ncount is reported as `ignored_pixels`. Combines with `region`: rects are\nalways window-relative and are intersected with it."New value: +"Window-relative excluded rects intersected with region. changed_pct uses remaining pixels; ignored_pixels reports the excluded count."
      • addedInput schema / properties / max_height
        Added value: +{
        +  "description": "Maximum returned image height; omit both limits for native output.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / max_width
        Added value: +{
        +  "description": "Maximum returned image width; shrinks after crop, preserving native comparison pixels.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / name / description
        Previous value: -"Name of a baseline saved by `glass_baseline_save`; an unsaved name errors\nrather than reporting no change."New value: +"Saved baseline name; unsaved names error."
      • changedInput schema / properties / region / description
        Previous value: -"Optional window-relative sub-rectangle to diff; omit to diff the whole\nwindow. Scopes the comparison (and the reported `bbox`, which becomes\nregion-relative) to just this area — the way to ask \"did *only* this part\nchange?\" Mirrors `glass_wait_for_region`'s `region`."New value: +"Window-relative comparison area; omit for whole window. Returned bbox is region-relative."
    • Changedglass_do74 fields changed
      • addedInput schema / $defs / ActionModeArg
        Added value: +{
        +  "enum": [
        +    "auto",
        +    "native",
        +    "pointer"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / $defs / ActionScopeArgs
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "query": {
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "role": {
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "states": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": [
        +        "array",
        +        "null"
        +      ]
        +    }
        +  },
        +  "type": "object"
        +}
      • addedInput schema / $defs / ActionTargetArgs
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "query": {
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "role": {
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "states": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": [
        +        "array",
        +        "null"
        +      ]
        +    },
        +    "within": {
        +      "anyOf": [
        +        {
        +          "$ref": "#/$defs/ActionScopeArgs"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ]
        +    }
        +  },
        +  "type": "object"
        +}
      • changedInput schema / $defs / ClickArgs / properties / count / description
        Previous value: -"Consecutive clicks at this point (default 1, valid range 1 through 10); pass 2 for a\ndouble-click."New value: +"Click count (default 1, range 1 through 10); 2 double-clicks."
      • changedInput schema / $defs / ClickArgs / properties / modifiers / description
        Previous value: -"Modifier keys to hold during the action, e.g. [\"ctrl\"] or [\"ctrl\",\"shift\"] for multi/range-select."New value: +"Held modifiers, e.g. [\"ctrl\", \"shift\"]."
      • changedInput schema / $defs / ClickArgs / properties / x / description
        Previous value: -"Click point x, window-relative — 0 is the window's left edge, not the screen's."New value: +"Click x in window-relative pixels."
      • changedInput schema / $defs / ClickArgs / properties / y / description
        Previous value: -"Click point y, window-relative — 0 is the window's top edge, not the screen's."New value: +"Click y in window-relative pixels."
      • changedInput schema / $defs / ClickElementArgs / properties / id / description
        Previous value: -"Element `#id` from the latest `glass_a11y_snapshot`.\nRe-snapshot after UI changes.\nPopover-owned targets route to their window and restore the prior active window.\nThe role-appropriate native accessibility operation handles occluded or off-screen targets\nand separate labels through their enclosing control.\nUnavailable native operations fall back to a pointer click at the target center.\nText editors may receive focus without activation.\n`method:\"native-action\"` labels any native path.\n`native_fallback` explains pointer fallback.\n`actuated_id` identifies a substituted enclosing control."New value: +"Latest snapshot ID, exclusive with target; re-read after UI changes. Native action may focus text editors; fallback clicks the center. Popover actions restore the prior window."
      • changedInput schema / $defs / ClickElementArgs / properties / id / type
        Previous value: -"integer"New value: +[
        +  "integer",
        +  "null"
        +]
      • addedInput schema / $defs / ClickElementArgs / properties / max_nodes
        Added value: +{
        +  "format": "uint32",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / $defs / ClickElementArgs / properties / mode
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionModeArg"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • changedInput schema / $defs / ClickElementArgs / properties / return / description
        Previous value: -"Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default)."New value: +"none (default): no observation; settle: text-only visual stability; snapshot: settle and refresh/fold the accessibility tree."
      • addedInput schema / $defs / ClickElementArgs / properties / target
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionTargetArgs"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • addedInput schema / $defs / ClickElementArgs / properties / timeout_ms
        Added value: +{
        +  "format": "uint64",
        +  "maximum": 120000,
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • removedInput schema / $defs / ClickElementArgs / required
        Removed value: -[
        -  "id"
        -]
      • changedInput schema / $defs / DiffArgs / properties / ignore / description
        Previous value: -"Window-relative rectangles to exclude from the comparison. Use for\nperpetually animating content — a blinking text caret, a clock, a\nspinner — which otherwise keeps `changed_pct` permanently non-zero.\n`changed_pct` is measured over the pixels that remain; the excluded\ncount is reported as `ignored_pixels`. Combines with `region`: rects are\nalways window-relative and are intersected with it."New value: +"Window-relative excluded rects intersected with region. changed_pct uses remaining pixels; ignored_pixels reports the excluded count."
      • addedInput schema / $defs / DiffArgs / properties / max_height
        Added value: +{
        +  "description": "Maximum returned image height; omit both limits for native output.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / $defs / DiffArgs / properties / max_width
        Added value: +{
        +  "description": "Maximum returned image width; shrinks after crop, preserving native comparison pixels.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / $defs / DiffArgs / properties / name / description
        Previous value: -"Name of a baseline saved by `glass_baseline_save`; an unsaved name errors\nrather than reporting no change."New value: +"Saved baseline name; unsaved names error."
      • changedInput schema / $defs / DiffArgs / properties / region / description
        Previous value: -"Optional window-relative sub-rectangle to diff; omit to diff the whole\nwindow. Scopes the comparison (and the reported `bbox`, which becomes\nregion-relative) to just this area — the way to ask \"did *only* this part\nchange?\" Mirrors `glass_wait_for_region`'s `region`."New value: +"Window-relative comparison area; omit for whole window. Returned bbox is region-relative."
      • changedInput schema / $defs / DragArgs / properties / duration_ms / description
        Previous value: -"Span the drag's motion over this many milliseconds so a frame-based GUI\nsamples the path across multiple frames (and registers the drag even while\nit repaints). Default 200. Lower = faster but coarser."New value: +"Motion duration in ms (default 200). Faster motion gives the app fewer sampled frames."
      • changedInput schema / $defs / DragArgs / properties / modifiers / description
        Previous value: -"Modifier keys to hold during the action, e.g. [\"ctrl\"] or [\"ctrl\",\"shift\"] for multi/range-select."New value: +"Held modifiers, e.g. [\"ctrl\", \"shift\"]."
      • changedInput schema / $defs / DragArgs / properties / x1 / description
        Previous value: -"Press-point x, window-relative — 0 is the window's left edge, not the screen's."New value: +"Press x in window-relative pixels."
      • changedInput schema / $defs / DragArgs / properties / x2 / description
        Previous value: -"Release-point x, window-relative."New value: +"Release x in window-relative pixels."
      • changedInput schema / $defs / DragArgs / properties / y1 / description
        Previous value: -"Press-point y, window-relative — 0 is the window's top edge, not the screen's."New value: +"Press y in window-relative pixels."
      • changedInput schema / $defs / DragArgs / properties / y2 / description
        Previous value: -"Release-point y, window-relative."New value: +"Release y in window-relative pixels."
      • changedInput schema / $defs / KeyArgs / properties / chord / description
        Previous value: -"A chord like \"ctrl+s\", \"Return\", \"alt+F4\"."New value: +"Modifiers plus one key: ctrl+s, Return, alt+F4. Modifiers: ctrl/shift/alt/super (cmd/win/meta aliases). Key: named key, F1-F12 or printable ASCII; case-insensitive names."
      • changedInput schema / $defs / MoveArgs / properties / x / description
        Previous value: -"Destination x, window-relative — 0 is the window's left edge, not the screen's."New value: +"Destination x in window-relative pixels."
      • changedInput schema / $defs / MoveArgs / properties / y / description
        Previous value: -"Destination y, window-relative — 0 is the window's top edge, not the screen's."New value: +"Destination y in window-relative pixels."
      • changedInput schema / $defs / RegionArgs / properties / height / description
        Previous value: -"Height in pixels, extending down from `y`."New value: +"Height in pixels."
      • changedInput schema / $defs / RegionArgs / properties / width / description
        Previous value: -"Width in pixels, extending right from `x`."New value: +"Width in pixels."
      • changedInput schema / $defs / RegionArgs / properties / x / description
        Previous value: -"Left edge in pixels, window-relative — 0 is the window's left edge, not the screen's."New value: +"Left edge in window-relative pixels."
      • changedInput schema / $defs / RegionArgs / properties / y / description
        Previous value: -"Top edge in pixels, window-relative — 0 is the window's top edge, not the screen's."New value: +"Top edge in window-relative pixels."
      • addedInput schema / $defs / ScreenshotArgs / properties / max_height
        Added value: +{
        +  "description": "Maximum returned image height; omit both limits for native output.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / $defs / ScreenshotArgs / properties / max_width
        Added value: +{
        +  "description": "Maximum returned image width; shrinks after crop, preserving native comparison pixels.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / $defs / ScreenshotArgs / properties / window_id / description
        Previous value: -"Capture/observe this window (id from `glass_list_windows`) instead of the\nactive one, without changing which window subsequent ops target. Omit for\nthe active window."New value: +"Observe this current glass_list_windows ID without selecting it; omit for active window."
      • changedInput schema / $defs / ScrollArgs / properties / dx / description
        Previous value: -"Horizontal wheel notches from -100 through 100, not pixels.\nPositive is right and negative is left, repeated `|dx|` times.\nTypical values are 1–5."New value: +"Horizontal wheel notches, -100 through 100 (positive right, negative left); typically 1-5, not pixels."
      • changedInput schema / $defs / ScrollArgs / properties / dy / description
        Previous value: -"Vertical wheel notches from -100 through 100, not pixels.\nPositive is down and negative is up, repeated `|dy|` times.\nTypical values are 1–5.\nApps choose how a notch maps to lines, pixels, or zoom."New value: +"Vertical wheel notches, -100 through 100 (positive down, negative up); typically 1-5. App determines distance/zoom."
      • changedInput schema / $defs / ScrollArgs / properties / modifiers / description
        Previous value: -"Modifier keys to hold during the action, e.g. [\"ctrl\"] or [\"ctrl\",\"shift\"] for multi/range-select."New value: +"Held modifiers, e.g. [\"ctrl\", \"shift\"]."
      • changedInput schema / $defs / ScrollArgs / properties / x / description
        Previous value: -"Pointer x the wheel is aimed at, window-relative — apps scroll the container\nunder this point, so it selects which pane moves."New value: +"Window-relative anchor x; selects the container under the pointer."
      • changedInput schema / $defs / ScrollArgs / properties / y / description
        Previous value: -"Pointer y the wheel is aimed at, window-relative. See `x`."New value: +"Window-relative anchor y."
      • changedInput schema / $defs / ScrollToElementArgs / properties / description / description
        Previous value: -"Substring of the target element's accessible description (selector). This can select\nan unnamed control, including an Android text field labelled only by its hint."New value: +"Accessible-description substring; can select unnamed controls."
      • changedInput schema / $defs / ScrollToElementArgs / properties / direction / description
        Previous value: -"Sweep direction: \"up\"/\"down\" (vertical) or \"left\"/\"right\" (horizontal).\nOmit to infer it from the target's off-screen position (falls back to a\nvertical down→up sweep when the target isn't in the a11y tree yet). The\nsearch reverses to the other end if the target isn't found first."New value: +"up/down/left/right. Default infers off-screen direction, or down then up if absent. Reverses at the first end."
      • changedInput schema / $defs / ScrollToElementArgs / properties / step / description
        Previous value: -"Wheel notches per scroll step (default 3). A calibration escape hatch — larger\ncovers distance faster but risks stepping past a row's/column's realized band."New value: +"Wheel notches per step (default 3); large steps can skip realized rows."
      • changedInput schema / $defs / ScrollToElementArgs / properties / timeout_ms / description
        Previous value: -"Give up after this long (default 20000ms); returns `{matched:false}`."New value: +"Timeout in ms (default 20000). Standalone: matched:false; batched: fails the sequence."
      • changedInput schema / $defs / ScrollToElementArgs / properties / value_contains / description
        Previous value: -"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required."New value: +"Value substring; requires name, description and/or role."
      • changedInput schema / $defs / ScrollToElementArgs / properties / x / description
        Previous value: -"Scroll anchor x (window-relative). By default the swipe anchors on the target's\nown row/column (falling back to the window center if it isn't in the a11y tree\nyet); set both `x` and `y` to point the wheel at a specific container instead."New value: +"Window-relative anchor x; supply both x/y for a container. Default target row/column, or window center if absent."
      • changedInput schema / $defs / SetValueArgs / properties / id / description
        Previous value: -"The element `#id` from `glass_a11y_snapshot`."New value: +"Latest snapshot ID, exclusive with target."
      • changedInput schema / $defs / SetValueArgs / properties / id / type
        Previous value: -"integer"New value: +[
        +  "integer",
        +  "null"
        +]
      • addedInput schema / $defs / SetValueArgs / properties / max_nodes
        Added value: +{
        +  "format": "uint32",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / $defs / SetValueArgs / properties / return / description
        Previous value: -"Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (waits for visual\nstability and returns text-only metadata), or \"none\" (default)."New value: +"none (default): no observation; settle: text-only visual stability; snapshot: settle and refresh/fold the accessibility tree."
      • addedInput schema / $defs / SetValueArgs / properties / target
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionTargetArgs"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • changedInput schema / $defs / SetValueArgs / properties / text / description
        Previous value: -"The value to set. For a text field, the text. For a spin/slider, a number.\nFor a switch/checkbox/toggle, a boolean (`\"true\"`/`\"false\"`/`\"on\"`/`\"off\"`/\n`\"1\"`/`\"0\"`) — idempotent. For a dropdown/combo box, an option label\n(case-insensitive); glass opens it and picks that option."New value: +"Text, slider/spin number, toggle boolean (true/false/on/off/1/0), or case-insensitive combo option label. Toggle writes are idempotent; combos open and choose the option."
      • addedInput schema / $defs / SetValueArgs / properties / timeout_ms
        Added value: +{
        +  "format": "uint64",
        +  "maximum": 120000,
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / $defs / SetValueArgs / required
        Previous value: -[
        -  "id",
        -  "text"
        -]New value: +[
        +  "text"
        +]
      • changedInput schema / $defs / SettleArgs / properties / ignore / description
        Previous value: -"Window-relative rectangles to exclude from the settle comparison. Use for\nperpetually animating content — a blinking text caret, a clock, a\nspinner — which otherwise keeps the window from ever settling.\nCombines with `stability_region`: rects are always window-relative and\nare intersected with it. A rect that falls partially or entirely\noutside the compared area is silently clamped or dropped, masking less\nthan requested or nothing at all — double-check placement."New value: +"Window-relative excluded rects intersected with stability_region. Off-area rects clamp/drop silently; check placement."
      • changedInput schema / $defs / SettleArgs / properties / stability_region / description
        Previous value: -"Window-relative sub-rectangle to watch for settling; when set, changes outside\nit are ignored."New value: +"Window-relative area watched for settling."
      • changedInput schema / $defs / SettleArgs / properties / timeout_ms / description
        Previous value: -"Give up after this long (default 5000ms).\nThis settle's timeout returns settled:false and completes the step.\nThe enclosing glass_do deadline fails the sequence."New value: +"Timeout in ms (default 5000) completes with settled:false. The overall glass_do deadline instead fails the sequence."
      • changedInput schema / $defs / SettleArgs / properties / tolerance / description
        Previous value: -"Per-channel difference (0–255) two frames may have and still count as\nunchanged (default 0, exact match)."New value: +"Per-channel difference allowed (0-255, default 0)."
      • changedInput schema / $defs / ThenArgs / properties / screenshot / description
        Previous value: -"Capture the window as an image — the only field here that always spends image\ntokens; prefer `diff` when you just need to know whether something changed."New value: +"Return a screenshot. Prefer diff for change detection without image tokens."
      • changedInput schema / $defs / ThenArgs / properties / settle / description
        Previous value: -"Wait for the UI to stop changing first. Set this whenever `diff` or\n`screenshot` follows, or they observe a half-drawn frame."New value: +"Wait for quiescence before diff/screenshot; otherwise they may observe a half-drawn frame."
      • addedInput schema / $defs / TypeArgs / properties / focus_mode
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionModeArg"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • addedInput schema / $defs / TypeArgs / properties / max_nodes
        Added value: +{
        +  "format": "uint32",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / $defs / TypeArgs / properties / return / description
        Previous value: -"Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default)."New value: +"none (default): no observation; settle: text-only visual stability; snapshot: settle and read the current accessibility tree, including app-visible text, with secure values redacted."
      • addedInput schema / $defs / TypeArgs / properties / target
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionTargetArgs"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • changedInput schema / $defs / TypeArgs / properties / text / description
        Previous value: -"Text to type into whatever currently has keyboard focus — this tool does not\nfocus a field, so click or `glass_click_element` one first. Sent as synthetic\nkey events, not pasted, so an app's per-keystroke handlers run."New value: +"Synthetic keystrokes, not paste. When `target` is omitted: current keyboard focus. When `target` is supplied: resolves and focuses the target, confirms focus, then types. Newlines do not press Return."
      • addedInput schema / $defs / TypeArgs / properties / timeout_ms
        Added value: +{
        +  "format": "uint64",
        +  "maximum": 120000,
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / $defs / WaitForElementArgs / properties / condition / description
        Previous value: -"What to wait for (default \"appears\"): appears|disappears|enabled|disabled|\nchecked|unchecked|selected|unselected|expanded|collapsed|focused|visible|hidden.\n`checked`/`unchecked` only match a checkable element (one exposing a real toggle\nstate) — a non-toggle element matches neither."New value: +"Default appears. appears|disappears|enabled|disabled|checked|unchecked|selected|unselected|expanded|collapsed|focused|visible|hidden. checked/unchecked require a real checkable state."
      • changedInput schema / $defs / WaitForElementArgs / properties / description / description
        Previous value: -"Substring of the element's accessible description (selector). Useful for unnamed\ncontrols whose platform label is exposed as a hint, help text, or description."New value: +"Accessible-description substring; can select unnamed controls."
      • changedInput schema / $defs / WaitForElementArgs / properties / timeout_ms / description
        Previous value: -"Give up after this long (default 10000ms); returns `{matched:false}`."New value: +"Timeout in ms (default 10000). Standalone: matched:false; batched: fails the sequence."
      • changedInput schema / $defs / WaitForElementArgs / properties / value_contains / description
        Previous value: -"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required."New value: +"Value substring; requires name, description and/or role; exclusive with value."
      • changedInput schema / properties / actions / description
        Previous value: -"Actions to run in order; must be non-empty. Fail-fast — the first failing\naction aborts the rest and reports its index, so a partial sequence may\nalready have landed."New value: +"Non-empty ordered actions; failure stops remaining steps. Completed mutations may already have landed."
      • changedInput schema / properties / then / description
        Previous value: -"Optional observe run once after the last action, in the order settle → diff\n→ screenshot."New value: +"Terminal observation in order: settle, diff, screenshot."
      • changedInput schema / properties / timeout_ms / description
        Previous value: -"Overall sequence budget in milliseconds. Omit for 30000; valid range\n1..=120000. One absolute deadline is shared by all actions and terminal\nobservations."New value: +"Overall budget in ms (default 30000, range 1..120000), shared by all actions and terminal observations."
    • Changedglass_doctor1 field changed
      • changedInput schema / properties / deep / description
        Previous value: -"Also spawn and tear down the default backend's headless display to verify it\nactually starts (slower). Default false."New value: +"Start and tear down the default display to prove it starts (default false)."
    • Changedglass_drag6 fields changed
      • changedInput schema / properties / duration_ms / description
        Previous value: -"Span the drag's motion over this many milliseconds so a frame-based GUI\nsamples the path across multiple frames (and registers the drag even while\nit repaints). Default 200. Lower = faster but coarser."New value: +"Motion duration in ms (default 200). Faster motion gives the app fewer sampled frames."
      • changedInput schema / properties / modifiers / description
        Previous value: -"Modifier keys to hold during the action, e.g. [\"ctrl\"] or [\"ctrl\",\"shift\"] for multi/range-select."New value: +"Held modifiers, e.g. [\"ctrl\", \"shift\"]."
      • changedInput schema / properties / x1 / description
        Previous value: -"Press-point x, window-relative — 0 is the window's left edge, not the screen's."New value: +"Press x in window-relative pixels."
      • changedInput schema / properties / x2 / description
        Previous value: -"Release-point x, window-relative."New value: +"Release x in window-relative pixels."
      • changedInput schema / properties / y1 / description
        Previous value: -"Press-point y, window-relative — 0 is the window's top edge, not the screen's."New value: +"Press y in window-relative pixels."
      • changedInput schema / properties / y2 / description
        Previous value: -"Release-point y, window-relative."New value: +"Release y in window-relative pixels."
    • Changedglass_gesture2 fields changed
      • changedInput schema / $defs / PointArg / properties / x / description
        Previous value: -"Window-relative x — 0 is the window's left edge, not the screen's."New value: +"Window-relative x in pixels."
      • changedInput schema / $defs / PointArg / properties / y / description
        Previous value: -"Window-relative y — 0 is the window's top edge, not the screen's."New value: +"Window-relative y in pixels."
    • Changedglass_key1 field changed
      • changedInput schema / properties / chord / description
        Previous value: -"A chord like \"ctrl+s\", \"Return\", \"alt+F4\"."New value: +"Modifiers plus one key: ctrl+s, Return, alt+F4. Modifiers: ctrl/shift/alt/super (cmd/win/meta aliases). Key: named key, F1-F12 or printable ASCII; case-insensitive names."
    • Changedglass_logs3 fields changed
      • changedInput schema / properties / contains / description
        Previous value: -"Return only lines containing this substring (case-sensitive). Filtering\nhappens server-side, so it narrows what the cap applies to."New value: +"Case-sensitive substring filter applied before the line cap."
      • changedInput schema / properties / cursor / description
        Previous value: -"Resume point — the `cursor` a previous call returned, to read only what has\nbeen logged since. Omit to read from the oldest buffered line."New value: +"Resume from a prior returned cursor; omit for oldest buffered line."
      • changedInput schema / properties / max_lines / description
        Previous value: -"Cap on lines returned (default 200); the returned `cursor` resumes at the\nfirst line left unread, so a capped read is not a lost one."New value: +"Line cap (default 200). Returned cursor resumes at the first unread line."
    • Changedglass_move2 fields changed
      • changedInput schema / properties / x / description
        Previous value: -"Destination x, window-relative — 0 is the window's left edge, not the screen's."New value: +"Destination x in window-relative pixels."
      • changedInput schema / properties / y / description
        Previous value: -"Destination y, window-relative — 0 is the window's top edge, not the screen's."New value: +"Destination y in window-relative pixels."
    • Changedglass_screenshot7 fields changed
      • changedInput schema / $defs / RegionArgs / properties / height / description
        Previous value: -"Height in pixels, extending down from `y`."New value: +"Height in pixels."
      • changedInput schema / $defs / RegionArgs / properties / width / description
        Previous value: -"Width in pixels, extending right from `x`."New value: +"Width in pixels."
      • changedInput schema / $defs / RegionArgs / properties / x / description
        Previous value: -"Left edge in pixels, window-relative — 0 is the window's left edge, not the screen's."New value: +"Left edge in window-relative pixels."
      • changedInput schema / $defs / RegionArgs / properties / y / description
        Previous value: -"Top edge in pixels, window-relative — 0 is the window's top edge, not the screen's."New value: +"Top edge in window-relative pixels."
      • addedInput schema / properties / max_height
        Added value: +{
        +  "description": "Maximum returned image height; omit both limits for native output.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / max_width
        Added value: +{
        +  "description": "Maximum returned image width; shrinks after crop, preserving native comparison pixels.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / window_id / description
        Previous value: -"Capture/observe this window (id from `glass_list_windows`) instead of the\nactive one, without changing which window subsequent ops target. Omit for\nthe active window."New value: +"Observe this current glass_list_windows ID without selecting it; omit for active window."
    • Changedglass_scroll5 fields changed
      • changedInput schema / properties / dx / description
        Previous value: -"Horizontal wheel notches from -100 through 100, not pixels.\nPositive is right and negative is left, repeated `|dx|` times.\nTypical values are 1–5."New value: +"Horizontal wheel notches, -100 through 100 (positive right, negative left); typically 1-5, not pixels."
      • changedInput schema / properties / dy / description
        Previous value: -"Vertical wheel notches from -100 through 100, not pixels.\nPositive is down and negative is up, repeated `|dy|` times.\nTypical values are 1–5.\nApps choose how a notch maps to lines, pixels, or zoom."New value: +"Vertical wheel notches, -100 through 100 (positive down, negative up); typically 1-5. App determines distance/zoom."
      • changedInput schema / properties / modifiers / description
        Previous value: -"Modifier keys to hold during the action, e.g. [\"ctrl\"] or [\"ctrl\",\"shift\"] for multi/range-select."New value: +"Held modifiers, e.g. [\"ctrl\", \"shift\"]."
      • changedInput schema / properties / x / description
        Previous value: -"Pointer x the wheel is aimed at, window-relative — apps scroll the container\nunder this point, so it selects which pane moves."New value: +"Window-relative anchor x; selects the container under the pointer."
      • changedInput schema / properties / y / description
        Previous value: -"Pointer y the wheel is aimed at, window-relative. See `x`."New value: +"Window-relative anchor y."
    • Changedglass_scroll_to_element6 fields changed
      • changedInput schema / properties / description / description
        Previous value: -"Substring of the target element's accessible description (selector). This can select\nan unnamed control, including an Android text field labelled only by its hint."New value: +"Accessible-description substring; can select unnamed controls."
      • changedInput schema / properties / direction / description
        Previous value: -"Sweep direction: \"up\"/\"down\" (vertical) or \"left\"/\"right\" (horizontal).\nOmit to infer it from the target's off-screen position (falls back to a\nvertical down→up sweep when the target isn't in the a11y tree yet). The\nsearch reverses to the other end if the target isn't found first."New value: +"up/down/left/right. Default infers off-screen direction, or down then up if absent. Reverses at the first end."
      • changedInput schema / properties / step / description
        Previous value: -"Wheel notches per scroll step (default 3). A calibration escape hatch — larger\ncovers distance faster but risks stepping past a row's/column's realized band."New value: +"Wheel notches per step (default 3); large steps can skip realized rows."
      • changedInput schema / properties / timeout_ms / description
        Previous value: -"Give up after this long (default 20000ms); returns `{matched:false}`."New value: +"Timeout in ms (default 20000). Standalone: matched:false; batched: fails the sequence."
      • changedInput schema / properties / value_contains / description
        Previous value: -"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required."New value: +"Value substring; requires name, description and/or role."
      • changedInput schema / properties / x / description
        Previous value: -"Scroll anchor x (window-relative). By default the swipe anchors on the target's\nown row/column (falling back to the window center if it isn't in the a11y tree\nyet); set both `x` and `y` to point the wheel at a specific container instead."New value: +"Window-relative anchor x; supply both x/y for a container. Default target row/column, or window center if absent."
    • Changedglass_select_window1 field changed
      • changedInput schema / properties / id / description
        Previous value: -"The window id from `glass_list_windows`. Ids are not stable across calls —\nre-list rather than caching them."New value: +"Current ID from glass_list_windows; re-list after window changes."
    • Changedglass_set_value9 fields changed
      • addedInput schema / $defs
        Added value: +{
        +  "ActionScopeArgs": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "query": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "role": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "states": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": [
        +          "array",
        +          "null"
        +        ]
        +      }
        +    },
        +    "type": "object"
        +  },
        +  "ActionTargetArgs": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "query": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "role": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "states": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": [
        +          "array",
        +          "null"
        +        ]
        +      },
        +      "within": {
        +        "anyOf": [
        +          {
        +            "$ref": "#/$defs/ActionScopeArgs"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ]
        +      }
        +    },
        +    "type": "object"
        +  }
        +}
      • changedInput schema / properties / id / description
        Previous value: -"The element `#id` from `glass_a11y_snapshot`."New value: +"Latest snapshot ID, exclusive with target."
      • changedInput schema / properties / id / type
        Previous value: -"integer"New value: +[
        +  "integer",
        +  "null"
        +]
      • addedInput schema / properties / max_nodes
        Added value: +{
        +  "format": "uint32",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / return / description
        Previous value: -"Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (waits for visual\nstability and returns text-only metadata), or \"none\" (default)."New value: +"none (default): no observation; settle: text-only visual stability; snapshot: settle and refresh/fold the accessibility tree."
      • addedInput schema / properties / target
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionTargetArgs"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • changedInput schema / properties / text / description
        Previous value: -"The value to set. For a text field, the text. For a spin/slider, a number.\nFor a switch/checkbox/toggle, a boolean (`\"true\"`/`\"false\"`/`\"on\"`/`\"off\"`/\n`\"1\"`/`\"0\"`) — idempotent. For a dropdown/combo box, an option label\n(case-insensitive); glass opens it and picks that option."New value: +"Text, slider/spin number, toggle boolean (true/false/on/off/1/0), or case-insensitive combo option label. Toggle writes are idempotent; combos open and choose the option."
      • addedInput schema / properties / timeout_ms
        Added value: +{
        +  "format": "uint64",
        +  "maximum": 120000,
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / required
        Previous value: -[
        -  "id",
        -  "text"
        -]New value: +[
        +  "text"
        +]
    • Changedglass_start10 fields changed
      • changedInput schema / $defs / WindowHintArgs / properties / class / description
        Previous value: -"Exact window-class match. Same purpose as `title` but more stable, since\nclass names rarely carry the dynamic prefixes/suffixes that titles do."New value: +"Exact window-class match."
      • changedInput schema / $defs / WindowHintArgs / properties / title / description
        Previous value: -"Case-insensitive substring matched against window titles. Used to pick the\nright window when several appear, and — since it ignores the process tree —\nto locate a window the launched process hands off to an unrelated process."New value: +"Case-insensitive title substring; can locate a window handed off to an unrelated process."
      • changedInput schema / properties / a11y / description
        Previous value: -"Spawn a private accessibility (AT-SPI) bus so `glass_a11y_snapshot` / `marks` /\n`set_value` / `click_element` / `wait_for_element` work against this app. **On by\ndefault** — the accessibility path is the cheap, low-token way to drive a UI, so it\nis available unless you opt out. Pass `false` to skip the bus for canvas/pixel-only\napps (it spawns extra processes). Effective on Linux only; other backends read\naccessibility ambiently and ignore this flag."New value: +"Enable the private accessibility bus (default true). False skips it for canvas-only apps. Linux only; other backends read accessibility ambiently."
      • changedInput schema / properties / backend / description
        Previous value: -"Backend to launch under: `\"x11\"` or `\"wayland\"` (Linux), `\"windows\"` (on a\nWindows host), `\"macos\"` (on a macOS host), `\"android\"` (an AVD emulator, any\nhost), or `\"ios\"` (an iOS Simulator, macOS host). Omit for the server default\n(`GLASS_BACKEND`, else `windows` on Windows, `macos` on macOS, else x11)."New value: +"x11/wayland (Linux), windows (Windows), macos (macOS), android (any host), ios (macOS). Default: GLASS_BACKEND, else host default (x11 on Linux)."
      • changedInput schema / properties / cwd / description
        Previous value: -"Working directory for both `build` and the launched app; omit to inherit the\nserver's own."New value: +"Working directory for build and app; defaults to the server directory."
      • changedInput schema / properties / env / description
        Previous value: -"Extra environment variables, as a `{ \"KEY\": \"VALUE\" }` object. They reach the launched app\non the desktop backends and on `ios`; on `android` they configure the `build` command on\nthe host only, since an app launched by `am start` is forked from zygote and never sees\nthe shell's environment."New value: +"Extra {KEY: VALUE} environment for build and app. Android applies it only to the host build, not the app."
      • changedInput schema / properties / run / description
        Previous value: -"What to launch: desktop `[executable, args...]`; iOS `[.app-or-bundle-id, args...]`;\nAndroid `[apk?, package/.Activity]` in either order, for example\n`[\"/absolute/path/app.apk\", \"com.example.app/.MainActivity\"]`."New value: +"Desktop: [executable, args...]. iOS: [.app-or-bundle-id, args...]. Android: [apk?, package/.Activity] in either order, e.g. [\"/absolute/path/app.apk\", \"com.example.app/.MainActivity\"]."
      • changedInput schema / properties / sandbox / description
        Previous value: -"Containment level for the launched app: `\"default\"` (filesystem/process\ncontainment, network on), `\"strict\"` (also no network), or `\"off\"` (no\ncontainment). Omit for the server default (`GLASS_SANDBOX`, else `default`).\nAn operator-set floor (`GLASS_SANDBOX_FLOOR`) may raise an omitted level, and\nrefuses an explicit level requested below it."New value: +"default: filesystem/process containment, network on; strict: also no network; off: uncontained. Default GLASS_SANDBOX or default. GLASS_SANDBOX_FLOOR raises omitted levels and refuses explicit lower levels."
      • changedInput schema / properties / timeout_ms / description
        Previous value: -"How long to wait for the app's window to appear before failing the launch\n(default 10000ms). Does not bound `build`."New value: +"Window-publication timeout in ms (default 10000); does not bound build."
      • changedInput schema / properties / window_hint / description
        Previous value: -"Optional `{ title?, class? }` to disambiguate which window is the app's when\nmore than one appears, or to find a window the launched process hands off to\nan unrelated process. Omit to take the first window owned by the launched\nprocess or a descendant it can follow."New value: +"Select a window by title/class, including process handoffs. Omit for the first window owned by the process or a followable descendant."
    • Changedglass_type7 fields changed
      • addedInput schema / $defs
        Added value: +{
        +  "ActionModeArg": {
        +    "enum": [
        +      "auto",
        +      "native",
        +      "pointer"
        +    ],
        +    "type": "string"
        +  },
        +  "ActionScopeArgs": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "query": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "role": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "states": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": [
        +          "array",
        +          "null"
        +        ]
        +      }
        +    },
        +    "type": "object"
        +  },
        +  "ActionTargetArgs": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "query": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "role": {
        +        "type": [
        +          "string",
        +          "null"
        +        ]
        +      },
        +      "states": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "type": [
        +          "array",
        +          "null"
        +        ]
        +      },
        +      "within": {
        +        "anyOf": [
        +          {
        +            "$ref": "#/$defs/ActionScopeArgs"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ]
        +      }
        +    },
        +    "type": "object"
        +  }
        +}
      • addedInput schema / properties / focus_mode
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionModeArg"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • addedInput schema / properties / max_nodes
        Added value: +{
        +  "format": "uint32",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / return / description
        Previous value: -"Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default)."New value: +"none (default): no observation; settle: text-only visual stability; snapshot: settle and read the current accessibility tree, including app-visible text, with secure values redacted."
      • addedInput schema / properties / target
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ActionTargetArgs"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ]
        +}
      • changedInput schema / properties / text / description
        Previous value: -"Text to type into whatever currently has keyboard focus — this tool does not\nfocus a field, so click or `glass_click_element` one first. Sent as synthetic\nkey events, not pasted, so an app's per-keystroke handlers run."New value: +"Synthetic keystrokes, not paste. When `target` is omitted: current keyboard focus. When `target` is supplied: resolves and focuses the target, confirms focus, then types. Newlines do not press Return."
      • addedInput schema / properties / timeout_ms
        Added value: +{
        +  "format": "uint64",
        +  "maximum": 120000,
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
    • Changedglass_wait_for_element4 fields changed
      • changedInput schema / properties / condition / description
        Previous value: -"What to wait for (default \"appears\"): appears|disappears|enabled|disabled|\nchecked|unchecked|selected|unselected|expanded|collapsed|focused|visible|hidden.\n`checked`/`unchecked` only match a checkable element (one exposing a real toggle\nstate) — a non-toggle element matches neither."New value: +"Default appears. appears|disappears|enabled|disabled|checked|unchecked|selected|unselected|expanded|collapsed|focused|visible|hidden. checked/unchecked require a real checkable state."
      • changedInput schema / properties / description / description
        Previous value: -"Substring of the element's accessible description (selector). Useful for unnamed\ncontrols whose platform label is exposed as a hint, help text, or description."New value: +"Accessible-description substring; can select unnamed controls."
      • changedInput schema / properties / timeout_ms / description
        Previous value: -"Give up after this long (default 10000ms); returns `{matched:false}`."New value: +"Timeout in ms (default 10000). Standalone: matched:false; batched: fails the sequence."
      • changedInput schema / properties / value_contains / description
        Previous value: -"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required."New value: +"Value substring; requires name, description and/or role; exclusive with value."
    • Changedglass_wait_for_log1 field changed
      • changedInput schema / properties / cursor / description
        Previous value: -"Start scanning from this cursor (from a prior glass_logs). Omit to match\nonly lines emitted after this call."New value: +"Cursor after draining logs before the action; 0 includes retained history. Omit for future lines only."
    • Changedglass_wait_for_region8 fields changed
      • changedInput schema / $defs / RegionArgs / properties / height / description
        Previous value: -"Height in pixels, extending down from `y`."New value: +"Height in pixels."
      • changedInput schema / $defs / RegionArgs / properties / width / description
        Previous value: -"Width in pixels, extending right from `x`."New value: +"Width in pixels."
      • changedInput schema / $defs / RegionArgs / properties / x / description
        Previous value: -"Left edge in pixels, window-relative — 0 is the window's left edge, not the screen's."New value: +"Left edge in window-relative pixels."
      • changedInput schema / $defs / RegionArgs / properties / y / description
        Previous value: -"Top edge in pixels, window-relative — 0 is the window's top edge, not the screen's."New value: +"Top edge in window-relative pixels."
      • changedInput schema / properties / ignore / description
        Previous value: -"Window-relative rectangles to exclude from the comparison. Use for\nperpetually animating content — a blinking text caret, a clock, a\nspinner — which otherwise keeps `changed_pct` permanently non-zero.\n`changed_pct` is measured over the pixels that remain. Combines with\n`region`: rects are always window-relative and are intersected with it.\nA rect that falls partially or entirely outside the compared area —\nthe frame, or the `region` sub-rectangle when one is set — is silently\nclamped or dropped, masking less than requested or nothing at all; the\nexcluded count is reported as `ignored_pixels`, so a smaller-than-\nexpected value flags a misplaced rect."New value: +"Window-relative excluded rects intersected with region. Off-area rects clamp/drop silently; check ignored_pixels. changed_pct uses remaining pixels."
      • addedInput schema / properties / max_height
        Added value: +{
        +  "description": "Maximum returned image height; omit both limits for native output.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / max_width
        Added value: +{
        +  "description": "Maximum returned image width; shrinks after crop, preserving native comparison pixels.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / window_id / description
        Previous value: -"Capture/observe this window (id from `glass_list_windows`) instead of the\nactive one, without changing which window subsequent ops target. Omit for\nthe active window."New value: +"Observe this current glass_list_windows ID without selecting it; omit for active window."
    • Changedglass_wait_stable12 fields changed
      • changedInput schema / $defs / RegionArgs / properties / height / description
        Previous value: -"Height in pixels, extending down from `y`."New value: +"Height in pixels."
      • changedInput schema / $defs / RegionArgs / properties / width / description
        Previous value: -"Width in pixels, extending right from `x`."New value: +"Width in pixels."
      • changedInput schema / $defs / RegionArgs / properties / x / description
        Previous value: -"Left edge in pixels, window-relative — 0 is the window's left edge, not the screen's."New value: +"Left edge in window-relative pixels."
      • changedInput schema / $defs / RegionArgs / properties / y / description
        Previous value: -"Top edge in pixels, window-relative — 0 is the window's top edge, not the screen's."New value: +"Top edge in window-relative pixels."
      • changedInput schema / properties / ignore / description
        Previous value: -"Window-relative rectangles to exclude from the settle comparison. Use for\nperpetually animating content — a blinking text caret, a clock, a\nspinner — which otherwise keeps the window from ever settling. Pixels\ninside a rect never count as changed and never set `saw_motion`.\nCombines with `stability_region`: rects are always window-relative and\nare intersected with it. Independent of `region`, which only crops the\nreturned image. A rect that falls partially or entirely outside the\ncompared area — the frame, or the `stability_region` sub-rectangle when\none is set — is silently clamped or dropped, masking less than\nrequested or nothing at all; the excluded count is reported as\n`ignored_pixels`, so a smaller-than-expected value flags a misplaced rect."New value: +"Window-relative rects excluded from comparison and saw_motion, intersected with stability_region. Off-area rects clamp/drop silently; check ignored_pixels for misplaced masks."
      • changedInput schema / properties / include_image / description
        Previous value: -"Return the settled frame as an image (default true). Set false for a\ntext-only `{settled, saw_motion, observed_ms, ignored_pixels, width,\nheight}` result with no WebP — cheap when the next step is a text\n`glass_diff`. `region` is ignored when false."New value: +"Return image (default true). False returns settled/saw_motion/observed_ms/ignored_pixels/dimensions as text; region then has no effect."
      • addedInput schema / properties / max_height
        Added value: +{
        +  "description": "Maximum returned image height; omit both limits for native output.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / max_width
        Added value: +{
        +  "description": "Maximum returned image width; shrinks after crop, preserving native comparison pixels.",
        +  "format": "uint32",
        +  "maximum": 4294967295,
        +  "minimum": 1,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / settle_frames / description
        Previous value: -"Consecutive unchanged frames required before the UI counts as settled\n(default 3). Raise it for an app that pauses mid-animation."New value: +"Consecutive unchanged frames required (default 3)."
      • changedInput schema / properties / stability_region / description
        Previous value: -"Optional window-relative sub-rectangle to watch for settling; when set,\nthe settle decision ignores changes outside it. Independent of `region`."New value: +"Window-relative area watched for settling, independent of returned-image region."
      • changedInput schema / properties / tolerance / description
        Previous value: -"Per-channel difference (0–255) two frames may have and still count as\nunchanged (default 0, exact match). Raise it for a backend with dithering\nor compression noise."New value: +"Per-channel difference allowed (0-255, default 0)."
      • changedInput schema / properties / window_id / description
        Previous value: -"Capture/observe this window (id from `glass_list_windows`) instead of the\nactive one, without changing which window subsequent ops target. Omit for\nthe active window."New value: +"Observe this current glass_list_windows ID without selecting it; omit for active window."
    • Changedglass_window4 fields changed
      • changedInput schema / properties / height / description
        Previous value: -"New height in pixels for `op: \"resize\"`, required there and ignored otherwise."New value: +"Height in pixels; required only for resize."
      • changedInput schema / properties / width / description
        Previous value: -"New width in pixels for `op: \"resize\"`, required there and ignored otherwise."New value: +"Width in pixels; required only for resize."
      • changedInput schema / properties / x / description
        Previous value: -"New left edge for `op: \"move\"`, required there and ignored otherwise. Screen\ncoordinates — the one place in this API that is not window-relative, since a\nwindow cannot be positioned relative to itself."New value: +"Screen-relative left edge; required only for move."
      • changedInput schema / properties / y / description
        Previous value: -"New top edge for `op: \"move\"`, required there and ignored otherwise. Screen\ncoordinates; see `x`."New value: +"Screen-relative top edge; required only for move."
  2. 10 tool updatesv1.5.1
    • Changedglass_click3 fields changed
      • changedInput schema / properties / count / description
        Previous value: -"Consecutive clicks at this point (default 1); pass 2 for a double-click."New value: +"Consecutive clicks at this point (default 1, valid range 1 through 10); pass 2 for a\ndouble-click."
      • addedInput schema / properties / count / maximum
        Added value: +10
      • changedInput schema / properties / count / minimum
        Previous value: -0New value: +1
    • Changedglass_click_element2 fields changed
      • changedInput schema / properties / id / description
        Previous value: -"The element `#id` from `glass_a11y_snapshot`. Valid only within the latest\nsnapshot — re-snapshot if the UI changed. If the element actually renders in a\npopover owned by a different window than the active one (e.g. an open\ndropdown's option row), the click is automatically routed into that popover\nwindow and the previously-active window is restored afterward — no extra step\nneeded.\n\nClicks via the platform's native accessibility action when the element exposes\none (works even when the element is occluded or scrolled off-screen), falling\nback to a synthetic pointer click at the element's center; the result's\n`method` field says which path ran, and `native_fallback` says why when the\npointer path was used. Where a control's label is a separate element from the\ncontrol itself, the native action fires on the enclosing control and the\nresult carries `actuated_id` — the element actually clicked."New value: +"Element `#id` from the latest `glass_a11y_snapshot`.\nRe-snapshot after UI changes.\nPopover-owned targets route to their window and restore the prior active window.\nThe role-appropriate native accessibility operation handles occluded or off-screen targets\nand separate labels through their enclosing control.\nUnavailable native operations fall back to a pointer click at the target center.\nText editors may receive focus without activation.\n`method:\"native-action\"` labels any native path.\n`native_fallback` explains pointer fallback.\n`actuated_id` identifies a substituted enclosing control."
      • changedInput schema / properties / return / description
        Previous value: -"Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (wait for the UI\nto stop changing, text-only), or \"none\" (default)."New value: +"Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default)."
    • Changedglass_do17 fields changed
      • changedInput schema / $defs / Action / oneOf
        Previous value: -[
        -  {
        -    "$ref": "#/$defs/ClickArgs",
        -    "properties": {
        -      "action": {
        -        "const": "click",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "action"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "$ref": "#/$defs/MoveArgs",
        -    "properties": {
        -      "action": {
        -        "const": "move",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "action"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "$ref": "#/$defs/DragArgs",
        -    "properties": {
        -      "action": {
        -        "const": "drag",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "action"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "$ref": "#/$defs/ScrollArgs",
        -    "properties": {
        -      "action": {
        -        "const": "scroll",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "action"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "$ref": "#/$defs/TypeArgs",
        -    "properties": {
        -      "action": {
        -        "const": "type",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "action"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "$ref": "#/$defs/KeyArgs",
        -    "properties": {
        -      "action": {
        -        "const": "key",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "action"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "$ref": "#/$defs/SettleArgs",
        -    "properties": {
        -      "action": {
        -        "const": "settle",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "action"
        -    ],
        -    "type": "object"
        -  }
        -]New value: +[
        +  {
        +    "$ref": "#/$defs/ClickArgs",
        +    "properties": {
        +      "action": {
        +        "const": "click",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/MoveArgs",
        +    "properties": {
        +      "action": {
        +        "const": "move",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/DragArgs",
        +    "properties": {
        +      "action": {
        +        "const": "drag",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/ScrollArgs",
        +    "properties": {
        +      "action": {
        +        "const": "scroll",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/TypeArgs",
        +    "properties": {
        +      "action": {
        +        "const": "type",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/KeyArgs",
        +    "properties": {
        +      "action": {
        +        "const": "key",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/SettleArgs",
        +    "properties": {
        +      "action": {
        +        "const": "settle",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/ClickElementArgs",
        +    "properties": {
        +      "action": {
        +        "const": "click_element",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/SetValueArgs",
        +    "properties": {
        +      "action": {
        +        "const": "set_value",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/WaitForElementArgs",
        +    "properties": {
        +      "action": {
        +        "const": "wait_for_element",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "$ref": "#/$defs/ScrollToElementArgs",
        +    "properties": {
        +      "action": {
        +        "const": "scroll_to_element",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "action"
        +    ],
        +    "type": "object"
        +  }
        +]
      • changedInput schema / $defs / ClickArgs / properties / count / description
        Previous value: -"Consecutive clicks at this point (default 1); pass 2 for a double-click."New value: +"Consecutive clicks at this point (default 1, valid range 1 through 10); pass 2 for a\ndouble-click."
      • addedInput schema / $defs / ClickArgs / properties / count / maximum
        Added value: +10
      • changedInput schema / $defs / ClickArgs / properties / count / minimum
        Previous value: -0New value: +1
      • addedInput schema / $defs / ClickElementArgs
        Added value: +{
        +  "properties": {
        +    "id": {
        +      "description": "Element `#id` from the latest `glass_a11y_snapshot`.\nRe-snapshot after UI changes.\nPopover-owned targets route to their window and restore the prior active window.\nThe role-appropriate native accessibility operation handles occluded or off-screen targets\nand separate labels through their enclosing control.\nUnavailable native operations fall back to a pointer click at the target center.\nText editors may receive focus without activation.\n`method:\"native-action\"` labels any native path.\n`native_fallback` explains pointer fallback.\n`actuated_id` identifies a substituted enclosing control.",
        +      "format": "uint32",
        +      "minimum": 0,
        +      "type": "integer"
        +    },
        +    "return": {
        +      "description": "Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default).",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    }
        +  },
        +  "required": [
        +    "id"
        +  ],
        +  "type": "object"
        +}
      • changedInput schema / $defs / ScrollArgs / properties / dx / description
        Previous value: -"Horizontal scroll in **wheel notches** (discrete clicks — small integers like 1–5, NOT\npixels). Positive `dx` sends wheel-right, negative wheel-left; glass clicks `|dx|` times."New value: +"Horizontal wheel notches from -100 through 100, not pixels.\nPositive is right and negative is left, repeated `|dx|` times.\nTypical values are 1–5."
      • addedInput schema / $defs / ScrollArgs / properties / dx / maximum
        Added value: +100
      • addedInput schema / $defs / ScrollArgs / properties / dx / minimum
        Added value: +-100
      • changedInput schema / $defs / ScrollArgs / properties / dy / description
        Previous value: -"Vertical scroll in **wheel notches** (discrete clicks — small integers like 1–5, NOT\npixels). Positive `dy` sends wheel-down, negative wheel-up; glass clicks `|dy|` times. How\nan app maps a wheel notch to its view (lines, pixels, zoom) is the app's choice."New value: +"Vertical wheel notches from -100 through 100, not pixels.\nPositive is down and negative is up, repeated `|dy|` times.\nTypical values are 1–5.\nApps choose how a notch maps to lines, pixels, or zoom."
      • addedInput schema / $defs / ScrollArgs / properties / dy / maximum
        Added value: +100
      • addedInput schema / $defs / ScrollArgs / properties / dy / minimum
        Added value: +-100
      • addedInput schema / $defs / ScrollToElementArgs
        Added value: +{
        +  "properties": {
        +    "description": {
        +      "description": "Substring of the target element's accessible description (selector). This can select\nan unnamed control, including an Android text field labelled only by its hint.",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "direction": {
        +      "description": "Sweep direction: \"up\"/\"down\" (vertical) or \"left\"/\"right\" (horizontal).\nOmit to infer it from the target's off-screen position (falls back to a\nvertical down→up sweep when the target isn't in the a11y tree yet). The\nsearch reverses to the other end if the target isn't found first.",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "name": {
        +      "description": "Substring of the target element's accessible name (selector).",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "role": {
        +      "description": "Element role filter, e.g. \"ListItem\", \"Button\", \"Document\" (selector).",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "step": {
        +      "description": "Wheel notches per scroll step (default 3). A calibration escape hatch — larger\ncovers distance faster but risks stepping past a row's/column's realized band.",
        +      "format": "uint32",
        +      "minimum": 0,
        +      "type": [
        +        "integer",
        +        "null"
        +      ]
        +    },
        +    "timeout_ms": {
        +      "description": "Give up after this long (default 20000ms); returns `{matched:false}`.",
        +      "format": "uint64",
        +      "minimum": 0,
        +      "type": [
        +        "integer",
        +        "null"
        +      ]
        +    },
        +    "value_contains": {
        +      "description": "Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required.",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "x": {
        +      "description": "Scroll anchor x (window-relative). By default the swipe anchors on the target's\nown row/column (falling back to the window center if it isn't in the a11y tree\nyet); set both `x` and `y` to point the wheel at a specific container instead.",
        +      "format": "int32",
        +      "type": [
        +        "integer",
        +        "null"
        +      ]
        +    },
        +    "y": {
        +      "description": "Scroll anchor y (window-relative). See `x`.",
        +      "format": "int32",
        +      "type": [
        +        "integer",
        +        "null"
        +      ]
        +    }
        +  },
        +  "type": "object"
        +}
      • addedInput schema / $defs / SetValueArgs
        Added value: +{
        +  "properties": {
        +    "id": {
        +      "description": "The element `#id` from `glass_a11y_snapshot`.",
        +      "format": "uint32",
        +      "minimum": 0,
        +      "type": "integer"
        +    },
        +    "return": {
        +      "description": "Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (waits for visual\nstability and returns text-only metadata), or \"none\" (default).",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "text": {
        +      "description": "The value to set. For a text field, the text. For a spin/slider, a number.\nFor a switch/checkbox/toggle, a boolean (`\"true\"`/`\"false\"`/`\"on\"`/`\"off\"`/\n`\"1\"`/`\"0\"`) — idempotent. For a dropdown/combo box, an option label\n(case-insensitive); glass opens it and picks that option.",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "id",
        +    "text"
        +  ],
        +  "type": "object"
        +}
      • changedInput schema / $defs / SettleArgs / properties / timeout_ms / description
        Previous value: -"Give up after this long (default 5000ms); the sequence continues rather than\nfailing."New value: +"Give up after this long (default 5000ms).\nThis settle's timeout returns settled:false and completes the step.\nThe enclosing glass_do deadline fails the sequence."
      • changedInput schema / $defs / TypeArgs / properties / return / description
        Previous value: -"Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (wait for the UI\nto stop changing, text-only), or \"none\" (default). Not accepted inside a `glass_do`\n`type` action — use a `settle` action or the terminal `then` observe there."New value: +"Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default)."
      • addedInput schema / $defs / WaitForElementArgs
        Added value: +{
        +  "properties": {
        +    "condition": {
        +      "description": "What to wait for (default \"appears\"): appears|disappears|enabled|disabled|\nchecked|unchecked|selected|unselected|expanded|collapsed|focused|visible|hidden.\n`checked`/`unchecked` only match a checkable element (one exposing a real toggle\nstate) — a non-toggle element matches neither.",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "description": {
        +      "description": "Substring of the element's accessible description (selector). Useful for unnamed\ncontrols whose platform label is exposed as a hint, help text, or description.",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "interval_ms": {
        +      "description": "Poll interval (default 200ms — an a11y snapshot per tick).",
        +      "format": "uint64",
        +      "minimum": 0,
        +      "type": [
        +        "integer",
        +        "null"
        +      ]
        +    },
        +    "name": {
        +      "description": "Substring of the element's accessible name (selector).",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "role": {
        +      "description": "Element role filter, e.g. \"Button\", \"ProgressBar\", \"Document\" (selector).",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "timeout_ms": {
        +      "description": "Give up after this long (default 10000ms); returns `{matched:false}`.",
        +      "format": "uint64",
        +      "minimum": 0,
        +      "type": [
        +        "integer",
        +        "null"
        +      ]
        +    },
        +    "value": {
        +      "description": "Exact case-sensitive accessible value, requiring another selector and excluding `value_contains`.",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    },
        +    "value_contains": {
        +      "description": "Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required.",
        +      "type": [
        +        "string",
        +        "null"
        +      ]
        +    }
        +  },
        +  "type": "object"
        +}
      • addedInput schema / properties / timeout_ms
        Added value: +{
        +  "description": "Overall sequence budget in milliseconds. Omit for 30000; valid range\n1..=120000. One absolute deadline is shared by all actions and terminal\nobservations.",
        +  "format": "uint64",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
    • Addedglass_find_elements
    • Changedglass_scroll6 fields changed
      • changedInput schema / properties / dx / description
        Previous value: -"Horizontal scroll in **wheel notches** (discrete clicks — small integers like 1–5, NOT\npixels). Positive `dx` sends wheel-right, negative wheel-left; glass clicks `|dx|` times."New value: +"Horizontal wheel notches from -100 through 100, not pixels.\nPositive is right and negative is left, repeated `|dx|` times.\nTypical values are 1–5."
      • addedInput schema / properties / dx / maximum
        Added value: +100
      • addedInput schema / properties / dx / minimum
        Added value: +-100
      • changedInput schema / properties / dy / description
        Previous value: -"Vertical scroll in **wheel notches** (discrete clicks — small integers like 1–5, NOT\npixels). Positive `dy` sends wheel-down, negative wheel-up; glass clicks `|dy|` times. How\nan app maps a wheel notch to its view (lines, pixels, zoom) is the app's choice."New value: +"Vertical wheel notches from -100 through 100, not pixels.\nPositive is down and negative is up, repeated `|dy|` times.\nTypical values are 1–5.\nApps choose how a notch maps to lines, pixels, or zoom."
      • addedInput schema / properties / dy / maximum
        Added value: +100
      • addedInput schema / properties / dy / minimum
        Added value: +-100
    • Changedglass_scroll_to_element4 fields changed
      • addedInput schema / properties / description
        Added value: +{
        +  "description": "Substring of the target element's accessible description (selector). This can select\nan unnamed control, including an Android text field labelled only by its hint.",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / name / description
        Previous value: -"Substring of the target element's accessible name (selector). `name` and/or\n`role` is required."New value: +"Substring of the target element's accessible name (selector)."
      • changedInput schema / properties / role / description
        Previous value: -"Element role filter, e.g. \"ListItem\", \"Button\" (selector)."New value: +"Element role filter, e.g. \"ListItem\", \"Button\", \"Document\" (selector)."
      • changedInput schema / properties / value_contains / description
        Previous value: -"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name` and/or `role` is still required."New value: +"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required."
    • Changedglass_set_value1 field changed
      • changedInput schema / properties / return / description
        Previous value: -"Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (wait for the UI\nto stop changing, text-only), or \"none\" (default)."New value: +"Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (waits for visual\nstability and returns text-only metadata), or \"none\" (default)."
    • Changedglass_start1 field changed
      • changedInput schema / properties / run / description
        Previous value: -"What to launch, then its arguments. `run[0]` is the executable on a desktop backend, an\n`.app` path or bundle id on `ios`, and a `package/.Activity` component — optionally with\nan `.apk` to install first — on `android`. `run[1..]` are the app's own arguments;\n`android` has no argument vector to put them in and returns an error rather than\nignoring them."New value: +"What to launch: desktop `[executable, args...]`; iOS `[.app-or-bundle-id, args...]`;\nAndroid `[apk?, package/.Activity]` in either order, for example\n`[\"/absolute/path/app.apk\", \"com.example.app/.MainActivity\"]`."
    • Changedglass_type1 field changed
      • changedInput schema / properties / return / description
        Previous value: -"Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (wait for the UI\nto stop changing, text-only), or \"none\" (default). Not accepted inside a `glass_do`\n`type` action — use a `settle` action or the terminal `then` observe there."New value: +"Terminal observation: \"snapshot\" settles and refreshes/folds a11y, \"settle\" waits for\nvisual stability and returns text-only metadata, and \"none\" skips observation (default)."
    • Changedglass_wait_for_element4 fields changed
      • addedInput schema / properties / description
        Added value: +{
        +  "description": "Substring of the element's accessible description (selector). Useful for unnamed\ncontrols whose platform label is exposed as a hint, help text, or description.",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / role / description
        Previous value: -"Element role filter, e.g. \"Button\", \"ProgressBar\" (selector)."New value: +"Element role filter, e.g. \"Button\", \"ProgressBar\", \"Document\" (selector)."
      • addedInput schema / properties / value
        Added value: +{
        +  "description": "Exact case-sensitive accessible value, requiring another selector and excluding `value_contains`.",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • changedInput schema / properties / value_contains / description
        Previous value: -"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name` and/or `role` is still required."New value: +"Additionally require the matched element's `value` to contain this substring.\nNot a standalone selector — `name`, `description`, and/or `role` is still required."
  3. 26 tool updatesv1.2.0
    • Changedglass_a11y_snapshot2 fields changed
      • addedInput schema / $schema
        Added value: +"https://json-schema.org/draft/2020-12/schema"
      • addedInput schema / properties / max_nodes
        Added value: +{
        +  "description": "Maximum number of elements to include. Omit for the default cap (protects the token\nbudget). Pass a larger number to raise it, or `0` for the full tree (no limit). A\nsnapshot renumbers ids, so re-read them after changing this.",
        +  "format": "uint32",
        +  "minimum": 0,
        +  "type": [
        +    "integer",
        +    "null"
        +  ]
        +}
    • Changedglass_baseline_save2 fields changed
      • addedInput schema / properties / name / description
        Added value: +"Name to file this baseline under, reused by `glass_diff` and\n`glass_wait_for_region`. ASCII letters, digits, `-` and `_` only; saving over\nan existing name replaces it without warning."
      • removedInput schema / title
        Removed value: -"BaselineSaveArgs"
    • Addedglass_capabilities
    • Changedglass_click4 fields changed
      • addedInput schema / properties / count / description
        Added value: +"Consecutive clicks at this point (default 1); pass 2 for a double-click."
      • addedInput schema / properties / x / description
        Added value: +"Click point x, window-relative — 0 is the window's left edge, not the screen's."
      • addedInput schema / properties / y / description
        Added value: +"Click point y, window-relative — 0 is the window's top edge, not the screen's."
      • removedInput schema / title
        Removed value: -"ClickArgs"
    • Changedglass_click_element2 fields changed
      • changedInput schema / properties / id / description
        Previous value: -"The element `#id` from `glass_a11y_snapshot`. Valid only within the latest\nsnapshot — re-snapshot if the UI changed. If the element actually renders in a\npopover owned by a different window than the active one (e.g. an open\ndropdown's option row), the click is automatically routed into that popover\nwindow and the previously-active window is restored afterward — no extra step\nneeded."New value: +"The element `#id` from `glass_a11y_snapshot`. Valid only within the latest\nsnapshot — re-snapshot if the UI changed. If the element actually renders in a\npopover owned by a different window than the active one (e.g. an open\ndropdown's option row), the click is automatically routed into that popover\nwindow and the previously-active window is restored afterward — no extra step\nneeded.\n\nClicks via the platform's native accessibility action when the element exposes\none (works even when the element is occluded or scrolled off-screen), falling\nback to a synthetic pointer click at the element's center; the result's\n`method` field says which path ran, and `native_fallback` says why when the\npointer path was used. Where a control's label is a separate element from the\ncontrol itself, the native action fires on the enclosing control and the\nresult carries `actuated_id` — the element actually clicked."
      • removedInput schema / title
        Removed value: -"ClickElementArgs"
    • Changedglass_clipboard_set1 field changed
      • removedInput schema / title
        Removed value: -"ClipboardSetArgs"
    • Changedglass_diff6 fields changed
      • addedInput schema / $defs / RegionArgs / properties / height / description
        Added value: +"Height in pixels, extending down from `y`."
      • addedInput schema / $defs / RegionArgs / properties / width / description
        Added value: +"Width in pixels, extending right from `x`."
      • addedInput schema / $defs / RegionArgs / properties / x / description
        Added value: +"Left edge in pixels, window-relative — 0 is the window's left edge, not the screen's."
      • addedInput schema / $defs / RegionArgs / properties / y / description
        Added value: +"Top edge in pixels, window-relative — 0 is the window's top edge, not the screen's."
      • addedInput schema / properties / name / description
        Added value: +"Name of a baseline saved by `glass_baseline_save`; an unsaved name errors\nrather than reporting no change."
      • removedInput schema / title
        Removed value: -"DiffArgs"
    • Changedglass_do31 fields changed
      • addedInput schema / $defs / ClickArgs / properties / count / description
        Added value: +"Consecutive clicks at this point (default 1); pass 2 for a double-click."
      • addedInput schema / $defs / ClickArgs / properties / x / description
        Added value: +"Click point x, window-relative — 0 is the window's left edge, not the screen's."
      • addedInput schema / $defs / ClickArgs / properties / y / description
        Added value: +"Click point y, window-relative — 0 is the window's top edge, not the screen's."
      • addedInput schema / $defs / DiffArgs / properties / name / description
        Added value: +"Name of a baseline saved by `glass_baseline_save`; an unsaved name errors\nrather than reporting no change."
      • addedInput schema / $defs / DragArgs / properties / button / description
        Added value: +"Button held for the drag: \"left\" (default), \"right\", or \"middle\"."
      • addedInput schema / $defs / DragArgs / properties / x1 / description
        Added value: +"Press-point x, window-relative — 0 is the window's left edge, not the screen's."
      • addedInput schema / $defs / DragArgs / properties / x2 / description
        Added value: +"Release-point x, window-relative."
      • addedInput schema / $defs / DragArgs / properties / y1 / description
        Added value: +"Press-point y, window-relative — 0 is the window's top edge, not the screen's."
      • addedInput schema / $defs / DragArgs / properties / y2 / description
        Added value: +"Release-point y, window-relative."
      • addedInput schema / $defs / MoveArgs / properties / x / description
        Added value: +"Destination x, window-relative — 0 is the window's left edge, not the screen's."
      • addedInput schema / $defs / MoveArgs / properties / y / description
        Added value: +"Destination y, window-relative — 0 is the window's top edge, not the screen's."
      • addedInput schema / $defs / RegionArgs / properties / height / description
        Added value: +"Height in pixels, extending down from `y`."
      • addedInput schema / $defs / RegionArgs / properties / width / description
        Added value: +"Width in pixels, extending right from `x`."
      • addedInput schema / $defs / RegionArgs / properties / x / description
        Added value: +"Left edge in pixels, window-relative — 0 is the window's left edge, not the screen's."
      • addedInput schema / $defs / RegionArgs / properties / y / description
        Added value: +"Top edge in pixels, window-relative — 0 is the window's top edge, not the screen's."
      • addedInput schema / $defs / ScrollArgs / properties / x / description
        Added value: +"Pointer x the wheel is aimed at, window-relative — apps scroll the container\nunder this point, so it selects which pane moves."
      • addedInput schema / $defs / ScrollArgs / properties / y / description
        Added value: +"Pointer y the wheel is aimed at, window-relative. See `x`."
      • addedInput schema / $defs / SettleArgs / properties / interval_ms / description
        Added value: +"How long to wait between capture ticks (default 100ms)."
      • addedInput schema / $defs / SettleArgs / properties / settle_frames / description
        Added value: +"Consecutive unchanged frames required before the UI counts as settled (default 3)."
      • addedInput schema / $defs / SettleArgs / properties / stability_region / description
        Added value: +"Window-relative sub-rectangle to watch for settling; when set, changes outside\nit are ignored."
      • addedInput schema / $defs / SettleArgs / properties / timeout_ms / description
        Added value: +"Give up after this long (default 5000ms); the sequence continues rather than\nfailing."
      • addedInput schema / $defs / SettleArgs / properties / tolerance / description
        Added value: +"Per-channel difference (0–255) two frames may have and still count as\nunchanged (default 0, exact match)."
      • addedInput schema / $defs / ThenArgs / properties / diff / description
        Added value: +"Compare against a saved baseline and return change stats as text."
      • addedInput schema / $defs / ThenArgs / properties / screenshot / description
        Added value: +"Capture the window as an image — the only field here that always spends image\ntokens; prefer `diff` when you just need to know whether something changed."
      • addedInput schema / $defs / ThenArgs / properties / settle / description
        Added value: +"Wait for the UI to stop changing first. Set this whenever `diff` or\n`screenshot` follows, or they observe a half-drawn frame."
      • addedInput schema / $defs / TypeArgs / properties / return
        Added value: +{
        +  "description": "Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (wait for the UI\nto stop changing, text-only), or \"none\" (default). Not accepted inside a `glass_do`\n`type` action — use a `settle` action or the terminal `then` observe there.",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • addedInput schema / $defs / TypeArgs / properties / text / description
        Added value: +"Text to type into whatever currently has keyboard focus — this tool does not\nfocus a field, so click or `glass_click_element` one first. Sent as synthetic\nkey events, not pasted, so an app's per-keystroke handlers run."
      • removedInput schema / description
        Removed value: -"Arguments for `glass_do`: an ordered, non-empty action sequence + optional observe."
      • addedInput schema / properties / actions / description
        Added value: +"Actions to run in order; must be non-empty. Fail-fast — the first failing\naction aborts the rest and reports its index, so a partial sequence may\nalready have landed."
      • addedInput schema / properties / then / description
        Added value: +"Optional observe run once after the last action, in the order settle → diff\n→ screenshot."
      • removedInput schema / title
        Removed value: -"DoArgs"
    • Changedglass_doctor1 field changed
      • removedInput schema / title
        Removed value: -"DoctorArgs"
    • Changedglass_drag6 fields changed
      • addedInput schema / properties / button / description
        Added value: +"Button held for the drag: \"left\" (default), \"right\", or \"middle\"."
      • addedInput schema / properties / x1 / description
        Added value: +"Press-point x, window-relative — 0 is the window's left edge, not the screen's."
      • addedInput schema / properties / x2 / description
        Added value: +"Release-point x, window-relative."
      • addedInput schema / properties / y1 / description
        Added value: +"Press-point y, window-relative — 0 is the window's top edge, not the screen's."
      • addedInput schema / properties / y2 / description
        Added value: +"Release-point y, window-relative."
      • removedInput schema / title
        Removed value: -"DragArgs"
    • Changedglass_gesture3 fields changed
      • addedInput schema / $defs / PointArg / properties / x / description
        Added value: +"Window-relative x — 0 is the window's left edge, not the screen's."
      • addedInput schema / $defs / PointArg / properties / y / description
        Added value: +"Window-relative y — 0 is the window's top edge, not the screen's."
      • removedInput schema / title
        Removed value: -"GestureArgs"
    • Changedglass_key1 field changed
      • removedInput schema / title
        Removed value: -"KeyArgs"
    • Changedglass_logs4 fields changed
      • addedInput schema / properties / contains / description
        Added value: +"Return only lines containing this substring (case-sensitive). Filtering\nhappens server-side, so it narrows what the cap applies to."
      • addedInput schema / properties / cursor / description
        Added value: +"Resume point — the `cursor` a previous call returned, to read only what has\nbeen logged since. Omit to read from the oldest buffered line."
      • addedInput schema / properties / max_lines / description
        Added value: +"Cap on lines returned (default 200); the returned `cursor` resumes at the\nfirst line left unread, so a capped read is not a lost one."
      • removedInput schema / title
        Removed value: -"LogsArgs"
    • Addedglass_move
    • Addedglass_screenshot
    • Changedglass_scroll3 fields changed
      • addedInput schema / properties / x / description
        Added value: +"Pointer x the wheel is aimed at, window-relative — apps scroll the container\nunder this point, so it selects which pane moves."
      • addedInput schema / properties / y / description
        Added value: +"Pointer y the wheel is aimed at, window-relative. See `x`."
      • removedInput schema / title
        Removed value: -"ScrollArgs"
    • Changedglass_scroll_to_element1 field changed
      • removedInput schema / title
        Removed value: -"ScrollToElementArgs"
    • Changedglass_select_window1 field changed
      • removedInput schema / title
        Removed value: -"SelectWindowArgs"
    • Addedglass_set_value
    • Changedglass_start5 fields changed
      • addedInput schema / properties / cwd / description
        Added value: +"Working directory for both `build` and the launched app; omit to inherit the\nserver's own."
      • changedInput schema / properties / env / description
        Previous value: -"Extra environment variables for the launched app, as a `{ \"KEY\": \"VALUE\" }` object."New value: +"Extra environment variables, as a `{ \"KEY\": \"VALUE\" }` object. They reach the launched app\non the desktop backends and on `ios`; on `android` they configure the `build` command on\nthe host only, since an app launched by `am start` is forked from zygote and never sees\nthe shell's environment."
      • changedInput schema / properties / run / description
        Previous value: -"Program and arguments to launch; `run[0]` is the executable."New value: +"What to launch, then its arguments. `run[0]` is the executable on a desktop backend, an\n`.app` path or bundle id on `ios`, and a `package/.Activity` component — optionally with\nan `.apk` to install first — on `android`. `run[1..]` are the app's own arguments;\n`android` has no argument vector to put them in and returns an error rather than\nignoring them."
      • addedInput schema / properties / timeout_ms / description
        Added value: +"How long to wait for the app's window to appear before failing the launch\n(default 10000ms). Does not bound `build`."
      • removedInput schema / title
        Removed value: -"StartArgs"
    • Changedglass_type3 fields changed
      • addedInput schema / properties / return
        Added value: +{
        +  "description": "Optional observe folded into the result: \"snapshot\" (wait for the UI to settle, then\nfold a fresh a11y tree, also refreshing the snapshot cache), \"settle\" (wait for the UI\nto stop changing, text-only), or \"none\" (default). Not accepted inside a `glass_do`\n`type` action — use a `settle` action or the terminal `then` observe there.",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / text / description
        Added value: +"Text to type into whatever currently has keyboard focus — this tool does not\nfocus a field, so click or `glass_click_element` one first. Sent as synthetic\nkey events, not pasted, so an app's per-keystroke handlers run."
      • removedInput schema / title
        Removed value: -"TypeArgs"
    • Addedglass_wait_for_element
    • Addedglass_wait_for_log
    • Addedglass_wait_for_region
    • Addedglass_wait_stable
    • Addedglass_window
  4. 13 tool updatesv1.1.0
    • Addedglass_a11y_marks
    • Removedglass_capabilities
    • Addedglass_clipboard_get
    • Addedglass_diff
    • Addedglass_do
    • Addedglass_key
    • Removedglass_screenshot
    • Addedglass_scroll_to_element
    • Addedglass_select_window
    • Removedglass_set_value
    • Addedglass_start
    • Removedglass_wait_for_log
    • Removedglass_wait_for_region
  5. 20 tool updatesv1.0.3
    • Addedglass_a11y_snapshot
    • Addedglass_baseline_save
    • Addedglass_capabilities
    • Addedglass_click
    • Addedglass_click_element
    • Removedglass_clipboard_get
    • Removedglass_diff
    • Removedglass_do
    • Removedglass_key
    • Addedglass_list_windows
    • Removedglass_move
    • Addedglass_screenshot
    • Addedglass_scroll
    • Removedglass_scroll_to_element
    • Removedglass_select_window
    • Addedglass_stop
    • Addedglass_type
    • Removedglass_wait_for_element
    • Addedglass_wait_for_log
    • Addedglass_wait_for_region
  6. 9 tool updatesv1.0.3
    • Addedglass_clipboard_get
    • Addedglass_clipboard_set
    • Removedglass_list_windows
    • Addedglass_logs
    • Removedglass_screenshot
    • Addedglass_scroll_to_element
    • Removedglass_start
    • Removedglass_wait_for_log
    • Removedglass_window
  7. 12 tool updatesv1.0.2
    • Addedglass_diff
    • Addedglass_do
    • Addedglass_gesture
    • Addedglass_key
    • Addedglass_list_windows
    • Removedglass_logs
    • Addedglass_move
    • Removedglass_scroll_to_element
    • Addedglass_select_window
    • Addedglass_set_value
    • Addedglass_start
    • Removedglass_type
  8. 9 tool updatesv1.0.1
    • First observedglass_doctor
    • First observedglass_drag
    • First observedglass_logs
    • First observedglass_screenshot
    • First observedglass_scroll_to_element
    • First observedglass_type
    • First observedglass_wait_for_element
    • First observedglass_wait_for_log
    • First observedglass_window

TDQS

A3.7/5.0

Scored across 31 tools

Disambiguation4/5

Most tools are clearly separated by target and action: click vs click_element, type vs set_value, and the wait_for_* variants are distinguishable. A few pairs (screenshot/a11y_snapshot, click/click_element, wait_stable/wait_for_region) could still be mis-selected by an agent skimming descriptions, but each has a precise purpose.

Naming Consistency3/5

All names share the glass_ prefix and snake_case, but the pattern is mixed: many are verb_noun (select_window, wait_for_element) while others are bare nouns (logs, doctor), verbs (click, type), or reversed noun_verb (baseline_save, clipboard_get). Still readable and navigable, but not a consistent verb_noun convention.

Tool Count2/5

31 tools crosses the 25+ threshold and is more than double the well-scoped range. The granularity (several wait variants, clipboard get/set, separate click vs click_element) may be useful, but the surface is heavy and could be consolidated.

Completeness4/5

The set covers the full UI-automation lifecycle: start/stop, window management, input, semantic and pixel observation, baselines, logs, clipboard, and diagnostics. Minor gaps like no direct hover/double-click convenience or element-attribute introspection are easily worked around with move/click/wait, so there are no dead ends.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers