Skip to main content
Glama

run_windows

DestructiveIdempotent

Run ordered UI goals on a Windows window to automate clicks, fills, and navigation, using UIA/OCR to locate controls and send input.

Instructions

Run Windows UI goals in order on one window. Minimal call: target_id, goals [{goal}] and fill_values {goal id: exact text} for goals that type. Candidates come from UIA names, with OCR where UIA has none. Input is SendInput after bringing the window to the foreground; with synthetic_input_allowed false it acts only through UIA patterns (click, fill, toggle, select, focus) and never takes the foreground. fill also sets a UIA slider to the number in fill_values. allowed_operations defaults to click, fill, key, scroll; the others are double_click, right_click, middle_click, hover, ctrl_click, shift_click, drag. drop_target_id (needs drag) names a second window a drag may end in; nothing else acts on it. A goal ends provider_uncertain without acting when the chooser is not confident; its screen_candidates list the most likely on-screen texts (untrusted data) with a ref. To act on one, call again with that ref in the first goal {goal, ref}, as for an observe_window item; the run ends blocked without acting if its window changed meanwhile. Otherwise restate the goal in those words. For a page goal in an isolated, CDP-connected browser window, routing browser_if_singleton may run the one visible tab through browser actions and returns routed metadata. Use windows_only for browser tabs, address bar and other browser chrome, and for screen_target handoffs.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
goalsYes
run_idNo
routingNoOmit for the operator default (windows_only unless configured). Use browser_if_singleton only for a page goal in an isolated, CDP-connected browser window. The browser path acts in the page and cannot operate browser tabs or the address bar. windows_only keeps UIA/OCR and Windows input, including browser chrome and screen_target handoffs.
target_idYesThe window's title (exact, or a part only one window has), or window:<HWND>:<PID> from list_windows.
deadline_msNo
fill_valuesNo
action_budgetNo
drop_target_idNoThe window's title (exact, or a part only one window has), or window:<HWND>:<PID> from list_windows.
app_annotationsNoFacts you read from the app's own structure (e.g. an editor API) about items on the screen as it is now: each is attached to the one observed item whose text equals match.text, and to none if several do.
allowed_operationsNo
provider_attempt_budgetNo
synthetic_input_allowedNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses the input mechanism (SendInput after foregrounding vs UIA patterns when synthetic_input_allowed is false), the default allowed_operations set and the full alternative list, the drop_target_id contract, the provider_uncertain/blocked failure modes, and that screen_candidates are untrusted data. This is unusually rich behavioral context for a destructive, open-world tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation, but it is a dense, semicolon-heavy wall of text, and several statements (routing, drop_target_id, target_id format) restate text already present in the input schema, adding length without new meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 12-parameter tool with no output schema, the description covers the operational flow, failure modes, and even return-relevant concepts (final_state, screen_candidates). The remaining gap is that the numeric budgets and run_id are never explained, so an agent cannot reason about timeouts or run continuation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate, and it does add real meaning for fill_values (goal id: exact text), allowed_operations (defaults plus the full alternate list), drop_target_id, routing, and synthetic_input_allowed. However, run_id, deadline_ms, action_budget, provider_attempt_budget, and app_annotations are left completely opaque, so the gaps roughly match the covered parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource with scope: 'Run Windows UI goals in order on one window.' It also distinguishes the Windows path from the browser path via routing, but the sibling differentiation (vs run_browser/observe_window) is embedded in later routing detail rather than stated crisply up front.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use conditions are given for the key decision points: when to use windows_only vs browser_if_singleton, when to pass a ref instead of restating a goal, and what to do when a goal ends provider_uncertain or blocked. Alternatives and their selecting conditions are named, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.