grok-computer-mcp
Enables computer use on Linux desktops: the server observes the screen (accessibility tree first, Set-of-Mark screenshots as fallback), and clicks, types, scrolls and drags in native desktop apps and browsers, with window offsets and multiple displays handled internally.
Enables computer use on macOS desktops: the server observes the screen (accessibility tree first, Set-of-Mark screenshots as fallback), and clicks, types, scrolls and drags in native desktop apps and browsers, with HiDPI scaling and multi-display coordinates handled internally.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@grok-computer-mcpopen Finder and verify my Downloads folder shows the latest file"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Grok Computer
English | 简体中文
Computer use for Grok Build: a Grok Build plugin (computer-use) plus a purpose-built MCP server (the grok-computer-mcp facade) that let Grok Build see the screen, click and type in desktop apps and browsers on macOS, Windows and Linux. The main use is checking UI inside the development loop (change code → launch the app → look at the UI → change again), and operating desktop and web tools that have no API.
Design document: docs/goal.md. Installation and configuration: docs/install.md. The documents under docs/ are written in Chinese.
How it works
main agent ──delegates──▶ computer subagent (own context; only reports go back)
│
├─▶ computer: grok-computer-mcp facade ──▶ Cua Driver ──▶ desktop (host / sandbox)
└─▶ browser: Playwright MCP (--isolated)
hooks: guard (risk grading and confirmation) · audit · verify_gate (report and final check) · selfcheckScreenshots stay out of the main session: all GUI work goes to the
computer-use:computersubagent; the main agent only receives a structured report of at most 2 KB (STATUS / SUMMARY / EVIDENCE / BLOCKERS / NEXT).Accessibility tree first, pixels last: an observation is a compact list of elements by default (one interactive element per line, with a stable ref). When the tree is not enough it falls back to a numbered Set-of-Mark screenshot, and only then to a grounding model or raw coordinates.
One coordinate space: every coordinate the model sees or sends is a pixel of the current observation image (long edge 1280 px). HiDPI scaling, window offsets and multiple displays are all handled inside the facade, and an action based on an observation of a screen that has since changed is refused (
STALE_OBSERVATION).Two safety layers: the hooks ask you before high-risk actions (send, delete, pay, ...), and the facade enforces hard rules itself and refuses everything when its policy cannot be loaded (deny-listed apps, secure fields, dangerous key chords, a session lock, an emergency stop).
The facade exposes nine tools: observe, click, type_text, press_keys, scroll, drag, apps, wait_for and locate.
Related MCP server: Universal Computer Control MCP
Quick start
curl -fsSL https://cua.ai/driver/install.sh | bash # desktop driver (macOS also needs permissions, see the install guide)
curl -LsSf https://astral.sh/uv/install.sh | sh # runs the facade
grok plugin marketplace add bo-516/computer-use
grok plugin install computer-use --trust
grok plugin enable computer-use
uvx grok-computer-mcp@0.2.0 doctor # self-checkThen describe the task to grok in plain language, or use /computer open Settings, turn on dark mode and confirm it took effect.
Privacy: screenshots are sent to the model provider. Deny-listed apps (password managers, keychains, terminals, banking apps, ...) are never observed; traces stay on your machine and are deleted after 7 days by default. See section 6 of the install guide.
Repository layout
Path | Contents |
| Marketplace index and component catalog |
| The plugin: subagent, skill and platform notes, |
| The facade MCP server (Python 3.11+, published on PyPI, run with |
| Task set (60 Cua Bench variants), host-mode runner and metrics, grounding benchmark, Phase 0 probes; see eval/README.md |
| Design document, install guide, Phase 0 verification record, state file schema |
| hooklib sync, plugin index generator |
| Hook fixture tests and repository contract tests |
Development
Requires uv (it installs Python 3.11 when needed); the dev-loop task tests also need Node.js.
uv sync
uv run pytest # all tests (fake backend; no desktop or account needed)
uv run ruff check
uv run pyright
uv run --python 3.8 --no-project --with pytest==8.3.5 pytest tests/hooks # hooks on Python 3.8
uv run python scripts/sync_hooklib.py --check # the facade's hooklib copy matches the plugin
uv run python scripts/plugin_index.py --check # the component catalog is current
grok plugin validate plugins/computer-useConventions are in AGENTS.md. The plugin directory ships files only and the hooks use the standard library only; the facade is the only code that talks to Cua Driver; the safety rules (fail-closed facade, R3 never loosened, typed text redacted, ...) must not regress.
Status
The plugin, the facade, the task set, the grounding benchmark and the Phase 0 probe kit are implemented and tested against a fake backend and a fake Cua Driver. What still needs a real environment (running the Phase 0 probes and the grounding benchmark, running the task set on macOS, Windows and Linux machines) is listed in docs/goal.md §10.1 and docs/phase0-verification.md.
License
Available Tools
9 toolsappsApps and windowsBIdempotent
list apps, list windows (optionally of one app), launch an app in the background, or focus (bring to front) an app.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| action | Yes | ||
| observation_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=true, openWorldHint=false, and destructiveHint=false, painting a mixed but clear safety profile. The description adds that launch runs 'in the background' and focus brings to front, which is useful behavioral context beyond the annotations. However, it doesn't disclose side effects like whether launch can fail if already running or what states are required for focus.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence front-loads the action list without extraneous detail. It is efficient but could be better structured as bullet points for clarity with four distinct operations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained. But with 0% schema description coverage and three parameters, the description leaves gaps about when each parameter applies and the observation_id's role. It's minimally adequate for an agent that already knows the action enum, but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It says windows can be 'optionally of one app', implying the 'name' parameter filters results for the windows action, but doesn't clarify that name is ignored for list or required for focus/launch. With three parameters and no schema descriptions, the description adds some mapping but leaves ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names four specific operations (list apps, list windows, launch, focus) mapping to the action enum, clearly distinguishing this tool from siblings like click or type_text. It's a clear verb+resource statement, though it doesn't explicitly contrast with any specific sibling beyond what the domain implies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to choose this tool over observe (which likely also lists windows) or when to use launch vs focus. The parenthetical '(optionally of one app)' hints at the windows filter but doesn't state the required parameters or conditions for each action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickClickADestructive
Click an element: prefer ref, then mark (from a som observation), then point (image pixels, last resort). Pass the observation_id you act on. Returns what changed and the next observation_id.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| mark | No | ||
| count | No | ||
| point | No | ||
| button | No | left | |
| modifiers | No | ||
| observation_id | Yes | observation_id of the observation you act on |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and non-idempotent, so the agent knows this mutates state. The description adds that it returns what changed and the next observation_id, which is behavioral context beyond the annotations. However, it doesn't elaborate on side effects such as whether the click might trigger navigation, downloads, or other destructive results beyond the generic hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the action and target selection rule, followed by a required argument and return value. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A complex, destructive, multi-mode tool with an output schema and seven parameters. The description covers the essential selection logic and the crucial observation_id requirement, and defers return details to the output schema. It omits behavior for count/button/modifiers, but for a tool at this complexity level, most of the critical guidance is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 14%, so the description must compensate. It explains the priority and origin of ref, mark, and point ('from a som observation', 'image pixels, last resort'), and says to pass the observation_id you act on. It doesn't mention count, button, or modifiers, which is a gap, but the most critical mutually-exclusive parameters are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Click) on a specific resource (an element) and immediately lists the three targeting modes with a priority order. It's distinguishable from siblings like type_text or drag because it names the element-targeting mechanisms.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent the preference order: 'prefer ref, then mark (from a som observation), then point (image pixels, last resort).' This is exactly the kind of when-to-use-which guidance that helps an agent choose among the mutually exclusive targeting parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dragDragADestructive
Drag from one ref/mark/point to another in the same observation.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Exactly one of ref, mark or point. | |
| from | Yes | Exactly one of ref, mark or point. | |
| observation_id | Yes | observation_id of the observation you act on |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the agent knows this mutates state. The description adds the same-observation scoping constraint but says nothing about reversibility, whether a drag triggers intermediate hover events, or what the side effects are beyond what annotations cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the action, both endpoints, and the scope constraint with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full output schema and fully documented parameters, the description only needs to frame the action and its scope, which it does. A note on drag side effects or timing would make it fully complete for a destructive gesture tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the nested TargetIn descriptions already explain ref, mark, point, and the 'exactly one of' rule. The description only echoes the accepted forms, adding no format or precedence guidance (e.g., which of ref/mark/point to prefer) beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (drag) and names the source/target forms (ref/mark/point), which clearly separates it from siblings like click or scroll. It does not explicitly name an alternative tool, so it falls short of the 5-level sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in the same observation' implies a constraint (no cross-observation drags) but never states when to prefer drag over click/scroll or what preconditions apply. Usage is only implied, not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
locateLocate by descriptionARead-onlyIdempotent
Find a described target in the current screenshot with a grounding model when refs and marks fail. Returns candidate points in image pixels; when not confident, up to 3 candidates and an annotated screenshot.
| Name | Required | Description | Default |
|---|---|---|---|
| description | Yes | What to find, e.g. 'the blue Export button' | |
| observation_id | Yes | observation_id of the observation you act on |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and open-world behavior, so the bar is lower. The description adds genuinely useful behavioral context beyond that: results are image-pixel points, and on low confidence it returns up to 3 candidates plus an annotated screenshot, which tells the agent how to interpret ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, no filler, with the primary action front-loaded and the fallback trigger and return behavior following. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though an output schema exists (so return values need not be explained), the description still covers the ambiguity path an agent needs to reason about. Combined with annotations covering safety, nothing required to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with only 2 required parameters, so the schema documents both fully and the per-parameter mentions in the description ('described target') add no syntax or format detail. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Find a described target in the current screenshot') and immediately scopes the mechanism ('with a grounding model') and its fallback role ('when refs and marks fail'). An agent can distinguish this from observe and click without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit usage condition — invoke when refs and marks fail — which implicitly positions it as the fallback to observe. It does not name the alternative tool by name, so it stops short of a full when/when-not/alternatives statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
observeObserve the screenARead-onlyIdempotent
Observe the GUI: interactive elements with refs (e7) and [x,y,w,h] boxes in image pixels. mode auto returns the tree, or a Set-of-Mark screenshot when the tree is thin; screenshot/som attach an image. root_ref expands a container; scope='screen' lists windows for cross-app work. Start every step with this.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App to observe; default: the last observed window or the frontmost app | |
| mode | No | auto: tree, or Set-of-Mark screenshot when the tree is thin | auto |
| page | No | Set-of-Mark page (80 marks each) | |
| scope | No | screen only for cross-app work (deny-listed apps masked) | window |
| root_ref | No | Expand this container: list all its descendants | |
| window_id | No | Window from apps(action='windows') | |
| max_elements | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, so the bar is lower. The description adds real behavior: auto mode falls back to a Set-of-Mark screenshot when the tree is thin, screenshot/som attach an image, and scope='screen' masks deny-listed apps. It does not mention pagination limits or output size caps, which would have pushed this higher.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences with no filler, and the core purpose plus the 'start every step with this' directive are front-loaded. The middle clauses about mode and refs are packed tightly enough to be slightly hard to scan, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the essential workflow positioning plus mode/scope/root_ref semantics for a 7-parameter tool. It leaves minor gaps around default app resolution and result size limits, but an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the baseline is 3, but the description adds coordinate-frame semantics ([x,y,w,h] in image pixels) and ref format (e7) that the schema does not state. Mode, scope, and root_ref behavior are largely restated from the schema, so the gain is incremental rather than comprehensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (observe) and resource (the GUI), then describes exactly what is returned: interactive elements with refs like e7 and [x,y,w,h] boxes in image pixels. This clearly separates it from action siblings such as click, type_text, and scroll, which operate on refs that observe produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit workflow directive ('Start every step with this') and conditions for specific modes (scope='screen' for cross-app work, root_ref to expand a container). It does not mention the locate sibling or explain when to prefer locate over observe, so it falls short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keysPress keysADestructive
Press a key chord such as 'cmd+s', 'enter' or 'tab', or a list of chords in order. Quit, log-out, lock, task-manager and run-dialog shortcuts are refused.
| Name | Required | Description | Default |
|---|---|---|---|
| keys | Yes | One chord like 'cmd+s' or 'enter', or a list of chords pressed in order | |
| observation_id | Yes | observation_id of the observation you act on |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and non-idempotent, so the safety profile is covered. The description adds genuinely new behavioral context beyond the annotations: the explicit set of refused shortcut classes, which tells the agent some invocations will be blocked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no waste; the core action and its examples are front-loaded and the constraint follows. Nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the annotations carry the mutation semantics. The description is complete enough for a 2-parameter action tool, with the only gap being no positive routing guidance versus typing tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents both the chord format and the in-order list behavior, so the description adds no syntax or format detail beyond it. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('press a key chord') with concrete examples ('cmd+s', 'enter', 'tab') and covers the multi-chord case. It is clearly distinct from a text-typing tool, though it never names the sibling (e.g. type_text) it could be confused with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The refusal list ('quit, log-out, lock, task-manager and run-dialog shortcuts are refused') gives an implicit when-not boundary, which is useful. However, there is no positive guidance on when to choose this over click/type_text, and no mention of prerequisites such as needing a prior observation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollScrollB
Scroll the window, or at a ref/mark/point, by wheel notches.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| mark | No | ||
| point | No | ||
| amount | No | Wheel notches | |
| direction | Yes | ||
| observation_id | Yes | observation_id of the observation you act on |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. 'Scroll' implies a write-like action (readOnlyHint=false), but the description does not explain side effects, required state, observation dependencies, or what happens if scroll bounds are reached. This is a significant gap for a UI interaction tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that is efficient and non-redundant. It states the core action and target options without waste, though it could benefit from slightly more structure for a multi-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters, 33% schema coverage, no annotations, and a UI interaction context, the description is inadequate. It omits direction, amount, observation_id semantics, default targeting logic, and behavioral constraints. An agent would need to infer much from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 33%, so many parameters lack formal descriptions. The description mentions 'ref/mark/point' and 'wheel notches', which maps to some parameters but omits direction, amount, observation_id, and the default behavior (window). It adds marginal value but leaves key parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource combination ('Scroll the window') and describes scope ('or at a ref/mark/point') and method ('by wheel notches'). It clearly differentiates from siblings like click or type_text. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied by the object/target description. However, there are no explicit when-to-use or when-not-to-use guidelines, and no alternatives are named. The schema defaults (window if no ref/mark/point) are only in the schema description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
type_textType textADestructive
Type text into the field ref, or the focused field. Refused for password and credential fields: never type secrets. clear_first empties the field first.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Field to type into; default: the focused element | |
| text | Yes | ||
| clear_first | No | Empty the field before typing | |
| observation_id | Yes | observation_id of the observation you act on |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavior beyond the annotations: a hard refusal on password/credential fields, the fact that clear_first destroys existing field content, and the focused-element default. The destructiveHint and non-idempotent hints are consistent with this, but return/error behavior is left to the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the safety constraint, then the modifier. Tight overall, though the clear_first sentence duplicates the schema description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with an output schema and 75% schema coverage, the description covers the essential risk (secrets refusal) and the overwrite behavior. The required observation_id is undocumented in prose but is fully specified in the schema, so no critical gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, with ref, clear_first, and observation_id already described in the schema. The description largely restates those (ref default = focused field, clear_first empties the field) and adds no new syntax or constraints, so it sits at the baseline for a well-covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (type) and resource (text into a field ref or the focused field), which is enough to separate it from click, press_keys, and drag. It does not explicitly name a sibling to contrast against, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the default target (focused field when ref is omitted) and gives a clear exclusion: it is refused for password and credential fields. No named alternatives (e.g. press_keys for special keys) are offered, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forWait for an elementARead-onlyIdempotent
Wait until an element whose label or value contains text (optionally of a role) appears, or disappears with gone=true. Returns a fresh observation.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | ||
| gone | No | Wait until the element disappears instead | |
| text | No | Wait for an element whose label or value contains this | |
| ref_role | No | Only elements of this role (e.g. 'button') | |
| window_id | No | ||
| timeout_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that the call yields a fresh observation and the gone=true semantics, but says nothing about blocking behavior, timeout expiry (the schema caps timeout_ms at 30000), or what happens when the wait fails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the primary appear-mode is front-loaded and the gone=true variant follows immediately, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return value need not be spelled out, and the description still notes it returns a fresh observation. Combined with annotations covering the safety profile, the definition is sufficient for an agent to call this correctly, with only minor gaps around timeout failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: text, gone and ref_role are documented in the schema, and the description reinforces text/role/gone. However, app, window_id and timeout_ms carry no descriptions anywhere, and the description does not compensate by explaining scoping or timeout semantics beyond the default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('wait until an element ... appears') plus the inverted mode via gone=true, so the core purpose is unambiguous. It does not differentiate itself from siblings like locate or observe, even though the 'returns a fresh observation' phrasing hints at overlap with observe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the wording: use it to wait for an element to appear or, with gone=true, to disappear. There is no explicit when-not guidance, no mention of locate/observe as alternatives, and no statement about what to do if the element never arrives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.2.0- First observed
apps - First observed
click - First observed
drag - First observed
locate - First observed
observe - First observed
press_keys - First observed
scroll - First observed
type_text - First observed
wait_for
TDQS
Scored across 9 tools
Each tool targets a clearly distinct action: observe for perception, wait_for for synchronization, locate for model-based grounding, and click/type_text/press_keys/scroll/drag for distinct input primitives. The observe/wait_for/locate trio is well-differentiated by intent and fallback role, so an agent can reliably pick the right one.
Mostly consistent lower_snake_case, with verb_noun forms (wait_for, type_text, press_keys) and clean single verbs (observe, click, scroll, drag, locate). The lone noun-style name 'apps' deviates slightly, but it is readable and unambiguous.
Nine tools is well-scoped for a GUI automation server, with each input modality (click, type, keys, scroll, drag) and perception path (observe, wait_for, locate) earning its place. No redundant or filler tools.
The surface covers perception, waiting, all major input primitives, app lifecycle, and a grounding fallback, which is strong for GUI control. Minor gaps remain, such as explicit double-click/hover, clipboard, or screenshot-only capture, though observe's som/screenshot modes partially mitigate this.
Maintenance
Related MCP Connectors
Securely control computers you explicitly pair through files, terminals, processes, screenshots, desktop UI/input, clipboard, browser automation, diagnostics, and document tools.
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables the model to observe and control the live desktop via accessibility trees and screenshots, performing actions like clicking, typing, scrolling, dragging, and setting values.MIT
- AlicenseBqualityAmaintenanceThis server enables AI agents to operate Windows and Linux computers by observing the screen, understanding the UI through accessibility, OCR, and vision, and performing mouse, keyboard, window, clipboard, and system actions with verification and recovery.291MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to observe and control the user's live desktop by listing apps, reading accessibility trees, and performing clicks, typing, scrolling, dragging, and other input actions with per-action approvals and local audit archives.1MIT
- AlicenseCqualityCmaintenanceEnables agent-native desktop control by driving mouse, keyboard, and interface elements through accessibility-first semantic operations, with fallback to screen coordinates when accessibility cannot reach a target.30MIT