grok-computer-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| observeA | Observe the GUI: interactive elements with refs (e7) and [x,y,w,h] boxes in image pixels. mode auto returns the tree, or a Set-of-Mark screenshot when the tree is thin; screenshot/som attach an image. root_ref expands a container; scope='screen' lists windows for cross-app work. Start every step with this. |
| clickA | Click an element: prefer ref, then mark (from a som observation), then point (image pixels, last resort). Pass the observation_id you act on. Returns what changed and the next observation_id. |
| type_textA | Type text into the field ref, or the focused field. Refused for password and credential fields: never type secrets. clear_first empties the field first. |
| press_keysA | Press a key chord such as 'cmd+s', 'enter' or 'tab', or a list of chords in order. Quit, log-out, lock, task-manager and run-dialog shortcuts are refused. |
| scrollB | Scroll the window, or at a ref/mark/point, by wheel notches. |
| dragA | Drag from one ref/mark/point to another in the same observation. |
| appsB | list apps, list windows (optionally of one app), launch an app in the background, or focus (bring to front) an app. |
| wait_forA | Wait until an element whose label or value contains text (optionally of a role) appears, or disappears with gone=true. Returns a fresh observation. |
| locateA | Find a described target in the current screenshot with a grounding model when refs and marks fail. Returns candidate points in image pixels; when not confident, up to 3 candidates and an annotated screenshot. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 9 tools
Each tool targets a clearly distinct action: observe for perception, wait_for for synchronization, locate for model-based grounding, and click/type_text/press_keys/scroll/drag for distinct input primitives. The observe/wait_for/locate trio is well-differentiated by intent and fallback role, so an agent can reliably pick the right one.
Mostly consistent lower_snake_case, with verb_noun forms (wait_for, type_text, press_keys) and clean single verbs (observe, click, scroll, drag, locate). The lone noun-style name 'apps' deviates slightly, but it is readable and unambiguous.
Nine tools is well-scoped for a GUI automation server, with each input modality (click, type, keys, scroll, drag) and perception path (observe, wait_for, locate) earning its place. No redundant or filler tools.
The surface covers perception, waiting, all major input primitives, app lifecycle, and a grounding fallback, which is strong for GUI control. Minor gaps remain, such as explicit double-click/hover, clipboard, or screenshot-only capture, though observe's som/screenshot modes partially mitigate this.