linux-cu
Cinnamon is one of the desktop environments explicitly supported and tested as the target session the agent can control on X11.
Works with GTK applications on the X11 desktop: it reads their widgets via AT-SPI (buttons, fields and their coordinates) and draws the agent's pointer with a click-through GTK overlay.
Enables an AI agent to see and drive a Linux X11 desktop: taking screenshots, moving the mouse, clicking, dragging, scrolling, typing Unicode text and sending key combinations, and inspecting/acting on native widgets through the AT-SPI accessibility tree (list_windows, ui_tree, click_element, set_text). The agent works with its own virtual pointer and keyboard, so the user's cursor and focus are untouched, and it renders its own click-through agent cursor overlay.
Ubuntu (24.04) is one of the Linux distributions explicitly supported and tested as the target desktop environment the agent can control.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@linux-cuTake a screenshot, then click the File menu and select New Window"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
linux-computer-use
English | 한국어 | 日本語 | 简体中文 | Español
Computer use for Linux. An MCP server and skill that let Claude Code and Codex see and drive an X11 desktop: screenshots, mouse, keyboard and the accessibility tree. The agent gets its own virtual pointer and keyboard, so your mouse and focus stay yours while it works.

The agent (blue cursor) works in the left window through the MCP tools while you keep typing on the right. Recorded on a throwaway X display with scripts/record_demo.sh.
git clone https://github.com/flyingsquirrel0419/linux-computer-use ~/Documents/linux-computer-use
~/Documents/linux-computer-use/install.sh # registers the server + skill for Claude Code and Codex
claude mcp list # → linux-cu: … run lcu-supervised - ✔ ConnectedNo extra binaries. No xdotool or scrot: pure Python over XTEST, X Input 2 and AT-SPI.
Own pointer. A second X master pointer/keyboard per agent. Your cursor doesn't move and your focused window doesn't change.
Native widgets.
ui_treelists buttons and fields with coordinates;click_element/set_textact on them directly.Unicode typing. Korean, CJK and emoji work, including with an ibus Hangul input method active.
Codex-style cursor. A click-through overlay draws the agent's pointer with the Codex computer-use glyph and motion.
Survives crashes. A supervisor restarts the server and restores the MCP session, so the tools don't vanish mid-session.
Contents
Related MCP server: claude-linux-mcp
How it works
flowchart LR
host["Claude Code / Codex"] -- "MCP stdio" --> sup["lcu-supervised<br/>(restarts, replays init)"]
sup --> srv["lcu.server<br/>(18 tools)"]
srv -- "XTEST via own<br/>master pointer/keyboard" --> x11["X11 desktop"]
srv -- "AT-SPI" --> apps["GTK / Qt / Chromium apps"]
srv -- "positions, clicks" --> ov["cursor overlay<br/>(GTK, click-through)"]
ov --> x11Every coordinate the agent sends or receives is in the pixel space of the screenshot it was given (long edge 1280 px by default). The server converts to real pixels, so the model never rescales anything.
Requirements
OS / session | Linux with an X11 session (tested: Ubuntu 24.04, Cinnamon). Wayland is not supported. |
Python | System |
Tools |
|
Agents | Claude Code and/or Codex CLI; |
Quick start
Install (idempotent, safe to re-run after
git pull):git clone https://github.com/flyingsquirrel0419/linux-computer-use ~/Documents/linux-computer-use ~/Documents/linux-computer-use/install.shIt creates a venv that uses the system PyGObject, registers the
linux-cuMCP server (throughlcu-supervised) with Claude Code (user scope) and Codex (~/.codex/config.toml, backup inconfig.toml.bak-lcu), and symlinks the skill into~/.claude/skillsand~/.codex/skills.Verify:
claude mcp list | grep linux-cu # ✔ Connected codex mcp list | grep linux-cu # enabledUse it. Start a new Claude Code or Codex session (running sessions don't pick up new servers) and ask for something on screen, for example "Open gedit and write a short note in Korean." The agent takes a screenshot, finds the widgets, clicks, types and checks the result.
Tools
Group | Tools |
See |
|
Mouse |
|
Keyboard |
|
Accessibility |
|
Other |
|
Action tools accept
screenshot_after: trueto return the resulting screen in the same call.type_textandkeyacceptexpect_window, a substring of the target window's title or WM class. If the window that would receive the keys doesn't match, nothing is sent.ui_treereturns lines like[9] push button "Save" @(756,21). The ids stay valid until the nextui_treecall.
Virtual pointer
Like Codex computer use, the agent works with its own input devices. On its first action the server creates an X Input 2 master pointer/keyboard pair named lcu-<pid> and routes all of its XTEST input through it.
Your mouse never moves and your keyboard focus never changes, so you can keep working.
Agent clicks don't raise or activate windows. Agent keystrokes go to the window under the agent's pointer, which is why
click_elementandset_textfirst move the pointer onto the element.Each agent gets its own pair: run Claude Code and Codex together and you'll see two cursors. The pair is removed when the server exits, and pairs left by crashed servers are cleaned up on the next start.
screen_inforeports"virtual_pointer": truewhen this mode is on.
Limit: the agent can only act on what is visible on screen. Clicking where a window is covered hits the window on top.
Agent cursor
A click-through GTK overlay draws the agent's pointer:
Glyph: the Codex
AgentCursoroutline, 14 px. The hotspot is the glyph's centre, not its tip. It has a translucent gradient fill, a 1.55 px rim and a soft glow.Motion: a cubic path picked from 20 candidates, driven by a damped spring (damping 0.9). On moves of 196 px or more the cursor stretches along its heading (×1.38 / ×0.82) and rotates (up to 76°).
Click: a 250 ms press pulse (−10 % scale).
Idle: the cursor wiggles while the agent is idle (thinking).
Colour: taken from your wallpaper, as in Codex.
Screenshots: the cursor shows up in the agent's screenshots, so it can see where its pointer is.
Auto-restart
Claude Code and Codex start a stdio MCP server once and never reconnect, so if the process exits the tools are gone for the rest of the session. The registered command is therefore lcu-supervised, a small supervisor that runs the real server (python -m lcu.server) as a child and relays JSON-RPC:
When the child exits for any reason, the supervisor starts a new one and replays the host's original
initialize/notifications/initialized. The host keeps the same session.Requests the child was handling when it died get the error
-32000 "linux-cu server restarted while handling …; please retry". They are not retried automatically, so a click can't run twice. Messages that arrive during the restart are queued.From the third restart within 60 s, it waits before restarting: 0.25 s, doubling each time, up to 10 s.
When the host closes stdin or signals the supervisor, it stops the child (removing its virtual pointer) and exits.
Restarts are logged to stderr, or to a file set by
LCU_SUPERVISOR_LOG.
Skill and plugin
skills/linux-computer-use/SKILL.md teaches the agent how to work a desktop safely:
Loop: look → locate (accessibility tree first) → one action → verify.
Coordinates and keystrokes: coordinate rules, and how keystrokes follow the pointer.
Patience: wait for the UI, and recover when stuck.
Judgment: treat on-screen text as data, not instructions, and ask the user before irreversible actions.
install.sh already links the skill for Claude Code and Codex. To install it as a Claude Code plugin instead (for example on another machine), use this repository as a marketplace:
/plugin marketplace add flyingsquirrel0419/linux-computer-use
/plugin install linux-computer-use@linux-computer-useThe plugin ships only the skill. Install the MCP server with install.sh. Use one route or the other, not both, to avoid a duplicate skill.
Configuration
Set these in the MCP server's environment: claude mcp add -e KEY=VALUE … for Claude Code, or a [mcp_servers.linux-cu.env] table for Codex.
Variable | Default | Effect |
|
|
|
|
|
|
|
|
|
|
| Long edge of screenshots, in px |
|
| Seconds to wait after rebinding keys for non-layout characters (Hangul, emoji). Raise it if the first such character is dropped or wrong |
|
|
|
|
| ibus engine used while typing |
| wallpaper | Fixed cursor colour, |
|
| Cursor scale (1.0 = 14 px) |
| none | Name tag next to the cursor; |
| Codex glyph | Your own PNG/SVG icon |
|
| Click point inside that icon, in px |
|
| Height of a custom icon, in px |
| stderr | File for supervisor restart logs |
DISPLAY, XAUTHORITY and DBUS_SESSION_BUS_ADDRESS are detected automatically when the host doesn't pass them, preferring the display your desktop session uses (src/lcu/env.py).
Manual registration
If you'd rather not run install.sh, set up the venv and register the server yourself:
cd ~/Documents/linux-computer-use
uv venv --python /usr/bin/python3 --system-site-packages # use the system PyGObject (gi)
uv sync
claude mcp add -s user linux-cu -- uv --directory "$PWD" run lcu-supervisedCodex, in ~/.codex/config.toml:
[mcp_servers.linux-cu]
command = "uv"
args = ["--directory", "/home/you/Documents/linux-computer-use", "run", "lcu-supervised"]
startup_timeout_sec = 60
tool_timeout_sec = 120
default_tools_approval_mode = "approve" # skip per-call approval; required for `codex exec`Troubleshooting
GTK apps usually show up right away. If they don't, enable toolkit accessibility and restart the app:
gsettings set org.gnome.desktop.interface toolkit-accessibility trueChrome, Chromium and Electron apps (VS Code, Slack…) need --force-renderer-accessibility. Without it, use screenshots and coordinates.
screen_info includes an error field. Common causes: xinput isn't installed (sudo apt install xinput), or LCU_VIRTUAL_POINTER=0 is set. In this mode the agent moves your real mouse.
Characters missing from the keyboard layout are typed by briefly binding them to spare keycodes. If an app picks up keymap changes slowly, raise
LCU_REMAP_SETTLE(for example0.25).ibus Hangul mode: while typing, ibus is switched to
LCU_PLAIN_ENGINEand then back. ibus-hangul then starts again in itsinitial-input-mode, usually Latin.
Check that the registered command is lcu-supervised, not lcu (claude mcp get linux-cu). Re-running install.sh switches old registrations over. Sessions started before a change keep the old server until you start a new session.
Wayland blocks XTEST input and screen capture for other clients. Log in to an X11 session.
Uninstall
claude mcp remove -s user linux-cu
rm ~/.claude/skills/linux-computer-use ~/.codex/skills/linux-computer-use # symlinks onlyIn ~/.codex/config.toml, delete the [mcp_servers.linux-cu] table, or restore ~/.codex/config.toml.bak-lcu. If a crashed server left an agent pointer behind, remove it:
xinput list --short | grep lcu-
xinput remove-master "lcu-<pid> pointer"Safety
The agent operates your real desktop. Beyond the skill's guidance there are no built-in guardrails, and with
default_tools_approval_mode = "approve"Codex calls the tools without asking.To isolate the agent, run the server against a separate display (for example
Xvfb :99with a window manager) by settingDISPLAY=:99in its environment.Screenshots of your screen are sent to the model provider that the agent uses.
Development
uv run python scripts/smoke_mcp.py # read-only: lists tools, screen info, windows, saves a screenshot
DISPLAY=:99 uv run python scripts/smoke_mcp.py
uv run python scripts/smoke_mcp.py --direct # bypass the supervisor
dbus-run-session -- scripts/record_demo.sh # re-record docs/demo.gif on a throwaway displaySource layout: server.py (MCP tools), supervisor.py, vpointer.py (MPX pointer), input.py (XTEST, keymap), capture.py, a11y.py (AT-SPI), overlay.py / motion.py (cursor), ime.py, env.py.
Credits
The cursor glyph and motion model in src/lcu/motion.py are ported from maka-agent (Apache-2.0). That project recovered them from the Codex desktop app. See NOTICE. This project isn't affiliated with or endorsed by OpenAI or Anthropic.
License
Apache License 2.0. See NOTICE for attribution.
Available Tools
18 toolsactive_windowA
Where your keystrokes will go (keys_go_to: the window under your
pointer when you have your own pointer) and which window the user has
focused. Check this before typing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It helpfully explains the returned `keys_go_to` semantics (the window under your pointer when you own the pointer) and the focus state, implying a safe read query, but it never explicitly states read-only behavior, whether it has side effects, or any conditions under which the values are unavailable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence delivers the purpose, the meaning of the key return field, and the action guidance with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, yet the description usefully characterizes the primary field. With zero parameters and a focused, single-condition use case, the definition is nearly complete; only explicit behavioral traits are left uncovered by the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify; the baseline for a parameterless tool is 4. No parameter meaning is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (the active window) and clarifies two distinct facts it reports: where keystrokes go and which window the user has focused. It is distinguishable from a plain window enumerator like list_windows, though it never explicitly contrasts itself with that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Check this before typing" gives a clear, actionable usage condition tied to the type_text/key/hold_key siblings. It stops short of naming alternatives (e.g., list_windows) or stating when not to use it, so it is context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickB
Click at (x, y). count=2 for double-click, 3 for triple-click. modifiers e.g. ["ctrl"] or ["shift"] are held during the click.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| count | No | ||
| button | No | left | |
| modifiers | No | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose two useful traits beyond the schema: that count produces multi-click gestures and that modifiers are held during the click. It omits whether the cursor is moved to (x, y) first, whether the call blocks, and what screenshot_after actually does, so coverage is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and coordinate requirement, and every clause adds information. Nothing is padded, though it is too terse to cover the remaining undocumented parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must stand alone for a 6-parameter mutation tool. It adequately covers the click gesture but leaves screenshot_after, the button enum semantics, and whether coordinates are screen- or window-relative undefined, which matters in a multi-window desktop environment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must compensate. It meaningfully clarifies count (2=double, 3=triple) and modifiers (held during the click, with example values), but leaves button, screenshot_after, and the coordinate reference frame entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Click at (x, y)" names a specific action and its coordinate-based scoping, which separates it from element-based clicking. However it never names click_element (the obvious alternative for semantic element targeting) or explains how it differs from mouse_down/mouse_up, so the sibling differentiation is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains parameter values (count for double/triple click, modifiers) but gives no guidance on when to choose this tool over click_element, drag, or the mouse_down/mouse_up pair. With 17 siblings in a desktop-automation family, the absence of any routing or prerequisite guidance is a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
click_elementB
Activate a ui_tree element. auto = AT-SPI action if available, else a real mouse click at its center.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| method | No | auto | |
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the fallback chain for method=auto (AT-SPI action first, otherwise a real mouse click at the element center), which is real behavioral value. However it says nothing about side effects, focus changes, failure behavior when the element is not actionable, or what screenshot_after does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The core action is stated first and the method semantics follow immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter action tool with no annotations and no output schema, the definition covers the core mechanism but omits failure modes, permission/accessibility requirements, and the meaning of screenshot_after, leaving gaps an agent would need to resolve by trial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the 'auto' value of method but leaves 'action' and 'mouse' implicit, and never explains that 'id' is a ui_tree node id or what 'screenshot_after' does. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (activate) and resource (ui_tree element), and the parenthetical clarifies the mechanism. It does not explicitly contrast itself with the sibling 'click' tool, which operates on raw coordinates, so an agent must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool works on elements obtained from ui_tree but never states when to use it instead of 'click', 'set_text', or the mouse_* primitives. No prerequisites (e.g., 'id must come from ui_tree') or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cursor_positionA
Current mouse pointer position in screenshot coordinates.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It implies a read-only query by saying 'Current' and usefully discloses that coordinates are in the screenshot frame, but it does not state that it is side-effect-free or whether it requires a visible/active session.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, correctly sized for a zero-argument query tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and there are no parameters to document. The description is nearly complete for this simplicity, only missing explicit read-only framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty, so the baseline of 4 applies. There is nothing further the description could add on this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (mouse pointer position) and the coordinate frame it is reported in. That distinguishes it from siblings like mouse_move and click, though it does not name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given. An agent must infer that this is a read-side companion to mouse_move/click for verifying pointer location; no alternatives or conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dragC
Press at start, move smoothly to end, release.
| Name | Required | Description | Default |
|---|---|---|---|
| end_x | Yes | ||
| end_y | Yes | ||
| button | No | left | |
| start_x | Yes | ||
| start_y | Yes | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It hints that motion is interpolated ('move smoothly'), which is a genuine trait, but omits button state semantics, coordinate system/origin, timing/duration, and what screenshot_after does. For a low-level input-synthesis tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero filler that directly conveys the action sequence. It is well-sized for a terse gesture primitive, though its brevity overlaps with under-specification rather than pure economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter input tool with no annotations and no output schema, the description leaves too much unsaid: coordinate origin, units, button default behavior, and the side effect of screenshot_after. An agent lacks enough context to invoke it confidently without reasoning from sibling conventions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 6 parameters. The phrase 'start'/'end' loosely maps to the start_x/start_y and end_x/end_y pairs, but the description never mentions the button enum or screenshot_after, nor explains coordinates are integer screen pixels. It only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb sequence (press, move, release) that clearly defines a drag gesture on a resource (screen coordinates). An agent can tell it produces a drag rather than a click or a scroll, though it never explicitly distinguishes itself from the sibling composite tools mouse_down/mouse_up/mouse_move.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use drag versus composing mouse_down + mouse_move + mouse_up, nor versus click. No prerequisites, no context for choosing among the many pointer siblings. Usage must be fully inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hold_keyB
Hold key(s) down for seconds (max 10), e.g. for games or key repeat.
| Name | Required | Description | Default |
|---|---|---|---|
| combo | Yes | ||
| seconds | Yes | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the max duration (10 seconds) and implies automatic release after the given seconds, but does not explain failure modes, interruption behavior, or required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that includes the key constraint (max 10 seconds) and a usage example without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0% schema coverage for three parameters, the description is incomplete. It omits any explanation of `screenshot_after` and does not clarify whether keys are released automatically or what happens if the hold is interrupted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three parameters. It clarifies `seconds` (with a max of 10) and implies `combo` means key(s), but never mentions `screenshot_after`, leaving one parameter completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (hold key(s) down) and the duration parameter. Distinguishes from the sibling `key` implicitly by emphasizing holding rather than pressing, but does not explicitly name the alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides example contexts ('for games or key repeat') which imply when to use it, but does not explicitly compare with the sibling `key` or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
keyC
Press a key or chord: "Return", "ctrl+c", "ctrl+shift+t", "alt+F4", "super", "Escape", "Page_Down". X keysym names are accepted. expect_window works as in type_text.
| Name | Required | Description | Default |
|---|---|---|---|
| combo | Yes | ||
| repeat | No | ||
| expect_window | No | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. It does not say which window receives the keystroke, whether an app must be focused first, whether the call waits for the target to react, or what expect_window actually verifies — behavior is deferred to type_text's documentation rather than stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loads the core purpose before the example list, with no filler sentences. The example enumeration is long but each entry carries real formatting information, so it earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% schema coverage across four parameters, the description should explain repeat, screenshot_after, and the window-targeting precondition. It only partially covers combo and hand-waves expect_window, so an agent cannot call this confidently in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so all four parameters need the description's help. The description usefully documents the combo format with examples and mentions expect_window (though it outsources its meaning), but says nothing about 'repeat' or 'screenshot_after', leaving half the parameters undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Press a key or chord') and concrete examples ('ctrl+c', 'alt+F4', 'super') plus the accepted keysym format, which clearly delineates it from type_text and hold_key. It stops short of naming those siblings explicitly, so an agent must infer the boundary rather than being told.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance against siblings like type_text (for strings) or hold_key (for press-and-hold). The only routing signal is the implicit contrast between 'press a key' and typing text, plus a cross-reference that 'expect_window works as in type_text', which is a weak proxy for real guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_windowsB
Top-level windows known to AT-SPI: app, title, active/visible, box.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses the returned fields (app, title, active/visible, box), which tells the agent what each window entry contains, but says nothing about ordering, filtering, permissions, or whether the list can be empty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact clause with the resource front-loaded and the returned fields enumerated compactly. It is telegraphic but wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only listing tool with an output schema present, the description covers what is needed. The output schema handles the return structure, and the description adds the AT-SPI scope and a field preview, leaving only minor gaps like ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document beyond what the schema already covers; the baseline for a parameterless tool is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (top-level windows) plus the scope 'known to AT-SPI'. It implicitly distinguishes itself from the sibling active_window by returning all top-level windows rather than the focused one, but never names that alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives like active_window, ui_tree, or screen_info. The usage is only implied by the name and the word 'Top-level'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mouse_downA
Press and hold a mouse button at the current pointer position.
| Name | Required | Description | Default |
|---|---|---|---|
| button | No | left |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the core behavior (button stays held) but omits the key operational fact that a matching mouse_up is required to release, plus any error/state behavior if invoked repeatedly. Adequate for a simple primitive but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler; the action and the positional constraint are both stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and the input schema covers the button enum, so little is missing structurally. However, with no annotations and no mention of the required mouse_up pairing, the description under-specifies the lifecycle of the held button.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter ('button') with an enum and a default, and the description never mentions which button is pressed. The enum values and default are self-describing in the schema, so the description adds no meaning beyond it. Baseline 3 for an essentially self-documenting single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource: 'Press and hold a mouse button' with the scope qualifier 'at the current pointer position.' An agent can distinguish it from click (instant press+release) and mouse_up (the release half). It does not explicitly name or route against siblings, but the semantics are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'press and hold' hints that this is the opening half of a drag/selection sequence paired with mouse_up, and 'at the current pointer position' implies the cursor must already be placed (e.g., via mouse_move). No alternatives (click, drag) or when-not conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mouse_moveB
Move the pointer to (x, y) without clicking (e.g. to reveal hover menus).
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden but only adds 'without clicking' and a motivation. It omits the coordinate space (screen vs. window), whether motion is instant or animated, and any side effects. The screenshot_after parameter's behavior is not disclosed at all.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the action and destination and folds the usage example into the end. Zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter pointer tool with no output schema, the definition is minimally adequate, but the unexplained coordinate space and the entirely undocumented screenshot_after flag leave an agent guessing at correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only restates 'to (x, y)' with no units or coordinate-space meaning. The third parameter, screenshot_after, is never mentioned, leaving a third of the parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Move') and resource ('the pointer') with the destination coordinates, and the phrase 'without clicking' cleanly distinguishes it from the click sibling. It is clear enough to select against most siblings, though it does not explicitly contrast with drag or mouse_down/mouse_up.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(e.g. to reveal hover menus)' implies a use case for hover-style interaction but is not an explicit when-to-use rule. No alternatives are named and there are no exclusions or prerequisites, leaving the agent to infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mouse_upC
Release a mouse button.
| Name | Required | Description | Default |
|---|---|---|---|
| button | No | left | |
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, but it only restates the operation. It does not disclose state requirements (e.g., whether a button must already be held), what happens if no button is down, or what screenshot_after actually captures. This is thin for a stateful input tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with zero filler and the core operation stated first. The terseness borders on under-specification rather than being genuinely economical, but structurally it is clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description is still missing the pieces an agent needs: the relationship to mouse_down, and the meaning of screenshot_after. Since the schema itself documents nothing, the definition leaves the caller guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across two parameters, so the description must compensate and does not. The button enum (left/right/middle/back/forward, default left) and the screenshot_after flag are entirely undocumented in prose; screenshot_after in particular is not self-evident about timing or scope.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ("Release a mouse button"), so an agent understands the operation immediately and can distinguish it from unrelated siblings like screenshot or type_text. However, it does nothing to differentiate itself from its nearest sibling, mouse_down, which is the operation agents will most often confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus alternatives. In particular, the paired sibling mouse_down exists and the natural usage pattern (a mouse_down must precede a mouse_up) is never stated, nor is any context about drag versus click workflows.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screen_infoB
Screen size (real and screenshot space), scale, and pointer position.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, but it is inherently a non-mutating info query and the text discloses the returned dimensions including the real-vs-screenshot-space distinction, which is genuinely useful. It stops short of stating that it takes no arguments or is side-effect free.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, and the most important item (screen size in both spaces) is front-loaded. It could be marginally clearer about what 'scale' means, but it wastes nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not enumerate return fields, and it sensibly summarizes them. For a zero-argument info tool this is close to complete; only the relationship to cursor_position is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool applies. The description correctly focuses on the return payload instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact data the tool returns — screen size in both real and screenshot coordinate spaces, scale, and pointer position — so the agent knows this is a read-only environment-info query. It is clear but does not distinguish itself from the sibling cursor_position, which overlaps with the 'pointer position' element.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of the cursor_position sibling, which appears to expose overlapping information. The agent must infer that this is the tool to call before converting between screenshot and real coordinates.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotA
Capture the screen. Optionally pass x/y/width/height (screenshot coords) to zoom into a region for reading small text; coordinates you act on must still be full-screen screenshot coordinates.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| width | No | ||
| height | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose one important behavioral gotcha — that zoomed coordinates still map to full-screen coordinates for subsequent actions — which is real value beyond the schema. But it says nothing about the returned artifact (image format, resolution, how the agent consumes it), leaving a meaningful gap for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core action front-loaded and the coordinate caveat placed where it matters. Efficient, though the second clause is slightly run-on and could be split for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations mean the description is the only source of truth, and it omits any description of the return value or default full-screen behavior. Parameter semantics and the coordinate caveat are covered, but the definition is not fully self-sufficient for an agent deciding how to consume the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does name all four params, labels them as screenshot coordinates, and explains their combined effect (region zoom). It stops short of stating units or that the four should be supplied together, so not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Capture the screen" is a specific verb+resource that an agent can immediately match to a screenshot operation. It does not explicitly differentiate from siblings like ui_tree or screen_info, which also expose screen state, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an implied use case for the optional region args ("zoom into a region for reading small text"), which is genuinely useful. However it never states when to prefer this over ui_tree/screen_info or that the no-arg form is the default whole-screen capture, so usage remains inferential.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollC
Scroll the wheel at (x, y). amount = number of wheel clicks.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| amount | No | ||
| direction | No | down | |
| modifiers | No | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It says nothing about focus requirements, whether scrolling is instant or animated, what happens at scroll boundaries, or that modifiers (e.g., ctrl) change zoom rather than scroll. Only the meaning of 'amount' is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the action is front-loaded before the parameter note. The brevity is efficient, though it shades into under-specification rather than ideal economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, annotation-free, output-schema-free tool, the description covers only one of six parameters and no behavioral context. An agent cannot reliably know what modifiers or screenshot_after do, or what the side effects of scrolling are.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, and the description only clarifies 'amount = number of wheel clicks'. The x/y coordinate semantics, the direction enum values, the modifiers array contents, and screenshot_after are all undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Scroll the wheel') plus the coordinate anchor, which is enough to distinguish it from click, drag, and mouse_move siblings. It does not explicitly name those siblings or contrast its role with them, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use coordinate-based scrolling versus ui_tree, click_element, or set_text, nor when to prefer key (e.g., Page Down) instead. The agent must infer the context entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_textB
Replace the contents of an editable ui_tree element. Falls back to focusing it, select-all and typing when the widget isn't directly editable.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| text | Yes | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the fallback strategy (focus, select-all, typing) which implies existing content is overwritten, but it omits return behavior, error cases when the element isn't editable at all, and side effects like focus stealing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the primary action front-loaded and the fallback detail second. No wasted words, though the fallback clause is slightly run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A mutation tool with no annotations, no output schema, and 0% parameter coverage leaves meaningful gaps: 'screenshot_after' is unexplained and return/error behavior is unstated. The description covers the core action but not enough for an agent to call it confidently in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters. The description only obliquely covers 'id' (the ui_tree element) and 'text' (the contents), and says nothing at all about 'screenshot_after', which is entirely undocumented in both description and schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Replace the contents of an editable ui_tree element.' The verb 'replace' (versus the sibling type_text's append/type behavior) gives the agent a usable distinction, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage by targeting 'an editable ui_tree element' and describes a fallback path, but it never states when to choose this over type_text, click_element, or key, nor any prerequisite such as the element first being located via ui_tree.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
type_textA
Type text into the focused widget. Any Unicode works (Korean, emoji...). Newlines press Return. expect_window: substring of the focused window's title or class; if it doesn't match, nothing is typed.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| expect_window | No | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It discloses meaningful behavior: Unicode support, that newlines press Return, and that a mismatched expect_window results in no typing at all. It omits prerequisites such as focus requirements, accessibility permissions, and typing speed/delay behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, then edge-case behavior. Every sentence adds information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple input tool with no output schema, the description covers the main action, Unicode handling, newline semantics, and the guard parameter's failure mode. The undocumented screenshot_after parameter is the notable remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains expect_window in detail (substring match against title or class, with failure mode) and implies text semantics (Unicode, newlines), but the third parameter screenshot_after is never mentioned anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Type text into the focused widget') and adds scope detail (Unicode, newline handling). It does not differentiate itself from the sibling set_text or key, which an agent must choose between, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through 'focused widget' and the expect_window guard, but there is no explicit statement of when to prefer this over set_text or key. The conditional 'nothing is typed' clause hints at a safety use case without naming it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_treeA
List visible widgets as [id] role "name" @(x,y) with center coordinates
in screenshot space. Filter by app/window name substring, or
active_window_only=True. Ids are valid until the next ui_tree call.
Apps only appear if accessibility is enabled (Chrome/Electron need
--force-renderer-accessibility).
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | ||
| window | No | ||
| max_depth | No | ||
| only_interactive | No | ||
| active_window_only | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the id lifecycle ('Ids are valid until the next ui_tree call'), the coordinate space, and the accessibility prerequisite that silently hides apps. It does not state that this is a non-mutating read, but the list semantics make that reasonably clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the output format, then filters, then the id lifecycle and the accessibility caveat. Four dense sentences, no filler, and the most decision-relevant facts come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no further explanation, and the id-invalidation warning is valuable. However, for a 5-parameter tool with zero schema descriptions and zero annotations, the unexplained only_interactive default and max_depth leave real gaps an agent would hit when widgets it expects are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description must compensate. It explains app/window substring filtering and active_window_only, but says nothing about max_depth or only_interactive, whose default of true materially changes results (non-interactive widgets are silently omitted).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list visible widgets) and even pins down the exact output shape `[id] role "name" @(x,y)` in screenshot space. It is clearly distinguishable in substance from siblings like screenshot or list_windows, though it never names an alternative to route between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers concrete usage mechanics (filter by app/window substring, or active_window_only=True) and a prerequisite (accessibility must be enabled, with the Chrome/Electron flag), but gives no explicit when-to-use-vs-screenshot/list_windows guidance or when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
waitB
Wait for the UI (max 30s), then by default return a screenshot.
| Name | Required | Description | Default |
|---|---|---|---|
| seconds | No | ||
| screenshot_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the 30s cap and the default screenshot return. However it omits what happens on timeout (error, partial capture, or silent return) and whether waiting beyond 30s is clamped or rejected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the duration cap and default return front-loaded. Every clause carries information and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple zero-required-param tool, but with no output schema, no annotations, and 0% schema coverage, the description should have covered the 'seconds' cap semantics and the timeout outcome. It leaves the agent guessing about edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and neither parameter is named, but the description maps to both: 'max 30s' constrains 'seconds' and 'by default return a screenshot' explains the 'screenshot_after' default of true. It conveys intent but no units, clamping behavior, or how to suppress the screenshot.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (wait) and target (the UI) plus the two key behaviors: a 30s ceiling and screenshot-by-default. It is clearly distinguishable from the sibling 'screenshot', though the relationship between 'wait' and 'screenshot_after' is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to wait versus calling 'screenshot' directly, and no mention of typical placement (e.g., after click/type_text to let the UI settle). The default-return note hints at an alternative but never names a condition for choosing it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v0.1.0- First observed
active_window - First observed
click - First observed
click_element - First observed
cursor_position - First observed
drag - First observed
hold_key - First observed
key - First observed
list_windows - First observed
mouse_down - First observed
mouse_move - First observed
mouse_up - First observed
screen_info - First observed
screenshot - First observed
scroll - First observed
set_text - First observed
type_text - First observed
ui_tree - First observed
wait
TDQS
Scored across 18 tools
Most tools target clearly distinct primitives (coordinate click vs. semantic click_element, type_text vs. set_text, key vs. hold_key, drag vs. mouse_down/up). Minor overlap: screen_info already reports the pointer position that cursor_position returns, and active_window/list_windows/ui_tree all relate to window/widget state and could be briefly confusing. Descriptions largely resolve these ambiguities.
All names use consistent snake_case, which reads cleanly across the set. There is some variance between verb-based (mouse_move, list_windows, click_element, set_text) and noun-based (screenshot, cursor_position, active_window, ui_tree) names, but the vocabulary is predictable for a low-level GUI primitive API.
18 tools is slightly above the typical 3-15 range but justified for full GUI automation, since each covers a genuinely distinct primitive (mouse buttons, keys, accessibility, screen capture). No redundant tools inflate the count.
The surface covers screen capture and geometry, full mouse control (click, move, drag, down/up, scroll), keyboard input (type, chords, holds), accessibility inspection and activation, plus waiting and focus verification. This is a complete lifecycle for computer-use tasks with no obvious dead ends.
Maintenance
Related MCP Connectors
Use your Mac, Windows or Linux computer from ChatGPT, Claude or Codex: files, commands, documents.
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Securely control computers you explicitly pair through files, terminals, processes, screenshots, desktop UI/input, clipboard, browser automation, diagnostics, and document tools.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables Claude Code to interact with Linux X11 desktops by performing mouse actions, keyboard input, and capturing screenshots. It provides a suite of tools for automation tasks such as clicking, typing, and zooming on specific screen regions.16-
- AlicenseAqualityCmaintenanceGive Claude Desktop full desktop control on Linux/X11: screenshot, mouse, keyboard, windows, clipboard, app launch. Zero-dependency MCP extension, MIT-licensed.15MIT
- AlicenseNot gradedqualityDmaintenanceEnables Claude to control the local desktop via screenshot, mouse, keyboard, and clipboard operations.MIT
- AlicenseAqualityCmaintenanceEnables local X11 desktop control for MCP clients, including persistent Python execution, AT-SPI accessibility inspection, screenshots, and recoverable mouse/keyboard input for operating real applications.5MIT