Skip to main content
Glama

Introduction

Computer Use is an open-source MCP server — a computer-use agent (CUA) backend — that lets any coding agent use your computer the way a person does. It reads the screen through accessibility trees, clicks and types in the background so your mouse stays yours, shows an agent pointer where it is working, zooms in on small text, and drives tabs in your signed-in Chrome — on macOS, Windows and Linux.

Works with Claude Code, Codex, Cursor and MT Code, or any other MCP client, with any model — no vision model is required for interaction.

Built by Munim Technologies as the Computer Use engine of MT Code, and published here on its own.

Related MCP server: openowl

Table of contents

Quick start

  1. Download the latest binary for your platform from Releases (munim-computer-use-macos-universal.zip, munim-computer-use-windows-x64.zip) or build from source.

  2. Put it somewhere on your PATH (/usr/local/bin/munim-computer-use, or %LOCALAPPDATA%\Programs\munim-computer-use\munim-computer-use.exe).

  3. macOS only: run munim-computer-use request-permissions once to be prompted for Accessibility and Screen Recording.

  4. Register it with your agent:

# Claude Code — fastest: the npm launcher fetches the signed binary on first run.
# Add `--scope user` to register it for every project instead of just this one.
claude mcp add munim-computer-use -- npx -y munim-computer-use
# or point at a downloaded binary
claude mcp add munim-computer-use -- /usr/local/bin/munim-computer-use
# Codex — writes the entry below into ~/.codex/config.toml for you (needs a Codex CLI
# with `codex mcp`; check with `codex mcp --help`)
codex mcp add munim-computer-use -- npx -y munim-computer-use
# Codex — ~/.codex/config.toml, if you would rather edit it yourself
[mcp_servers.munim-computer-use]
command = "npx"
args = ["-y", "munim-computer-use"]
# Cursor — no MCP subcommand in its CLI, so write the config. ~/.cursor/mcp.json applies
# to every project; .cursor/mcp.json in a repo applies to that one.
mkdir -p ~/.cursor && [ -s ~/.cursor/mcp.json ] || echo '{}' > ~/.cursor/mcp.json
jq '.mcpServers["munim-computer-use"] = {"command":"npx","args":["-y","munim-computer-use"]}' \
  ~/.cursor/mcp.json > ~/.cursor/mcp.json.tmp && mv ~/.cursor/mcp.json.tmp ~/.cursor/mcp.json
// Cursor — the entry that produces, in .cursor/mcp.json
{ "mcpServers": { "munim-computer-use": { "command": "npx", "args": ["-y", "munim-computer-use"] } } }

Then ask: "Open Safari, find the cheapest flight to Denver on Tuesday and put it in a note." The agent reads the UI with get_app_state, acts by element id, and verifies with screenshot.

Capability matrix

Munim Computer Use is the highlighted first column; the others are the computer-use servers people reach for. each cell comes from that project's own README or docs in September 2026 (sources under Credits and license). ✅ present · ❌ absent or not documented · ⚠️ partial.

Capability

Munim Computer Use

Codex Computer Use

OpenAI Agents API

Anthropic reference demo

Windows-MCP

MacOS-MCP

open-computer-use

computer-use-mcp (zavora)

Notes

macOS

✅

✅

n/a

❌

❌

✅

✅

✅

The Anthropic demo drives a Linux desktop inside Docker, not your machine. The OpenAI Agents API (September 2026) runs a browser OpenAI hosts, so it never drives your machine either.

Windows

✅

✅

n/a

❌

✅

❌

✅

✅

Windows-MCP is Windows only; MacOS-MCP is macOS only.

Linux

✅

❌

n/a

✅ (sandbox)

❌

❌

✅

✅

Munim Computer Use uses AT-SPI + X11; native Wayland apps get element actions but not coordinate clicks.

Accessibility tree with element ids

✅

❌

❌

❌

✅

✅

✅

✅

Codex, the Agents API and the Anthropic demo are screenshot-driven. Ids let the agent press the button instead of a pixel.

Background input (your mouse never moves)

✅

✅

n/a

n/a

❌

❌

❌

❌

Munim Computer Use addresses events to the target window: always on macOS (SkyLight and per-process events), for UI Automation patterns and classic Win32 controls on Windows; other Windows UI and all Linux pointer input (XTEST) still move the real pointer. Codex does this too, with a second cursor of its own. Every other server in this table drives the real cursor.

Agent pointer overlay

✅

✅

❌

❌

⚠️

❌

❌

❌

Windows-MCP flashes a border around captures; Codex draws its own cursor on your screen.

Zoom into a region at full resolution

✅

❌

❌

✅

❌

❌

❌

❌

Anthropic's toolset has zoom; here it is a tool on every platform.

Screenshots carry screen-coordinate mapping

✅

n/a

n/a

n/a

❌

❌

❌

❌

Origin and pixels-per-point in every capture, so clicks from Retina or downscaled images land.

Hover, wait, label query

✅

⚠️

❌

⚠️

✅

⚠️

❌

⚠️

Windows-MCP has Wait/WaitFor; MacOS-MCP has Wait; Anthropic has wait/mouse_move.

Your signed-in Chrome, own tab group or yours

✅

⚠️

❌

❌

⚠️

❌

❌

❌

Codex uses its in-app browser; Windows-MCP reads the DOM of open browsers. Munim Computer Use opens its own labelled tab group in your real Chrome, and can also take over a tab you already have open when you ask it to.

Act and observe in one call

✅

n/a

n/a

✅

❌

❌

❌

❌

return_state on any action appends the fresh accessibility tree or page snapshot, so one call acts and checks. The Anthropic demo returns a screenshot after each action; Codex and the Agents API run the loop themselves.

Read a page's text in a signed-in tab

✅

❌

❌

❌

⚠️

❌

❌

❌

browser_read returns the page as text with headings, in chunks, with a query filter, from a background tab. Windows-MCP scrapes pages.

Sign-in without the model seeing the password

✅

❌

✅

❌

❌

❌

❌

❌

browser_request_credentials opens a Chrome window naming the site's real origin; what the user types goes into the page and never into the conversation. The Agents API hands sign-in to the developer's app the same way. Desktop password fields refuse typing by default.

Per-app and per-site allow / ask / block rules

✅

⚠️

⚠️

❌

❌

❌

❌

❌

A policy file the user controls; ask shows an approval prompt once per task. Codex asks before each new app and keeps an "Always allow" list; the Agents API asks the developer's app before each new website.

Works with any MCP client

✅

❌

❌

❌

✅

✅

✅

✅

Codex Computer Use is Codex only; the Anthropic demo is Claude only.

Identical tool surface on every platform

✅

n/a

n/a

n/a

n/a

n/a

⚠️

✅

32 tools with byte-identical schemas across the Swift and Rust servers.

Prebuilt signed binaries + npm launcher

✅

✅

n/a

❌

❌

❌

✅ (npm)

✅ (npm)

macOS universal (Developer ID signed) and Windows x64 on Releases.

Open source

✅ Apache-2.0

❌

❌

✅

✅ MIT

✅ MIT

✅ MIT

✅ MIT

Also looked at: mediar-ai/mcp-server-macos-use (macOS, accessibility, real input), deploymenttheory/windows-mcp-server (Windows, UIA Invoke patterns, WaitFor), nuphus-mcp (OCR + bring-your-own vision model, CDP Chrome), computer-control-mcp (PyAutoGUI + OCR), and microsoft/playwright-mcp (browser only). Corrections welcome — open an issue with a link.

Why it works well

  • Accessibility first, pixels second. get_app_state returns the app's accessibility tree with stable element ids, so the agent presses the button instead of guessing at a coordinate. It costs a fraction of the tokens of a screenshot and it is what scores highest on OSWorld-style tasks. Screenshots are for verifying and for content the tree cannot describe.

  • Background control. Events are addressed to the target window (SkyLight on macOS, UI Automation patterns and posted window messages on Windows). On macOS the agent never takes your mouse or keyboard; see Works alongside you for the exact guarantee on each platform.

  • Pointer overlay, not your pointer. A soft lavender agent pointer shows where the agent is acting. Your cursor is untouched.

  • Coordinates that land. Every screenshot and zoom carries its screen origin and pixels-per-point. zoom captures any region at full physical resolution.

  • Your browser, your logins. The Chrome extension gives the agent its own labelled tab group in your signed-in Chrome, and leaves your tabs alone unless you point it at one.

  • Or the tab you already have open. browser_list_tabs all=true shows every tab in the browser and browser_use_tab takes one over in place — useful when the page is already signed in or mid-flow and re-opening the URL would throw that away. An adopted tab is not moved into the agent's group, not activated and not reloaded; cleanup releases it rather than closing it, and browser_release_tab hands it back early.

  • Parallel tasks, one extension. Every MCP process gets its own tab group, and one process can run several tasks by passing a stable session_id on its browser calls. A task cannot drive or adopt another task's tabs, and cleanup (including a process exiting) closes only its own. Any number of MCP processes share the one extension: the first owns it and the rest go through it, and if the owner exits another takes over without closing anyone's tabs. Tasks share Chrome's cookies and logins, and desktop apps and the clipboard are not isolated.

  • Model-agnostic. No vision model is required for interaction; local models work too.

  • Look → act → verify. hover for mouse-over menus, wait for loads, query to find a control by label without reading a whole tree.

  • Act and look in one call. Pass return_state: true to any action (click, type_text, set_value, browser_click, browser_navigate, …) and the result carries the app's fresh accessibility tree, or the page's fresh snapshot, taken once the UI has settled. That halves the round trips of a look-act-verify loop. state_query narrows the desktop tree the same way query does.

  • Pages as text. browser_read returns a page's readable text with its headings, including what is scrolled out of view, from a tab in the background. It comes in chunks you can continue with offset, query keeps only the matching lines under their heading, and include_links lists the links.

  • Passwords stay out of the conversation. browser_request_credentials asks the user to sign in in a small Chrome window that shows the site's real origin. What they type goes straight into the page's fields and is never returned to the model. browser_snapshot never shows a password field's value, and on the desktop, typing into password fields is refused by default.

  • Pick up what you are looking at. get_app_state and screenshot take app: "frontmost" for the app in front of the user, so a client can hand the agent "this window" in one step.

Works alongside you

The agent has its own pointer; yours stays yours.

macOS — guaranteed. Every action goes through accessibility (press, set value, select text, show menu, scroll bars) or through events addressed to the target app's process and window. The server never moves your pointer, never posts into the system-wide input stream, and never holds or blocks your input, so you can keep clicking and typing in other apps while the agent works — even in the same app, on another window. Events come from a private source, so a modifier you are holding does not leak into the agent's clicks. The target app may be brought forward when that is the point of the step (activate_app, or handing you a password field), but not on every action.

The exceptions refuse instead of borrowing your pointer: a click, hover or scroll with no target app (coordinates over the desktop before any get_app_state), and drags that leave the source window (between apps, or onto the desktop). The error says what to pass instead.

Windows — best effort, reported. Element presses (Invoke), set_value, select_text and scrolling through UI Automation's ScrollPattern never touch your pointer, and classic Win32 controls also take clicks, hovers, wheel and drags as posted window messages. Other UI (Chromium, Electron, WPF, UWP) only reacts to real mouse input, so those clicks, hovers and drags move your pointer, and type_text/press_key go to the focused window. Whenever the real pointer was used, the result says via cursor.

Linux — pointer actions use it. XTEST input moves the real pointer and goes to the focused window, and results say via cursor. Element actions through AT-SPI (press, set value, insert text, select) do not move it.

Apps and sites the agent may use

You decide which apps and websites the agent may touch. Put a policy.json in the server's support directory (~/Library/Application Support/computer-use on macOS, %LOCALAPPDATA%\munim-computer-use on Windows, ~/.local/share/munim-computer-use on Linux; munim-computer-use identity prints it as supportDir), or point COMPUTER_USE_POLICY at a file anywhere. An app that embeds the server under its own profile reads the file from its own support directory:

{
  "apps": { "Keychain Access": "block", "com.apple.MobileSMS": "ask" },
  "sites": { "bank.example": "block", "mail.google.com": "ask", "*": "allow" }
}
  • allow, ask or block per app or site. ask shows the user a prompt the first time the agent reaches for it (a system dialog for apps, a small Chrome window for sites), and an approval lasts until that agent task ends. An unanswered prompt counts as no after two minutes.

  • Apps match their name or bundle id, ignoring case. The rule applies when the agent reads, captures or activates the app. Actions then target elements from a snapshot that was allowed.

  • Sites match a host and all its subdomains, and the most specific pattern wins. They are checked on every browser action against the page the tab is showing at that moment, so following a link into a blocked site does not get around it.

  • * sets the default, which is otherwise allow, so {"apps": {"*": "block", "Notes": "allow"}} is an allow-list.

  • Changes apply at once, with no restart. A file that is not valid blocks everything rather than being ignored, so a typo cannot quietly turn a block into an allow.

This is a guard rail for an agent that follows its instructions, not a sandbox. A coordinate click lands wherever it points, and a whole-display screenshot shows every window.

Tools (32)

Area

Tools

See

list_apps, get_app_state (with query), screenshot, zoom, list_displays

Act

click, right_click, hover, drag, scroll, type_text, set_value, select_text, press_key, activate_app, wait

Clipboard

clipboard_read, clipboard_write (plain text)

Browser

browser_open_tab, browser_list_tabs, browser_use_tab, browser_release_tab, browser_select_tab, browser_navigate, browser_snapshot, browser_read, browser_click, browser_type, browser_request_credentials, browser_press_key, browser_close_tab, browser_close_all_tabs

Every action tool also takes return_state, which returns the state after the action in the same call.

Names, argument shapes and descriptions are identical on every platform, and CI enforces it (node scripts/check-tool-parity.mjs); a model that learned them on a Mac needs nothing new on Windows. To change a tool, edit both literals at once with scripts/tool-defs.mjs rather than by hand.

Repository layout

Directory

What

Build

macos/

Swift server on the Accessibility API and ScreenCaptureKit (macOS 14+)

swift build -c release → .build/release/munim-computer-use

windows-linux/

Rust server: UI Automation on Windows, AT-SPI + X11 on Linux

cargo build --release → target/release/munim-computer-use

chrome-extension/

Chrome extension + native messaging host for the browser_* tools

Load unpacked; sh install.sh / install.ps1 registers the host; node background.test.mjs

Dockerfile

Headless Linux build of the Rust server for registry introspection (Glama and similar); no desktop control inside a container

docker build -t munim-computer-use .

Build from source

# macOS
cd macos && swift build -c release
# Windows / Linux
cd windows-linux && cargo build --release

Linux notes: element actions work everywhere; coordinate clicks need an X11 or XWayland client, since native Wayland apps do not expose absolute geometry.

Tests

cd windows-linux && cargo test                  # Rust server
node chrome-extension/background.test.mjs       # extension, against a fake Chrome
node scripts/check-tool-parity.mjs              # Swift and Rust tool lists match
# The browser tools end to end, in a throwaway Chrome for Testing profile that
# never touches your own Chrome (npx playwright install chromium to get one):
node scripts/e2e-browser.mjs --server <munim-computer-use binary> --chrome <Chrome for Testing binary>

Chrome extension (optional)

  1. chrome://extensions → Developer mode → Load unpacked → select chrome-extension/.

  2. Register the native messaging host: munim-computer-use install-native-host (for example npx -y munim-computer-use install-native-host), or from a checkout sh chrome-extension/install.sh (macOS/Linux) / powershell -File chrome-extension/install.ps1 (Windows), which find the build and run the same command. Point COMPUTER_USE_PATH at the binary if it is not in the default build location.

The standalone server's host is com.munimtech.computer_use.desktop; MT Code, which bundles this server, registers com.munim.mtcode.desktop. The extension connects to every host it knows at once and answers each on its own connection, so MT Code and a standalone server (npx, Claude Code, Cursor, a checkout) can both drive Chrome at the same time, each in its own tab groups.

Environment flags

Variable

Effect

COMPUTER_USE_BROWSER=0

Hide the browser_* tools

COMPUTER_USE_AGENT_CURSOR=0

Do not draw the agent pointer

COMPUTER_USE_AGENT_CURSOR_TASK_FADE_SECS

How long the pointer stays after the last tool call (default 8)

COMPUTER_USE_ALLOW_SECURE_FIELD_INPUT=1

Allow typing into password fields (refused by default)

COMPUTER_USE_REMOTE_CONTROL=1

Remote-desktop mode: input takes over the real pointer

COMPUTER_USE_POLICY=<file>

Where to read the app and site policy (default policy.json in the support directory)

Remote control

Normally this server never touches the pointer: coordinate clicks are routed to a specific window, keystrokes are posted to a specific process, and an action that cannot be targeted is refused rather than taking over the machine. That is what lets an agent work while the user keeps using their computer.

COMPUTER_USE_REMOTE_CONTROL=1 inverts that contract for one process, for the case where a person is watching this machine's screen from another one and is steering it themselves. Then click, right_click, drag, hover and scroll move the real cursor and type_text and press_key go to whatever is focused, the way Chrome Remote Desktop or Screen Sharing behave. scroll also accepts x/y so the wheel acts over the point the viewer scrolled at, and the agent-cursor overlay stays hidden — there is only one pointer now.

A host that runs both an agent and a viewer runs them as two processes, so turning this on for the viewer never takes the pointer away from the user on the agent's behalf.

Embedding in an app

An app can ship this binary inside its own bundle and run it under its own identity, so it never shares a browser bridge, agent-cursor app or native-messaging host with a standalone install on the same machine. MT Code does exactly this. Pass a profile, either as a JSON object or a path to a JSON file, with --profile <json|file> (any position) or COMPUTER_USE_PROFILE:

{
  "name": "example-desktop",
  "envPrefix": "EXAMPLE_DESKTOP_",
  "agentCursorName": "ExampleAgentCursor",
  "agentCursorBundleId": "com.example.agent-cursor",
  "nativeHostNames": ["com.example.desktop"],
  "extensionIds": ["abcdefghijklmnopabcdefghijklmnop"],
  "nativeHostDescription": "Example desktop control bridge"
}

Key

Default

Meaning

name

none (standalone paths)

Moves the support dir, bridge socket and Windows pipe under this name

supportDir

~/Library/Application Support/computer-use, $XDG_DATA_HOME/munim-computer-use, %LOCALAPPDATA%\munim-computer-use

Native-host wrapper, profile copy; on macOS also the bridge socket and a bare build's overlay app

bridgeSocket

<supportDir>/bridge.sock (macOS), $XDG_RUNTIME_DIR/<name>/bridge.sock (Linux), <name>-bridge-<user> pipe (Windows)

Where the MCP server and the Chrome relay meet

envPrefix

none

Tunables are read as <prefix>BROWSER, <prefix>AGENT_CURSOR, … before COMPUTER_USE_*

agentCursorName / agentCursorBundleId

MunimAgentCursor / com.munimtech.computer-use.agent-cursor

The pointer overlay's app, executable and window-class name, and its macOS bundle id

historyDir

none

Default --root for computer-history

nativeHostNames / extensionIds

com.munimtech.computer_use.desktop, com.munim.mtcode.desktop / kgdolgnijopbghhomnblabjkmjhnoage

What install-native-host registers, and for which extension. The first name is this identity's own; later ones are aliases, written only if no other installed app owns them

Each of supportDir, bridgeSocket, envPrefix, agentCursorName, agentCursorBundleId and historyDir can also be overridden by COMPUTER_USE_<SNAKE_CASE> (for example COMPUTER_USE_SUPPORT_DIR), which wins over the profile. munim-computer-use identity prints the resolved values.

  • Browser bridge. Run munim-computer-use install-native-host with the same profile. It writes a wrapper that relays Chrome into this identity's bridge (replaying the profile), and a host manifest for each name in every Chrome/Chromium profile directory (the registry on Windows). It rewrites nothing that is already current, so an app can call it on every launch.

  • Extension. Use the stock extension (it connects to com.munim.mtcode.desktop and com.munimtech.computer_use.desktop), or build a variant with its own host names, tab-group title and key: node scripts/build-extension.mjs --out <dir> --host com.example.desktop --group-title "Example" --key <base64>. It prints the variant's extension id for extensionIds.

  • macOS permissions. The MCP server is a bare executable, so Accessibility and Screen Recording are granted to the app that spawns it. Only the agent-cursor overlay has a bundle of its own; ship <agentCursorName>.app (a copy of the binary plus an LSUIElement Info.plist) beside the binary, or it is materialised under supportDir on first use.

Prompting your agent

Look → act → verify. get_app_state for ids, act by id with return_state: true to see the result in the same call, and screenshot when the tree cannot show it. Read pages with browser_read, and let the user type passwords through browser_request_credentials. Use zoom for small text, hover for menus that appear on mouse-over, wait after loads, keyboard shortcuts for stubborn widgets. The system-prompt text MT Code gives its agents lives in CodexDeveloperInstructions.ts and is a good starting point.

Contributing

This repository mirrors the native/ tree of munimtechnologies/mtcode, where the server is developed and shipped inside MT Code. Issues and discussions are welcome here; code changes land in mtcode first and are synced.

Credits and license

Designed and built by Munim Technologies (Munim, Inc.) for MT Code. Copyright 2026 Munim, Inc. Licensed under the Apache License 2.0; see LICENSE.

Comparison sources: Codex Computer Use and its docs · OpenAI Agents API computer use · Anthropic computer-use demo · CursorTouch/Windows-MCP · CursorTouch/MacOS-MCP · QwenLM/open-computer-use · zavora-ai/computer-use-mcp · mediar-ai/mcp-server-macos-use · deploymenttheory/windows-mcp-server · nuphus-mcp · computer-control-mcp · microsoft/playwright-mcp

Available Tools

32 tools
activate_appActivate appA
Idempotent

Bring an app's windows to the foreground and give it keyboard focus. Call it before press_key or type_text when the target app is not frontmost; element-id actions such as click and set_value do not need it. Side effect: the window the user was working in loses focus.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesApp name, bundle id, or pid exactly as reported by list_apps

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description adds the concrete side effect that the user's current window loses focus, and clarifies that it gives keyboard focus. It doesn't contradict annotations and provides useful behavioral context beyond the structured flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler. The primary purpose is in the first sentence, and the second sentence packs usage guidance and side effect efficiently. All content earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, side-effectful operation with no output schema, the description covers the action, the trigger conditions, the side effect, and the parameter format (via schema). Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter 'app' is fully documented with format ('App name, bundle id, or pid exactly as reported by list_apps'). The description adds no additional parameter-specific details, which is acceptable given the high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('bring') and resource ('an app's windows'), and clearly defines the outcome: foreground and keyboard focus. It also distinguishes from siblings by noting that element-id actions like click and set_value do not need it, which helps an agent pick the right tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to call it: before press_key or type_text when the target app is not frontmost, and explicitly says when not needed. It also mentions the side effect of losing focus in the previous window, giving clear situational guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_clickClick in browserA
Destructive

Click in one of the agent's tabs, either an element by its index from browser_snapshot (preferred) or a point given in page coordinates. Pass index or x and y, not both. Works on a background tab. Use click for native app windows. A click can submit forms or follow links, so snapshot first.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoPage x coordinate in CSS pixels, used together with y when no index is given
yNoPage y coordinate in CSS pixels, used together with x when no index is given
indexNoElement index from the latest browser_snapshot of this tab. Preferred over coordinates.
tab_idYestab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.
return_stateNoAfter acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the risk profile (destructiveHint=true, openWorldHint=true, non-idempotent, not read-only), so the bar is lower. The description still adds real context beyond them: it works on background tabs, clicks can cause navigation/form submission, and the snapshot indices get replaced after a state refresh.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the primary addressing mode (index) before the fallback (coordinates), then the exclusivity rule, then the side-effect warning. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating, open-world click tool with no output schema, the description covers targeting, exclusivity, background-tab behavior, the sibling alternative, and the side-effect hazard. It could say slightly more about what return_state yields, though the schema already explains that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so every parameter is already documented in the schema, giving a baseline of 3. The description adds one genuinely new constraint the schema lacks — 'Pass index or x and y, not both' — but otherwise restates what the schema already says about index vs. coordinates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (click) and resource (agent's browser tab), and explicitly delimits itself from the sibling `click` tool ('Use click for native app windows'). An agent can distinguish it from right_click/hover and from the native-app click without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: prefer the `index` from browser_snapshot over raw coordinates, use the sibling `click` for native app windows, and 'snapshot first' because a click can submit forms or follow links. It also notes it works on a background tab, which is a non-obvious usability fact.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_close_all_tabsClose all agent tabsA
DestructiveIdempotent

Close every tab the agent opened and remove its tab group. Tabs taken over with browser_use_tab are released back to the user, not closed. Call this when finished with the browser so no empty group is left in the user's tab strip. The MCP process also runs this automatically when the Computer Use session ends. Unsaved state in the agent's tabs is lost.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral detail beyond the annotations: it clarifies that user-taken-over tabs are released rather than closed, warns that unsaved state is lost, and discloses automatic execution when the Computer Use session ends. These are meaningful side effects not captured by the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Each sentence serves a distinct purpose: core action, exception handling, when to call, automatic behavior, and data-loss warning. The description is front-loaded with the main operation and remains appropriately sized for a destructive cleanup tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the action, side effects, user-impact, automatic invocation, and unsaved-state warning. For a destructive tool with a single documented parameter and no output schema, this provides everything an agent needs to call it correctly and understand consequences.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the session_id parameter is fully documented in the schema. The description does not add parameter-level detail, but the baseline of 3 applies because the schema handles the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Close every tab the agent opened and remove its tab group') and distinguishes it from sibling tools like browser_close_tab by specifying scope (all agent tabs) and the special handling of tabs taken over via browser_use_tab. This makes the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: 'Call this when finished with the browser' and notes the automatic execution at session end. It does not explicitly name an alternative for closing a single tab, but the overall context and sibling list make the intended condition clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_close_tabClose browser tabA
DestructiveIdempotent

Close one of the agent's tabs, discarding any unsaved page state. A tab taken over with browser_use_tab is released rather than closed — it belongs to the user. Use browser_close_all_tabs to clean up everything at the end of a task.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNo1-based position within the agent's tabs; fallback mode only, when tab_id is unavailable
tab_idNotab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds important behavioral context beyond the annotations: it discloses that closing a tab discards unsaved page state, and that tabs taken over with browser_use_tab are released rather than closed because they belong to the user. The annotations already declare destructiveHint=true and idempotentHint=true, and the description aligns with these. It doesn't contradict the annotations, and it adds meaningful context about the tool's side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with zero waste. It front-loads the core action and side effect, then provides the sibling distinction and the special case. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a destructive tab-closing tool. It covers the core behavior, the side effect (discarding unsaved state), the special case for user-owned tabs, and the alternative for bulk cleanup. The only minor gap is that it doesn't explicitly state what happens if both index and tab_id are provided, but the schema's 'fallback mode only' note covers this. With no output schema, the description doesn't need to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all three parameters (index, tab_id, session_id). The description doesn't add much beyond what the schema provides, but it does clarify the fallback relationship between index and tab_id ('fallback mode only, when tab_id is unavailable'). This is a baseline 3 because the schema carries the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: closing one of the agent's tabs and discarding unsaved page state. It also distinguishes itself from the sibling tool browser_close_all_tabs, which is for cleaning up everything at the end of a task. The verb 'close' and resource 'tab' are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool versus alternatives: use browser_close_all_tabs to clean up everything at the end of a task. It also clarifies the special case of tabs taken over with browser_use_tab, which should be released rather than closed. This is clear, actionable guidance for an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_list_tabsList browser tabsA
Read-onlyIdempotent

List the tabs in the agent's own Chrome tab group, marking the active one, with the tab_id each other browser tool needs. Pass all=true to see every tab open in the browser, including the user's, so you can pick one to drive with browser_use_tab. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
allNoList every tab in the browser, not just the agent's (default false)
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description reinforces 'Read-only'. It adds meaningful behavioral context beyond annotations: the agent's own tab group is the default scope, all=true expands to every tab, the active tab is marked, and the returned tab_id is needed by other browser tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The core behavior is front-loaded, the all=true variant is explained exactly where needed, and the safety note is appended last. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only listing tool with no output schema, the description adequately conveys the key returned data (tab list, active marker, tab_id) and the parameter behavior that matters for invocation. It could be slightly richer about the exact output shape, but nothing critical is missing for selecting and calling the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters. The description adds prose meaning for all ('Pass all=true to see every tab...') but does not discuss session_id beyond what the schema provides. This matches the baseline for fully documented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'List the tabs in the agent's own Chrome tab group', and clarifies the output includes active marking and tab_id. This clearly distinguishes it from sibling tools like browser_snapshot or browser_select_tab while positioning it as the listing companion to browser_use_tab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete guidance on when to use all=true (to see the user's tabs and pick one to drive with browser_use_tab) and clarifies the default scope. It does not explicitly state when not to use the tool, but the context is clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_navigateNavigate browser tabA
Destructive

Point one of the agent's tabs at a different URL, replacing the current page; unsaved page state is lost. Use browser_open_tab to keep the current page and open another. Follow with browser_snapshot, since element indices reset after navigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute URL to load in the tab
tab_idYestab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.
return_stateNoAfter acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false; the description adds the concrete consequence ('unsaved page state is lost') and a non-obvious post-condition ('element indices reset after navigation') that the structured fields cannot express. This materially affects how an agent sequences its next call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the destructive action, the alternative, and the required follow-up. The consequence is front-loaded before the routing advice.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description tells the agent what happens to prior state and what call to make next, covering the gaps a navigation tool would otherwise leave. Annotations plus the 100%-documented schema leave nothing important unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for all four parameters, so the schema already documents tab_id, url, session_id and return_state semantics. The description adds no parameter-level syntax or format detail, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Point one of the agent's tabs at a different URL, replacing the current page') and immediately frames it against the sibling browser_open_tab. An agent can distinguish it from every other tab tool without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the alternative and its selecting condition ('Use browser_open_tab to keep the current page and open another'), plus prescribes the follow-up call ('Follow with browser_snapshot'). Both when-to-use and what-to-do-next are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_open_tabOpen browser tabA

Open a URL in a new background tab inside the agent's own labelled tab group in the user's signed-in Chrome, and return its tab_id for browser_snapshot, browser_click, browser_type and browser_navigate. The tab opens in the background, so the user's browsing is not interrupted. Use browser_use_tab instead when the user already has the page open and signed in. Requires the Computer Use Chrome extension; a limited fallback mode applies without it.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoAbsolute URL to open (default about:blank)
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that the tab opens in the background, that it lives in the agent's own tab group, that it returns a tab_id for other browser tools, and that it requires the Computer Use Chrome extension with a limited fallback mode. These are material behavioral traits not present in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the core action and return value, the background behavior, and the alternative tool with extension requirement. The most important operational detail is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter optional tool with no output schema, the description is complete enough: it states what the tool returns, how the tab behaves, when to use a sibling instead, and the extension prerequisite. The agent has what it needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema fully documents url and session_id. The description adds context about the returned tab_id and its downstream use, but does not add parameter-level semantics beyond the schema. This matches the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Open a URL in a new background tab'), clarifies the tool's distinguishing behavior (background tab, agent's own labelled tab group), and names the return value (tab_id). This clearly differentiates it from sibling tools like browser_navigate and browser_use_tab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: to open a URL in a new background tab without interrupting the user. It also names the alternative and the exact condition for choosing it: 'Use browser_use_tab instead when the user already has the page open and signed in.' This is direct usage guidance, not just implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_press_keyPress key in browserA
Destructive

Press Enter, Tab, Escape or Backspace in one of the agent's tabs, for example Enter to submit a form after browser_type. Only these four keys are supported; use browser_type for characters. Enter can submit forms and Backspace deletes, so check the page state with browser_snapshot first.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesKey to press
tab_idYestab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.
return_stateNoAfter acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and idempotentHint=false, so the mutation risk is covered structurally. The description goes beyond that by naming the concrete side effects ('Enter can submit forms and Backspace deletes') and the key limitation, though it says nothing about failure modes or what happens on an invalid tab_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, all earning their place: capability first, then the exclusivity constraint, then the caution and precondition. No filler or repetition of name/title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with full schema coverage and annotations, the description supplies the essentials: supported keys, side effects, the sibling route for typing, and the snapshot precondition. It does not mention session_id isolation or the return_state snapshot behavior, but the schema covers both, so the gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents tab_id, session_id, and return_state in detail, and the enum defines the allowed keys. The description only restates the key restriction already encoded in the enum, adding no new per-parameter meaning; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Press Enter, Tab, Escape or Backspace in one of the agent's tabs') and enumerates the exact supported key set. It explicitly distinguishes itself from browser_type, so an agent can pick the right sibling without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives when-to-use with a concrete example ('Enter to submit a form after browser_type'), names the alternative and its exclusion condition ('use browser_type for characters'), and prescribes a precondition ('check the page state with browser_snapshot first').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_readRead page textA
Read-onlyIdempotent

Read the text of the page in one of the agent's tabs: headings marked with #, then paragraphs, lists and table text in reading order, including content scrolled out of view. Use it to read an article, results or documentation; use browser_snapshot to find something to click. Works on a background tab. Long pages come back in chunks: the result says where it stopped, and offset continues from there. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoOnly return lines containing this text (case-insensitive), each with its neighbouring lines and the heading above it
offsetNoCharacter offset to start from, to continue a page that was cut off (default 0)
tab_idYestab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
max_charsNoCharacters to return (default 20000, at most 200000)
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.
include_linksNoAlso list the page's links as text and URL (up to 150)

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only/idempotent/open-world safety, but the description adds genuine behavioral context beyond them: it works on a background tab, includes scrolled-out content, and returns long pages in chunks with a stated stop point. The trailing 'Read-only' merely restates readOnlyHint, keeping this short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what is read, then format, then usage/alternatives, then background-tab and chunking behavior. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a full-coverage, fully-annotated read tool with no output schema, the description covers purpose, alternative routing, return shape (chunking + offset continuation) and safe concurrent use (background tab). Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and all six parameters are documented in-schema, so the schema does the heavy lifting. The description only reinforces offset-based continuation, adding no syntax or semantics beyond what the schema already states, which is the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) plus resource (the page text in one of the agent's tabs) and even the output format (headings with #, paragraphs, lists, table text in reading order). It explicitly distinguishes itself from browser_snapshot, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use cases (read an article, results or documentation) and names the alternative with its own use case (browser_snapshot to find something to click). The selection condition between the two read-oriented siblings is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_release_tabRelease adopted tabA
Idempotent

Hand a tab taken over with browser_use_tab back to the user: the agent stops driving it and the page is left exactly as it is. Call it as soon as you are done with an adopted tab. Use browser_close_tab for tabs the agent opened itself.

ParametersJSON Schema
NameRequiredDescriptionDefault
tab_idYestab_id of a tab previously adopted with browser_use_tab
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral context beyond the annotations: the agent stops driving the tab and the page remains unchanged. This clarifies that the operation releases control rather than closing or modifying the page. It does not describe error/ownership edge cases, but the core behavioral contract is well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences with no filler: the first defines the action, the second states when to call it, and the third points to the relevant alternative. Key information is front-loaded and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with full schema coverage and helpful annotations, the description is complete enough. It conveys purpose, timing, behavior, and the key alternative without requiring an output schema or further elaboration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are well documented in the schema. The description reinforces that tab_id must refer to a previously adopted tab but does not add new semantic details beyond that. This meets the baseline for well-covered parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: it hands an adopted tab back to the user and stops agent control. It also clearly distinguishes from browser_close_tab by stating the page is left exactly as it is, so the agent can tell this tool apart from its sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit timing rule: 'Call it as soon as you are done with an adopted tab.' It also names the alternative, browser_close_tab, for tabs the agent opened itself, making when-to-use and when-not-to-use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_request_credentialsAsk the user to sign inA

Ask the user to sign in on the page in one of the agent's tabs without the values passing through you. Chrome shows the user a small window naming the site's real origin, with one field per input you list; what they type goes straight into those inputs and is never returned to you. Use it for passwords, verification codes and any secret instead of browser_type, and never ask the user to paste a secret into the conversation. Take the input indices from browser_snapshot (password inputs show as input[password]), then click the page's sign-in button yourself. Waits until the user answers, cancels, or the time runs out.

ParametersJSON Schema
NameRequiredDescriptionDefault
fieldsYesThe inputs to fill, in the order the user should see them (1 to 6)
reasonNoOne sentence the user sees explaining why you need to sign in
tab_idYestab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.
timeout_secondsNoHow long to wait for the user (default 180, at most 600)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: the values never return to the agent, Chrome names the site's real origin in its own window, one field is shown per listed input, and the call blocks until the user answers, cancels, or times out. None of that is derivable from readOnlyHint/openWorldHint/idempotentHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded with purpose and the safety rule; every sentence carries an instruction the agent needs. It runs a bit long for the number of concepts, but there is little pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the privacy model, the input-source workflow, the blocking behavior, and the tool's relationship to browser_type. It stops short of saying what the call yields on cancel or timeout, which matters since there is no output schema, but the coverage is otherwise strong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline would be 3, but the description adds real semantics: fields are shown in the order given, indices come from the latest browser_snapshot, and password inputs surface as input[password]. It does not explain the session_id or timeout_seconds behavior beyond the schema, which is fine.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Ask the user to sign in on the page in one of the agent's tabs') and explicitly distinguishes itself from the closest sibling: use this 'instead of browser_type' for secrets. An agent can route between the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (passwords, verification codes, any secret), explicit when-not (never ask the user to paste a secret into the conversation), and names the alternative (browser_type). It also supplies the surrounding workflow: pull indices from browser_snapshot, then click the sign-in button yourself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_select_tabSelect browser tabA
Idempotent

Make one of the agent's tabs the visible one in its window, for example before capturing it with screenshot. browser_snapshot, browser_click and browser_type work on background tabs, so most tasks never need this. The agent's group lives in the user's Chrome window, so this changes which tab that window shows; use it sparingly. The user's own tabs are never selected.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNo1-based position within the agent's tabs; fallback mode only, when tab_id is unavailable
tab_idNotab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses the real-world side effect: the agent's tab group lives in the user's Chrome window, so selecting a tab changes what the user sees. It also clarifies that the user's own tabs are never selectedholistically and that the action is idempotent, matching idempotentHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no filler. It front-loads the core purpose, then adds usage guidance and side-effect warnings, with every sentence earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tab-selection tool, the description covers purpose, when to use it, when not to use it, side effects, and scope boundaries. The input schema fully documents all parametersable, and there is no output schema, so no return-value explanation is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-level meaning beyond what the schema already provides, but it does reinforce the fallback relationship for index and the source of tab_id indirectly through the tool description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Make one of the agent's tabs the visible one in its window.' It also differentiates from related browser tools by noting that browser_snapshot, browser_click, and browser_type work on background tabs, clarifying why this tool is rarely needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly gives a usage example ('before capturing it with screenshot'), explains that most tasks never need it because other browser actions work on background tabs, and warns to use it sparingly because it changes the user's Chrome window. This is strong when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_snapshotSnapshot page elementsA
Read-onlyIdempotent

List the interactive elements (links, buttons, inputs) on the page in one of the agent's tabs, with the index each one has for browser_click, plus the page title and URL. Inputs show their type, such as input[password]. Works on a background tab, so the user can be looking at something else. Use it before every browser_click, because indices change when the page changes; use browser_read for the page's text. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
tab_idYestab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, but the description adds genuine context beyond them: it works on a background tab while the user views something else, indices become stale when the page changes, and inputs disclose their type (e.g. input[password]). It does not discuss pagination or element-count limits, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what the tool returns, then usage rules and the read-only flag. Every sentence carries distinct information (output contents, background-tab behavior, click prerequisite, sibling routing) with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema, the description enumerates the return shape (elements + index + title + URL) and the background-tab and index-volatility behaviors. An agent has everything needed to call it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so tab_id and session_id are already fully documented in the schema. The description only implies tab context ('one of the agent's tabs') without adding syntax or format detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the interactive elements (links, buttons, inputs)') and adds the key output detail that each element carries the index needed for browser_click, plus page title and URL. It explicitly distinguishes itself from browser_read by routing text extraction elsewhere.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a hard rule ('Use it before every browser_click, because indices change when the page changes') and names the alternative ('use browser_read for the page's text'). The when-to-use condition and the sibling routing are both explicit, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_typeType in browserA
Destructive

Type text into the field that currently has focus in one of the agent's tabs; browser_click the field first. Text is inserted at the caret without clearing existing content. Use browser_press_key for Enter, Tab, Escape or Backspace, and type_text for native apps.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesExact text to type into the focused field
tab_idYestab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.
return_stateNoAfter acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false and destructiveHint=true, so the safety profile is conveyed. The description adds genuinely new behavioral detail: text is inserted at the caret and does NOT clear existing content, which materially qualifies the destructiveHint and tells the agent the call is conditional on prior focus state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, zero filler. The core action and its focus precondition come first, alternatives second — correctly front-loaded for scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but for a mutation tool the agent needs the focus prerequisite, the non-clearing behavior, and the routing to sibling tools — all present. Only minor gaps remain (e.g. whether typed text fires input/keyboard events or how focus loss surfaces as an error).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (including tab_id provenance and return_state semantics) are already documented in the schema. The description adds no syntax for text or tab_id beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (type) and resource (focused field in one of the agent's tabs), and explicitly names the sibling tools it is not (type_text for native apps, browser_press_key for control keys). An agent can distinguish it from the other ~30 siblings immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition ('browser_click the field first') and routes two adjacent use cases to their correct alternatives: browser_press_key for Enter/Tab/Escape/Backspace and type_text for native apps. When and when-not are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_use_tabAdopt user's tabA
Idempotent

Take over a tab the user already has open, instead of opening a new one. Use this when the page is already signed in or mid-flow — a checkout, a draft, a dashboard behind SSO — and re-opening the URL would lose that state. Find the tab_id with browser_list_tabs all=true. The tab stays exactly where it is in the user's window; it is not moved into the agent's group, not activated, and not reloaded. It is never closed by cleanup — call browser_release_tab to hand it back. Ask the user before taking over a tab they are actively working in.

ParametersJSON Schema
NameRequiredDescriptionDefault
tab_idYestab_id of the user's tab, from browser_list_tabs all=true
session_idNoStable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false. The description adds substantial behavioral context: the tab stays in the user's window, is not moved, activated, or reloaded, is never closed by cleanup, and requires browser_release_tab to hand it back. This goes well beyond what annotations convey and is consistent with them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured: purpose, usage, behavior, and handoff are presented in a logical order. The description is slightly verbose but each sentence adds essential information—no redundancy. Could be trimmed slightly, but it's highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that adopts a user's tab, the description covers the crucial context: when to use it, how to obtain the tab ID, the exact behavioral guarantees (no movement/reload), and the required release step. It also covers the safety aspect of asking the user. No missing information for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters. The description repeats the tab_id source (from browser_list_tabs) but does not add new semantic meaning beyond what the schema already states. Baseline 3 is appropriate since the schema fully documents parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action: 'take over a tab the user already has open' rather than opening a new one. Distinguishes from siblings like browser_open_tab and browser_select_tab by specifying the unique behavior of adopting an existing tab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use: when page is signed in or mid-flow and re-opening would lose state. Provides the concrete step to find the tab_id via browser_list_tabs all=true. Also instructs to ask the user before taking over an active tab, which is a key usage constraint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickClickA
Destructive

Click an element by element_id (preferred: it uses the accessibility press action, so it works even when the element is scrolled out of view) or at absolute screen coordinates taken from a screenshot or zoom. Pass element_id or x and y, not both. Use browser_click for pages in the agent's Chrome tabs, right_click for context menus, and drag for press-move-release. The click reaches the target app for real and can trigger any action the user could, so read the target with get_app_state first. The agent's own pointer overlay moves to the target. On macOS the user's mouse pointer never moves; on Windows and Linux a target that ignores background input gets real mouse input, which moves the user's pointer, and the result then says via cursor.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoScreen x coordinate in points, used together with y when no element_id is given
yNoScreen y coordinate in points, used together with x when no element_id is given
element_idNoElement id from the most recent get_app_state snapshot, e.g. e12. Preferred over coordinates.
click_countNo1 for a single click (default), 2 for a double-click
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it explains that the click reaches the real app and can trigger any user-level action, that element_id uses the accessibility press action and works off-screen, that the agent's overlay moves to the target, and the platform-specific pointer side effects (macOS never moves the user's pointer; Windows/Linux may, signalled by 'via cursor' in the result). This is exactly the extra context annotations like destructiveHint cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and targeting modes, then folds routing, prerequisites, and platform caveats into a few dense sentences. Slightly long, but nearly every clause carries operational information; a minor tightening could help.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param, zero-required, destructive, annotation-covered tool with no output schema, the description covers targeting, prerequisites, exclusions, side effects, and even the shape of the result ('the result then says via cursor'). Nothing an agent needs to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, but the description still adds value: it states element_id is preferred and that element_id and x/y are mutually exclusive (0 required params means the schema does not enforce this). It does not explain click_count or state_query/return_state beyond what the schema says, so it stops short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Click an element') and immediately names the two accepted targeting modes (element_id vs x/y). It also explicitly differentiates itself from siblings (browser_click, right_click, drag), so an agent can route correctly without reading other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use routing ('Use browser_click for pages in the agent's Chrome tabs, right_click for context menus, drag for press-move-release'), a prerequisite ('read the target with get_app_state first'), and a parameter-level rule ('Pass element_id or x and y, not both').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clipboard_readRead clipboardA
Read-onlyIdempotent

Read the plain text on the system clipboard, for example after press_key cmd+c (ctrl+c off macOS) copied a selection, or when the user says they copied something for you. Prefer get_app_state, browser_snapshot or screenshot to read what is on screen; use this for text that was deliberately copied. Images and files on the clipboard are reported as no text. Read-only, but the clipboard can hold private data the user copied, such as passwords, so do not repeat it beyond what the task needs.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_charsNoMaximum characters to return (default 20000, maximum 200000). Longer text is cut off with a note giving its full length.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and idempotent behavior, so the description's safety profile is reinforced rather than duplicated. It adds useful behavioral context beyond annotations: images/files yield no text, and clipboard contents may contain private data like passwords, so output should not be repeated beyond task needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then provides essential usage guidance, exclusions, edge cases, and a privacy caveat. Every sentence earns its place; there is no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with one optional parameter and no output schema, the description covers the key conditions, alternatives, non-text edge cases, and privacy implications. Nothing critical is missing for an agent to decide when and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the sole parameter max_chars is already fully documented, including default, maximum, and truncation note. The description adds no additional parameter-specific semantics, matching the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Read the plain text on the system clipboard.' It also differentiates itself from siblings by explicitly saying to prefer get_app_state, browser_snapshot, or screenshot for on-screen content and to use this tool for text that was deliberately copied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete when-to-use examples: after press_key cmd+c / ctrl+c copies a selection, or when the user says they copied something. It also names alternatives and explicitly distinguishes on-screen reading from clipboard reading, and notes that images/files are reported as no text.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clipboard_writeWrite clipboardA
DestructiveIdempotent

Replace the system clipboard with plain text, typically so a long or multi-line value can be pasted with press_key cmd+v (ctrl+v off macOS) where set_value is rejected and type_text would be slow. This tool does not paste anything itself. Side effect: whatever the user had on the clipboard is overwritten and not restored, so tell the user when you use it. Prefer set_value or type_text when they work.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesExact text to place on the clipboard, replacing its current contents

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal destructiveness, but the description adds crucial behavior beyond them: it overwrites the user's clipboard and does not restore it, and it tells the agent to inform the user. It also clarifies that the tool only writes the clipboard and does not perform the paste, preventing a common misuse.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all necessary: the action, the use case, and the critical side-effect warning. The most decision-relevant facts are front-loaded, and there is no filler or repetition of schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers all an agent needs: what it does, when to use it, what it does not do, its side effects, and how it relates to sibling tools. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is already well described. The description adds a little extra meaning by specifying plain text and mentioning the typical use case of long or multi-line values, which helps the agent understand the character of the input beyond the schema's 'Exact text.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Replace the system clipboard with plain text.' It also explains the intended role in a larger workflow (enabling cmd+v paste) and explicitly differentiates itself from set_value and type_text, so an agent can distinguish it from siblings without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete when-to-use context: when a long or multi-line value must be pasted via press_key and set_value is rejected or type_text would be slow. It also states a clear exclusion: 'This tool does not paste anything itself' and explicitly recommends set_value or type_text when they work.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dragDragA
Destructive

Press at one point, move, and release at another to drag and drop, move a slider, or select a range. Give each end as an element id or as screen coordinates; the two ends may use different forms. A drop can move or reorder items in the app, so verify the result with get_app_state. On macOS a drag must stay inside one app window and never moves the user's pointer; on Windows and Linux most drags use real mouse input, which moves the user's pointer (the result then says via cursor).

ParametersJSON Schema
NameRequiredDescriptionDefault
to_xNoScreen x to release at, used with to_y when no to_element_id is given
to_yNoScreen y to release at, used with to_x when no to_element_id is given
from_xNoScreen x to start at, used with from_y when no from_element_id is given
from_yNoScreen y to start at, used with from_x when no from_element_id is given
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.
to_element_idNoElement to release on, from get_app_state
from_element_idNoElement to start the drag on, from get_app_state

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint/readOnlyHint) by disclosing consequences ('a drop can move or reorder items'), platform-specific input mechanics (macOS stays inside one window and never moves the pointer; Windows/Linux use real mouse input and report 'via cursor'), and that ids are invalidated by return_state. This is exactly the behavioral context annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the physical action and use cases, then the parameter-form rule, then verification and platform behavior. No filler sentences; each one changes how an agent would call the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-param, no-output-schema mutation tool, the description covers the interaction model, verification path (get_app_state / return_state), and platform caveats. Only the non-return_state response shape is left implicit, a minor gap given return_state is documented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the per-parameter docs carry the load and the baseline is 3. The description still adds non-obvious semantics: the two ends may use different forms (mixing element id and coordinates), exceeds what the schema descriptions alone convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and physical action ('press at one point, move, release at another') and enumerates the three distinct outcomes: drag-and-drop, slider movement, range selection. This is enough for an agent to separate 'drag' from click, hover, select_text, and press_key without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating conditions: each end may be an element id or coordinates and the two forms can be mixed, and the result should be verified with get_app_state. It also flags the platform constraint on macOS versus Windows/Linux. It stops short of explicit when-not-to-use guidance against siblings like click or select_text.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_app_stateRead accessibility treeA
Read-onlyIdempotent

Read an app's accessibility tree as an indented outline in which interactive elements carry ids like [e12] that click, type_text, set_value, scroll, hover and select_text accept. Use it instead of screenshot whenever you intend to act: it is far cheaper in tokens and gives exact targets. Call it before interacting and again after the UI changes, because ids are per-snapshot and a stale id fails. Read-only; it describes the app's visible windows and does not change focus.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesApp name, bundle id, or pid exactly as reported by list_apps, or "frontmost" for the app the user is looking at
queryNoOnly list elements whose role, label or value contains this text (case-insensitive). Ids stay valid. Use it instead of raising max_elements when you know what you are looking for.
max_depthNoMaximum nesting depth to descend (default 18). Lower it for a quick overview of a large window.
max_elementsNoMaximum elements to emit before the outline is truncated (default 800). Prefer `query` over raising this.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, and the description goes well beyond them: it discloses the id format ([e12]), that ids are per-snapshot and stale ids fail, that the output is a token-cheap outline, and that it does not change focus. This is exactly the extra behavioral context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tightly packed sentences, front-loaded with what the tool returns, then the screenshot comparison, then the staleness/sequencing rule, then the read-only guarantee. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining the return format (indented outline with [e12] ids) and it does so, plus it covers lifecycle (re-call after UI changes), safety (read-only, no focus change), and cost (token-cheap). Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: the [e12] id convention and which sibling tools accept those ids. It does not add much to the query/max_depth/max_elements tradeoff that the schema descriptions already state, so it lands just above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read an app's accessibility tree as an indented outline') and immediately distinguishes itself from the screenshot sibling by naming the alternative and the condition that selects this tool. An agent can tell exactly what it gets back and how it differs from the other read tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('Use it instead of screenshot whenever you intend to act'), explicit when-not/alternative, and a sequencing rule ('Call it before interacting and again after the UI changes'). It also names the tools whose parameters consume the returned ids, so the workflow is fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hoverHoverA
Idempotent

Move the agent pointer over an element or screen position without clicking, to reveal hover menus, toolbars, tooltips or drag handles. Follow with get_app_state or screenshot to see what appeared. Use click to activate. Pass element_id or x and y, not both. On macOS the user's own mouse pointer is not moved; on Windows and Linux it may be, and the result then says via cursor.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoScreen x coordinate in points, used together with y when no element_id is given
yNoScreen y coordinate in points, used together with x when no element_id is given
element_idNoElement id from the most recent get_app_state snapshot, e.g. e12
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes beyond the annotations by disclosing platform-specific side effects: on macOS the user's mouse pointer is not moved, while on Windows/Linux it may be and the result reports 'via cursor'. That is exactly the kind of environmental behavior an agent needs and that the schema cannot express. It does not, however, explain timing/settle semantics in any depth beyond the parameter descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with purpose and followed by the workflow and the platform caveat. No filler; each sentence carries a distinct piece of actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 100% schema coverage, no output schema, and annotations covering the safety profile, the description supplies the remaining gaps: expected follow-up tools, activation alternative, argument exclusivity, and platform pointer behavior. Nothing essential for a correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds a real constraint the schema only implies: element_id versus x/y are mutually exclusive, not to be combined. That extra coordination rule is meaningful beyond the per-parameter text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (move/hover) and resource (agent pointer over an element or screen position) and immediately scopes it against the sibling click with 'without clicking'. An agent can tell what this does and how it differs from click/right_click without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it (to reveal hover menus, toolbars, tooltips, drag handles), what to do next (get_app_state or screenshot), and what to use instead to activate (click). It also states the mutual-exclusivity rule: element_id or x and y, not both.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_appsList running appsA
Read-onlyIdempotent

List running applications with their bundle id, pid, window count, and which one is frontmost. Call it first to learn the exact app value that get_app_state, screenshot and activate_app accept. One app can have several running instances and only some own windows, so prefer the instance that has windows. Read-only: no window or input is touched.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation is read-only, idempotent, and non-destructive. The description adds meaningful behavioral context beyond annotations: it states no window or input is touched, and explains the multi-instance/window-ownership nuance that affects how an agent should use the results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states purpose and output fields, the second gives operational guidance with a concrete use-before pattern, and the third reassures about read-only safety. No filler or repeated annotation material.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only listing tool with no output schema, the description is complete: it enumerates the return fields, explains why the tool should be called first, and covers the instance-selection caveat. An agent has everything needed to call it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is trivially 100%, so there are no parameter details to explain. The description correctly focuses on the value of the tool's output rather than parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('running applications') and enumerates the exact fields returned (bundle id, pid, window count, frontmost). It also explains the tool's role as the discovery entry point for sibling tools, distinguishing it from the rest.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call this tool first to learn the exact 'app' value accepted by get_app_state, screenshot, and activate_app. It also provides concrete guidance on choosing the right instance when multiple exist and only some own windows.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_displaysList displaysA
Read-onlyIdempotent

List every attached display with its index, resolution and position, for use with screenshot(display: N) and for interpreting screen coordinates on multi-monitor setups. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds context beyond annotations by detailing the returned data (index, resolution, position) and its intended use. The 'Read-only' remark is redundant with annotations but not contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, front-loads the core action and result, and immediately states the use cases. Every clause adds value without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateless, parameterless read-only tool, the description is nearly complete. It states what is returned and why. A minor gap is not explicitly noting the zero-based nature of display indices or the exact output format, but these are inferable from the list of nouns and the reference to screenshot(display: N).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema coverage is 100% (empty schema fully describes the input). With no parameters to document, the baseline of 4 applies; the description correctly focuses on output and purpose rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('attached display'), and specifies what is returned: index, resolution, and position. It also explicitly ties the tool's purpose to downstream use with screenshot(display: N) and multi-monitor coordinate interpretation, making it clearly distinguishable from all sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: this is the tool to call before taking screenshots on a specific display or when working with screen coordinates on multi-monitor setups. It effectively tells an agent when to select this tool among the many siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_keyPress keyA
Destructive

Press one named key, optionally with modifiers held, e.g. key='s' modifiers=['cmd'] to save or key='return' to submit. Use it for shortcuts and navigation keys; use type_text for literal text and browser_press_key inside the agent's Chrome tabs. The key goes to the focused app, so call activate_app or click first when focus is uncertain. Shortcuts can close windows or delete content, so confirm the target before pressing.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesKey name: a single character such as 's' or '/' (resolved through the current keyboard layout), or a named key: return, tab, escape, space, delete, backspace, forwarddelete, up, down, left, right, home, end, pageup, pagedown, insert, f1 to f20, punctuation names (minus, equal, leftbracket, rightbracket, backslash, semicolon, quote, comma, period, slash, grave), numpad0 to numpad9, numpadadd, numpadsubtract, numpadmultiply, numpaddivide, numpaddecimal, numpadenter
modifiersNoModifier keys to hold while pressing: any of cmd, shift, alt, ctrl, fn. cmd maps to the Windows/Super key off macOS.
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is known; the description adds value on top by explaining that keystrokes are delivered to the focused app and that shortcuts can close windows or delete content, warranting confirmation of the target. It does not describe latency, error behavior, or what happens if the named key is unsupported, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each load-bearing: purpose+example, sibling routing, focus prerequisite, destructive warning. The primary action is front-loaded and nothing is repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does cover the feedback loop by referencing return_state and the fresh app state it appends. Combined with the focus and destructiveness guidance, an agent has enough to call this correctly, though a note on failure/unsupported-key behavior would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so key names, modifier names and return_state/state_query semantics are fully documented in the schema (baseline 3). The description adds marginal value by demonstrating a real invocation shape ('key=s', modifiers=['cmd']) and tying return_state to checking results and picking the next target, but does not add syntax beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Press one named key, optionally with modifiers held') and immediately grounds it with concrete examples ('s' with 'cmd' to save, 'return' to submit). It explicitly distinguishes itself from the two closest siblings, type_text and browser_press_key, so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use and when-not-to-use: shortcuts/navigation keys for this tool, literal text for type_text, and browser_press_key inside Chrome tabs. It also states the prerequisite for focus ('call activate_app or click first when focus is uncertain'), which is exactly the kind of routing condition an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

right_clickRight-clickA
Destructive

Right-click (secondary click) an element or screen position to open its context menu. Follow with get_app_state to read the menu items, then click one. Use click for normal activation. Pass element_id or x and y, not both. The user's pointer is treated as for click.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoScreen x coordinate in points, used together with y when no element_id is given
yNoScreen y coordinate in points, used together with x when no element_id is given
element_idNoElement id from the most recent get_app_state snapshot, e.g. e12
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false and openWorldHint=false, so the safety profile is covered structurally. The description adds the pointer-treatment behavior and the follow-up workflow, but never explains why a context-menu action is flagged destructive, which is the one non-obvious trait worth clarifying.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then the follow-up workflow, the alternative tool, and the parameter constraint. Four tight sentences with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety, a 100%-covered schema, and no output schema, the description supplies the essential workflow context an agent needs. It stops just short of explaining the destructive flag, but is otherwise complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds a genuinely useful constraint not stated in the schema: 'Pass element_id or x and y, not both.' It also implies the two targeting modes, giving the agent more than the raw property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Right-click (secondary click) an element or screen position to open its context menu') and names the distinguishing effect (opening a context menu) that separates it from the sibling click tool. An agent can identify this without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit workflow ('Follow with get_app_state to read the menu items, then click one') and names the alternative with its selection condition ('Use click for normal activation'), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotScreenshotA
Read-onlyIdempotent

Capture an app's largest window, or a whole display, as an image. The result text states the capture's screen origin and pixels-per-point so an image pixel can be converted into click or hover coordinates. Prefer get_app_state for interaction, which is cheaper and returns clickable element ids; use screenshot to verify an outcome or to see content the accessibility tree cannot describe (canvas, video, custom drawing), and zoom to read small text. Read-only; the captured window is not raised or focused.

ParametersJSON Schema
NameRequiredDescriptionDefault
appNoApp name, bundle id, or pid exactly as reported by list_apps. Captures that app's largest window. Provide either app or display, or "frontmost" for the app the user is looking at
formatNoImage encoding (default png). Use jpeg for live remote viewing.
displayNo0-based display index from list_displays. Captures the whole display instead of an app window.
qualityNoJPEG quality from 1 to 100 (default 55). Ignored for png. Raise it when small text must stay sharp.
max_widthNoDownscale the image to this width in pixels (default 1400). Lower it to save tokens.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive, but the description adds traits they do not cover: the captured window is not raised or focused, and the result text carries screen origin plus pixels-per-point so image pixels convert to click/hover coordinates. That is genuine extra behavioral context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three densely-packed sentences with the primary action, the alternative tool, and the side-effect disclosure front-loaded in that order. No filler and every clause carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by explaining the result text's coordinate-conversion data. Combined with the full schema for parameters and annotations for the safety profile, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents app/display/format/quality/max_width including defaults and valid values; that establishes the baseline of 3. The description adds only the loose hint that zoom helps read small text (related to quality/max_width) and otherwise does not extend parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: capture an app's largest window or a whole display as an image. It also names the output semantics (screen origin, pixels-per-point) and explicitly distinguishes itself from sibling get_app_state and zoom, so an agent can pick it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to prefer get_app_state for interaction because it is cheaper and returns clickable element ids, and defines exactly when screenshot is the right choice: verifying an outcome or seeing content the accessibility tree cannot describe, using zoom for small text. This is a full when/when-not/alternatives routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrollScrollA

Scroll up, down, left or right by a number of lines, over element_id when given so the right pane scrolls. Use it to bring off-screen content into view before get_app_state or screenshot. It only scrolls; nothing is clicked or selected. On macOS it scrolls the target app in the background without moving the user's pointer. On Windows and Linux, without element_id it scrolls whatever is under the user's pointer, and an element with no background route gets the real pointer moved onto it (the result then says via cursor).

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoScreen x to scroll over. Remote control only: the pointer moves there first. Ignored otherwise.
yNoScreen y to scroll over. Remote control only: the pointer moves there first. Ignored otherwise.
amountNoNumber of scroll lines, 1 to 100 (default 5)
directionNoScroll direction (default down)
element_idNoElement to scroll, from get_app_state; the nearest scrollable area around it moves. Omit to scroll the last inspected app (macOS) or whatever is under the pointer (Windows, Linux).
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only covering safety flags, the description carries the real behavioral burden and does so: platform-specific side effects (macOS scrolls the target app in the background without moving the pointer; Windows/Linux scroll under the pointer or move the real pointer onto elementless targets), plus the 'via cursor' result marker. This is exactly the side-effect disclosure an agent needs for a non-read-only tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and then usage and side effects; no filler. The final platform sentence is dense but every clause is load-bearing for correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, no-output-schema tool with platform-dependent behavior, the description covers the action, usage, and cross-platform side effects well. It leaves return_state/state_query semantics entirely to the schema, which is acceptable, but a note on what the scroll call returns by default would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents every parameter, including element_id fallback behavior and the x/y remote-control caveat. The description's parameter mentions (element_id 'so the right pane scrolls', x/y platform behavior) largely restate the schema rather than adding new syntax or constraints, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scroll) plus resource/axis and line count, and names the targeting parameter (element_id) that selects which pane moves. An agent can tell this apart from click, drag, and hover siblings immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it (bring off-screen content into view before get_app_state or screenshot) and draws a hard boundary against alternatives ('It only scrolls; nothing is clicked or selected'). This is direct routing guidance rather than implied usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_textSelect textA
Idempotent

Select a character range inside a text element through the accessibility API, for example to copy part of a value or to replace just that part with type_text. Defaults to selecting from start to the end of the value. Use set_value to replace the whole value instead. Only the selection changes; the text is not modified.

ParametersJSON Schema
NameRequiredDescriptionDefault
startNoZero-based character offset to start the selection at (default 0)
lengthNoNumber of characters to select (default: through the end of the value)
element_idYesText element to select in, from get_app_state
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare idempotentHint=true and destructiveHint=false, so safety is partly covered, but readOnlyHint=false is potentially confusing; the description resolves this well by clarifying 'Only the selection changes; the text is not modified.' It also documents the default selection range. It does not mention focus/in-view requirements or whether the target must already be focused, which is a real behavioral gap for a UI selection tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose first, then usage/alternative, then the non-mutation guarantee. No filler or repetition, and the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation-adjacent UI tool with no output schema, the description covers what the action does, its non-destructive effect, and its alternative. The main omission is focus/visibility prerequisites and what happens if element_id is stale, but nothing critical to invoking it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so start, length, element_id, state_query and return_state are already fully documented. The description's default note ('from start to the end of the value') merely restates what the schema says for those params. Baseline 3 is appropriate when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Select a character range inside a text element through the accessibility API.' It also explicitly distinguishes itself from set_value and type_text, so an agent can route correctly without opening sibling schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases ('to copy part of a value or to replace just that part with type_text') and names the alternative with its selection condition ('Use set_value to replace the whole value instead'). When-to-use and when-to-use-something-else are both explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_valueSet field valueA
DestructiveIdempotent

Replace a text field's entire contents in one step through the accessibility API, without keystrokes. Prefer it over type_text for long values or when the field already holds text; fall back to click plus type_text if the field rejects it, which the result reports. The previous value is discarded.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYesNew complete value for the field
element_idYesText field to set, from get_app_state
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is partly covered. The description adds meaningful context beyond them: the previous value is discarded, rejection is reported in the result, and no keystrokes are synthesized. A small gap remains on whether this requires focus or specific element state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with zero waste: mechanism first, then preference/alternative routing, then the destructive consequence. Each sentence earns its place and the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but the description notes that rejection is reported in the result and that return_state appends fresh state, which is enough for an agent to interpret outcomes. Minor gaps remain on prerequisites such as whether the field must be focused first.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters, including the return_state behavior. The description adds nothing about parameter formats beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (replace/Set) and resource (a text field's entire contents) plus the mechanism (accessibility API, no keystrokes). It explicitly distinguishes itself from the type_text sibling, so an agent can route correctly without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('prefer over type_text for long values or when the field already holds text') and a when-not/fallback path ('fall back to click plus type_text if the field rejects it'). Both the preferred case and the failure case are named with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

type_textType textA
Destructive

Type literal text as keystrokes into the field that currently has focus, optionally focusing element_id first. Use it for short entries and for fields that reject set_value; use set_value to replace a long value in one step, and press_key for shortcuts or keys such as return and tab. Text is inserted at the caret without clearing what is already there. Typing into password fields is refused by default (see COMPUTER_USE_ALLOW_SECURE_FIELD_INPUT).

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesExact text to type, character by character
element_idNoElement to focus before typing, from get_app_state. Omit to type into whatever currently has focus.
state_queryNoWith return_state, list only elements whose role, label or value contains this text, like get_app_state's query
return_stateNoAfter acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds real context beyond them: text is inserted at the caret without clearing existing content, and password fields are refused by default with a named env override. This is meaningful behavioral disclosure, though response/error behavior is not covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then when-to-use, then behavioral caveats. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, routing, and key behavioral caveats for a 4-param mutation tool with no output schema. Sufficient to invoke correctly, though error/return behavior is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including element_id focus and return_state. The description reinforces element_id focusing but adds little syntax beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Type literal text as keystrokes into the field that currently has focus') and explicitly distinguishes itself from set_value and press_key.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('short entries', 'fields that reject set_value') plus named alternatives for the other two cases (set_value for long replacements, press_key for shortcuts/return/tab). Routing is fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

waitWaitA
Read-onlyIdempotent

Pause before the next action so the UI can catch up: page loads, animations, dialogs opening, apps launching. Follow with get_app_state or screenshot to confirm the new state instead of guessing. Sends no input.

ParametersJSON Schema
NameRequiredDescriptionDefault
secondsNoSeconds to wait (default 1, maximum 30)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds useful behavioral context beyond that: the wait exists to let UI transitions settle, and it sends no input. This enriches the annotation profile without contradicting it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two purposeful sentences front-load the core purpose and usage context, then add the verification follow-up and the behavioral constraint. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple no-input pausing tool with strong annotations and a fully described schema, the description covers why, when, and how to confirm the result. Nothing an agent needs to correctly invoke this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the only parameter (seconds) is already fully documented with default and maximum. The description does not need to repeat this, but it also adds no additional nuance about how the parameter interacts with wait behavior. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Pause') and gives concrete use cases (page loads, animations, dialogs opening, apps launching), so an agent can tell exactly what the tool does. It also separates this action-oriented tool from siblings like get_app_state or screenshot by noting it 'sends no input.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when to use the tool: before the next action when the UI needs to catch up. It also advises following up with get_app_state or screenshot to confirm the new state instead of guessing, giving actionable post-wait guidance. It does not explicitly discuss when not to use it, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

zoomZoom into regionA
Read-onlyIdempotent

Capture one region of the screen at full resolution, to read small text, dense tables, file names or tiny controls that a normal screenshot blurs. Give the region as two corners in screen coordinates (the same space click uses); the result text explains how to map pixels in the zoomed image back to screen coordinates. Use screenshot for a whole window and get_app_state when the text is exposed by accessibility. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
x0YesLeft edge, screen coordinates
x1YesRight edge, screen coordinates
y0YesTop edge, screen coordinates
y1YesBottom edge, screen coordinates
max_widthNoDownscale the zoomed image to this width in pixels (default 1400)

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive behavior, so the description does not need to repeat that. It adds useful behavioral context: the coordinate space is the same as click, and the result text explains how to map pixels back to screen coordinates. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose/use cases, input format plus result behavior, and alternative routing. There is no filler or repetition; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only capture tool with fully documented parameters, the description covers why to use it, when to avoid it, how to specify input, and what the result will indicate. No output schema exists, but the description does disclose the essential result behavior, so the agent has what it needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds meaningful extra context by clarifying that the region is defined by two corners in the same coordinate space that click uses. It also explains why that matters by mentioning pixel-to-screen mapping in the result.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Capture one region of the screen at full resolution.' It also names concrete use cases (small text, dense tables, file names, tiny controls) and is clearly differentiated from the sibling screenshot tool by emphasizing regions at full resolution.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent when to use this tool versus alternatives: 'Use screenshot for a whole window and get_app_state when the text is exposed by accessibility.' This gives direct routing guidance beyond mere inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv0.5.2
    • Changedbrowser_click1 field changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.",
        +  "type": "boolean"
        +}
    • Changedbrowser_navigate1 field changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.",
        +  "type": "boolean"
        +}
    • Changedbrowser_press_key1 field changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.",
        +  "type": "boolean"
        +}
    • Addedbrowser_read
    • Addedbrowser_request_credentials
    • Changedbrowser_type1 field changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait for the page to settle (and finish loading, if the action navigated) and append a fresh browser_snapshot of this tab, so you can pick the next index in the same call. Its indices replace earlier ones.",
        +  "type": "boolean"
        +}
    • Changedclick2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changeddrag2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changedget_app_state1 field changed
      • changedInput schema / properties / app / description
        Previous value: -"App name, bundle id, or pid exactly as reported by list_apps"New value: +"App name, bundle id, or pid exactly as reported by list_apps, or \"frontmost\" for the app the user is looking at"
    • Changedhover2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changedpress_key2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changedright_click2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changedscreenshot2 fields changed
      • changedInput schema / properties / app / description
        Previous value: -"App name, bundle id, or pid exactly as reported by list_apps. Captures that app's largest window. Provide either app or display."New value: +"App name, bundle id, or pid exactly as reported by list_apps. Captures that app's largest window. Provide either app or display, or \"frontmost\" for the app the user is looking at"
      • addedInput schema / properties / quality
        Added value: +{
        +  "description": "JPEG quality from 1 to 100 (default 55). Ignored for png. Raise it when small text must stay sharp.",
        +  "maximum": 100,
        +  "minimum": 1,
        +  "type": "integer"
        +}
    • Changedscroll2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changedselect_text2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changedset_value2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
    • Changedtype_text2 fields changed
      • addedInput schema / properties / return_state
        Added value: +{
        +  "description": "After acting, wait briefly for the UI to settle and append a fresh get_app_state of the app you last read, so you can check the result and pick the next target in the same call. Its ids replace every earlier id.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / state_query
        Added value: +{
        +  "description": "With return_state, list only elements whose role, label or value contains this text, like get_app_state's query",
        +  "type": "string"
        +}
  2. 12 tool updatesv0.4.3
    • Changedbrowser_click1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_close_all_tabs1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_close_tab1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_list_tabs1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_navigate1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_open_tab1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_press_key1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_release_tab1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_select_tab1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_snapshot1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_type1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedbrowser_use_tab1 field changed
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Stable task or thread ID. Pass the same value on every browser call for this task. Different IDs isolate tabs and cleanup within one MCP process. Omit for the default process session.",
        +  "maxLength": 128,
        +  "minLength": 1,
        +  "type": "string"
        +}
  3. 1 tool updatev0.4.2
    • Changedscroll2 fields changed
      • addedInput schema / properties / x
        Added value: +{
        +  "description": "Screen x to scroll over. Remote control only: the pointer moves there first. Ignored otherwise.",
        +  "type": "number"
        +}
      • addedInput schema / properties / y
        Added value: +{
        +  "description": "Screen y to scroll over. Remote control only: the pointer moves there first. Ignored otherwise.",
        +  "type": "number"
        +}
  4. 4 tool updatesv0.4.0
    • Addedclipboard_read
    • Addedclipboard_write
    • Changedpress_key1 field changed
      • changedInput schema / properties / key / description
        Previous value: -"Key name: a single character such as 's', or a named key such as return, tab, escape, space, delete, backspace, up, down, left, right, home, end, pageup, pagedown"New value: +"Key name: a single character such as 's' or '/' (resolved through the current keyboard layout), or a named key: return, tab, escape, space, delete, backspace, forwarddelete, up, down, left, right, home, end, pageup, pagedown, insert, f1 to f20, punctuation names (minus, equal, leftbracket, rightbracket, backslash, semicolon, quote, comma, period, slash, grave), numpad0 to numpad9, numpadadd, numpadsubtract, numpadmultiply, numpaddivide, numpaddecimal, numpadenter"
    • Changedscroll2 fields changed
      • changedInput schema / properties / amount / description
        Previous value: -"Number of scroll lines (default 5)"New value: +"Number of scroll lines, 1 to 100 (default 5)"
      • changedInput schema / properties / element_id / description
        Previous value: -"Element to position the pointer over before scrolling, from get_app_state. Omit to scroll at the current pointer position."New value: +"Element to scroll, from get_app_state; the nearest scrollable area around it moves. Omit to scroll the last inspected app (macOS) or whatever is under the pointer (Windows, Linux)."
  5. 23 tool updatesv0.3.1
    • Changedactivate_app1 field changed
      • addedInput schema / properties / app / description
        Added value: +"App name, bundle id, or pid exactly as reported by list_apps"
    • Changedbrowser_click4 fields changed
      • changedInput schema / properties / index / description
        Previous value: -"Element index from browser_snapshot"New value: +"Element index from the latest browser_snapshot of this tab. Preferred over coordinates."
      • addedInput schema / properties / tab_id / description
        Added value: +"tab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs"
      • addedInput schema / properties / x / description
        Added value: +"Page x coordinate in CSS pixels, used together with y when no index is given"
      • addedInput schema / properties / y / description
        Added value: +"Page y coordinate in CSS pixels, used together with x when no index is given"
    • Changedbrowser_close_tab2 fields changed
      • changedInput schema / properties / index / description
        Previous value: -"1-based index, fallback mode only"New value: +"1-based position within the agent's tabs; fallback mode only, when tab_id is unavailable"
      • changedInput schema / properties / tab_id / description
        Previous value: -"From browser_list_tabs"New value: +"tab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs"
    • Changedbrowser_navigate2 fields changed
      • addedInput schema / properties / tab_id / description
        Added value: +"tab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs"
      • addedInput schema / properties / url / description
        Added value: +"Absolute URL to load in the tab"
    • Changedbrowser_open_tab1 field changed
      • changedInput schema / properties / url / description
        Previous value: -"URL to open (default about:blank)"New value: +"Absolute URL to open (default about:blank)"
    • Changedbrowser_press_key2 fields changed
      • addedInput schema / properties / key / description
        Added value: +"Key to press"
      • addedInput schema / properties / tab_id / description
        Added value: +"tab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs"
    • Changedbrowser_release_tab1 field changed
      • addedInput schema / properties / tab_id / description
        Added value: +"tab_id of a tab previously adopted with browser_use_tab"
    • Changedbrowser_select_tab2 fields changed
      • changedInput schema / properties / index / description
        Previous value: -"1-based index, fallback mode only"New value: +"1-based position within the agent's tabs; fallback mode only, when tab_id is unavailable"
      • changedInput schema / properties / tab_id / description
        Previous value: -"From browser_list_tabs"New value: +"tab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs"
    • Changedbrowser_snapshot1 field changed
      • changedInput schema / properties / tab_id / description
        Previous value: -"From browser_open_tab or browser_list_tabs"New value: +"tab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs"
    • Changedbrowser_type2 fields changed
      • addedInput schema / properties / tab_id / description
        Added value: +"tab_id of one of the agent's tabs, from browser_open_tab or browser_list_tabs"
      • addedInput schema / properties / text / description
        Added value: +"Exact text to type into the focused field"
    • Changedbrowser_use_tab1 field changed
      • changedInput schema / properties / tab_id / description
        Previous value: -"From browser_list_tabs all=true"New value: +"tab_id of the user's tab, from browser_list_tabs all=true"
    • Changedclick4 fields changed
      • changedInput schema / properties / click_count / description
        Previous value: -"1 for single, 2 for double-click"New value: +"1 for a single click (default), 2 for a double-click"
      • changedInput schema / properties / element_id / description
        Previous value: -"Element id from get_app_state, e.g. e12"New value: +"Element id from the most recent get_app_state snapshot, e.g. e12. Preferred over coordinates."
      • addedInput schema / properties / x / description
        Added value: +"Screen x coordinate in points, used together with y when no element_id is given"
      • addedInput schema / properties / y / description
        Added value: +"Screen y coordinate in points, used together with x when no element_id is given"
    • Changeddrag6 fields changed
      • addedInput schema / properties / from_element_id / description
        Added value: +"Element to start the drag on, from get_app_state"
      • addedInput schema / properties / from_x / description
        Added value: +"Screen x to start at, used with from_y when no from_element_id is given"
      • addedInput schema / properties / from_y / description
        Added value: +"Screen y to start at, used with from_x when no from_element_id is given"
      • addedInput schema / properties / to_element_id / description
        Added value: +"Element to release on, from get_app_state"
      • addedInput schema / properties / to_x / description
        Added value: +"Screen x to release at, used with to_y when no to_element_id is given"
      • addedInput schema / properties / to_y / description
        Added value: +"Screen y to release at, used with to_x when no to_element_id is given"
    • Changedget_app_state3 fields changed
      • changedInput schema / properties / app / description
        Previous value: -"App name, bundle id, or pid"New value: +"App name, bundle id, or pid exactly as reported by list_apps"
      • changedInput schema / properties / max_depth / description
        Previous value: -"Max tree depth (default 18)"New value: +"Maximum nesting depth to descend (default 18). Lower it for a quick overview of a large window."
      • changedInput schema / properties / max_elements / description
        Previous value: -"Max elements to emit (default 800)"New value: +"Maximum elements to emit before the outline is truncated (default 800). Prefer `query` over raising this."
    • Changedhover3 fields changed
      • changedInput schema / properties / element_id / description
        Previous value: -"Element id from get_app_state"New value: +"Element id from the most recent get_app_state snapshot, e.g. e12"
      • addedInput schema / properties / x / description
        Added value: +"Screen x coordinate in points, used together with y when no element_id is given"
      • addedInput schema / properties / y / description
        Added value: +"Screen y coordinate in points, used together with x when no element_id is given"
    • Changedpress_key2 fields changed
      • addedInput schema / properties / key / description
        Added value: +"Key name: a single character such as 's', or a named key such as return, tab, escape, space, delete, backspace, up, down, left, right, home, end, pageup, pagedown"
      • changedInput schema / properties / modifiers / description
        Previous value: -"Any of cmd, shift, alt, ctrl, fn. cmd maps to the Windows/Super key off macOS."New value: +"Modifier keys to hold while pressing: any of cmd, shift, alt, ctrl, fn. cmd maps to the Windows/Super key off macOS."
    • Changedright_click3 fields changed
      • addedInput schema / properties / element_id / description
        Added value: +"Element id from the most recent get_app_state snapshot, e.g. e12"
      • addedInput schema / properties / x / description
        Added value: +"Screen x coordinate in points, used together with y when no element_id is given"
      • addedInput schema / properties / y / description
        Added value: +"Screen y coordinate in points, used together with x when no element_id is given"
    • Changedscreenshot3 fields changed
      • changedInput schema / properties / app / description
        Previous value: -"App name, bundle id, or pid"New value: +"App name, bundle id, or pid exactly as reported by list_apps. Captures that app's largest window. Provide either app or display."
      • changedInput schema / properties / display / description
        Previous value: -"Capture a whole display by index (see list_displays) instead of an app window"New value: +"0-based display index from list_displays. Captures the whole display instead of an app window."
      • changedInput schema / properties / max_width / description
        Previous value: -"Downscale to this width in pixels (default 1400)"New value: +"Downscale the image to this width in pixels (default 1400). Lower it to save tokens."
    • Changedscroll3 fields changed
      • changedInput schema / properties / amount / description
        Previous value: -"Scroll lines (default 5)"New value: +"Number of scroll lines (default 5)"
      • addedInput schema / properties / direction / description
        Added value: +"Scroll direction (default down)"
      • addedInput schema / properties / element_id / description
        Added value: +"Element to position the pointer over before scrolling, from get_app_state. Omit to scroll at the current pointer position."
    • Changedselect_text3 fields changed
      • addedInput schema / properties / element_id / description
        Added value: +"Text element to select in, from get_app_state"
      • changedInput schema / properties / length / description
        Previous value: -"Characters to select (default: to end)"New value: +"Number of characters to select (default: through the end of the value)"
      • changedInput schema / properties / start / description
        Previous value: -"Start offset (default 0)"New value: +"Zero-based character offset to start the selection at (default 0)"
    • Changedset_value2 fields changed
      • addedInput schema / properties / element_id / description
        Added value: +"Text field to set, from get_app_state"
      • addedInput schema / properties / value / description
        Added value: +"New complete value for the field"
    • Changedtype_text2 fields changed
      • changedInput schema / properties / element_id / description
        Previous value: -"Focus this element before typing"New value: +"Element to focus before typing, from get_app_state. Omit to type into whatever currently has focus."
      • addedInput schema / properties / text / description
        Added value: +"Exact text to type, character by character"
    • Changedwait1 field changed
      • changedInput schema / properties / seconds / description
        Previous value: -"Seconds to wait (default 1, max 30)"New value: +"Seconds to wait (default 1, maximum 30)"
  6. 28 tool updatesv0.3.0
    • First observedactivate_app
    • First observedbrowser_click
    • First observedbrowser_close_all_tabs
    • First observedbrowser_close_tab
    • First observedbrowser_list_tabs
    • First observedbrowser_navigate
    • First observedbrowser_open_tab
    • First observedbrowser_press_key
    • First observedbrowser_release_tab
    • First observedbrowser_select_tab
    • First observedbrowser_snapshot
    • First observedbrowser_type
    • First observedbrowser_use_tab
    • First observedclick
    • First observeddrag
    • First observedget_app_state
    • First observedhover
    • First observedlist_apps
    • First observedlist_displays
    • First observedpress_key
    • First observedright_click
    • First observedscreenshot
    • First observedscroll
    • First observedselect_text
    • First observedset_value
    • First observedtype_text
    • First observedwait
    • First observedzoom

TDQS

A4.1/5.0

Scored across 32 tools

Disambiguation4/5

Tools are mostly distinct and descriptions explicitly separate native-app actions from browser actions, with cross-references like click vs browser_click and type_text vs browser_type. Some overlap remains among tab lifecycle and input tools, but an agent can generally tell them apart.

Naming Consistency4/5

Names use consistent snake_case and a clear browser_ prefix for browser-specific tools. Minor deviations like screenshot, zoom, wait, and clipboard_read/clipboard_write are still readable and predictable.

Tool Count2/5

At 32 tools, the set is heavy for a single server, exceeding the 25+ threshold in the rubric. While computer use is broad, several browser tab and input tools could be consolidated without losing core capability.

Completeness4/5

Coverage is strong across native app control, browser automation, tabs, clipboard, displays, and credential handling. Minor gaps include window management, browser drag/hover/select-text, double-click, and file upload/download workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    An open-source MCP server for macOS and Windows that provides native desktop control via Accessibility APIs, OCR, and Chrome CDP. It enables AI agents to interact with applications, manage browser sessions, and automate workflows with high-speed native UI actions.
    111 npm
    15
    AGPL 3.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that gives any AI assistant eyes and hands on your desktop — screenshots, clicking, typing, OCR, window management, accessibility-tree queries, workflow recording.
    5
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    An MCP server that gives AI agents real OS-level control of macOS, enabling them to click real buttons, type real keys, and observe rendered screens just like a human would.
    128 PyPI
    1
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    macOS MCP server that enables AI agents to directly control the host OS, including mouse, keyboard, windows, files, and accessibility automation for computer-use workflows.
    1
    -