Skip to main content
Glama

desktop-touch-mcp

desktop-touch-mcp MCP server

日本語

Computer-use MCP server for Windows. Lets Claude, Cursor, or any MCP client see and operate your Windows 10/11 desktop — screenshots, UI Automation, Chrome CDP, keyboard / mouse, terminal — with semantic discover-then-act targeting that avoids pixel-coordinate guessing, and per-action perception guards that catch wrong-window typing before it happens.

npx -y @harusame64/desktop-touch-mcp

32 tools, native Rust engine (UIA in 2 ms), zero-config PowerShell fallback, full CJK support, MIT licensed. Add the snippet above to your Claude / Cursor / VS Code Copilot config and Claude can drive Notepad, Excel, Chrome, Windows Terminal, and any other app on your machine.

Why this over pixel-clicking? Two ideas run through every tool: discover-then-act — desktop_discover returns interactive entities with short-lived leases instead of raw coordinates, so desktop_act operates on what you mean, not where it was — and per-action perception guards that verify the target window's identity and bounds before input lands, catching wrong-window typing and stale-coordinate clicks before they happen.

Under the hood: an 82× average speedup from the Rust native engine (UIA focus queries in 2 ms, SSE2-accelerated image diffing at 13–15×), with a transparent PowerShell fallback when the engine is absent. The npm launcher fetches only the GitHub Release tag matching the installed version and verifies the Windows runtime zip before extraction.


Features

  • ⚡ High-performance Rust Native Core — The UIA bridge and image-diff engine are written in Rust (napi-rs + windows-rs) and loaded as a native .node addon. Direct COM calls from a dedicated MTA thread eliminate PowerShell process spawning — getFocusedElement completes in 2 ms (160× faster), and getUiElements returns full trees in ~100 ms with a batch BFS algorithm that minimizes cross-process RPC. Image-diff operations use SSE2 SIMD for 13–15× throughput. When the native engine is unavailable, every function transparently falls back to PowerShell — zero config required.

  • 🎯 Set-of-Marks (SoM) visual fallback — Games, RDP sessions, and non-accessible Electron apps return clickable elements even when UIA is completely blind. screenshot(detail="text") automatically detects UIA sparsity and activates a Hybrid Non-CDP pipeline: Rust-powered grayscale + bilinear upscale → Windows OCR → clustering → red bounding-box annotation with numbered badges ([1], [2]…). Two parallel representations returned: a visual PNG for spatial orientation and a semantic elements[] list with clickAt coords — no CDP required.

  • 🔁 One-call confirmation on visual-only targets — On UIA-blind targets (Electron, PWAs, games, custom canvases, RDP windows), desktop_act can fold the post-action confirmation into its own response: an optional roiCapture carrying a PNG crop of just the region that changed plus a lease-less preview of the controls now visible there. The agent confirms what its click did and finds the next target without a separate desktop_state + screenshot. On visual-only targets it is on by default for a visible change (returnCapture:"on-change"); pass returnCapture:"never" to suppress it, or "always" to force it. Never attached on structured targets (browser/CDP, UIA-rich native), where desktop_state is cheaper and exact — so those responses are unchanged.

  • 🔐 Key Locker — the terminal autofills your SSH / sudo passwords — Save a credential once into the locker's own secure dialog (stored encrypted on your machine with Windows DPAPI; never shown to the assistant), then run ssh / sudo in a console opened by key_locker(action='launch_console') — the password is filled in automatically when the hidden prompt appears, with a per-fill confirmation prompt by default. See Key Locker.

  • LLM-native design — Built around how LLMs think, not how humans click. run_macro batches multiple operations into a single API call; diffMode sends only the windows that changed since the last frame. Minimal tokens, minimal round-trips.

  • Reactive Perception Graph — Register a lensId for a window or browser tab, pass it to action tools, and get guard-checked post.perception feedback after each action. It reduces repeated screenshot / desktop_state calls and prevents wrong-window typing or stale-coordinate clicks.

  • Full CJK support — Uses Win32 GetWindowTextW for window titles, avoiding nut-js garbling. IME bypass input supported for Japanese/Chinese/Korean environments.

  • 3-tier token reduction — detail="image" (~443 tok) / detail="text" (~100–300 tok) / diffMode=true (~160 tok). Send pixels only when you actually need to see them.

  • 1:1 coordinate mode — dotByDot=true captures at native resolution (WebP). Image pixel = screen coordinate — no scale math needed. With origin+scale passed to mouse_click, the server converts coords for you — eliminating off-by-one / scale bugs.

  • Browser capture data reduction — grayscale=true (~50% size), dotByDotMaxDimension=1280 (auto-scaled with coord preservation), and windowTitle + region sub-crops help exclude browser chrome and other irrelevant pixels. Typical reduction for heavy captures: 50–70%.

  • Chromium smart fallback — detail="text" on Chrome/Edge/Brave auto-skips UIA (prohibitively slow there) and runs Windows OCR. hints.chromiumGuard + hints.ocrFallbackFired flag the path taken.

  • UIA element extraction — detail="text" returns button names and clickAt coords as JSON. Claude can click the right element without ever looking at a screenshot.

  • Auto-dock CLI — window_dock(action='dock') snaps any window to a screen corner with always-on-top. Set DESKTOP_TOUCH_DOCK_TITLE='@parent' to auto-dock the terminal hosting Claude on MCP startup — the process-tree walker finds the right window regardless of title.

  • Emergency stop (Failsafe) — Park the mouse in the top-left corner of the primary monitor (within 10px of 0,0) for 500ms to trigger the emergency stop.


Related MCP server: pywinauto-mcp

Requirements

OS

Windows 10 / 11 (64-bit)

Node.js

v20+ recommended (tested on v22+) — **to develop or run the test suite, `^22.12

PowerShell

5.1+ (bundled with Windows) — used only as fallback when the Rust native engine is unavailable

Claude CLI

claude command must be available

Note: nut-js native bindings require the Visual C++ Redistributable. Download from Microsoft if not already installed.

Note (Key Locker): The credential helper Key Locker uses is an unsigned executable, so on some machines Windows SmartScreen or antivirus may show an "unknown publisher" warning the first time it runs. This is expected — the helper ships with desktop-touch-mcp and runs locally on your machine; you can allow it to proceed. Code signing is planned for a future release.


Installation

npx -y @harusame64/desktop-touch-mcp

The npm launcher resolves runtime strictly by npm package version. For package X.Y.Z, it fetches only GitHub Release tag vX.Y.Z, downloads desktop-touch-mcp-windows.zip, verifies its SHA256 digest, and only then expands it under %USERPROFILE%\.desktop-touch-mcp. Verified cached releases are reused on later runs.

Set DESKTOP_TOUCH_MCP_HOME to override the cache root directory.

On a shared or CI network? The first run reads the GitHub Releases API to locate the runtime zip. The anonymous limit is 60 requests/hour per IP, which a shared public address (CI runners, office NAT) can exhaust before your download even starts. Set GITHUB_TOKEN (or GH_TOKEN) in the environment and the launcher authenticates the request, raising the limit to 5,000 requests/hour. No token is needed on an ordinary home connection.

Running the launcher from a source checkout? A source build's bin/launcher.js carries a placeholder integrity hash (sha256: "PENDING") instead of a finalized one. Rather than download and run an unverified runtime, the launcher fails closed — this guard stops an accidentally published or unfinalized launcher from silently starting unverified code. Published npm releases always ship a real SHA256, so end users never see this. If you are intentionally running the launcher from source, set DESKTOP_TOUCH_MCP_ALLOW_UNVERIFIED=1 to skip integrity verification (development only).

Does your host give up before the launcher finishes? Some desktop hosts allow a plugin a fixed budget — 60 seconds is common — to become ready, and a launcher waiting on an unreachable GitHub can spend all of it. Two environment variables cover that case.

DESKTOP_TOUCH_MCP_FETCH_TIMEOUT_MS (default 15000) bounds how long the launcher waits without hearing from GitHub. It applies to the release lookup and to the download; for the download it counts silence rather than total time, so a large runtime still installs over a slow connection. A value that is not a positive number of milliseconds is ignored with a warning.

DESKTOP_TOUCH_MCP_OFFLINE_FALLBACK=1 lets the launcher start a release that is already installed when GitHub cannot be reached at all. It is off by default. GitHub is always contacted first, so a reachable network still re-downloads and repairs a damaged install; only a network failure reaches the fallback, which starts the copy of your version on disk — without re-verification — or, when that version was never installed, the newest older release that completed a verified install. Answers that are not network failures (a 404, the API rate limit, a mismatched integrity hash) still stop startup loudly. Leave it off unless a host timeout forces your hand: while it is set, a corrupted install of your current version is reused instead of being repaired.

The two work together: with the fallback on, startup still waits out the timeout before falling back, so lower DESKTOP_TOUCH_MCP_FETCH_TIMEOUT_MS if your host's budget is tight. Note also that a download which is still arriving, however slowly, is never interrupted — the fallback answers when the network has gone silent, not when it is merely slow.

Register with Claude CLI

Add to ~/.claude.json under mcpServers:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"]
    }
  }
}

No system prompt needed. The command reference is automatically injected into Claude via the MCP initialize response's instructions field.

Register with other clients (HTTP mode)

Clients that require an HTTP endpoint (GPT Desktop, VS Code Copilot, Cursor, etc.) can use the built-in Streamable HTTP transport:

npx -y @harusame64/desktop-touch-mcp --http
# or with a custom port:
npx -y @harusame64/desktop-touch-mcp --http --port 8080

The server starts at http://127.0.0.1:23847/mcp (localhost only). Register the URL in your MCP client settings. A health check is available at http://127.0.0.1:<port>/health.

In HTTP mode the system tray icon shows the active URL and provides quick-copy and open-in-browser shortcuts.

Development install

git clone https://github.com/Harusame64/desktop-touch-mcp.git
cd desktop-touch-mcp
npm install

Build after install:

npm run build

For a local checkout, register the built server directly:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "node",
      "args": ["D:/path/to/desktop-touch-mcp/dist/index.js"]
    }
  }
}

Note: Replace D:/path/to/desktop-touch-mcp with the actual path where you cloned this repository.


Tools (32 Optimized Tools)

📖 Full Reference: docs/system-overview.md — Exhaustive guide on parameters, return schemas, and coordinate math.

🌐 World-Graph V2 (Primary Path)

Tool

Description

desktop_discover

Observe the desktop. Returns interactive entities with leases (UIA, CDP, Terminal, Visual SoM).

desktop_act

Perform actions (click, type, drag) on entities via lease validation. Returns semantic diffs — plus an optional roiCapture (changed-region PNG + next-target preview) on visual-only targets.

👁️ Observation & State

Tool

Description

desktop_state

Lightweight check of focus, active window, cursor, and Auto-Perception attention signal.

screenshot

Multi-mode capture: detail='text' (UIA/OCR), diffMode (P-frame), dotByDot (1:1), and background. Returns a cheap screenshot://by-ref/{id} link to the saved image instead of inlining pixels every time.

screenshot_query / screenshot_gc

Inspect and prune the on-disk screenshot cache behind the by-ref links: screenshot_query lists saved captures without re-reading pixels; screenshot_gc reclaims space by retention policy (dry-run by default).

workspace_snapshot

Instant session orientation: all window thumbnails + UI summaries in one call.

server_status

Diagnostic check for native engine health and feature activation.

⌨️ Input & Control

Tool

Description

keyboard

Send keyboard input. Supports background input (WM_CHAR) and IME-safe clipboard bypass.

mouse_click / mouse_drag

Precision coordinate-based interaction with homing and force-focus protection.

scroll

Multi-strategy: raw (notches), to_element, smart (virtual lists), and capture (stitch).

click_element

Legacy UIA-based click by name/ID (fallback when entities are unavailable).

🌐 Browser CDP (Chrome/Edge/Brave)

Tool

Description

browser_open / browser_navigate

Idempotent debug-mode launch and reliable navigation.

browser_click / browser_fill / browser_form

High-level DOM interaction stable across repaints and framework re-renders.

browser_eval

Deep inspection via js (scripting), dom (HTML), and appState (SPA data extraction).

browser_overview / browser_search / browser_locate

Semantic discovery, grep-like DOM search, and pixel-accurate coordinate lookup.

🛠️ Utilities & Workflow

Tool

Description

terminal

Unified command execution: run (send + wait + read), read (OCR/UIA), and send. run completion modes: quiet, pattern, and exit (waits for the command to finish + returns its exit code — see Terminal command completion).

wait_until

Efficient server-side polling for window, focus, text, or URL state changes.

window_dock / focus_window

Window management: pin (always-on-top), unpin, dock (corner snap), and focus.

workspace_launch

Launch apps and auto-detect new HWNDs (supports localized titles).

run_macro

Batch up to 50 operations into a single round-trip for maximum efficiency.

clipboard / notification_show

System-level text exchange and user alerts.

key_locker

Manage credentials the terminal autofills for you (SSH key passphrases, sudo / login passwords). Secrets are entered once into the locker's own secure dialog and stored encrypted on this machine (Windows DPAPI); they are never shown to the assistant. action='launch_console' opens an autofill-capable console (returns a paneId to drive ssh/sudo into via terminal); save / list / forget / set_policy / status manage bindings. Autofill only fires in a console opened by launch_console. Disable with DESKTOP_TOUCH_DISABLE_KEY_LOCKER=1.

📊 Office (Excel)

Tool

Description

excel

Author and run Excel VBA macros via COM. action='run_vba' writes a macro into a managed Trusted Location and runs it; action='check_access_vbom' is a read-only preflight. Runs VBA where formula-only tools cannot. One-time setup: node scripts/enable-access-vbom.mjs.


Standard workflow (v1.0.0)

The v2 World-Graph surface (desktop_discover / desktop_act) is the recommended dispatch path. The four-call shape works for native apps, browsers, and terminals identically.

desktop_state          → orient: focused window/element, modal, attention signal
desktop_discover       → find actionable entities (returns lease + windows[])
desktop_act(lease, …)  → act on entity (returns attention + post.perception)
desktop_state          → confirm the world changed as expected

Clicking — priority order:

browser_click(selector)               → Chrome / Edge (CDP, stable across repaints)
desktop_act(lease, action='click')    → native / dialog / visual (entity-based; use after desktop_discover)
click_element(name | automationId)    → native UIA fallback if desktop_act returns ok:false
mouse_click(x, y, origin?, scale?)    → pixel last resort; origin+scale from dotByDot screenshots only

Recovery hints — read response.attention after every observation and response.warnings[] on desktop_discover / desktop_act. Common reasons:

  • lease_expired / lease_generation_mismatch / lease_digest_mismatch / entity_not_found → re-call desktop_discover

  • modal_blocking → response.blockingElement (when present) names the blocking modal. role: "dialog" means a separate dialog window has disabled the target's window: blockingElement.hwnd is that dialog — re-call desktop_discover with target.hwnd = blockingElement.hwnd, answer it there, then retry (name is its title, which may be empty or shared). Any other role: a window the desktop_discover snapshot holds, where the OS could not say whether it blocks this entity — with blockingElement.hwnd, re-call desktop_discover with target.hwnd = blockingElement.hwnd and answer it there; without it, dismiss via click_element(name=blockingElement.name). Then re-call desktop_discover on the original target and act on the new lease — this refusal came from that snapshot, so the same lease is refused again

  • entity_outside_viewport → the element moved off screen: scroll(action='to_element' | 'raw'), or re-call desktop_discover if its window moved or closed

  • origin_window_not_visible → the element's window is minimised or hidden, so nothing is drawn where it was found: focus_window(windowTitle) to restore it, then re-call desktop_discover

  • coordinate_outside_reachable_bounds → the coordinate is not on any connected monitor. Coordinate-based mouse input (mouse_click / mouse_drag / scroll / browser_click, and the mouse route inside desktop_act) now works on every monitor, including monitors placed left of or above the primary one, so this error normally means the coordinates are stale — the window moved or closed after they were read. Re-run desktop_discover and act on the new coordinates. If the server is running without its built-in Windows input module, mouse input falls back to the primary monitor only; the error message says so, and moving the window onto the primary monitor (or reinstalling the server) is the fix

  • cursor_placement_blocked → the coordinate is on a monitor, but the pointer could not be placed there, so nothing was clicked. This happens while another app confines the cursor to its own window (common in full-screen games), while a remote-desktop session is disconnected or locked, while another program keeps repositioning the pointer, or right after a monitor is added or removed. Leave the app holding the cursor, reconnect the session, or — after a monitor change — re-run desktop_discover, then retry. click_element acts through the accessibility API without moving the cursor and works meanwhile

  • keyboard_target_unsafe → a type was refused, because the characters would not have reached the field you named: the keyboard focus is on a different control or in a different window, the control that would receive them — or the field you named — is read-only, or the field you named — or its window — is disabled. Nothing was typed, and if_unexpected.detail says which. For a disabled field, answer or wait out whatever disabled it, then re-run desktop_discover and type again; clicking it does not help, and desktop_discover does not list a disabled field, so while it is missing there it is still disabled. A field you named that is read-only does not take text; typing again will not change that. For another control or window, put the focus on the field you named, then type again; if_unexpected.detail names the way back for the road the act took. On a window named by title, desktop_act with action='click' on the same entity does it. On a window named by handle nothing here moves the focus to a text field yet, so re-run desktop_discover by the window's title and click the field there — a common dialog's title resolves to a handle as well, so that road does not open there. For another window, bring the field's window forward first (focus_window): it comes forward with the focus it last had, and the window holding the focus is usually drawn over the field. Do not retry with a foreground keyboard type: whatever holds the focus would take the characters

  • executor_failed → fall back to click_element / mouse_click / browser_click

A successful type can carry landing: { confirmed: false, why }. The write took the background route, but the server could not confirm that it reached the field you named — for example, in a WPF window, whose fields have no window of their own. This is a report, not a state that can be resolved here: nothing in the response establishes whether the characters arrived, reading the field back does not settle it (desktop_state answers about the foreground, and may come back with no value at all — hints.focusedElementValueAbsent names the road that dropped it, view_road_has_no_value or masked_on_this_road, and no hint is not evidence a value was there — or name a field in another window with the same title), diff.value_changed is not delivery either, its baseline being your desktop_discover snapshot rather than the write, and retrying a nonempty write is not a repeat — a background write lands at the caret and replaces the selection, exactly as typing does.

Lease lifecycle:

  • Each desktop_discover response carries softExpiresAtMs (≈ 60 % of the TTL window). Past that timestamp the LLM should consider re-calling desktop_discover even though the lease is still technically valid — lease.expiresAtMs is the only correctness wall.

  • TTL adapts to view mode (action/explore/debug), entity count, and response payload size. Cap is 60 s.

  • Set DESKTOP_TOUCH_DISABLE_FUKUWARAI_V2=1 to fall back to the v1 tool surface (get_windows / get_ui_elements / set_element_value) for troubleshooting only — V2 is the recommended default.


Terminal command completion (until)

terminal(action='run') sends a command, waits for it to complete, and reads the output in one call. How it decides "complete" is controlled by until:

Mode

Waits for

Best for

quiet (default)

output to fall silent for quietMs

short interactive commands

pattern

a string/regex you expect in the output

long commands with a known final marker

exit

the command to actually finish

when you need completion or the exit code

Anchoring caveat (#384): a command whose final line has no trailing newline glues the marker to the next prompt with no line boundary (printf X → Xuser@host:~$), so an end-anchored pattern (X\s*\n / X$) can never bind. For completion use mode:'exit'; for content matching use a bare marker (no \n/$). mode:'pattern' also accepts an optional quietMs settle fallback: until:{mode:'pattern', pattern, quietMs:1000} completes with reason:'quiet' (no matchedPattern) once output is stable for that long without a match — instead of hanging until timeoutMs. It is opt-in (omit quietMs to keep waiting for the pattern; long commands with mid-run silent gaps are unaffected).

until:{mode:'exit'} — real completion + exit code

The heuristic modes can misfire on the common "append a sentinel" idiom (some-task; echo DONE matched by DONE): the sentinel also shows up in the echoed command line, and for multi-line commands there is no reliable way to tell that echo apart from real output. mode:'exit' removes the guesswork — the server appends its own completion marker whose printed form differs from its typed form, so it never matches the echoed command (even for multi-line input), and it returns the real process exit code:

terminal({
  action: 'run',
  windowTitle: 'pwsh',
  input: 'npm run build',
  until: { mode: 'exit', shell: 'powershell' },
})
// → completion: { reason: 'exited', exitCode: 0, elapsedMs: … }
//   output: just the command's real output (the injected marker is stripped)
  • Pass shell explicitly ('bash' or 'powershell'). shell:'auto' detects the shell from the terminal window, but it cannot see a shell running inside SSH or WSL — the window still looks like its local host — so for remote/nested sessions pass the remote side's shell (auto otherwise warns and may pick the outer shell). A window whose process is genuinely unidentifiable (e.g. Windows Terminal) returns ExitModeShellAmbiguous.

  • First-class shells: bash and powershell. cmd.exe is not supported yet (ExitModeShellUnsupported).

  • Unsafe input is rejected up front (ExitModeUnsafeInput) rather than hanging: a command ending mid-construct (unterminated quote, here-doc, $(…), a trailing \ or PowerShell backtick).

  • Exit mode controls its own delivery, so delivery-shaping sendOptions (method / preferClipboard / pressEnter / chunkSize / pasteKey) are rejected with InvalidArgs; focus options remain accepted.


Key Locker (terminal credential autofill)

Running ssh user@host or sudo … normally stops at a hidden password prompt an assistant can't safely type into. Key Locker stores your SSH key passphrases and sudo / login passwords encrypted on your machine (Windows DPAPI, current user) and fills them in automatically when a bound command reaches its prompt. The secret is typed once into the locker's own secure dialog — it is never shown to the assistant and never travels through the MCP channel.

// 1. Save the credential once — opens a secure dialog on your desktop
key_locker({ action:'save', uri:'ssh://user@host:22' })

// 2. Open an autofill-capable console (returns its paneId)
key_locker({ action:'launch_console' })   // → { paneId:'12345678', windowTitle:'…' }

// 3. Run the command through that pane — the password is filled at the prompt
terminal({ action:'send', paneId:'12345678', input:'ssh user@host' })
  • Autofill only fires in a console opened by launch_console — a pre-existing terminal is never autofilled. The console is a classic visible Windows console, so you can watch it and take over at any prompt yourself.

  • Every autofill asks you to confirm by default; opt a binding out with set_policy. list / status / forget manage saved credentials.

  • terminal read / send accept paneId as an alternative to windowTitle — it targets that exact window even after an ssh login renames its title.

  • Supported binding URIs: ssh://user@host:22, sudo://host/user, https-cred://host, and SSH key passphrases (sshkey:SHA256:…). An ssh save needs the host key already in known_hosts (connect to the host once first).

  • Windows only. Disable the whole feature with DESKTOP_TOUCH_DISABLE_KEY_LOCKER=1. The secure dialog is an unsigned helper executable — Windows SmartScreen may show an "unknown publisher" warning on first run (see the note under Requirements).


Browser CDP automation

For web automation, connect Chrome or Edge with the remote debugging port enabled — no Selenium or Playwright needed.

# Launch Chrome in CDP mode
chrome.exe --remote-debugging-port=9222 --user-data-dir=C:\tmp\cdp
browser_open({launch:{}})                          → spawn-if-needed Chrome in debug mode + list tabs (idempotent)
browser_open()                                     → connect-only (fail if no CDP endpoint live)
browser_locate({selector:"#submit"})               → CSS selector → physical screen coords
browser_click({selector:"#submit"})                → find + click in one step (auto-focuses browser)
browser_eval({action:"js", expression:"document.title"})  → evaluate JS, returns result
browser_eval({action:"dom", selector:"#main", maxLength:5000})  → outerHTML, truncated to maxLength chars
browser_eval({action:"appState"})                  → one-shot SPA state (Next/Nuxt/Remix/Apollo/GitHub react-app/Redux SSR)
browser_fill({selector:"#email", value:"user@example.com"})  → fill React/Vue/Svelte controlled input (state-safe)
browser_overview()                                 → links/buttons/inputs + ARIA toggles + viewportPosition per element
browser_search({by:"text", pattern:"..."})         → grep DOM with confidence ranking
browser_navigate({url:"https://example.com"})      → navigate via CDP (no address bar interaction)

For chained calls in the same tab, pass includeContext:false to omit the activeTab/readyState annotation (~150 tok/call saved). Boolean / object params accept the LLM-friendly string spellings ("true", "{}").

Coordinates returned by browser_locate account for the browser chrome (tab strip + address bar height) and devicePixelRatio, so they can be passed directly to mouse_click without any scaling.

Recommended web workflow:

browser_open({launch:{}}) → browser_eval({action:"dom"}) → browser_locate(selector) → browser_click(selector)

Auto-dock CLI on startup

Keep Claude CLI visible while operating other apps full-screen. Set env vars in your MCP config and the docked window auto-snaps into place every MCP startup.

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_DOCK_TITLE": "@parent",
        "DESKTOP_TOUCH_DOCK_CORNER": "bottom-right",
        "DESKTOP_TOUCH_DOCK_WIDTH": "480",
        "DESKTOP_TOUCH_DOCK_HEIGHT": "360",
        "DESKTOP_TOUCH_DOCK_PIN": "true"
      }
    }
  }
}

Env var

Default

Notes

DESKTOP_TOUCH_DOCK_TITLE

(unset = off)

@parent walks the MCP process tree to find the hosting terminal — immune to title / branch / project changes. Or use a literal substring.

DESKTOP_TOUCH_DOCK_CORNER

bottom-right

top-left / top-right / bottom-left / bottom-right

DESKTOP_TOUCH_DOCK_WIDTH / HEIGHT

480 / 360

px ("480") or ratio of work area ("25%") — 4K/8K auto-adapts

DESKTOP_TOUCH_DOCK_PIN

true

Always-on-top toggle

DESKTOP_TOUCH_DOCK_MONITOR

primary

Monitor id from desktop_state({includeScreen:true})

DESKTOP_TOUCH_DOCK_SCALE_DPI

false

If true, multiply px values by dpi / 96 (opt-in per-monitor scaling)

DESKTOP_TOUCH_DOCK_MARGIN

8

Screen-edge padding (px)

DESKTOP_TOUCH_DOCK_TIMEOUT_MS

5000

Max wait for the target window to appear

Input routing gotcha: when a pinned window is active (e.g. Claude CLI), keyboard(action='type') / keyboard(action='press') send keys to it, not the app you wanted to type into. Always call focus_window(title=...) before keyboard operations, then verify isActive=true via screenshot(detail='meta').

Screenshot cache (by-ref storage)

screenshot and the other visual results return a cheap screenshot://by-ref/{id} link to an image saved on disk instead of inlining the pixels every time, so routine look-act-confirm loops cost far fewer tokens. The cache bounds itself automatically and screenshot_query / screenshot_gc let you inspect and prune it. Tune the storage with:

Env var

Default

Notes

DESKTOP_TOUCH_SCREENSHOTS_DIR

(per-user cache dir)

Pin the cache to a specific folder. If the default folder can't be created or written (e.g. corporate policy blocking new folders under your profile), the server auto-probes this → the runtime dir → an OS temp folder and uses the first writable one instead of giving up on the cache.

DESKTOP_TOUCH_SCREENSHOT_MAX_COUNT

200

Keep at most this many captures in the cache.

DESKTOP_TOUCH_SCREENSHOT_MAX_BYTES

256 MiB

Cap the total cache size on disk.

DESKTOP_TOUCH_SCREENSHOT_MAX_AGE_MS

(off)

Drop captures older than this many milliseconds (opt-in).

DESKTOP_TOUCH_SCREENSHOT_AUTOPRUNE

on

Auto-trim the cache as new captures are saved. Set 0 to disable.

DESKTOP_TOUCH_SCREENSHOT_MIN_EVICT_AGE_MS

60000

Never auto-evict a capture younger than this (ms), so a by-ref link you were just handed survives long enough to open even when another AI/process on the same PC is also capturing. 0 disables.

Multi-monitor screenshots

screenshot(displayId=…) and screenshot(region=…) capture any monitor, including one placed left of or above the primary — those have negative desktop coordinates, and you pass them exactly as screenshot(detail='meta') reports them. screenshot() with no region is the primary monitor, as it has always been.

A region that cannot be captured comes back as RegionOutsideCapturableBounds rather than a raw Windows error, and the message says which of three things happened. The region may be on no monitor at all, which usually means the coordinates went stale because the window moved or closed — take a fresh screenshot and use the new numbers. It may overlap a monitor but stretch past the edge of the screen area, in which case the coordinates are fine and the region is simply too big: ask for a smaller one, or capture the window itself with screenshot(windowTitle=…). Or this server may be limited to the primary monitor, which the message says outright — along with why, because that decides the fix: if an env override pinned it, screenshot(windowTitle=…) normally still works on every monitor, whereas if the built-in capture module is missing then window capture usually needs that same module and fails too, so move the window onto the primary monitor or reinstall the server. Whole-screen capture and single-window capture are separate parts of that module, though, and a server can end up with one but not the other — so rather than working it out from the cause, read the message: it says plainly whether screenshot(windowTitle=…) is available on this server.

If Windows returns no pixels at all — a locked screen, a UAC prompt, a disconnected remote-desktop session — you get CaptureBackendFailed; capturing the window itself with screenshot(windowTitle=…) usually still works, because it reads through a different Windows API.

Env var

Default

Notes

DESKTOP_TOUCH_CAPTURE_BACKEND

(unset = automatic)

Diagnostic override for the screen-capture path. Set to nutjs to force the older capture backend, which can only read the primary monitor — useful for isolating a capture problem. The server picks the backend once at startup, so change this in your MCP client config and restart. Any other value is ignored.

Auto Perception (always-on)

Phase 4 privatizes the explicit perception_* tool family — the v0.12 Auto Perception layer attaches an attention signal to every desktop_state and desktop_act response automatically. Action tools also auto-guard when given a windowTitle. There is no longer a need to register / read / forget lenses manually.

# desktop_state always returns the attention signal
desktop_state() → {focusedWindow, focusedElement, modal, attention:"ok", ...}

# Action tools auto-guard when windowTitle is given:
keyboard({action:"type", text:"hello", windowTitle:"Notepad"})
→ post.perception:{status:"ok"}  // unsafe input blocked if guards fail

# When attention is dirty / stale / settling, refresh with desktop_state:
desktop_state()  // re-evaluates attention via Auto Perception

For advanced pinned-target workflows, the lensId parameter remains on action tools (keyboard, mouse_click, mouse_drag, click_element, browser_click, browser_navigate, browser_eval, desktop_act). Omit lensId for the normal Auto Perception path. The underlying registry, hot target cache, and sensor loop are unchanged; only the explicit perception_register / perception_read / perception_forget / perception_list tools were retired.


Mouse homing correction

When Claude calls screenshot(detail='text') to read coordinates and then mouse_click seconds later, the target window may have moved. The homing system corrects this automatically.

Tier

How to enable

Latency

What it does

1

Always-on (if cache exists)

<1ms

Applies (dx, dy) offset when window moved

2

Pass windowTitle hint

~100ms

Auto-focuses window if it went behind another

3

Pass elementName/elementId + windowTitle

1–3s

UIA re-query for fresh coords on resize

# Tier 1 only (automatic)
mouse_click(x=500, y=300)

# Tier 1 + 2: also bring window to front if hidden
mouse_click(x=500, y=300, windowTitle="Notepad")

# Tier 1 + 2 + 3: also re-query UIA if window resized
mouse_click(x=500, y=300, windowTitle="Notepad", elementName="Save")

# Traction control OFF — no correction
mouse_click(x=500, y=300, homing=false)

The homing parameter is available on mouse_click, mouse_drag, and scroll. The cache is updated automatically on every screenshot(), desktop_discover(), focus_window(), and workspace_snapshot() call.

mouse_click image-local coords (origin + scale)

When you take a dotByDot screenshot with dotByDotMaxDimension, the response prints the origin and scale values. Instead of computing screen coords manually, copy them into mouse_click:

# Screenshot response:
#   origin: (0, 120) | scale: 0.6667
#   To click image pixel (ix, iy): mouse_click(x=ix, y=iy, origin={x:0, y:120}, scale=0.6667)

mouse_click(x=640, y=300, origin={x:0, y:120}, scale=0.6667, windowTitle="Chrome")
# Server converts: screen = (0 + 640/0.6667, 120 + 300/0.6667) = (960, 570)

This eliminates a whole class of off-by-one and scale bugs. Without origin/scale, x/y remain absolute screen pixels (unchanged behavior).


screenshot key parameters

detail="image"          — PNG/WebP pixels (default)
detail="text"           — UIA element JSON + clickAt coords (no image, ~100–300 tok)
detail="meta"           — Title + region only (cheapest, ~20 tok/window)
dotByDot=true           — 1:1 WebP; image_px + origin = screen_px
dotByDotMaxDimension=N  — cap longest edge (response includes scale for coord math)
grayscale=true          — ~50% smaller for text-heavy captures (code/AWS console)
region={x,y,w,h}        — with windowTitle: window-local coords (exclude browser chrome)
                          without: virtual screen coords
diffMode=true           — I-frame first call, P-frame (changed windows only) after (~160 tok)
ocrFallback="auto"      — detail='text' auto-fires Windows OCR on uiaSparse or empty

Recommended Chrome combo (50–70% data reduction):

screenshot(windowTitle="Chrome",
           dotByDot=true, dotByDotMaxDimension=1280, grayscale=true,
           region={x:0, y:120, width:1920, height:900})  # skip browser chrome

Recommended workflow:

workspace_snapshot()                     → full orientation (resets diff buffer)
screenshot(detail="text", windowTitle=X) → get actionable[].clickAt coords
mouse_click(x, y)                        → click directly, no math needed
screenshot(diffMode=true)                → check only what changed (~160 tok)

Security

Emergency stop (Failsafe)

Park the mouse in the top-left corner of the primary monitor (within 10px of 0,0) for 500ms continuously to trigger the emergency stop.

  • The trigger corner is on the primary monitor only. Areas that used to trigger the stop in older versions (monitors left of or above the primary) no longer do; if the cursor dwells there, a one-time balloon notification points you to the right corner.

  • While a tool call is running: the server exits (exit code 1) — the runaway-automation brake. A balloon notification and a diagnostic log entry (with cursor coordinates) record why it stopped. Only a call that is actually mid-flight triggers the exit; in the rare case where that call finishes during the ~1 second the notification takes, the server stays up instead and a follow-up balloon corrects the first one.

  • While idle: the server stays up and refuses new tool calls until the cursor leaves the corner. Background credential autofill (key_locker) is cancelled before any of its dialogs open — while you hold the corner, no credential prompt dialog appears and no credential is typed. It does not pick up again by itself: move the cursor away from the corner and run the command again.

  • Per-tool check: runs before every tool handler. Background monitor: 500ms polling as a backup for long-running operations. Trigger radius: 10px.

  • DESKTOP_TOUCH_FAILSAFE_HOLD_MS — dwell time in ms before the stop fires (default 500; 0 = fire immediately on corner entry).

Blocked operations

workspace_launch blocklist: cmd.exe, powershell.exe, pwsh.exe, wscript.exe, cscript.exe, mshta.exe, regsvr32.exe, rundll32.exe, msiexec.exe, bash.exe, wsl.exe are blocked. Script extensions (.bat, .ps1, .vbs, etc.) are rejected. Arguments containing ;, &, |, `, $(, ${ are also rejected.

keyboard(action='press') blocklist: Win+R (Run dialog), Win+X (admin menu), Win+S (search), Win+L (lock screen) are blocked.

PowerShell injection protection

All -like patterns in the UIA bridge PowerShell fallback path are sanitized with escapeLike(), which escapes wildcard characters (*, ?, [, ]) before they reach PowerShell. When the Rust native engine is active, PowerShell is not invoked for UIA operations.

Allowlist for workspace_launch

Shell interpreters are blocked by default. To allow specific executables, create an allowlist file:

File locations (searched in order):

  1. Path in DESKTOP_TOUCH_ALLOWLIST environment variable

  2. ~/.claude/desktop-touch-allowlist.json

  3. desktop-touch-allowlist.json in the server's working directory

Format:

{
  "allowedExecutables": [
    "pwsh.exe",
    "C:\\Tools\\myapp.exe"
  ]
}

Changes take effect immediately — no restart needed.


Mouse movement speed

All mouse tools (mouse_click, mouse_drag, scroll) accept an optional speed parameter:

Value

Behavior

Omitted

Uses the configured default (see below)

0

Instant teleport — setPosition(), no animation

1–N

Animated movement at N px/sec

Default speed is 1500 px/sec. Change it permanently via the DESKTOP_TOUCH_MOUSE_SPEED environment variable:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_MOUSE_SPEED": "3000"
      }
    }
  }
}

Common values: 0 = teleport, 1500 = default gentle, 3000 = fast, 5000 = very fast.


Force-Focus (AttachThreadInput)

Windows foreground-stealing protection can prevent SetForegroundWindow from succeeding when another window (such as a pinned Claude CLI) is in the foreground. This causes subsequent keystrokes or clicks to land in the wrong window — a silent failure.

mouse_click, keyboard(action='type'), keyboard(action='press'), and terminal(action='send') all accept a forceFocus parameter that bypasses this protection using AttachThreadInput:

{
  "name": "mouse_click",
  "arguments": {
    "x": 500,
    "y": 300,
    "windowTitle": "Google Chrome",
    "forceFocus": true
  }
}

If the force attempt is refused despite AttachThreadInput, the response is ok:false with code: "ForegroundRestricted" (issue #202 unification — same shape as focus_window, keyboard, terminal_send, mouse_click). The action itself is suppressed so the keystrokes / click never land on the wrong window. Recover via focus_window's auto-escalate ladder before retrying. The legacy hints.warnings: ["ForceFocusRefused"] shape is no longer emitted.

Global default via environment variable:

{
  "mcpServers": {
    "desktop-touch": {
      "env": {
        "DESKTOP_TOUCH_FORCE_FOCUS": "1"
      }
    }
  }
}

Setting DESKTOP_TOUCH_FORCE_FOCUS=1 makes forceFocus: true the default for all four tools without changing each call.

Known tradeoffs:

  • During the ~10ms AttachThreadInput window, key state and mouse capture are shared between the two threads. In rapid macro sequences this can cause a race condition (rare in practice).

  • Disable forceFocus (or unset the env var) when the user is manually operating another app to avoid unexpected focus shifts.


Auto Guard

Action tools (mouse_click, mouse_drag, keyboard(action='type'/'press'), click_element, desktop_act, browser_click, browser_navigate) automatically guard each action when you pass windowTitle / tabId:

  • Verifies target window identity (process restart / HWND replacement detected)

  • Confirms click coordinates are inside the target window rect

  • Returns post.perception.status on every response — including failures — so the LLM can recover without a screenshot

Keyboard writes must name a destination. keyboard(action='type'/'press'/'sequence') requires either windowTitle or hwnd. Without one there is no target to guard, and the keys would land on whatever window is foreground at that instant — including one you just clicked into yourself. Such a call is refused with code:"DestinationRequired" before any key is sent, and a windowTitle that is empty or only spaces counts as no target at all. A window that has no title can be addressed by hwnd, but only while it is already the foreground window — keyboard focus and guarding cannot target a titleless window yet, so bring it forward with focus_window first if it is not in front. Such a call also comes back with a warning saying the input was delivered unguarded.

Variable

Default

Meaning

DESKTOP_TOUCH_REQUIRE_DESTINATION

(unset = required)

Set to 0 to type into the current foreground window on purpose. The refusal becomes a warning on the response instead of an error — never a silent pass.

DESKTOP_TOUCH_AUTO_GUARD

(unset = on)

Set to 0 to turn the whole guard layer off, the destination check included.

Disabling auto guard — set DESKTOP_TOUCH_AUTO_GUARD=0 to restore v0.11.12 behavior (no auto guard):

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_AUTO_GUARD": "0"
      }
    }
  }
}

When auto guard is enabled (default), post.perception.status will be one of:

Status

Meaning

ok

Guard passed — target verified

unguarded

windowTitle not provided; action ran without guard

ambiguous_target

Multiple windows matched; pass hwnd to name one exactly, or use a more specific title

target_not_found

No window matched the given title

identity_changed

Window was replaced (process restart / HWND change)

blocked_by_modal

A modal dialog is in the way — dismiss it, then retry

unsafe_coordinates

Click coordinates are outside the target window rect

browser_not_ready

The browser tab is still loading — wait, then retry

needs_escalation

Use browser_click or specify windowTitle

destination_required

A keyboard write named no target. Refused before the guard runs, so it arrives as code:"DestinationRequired" with this status under context.guard rather than in post.perception — pass windowTitle or hwnd

When unsafe_coordinates or identity_changed is returned, the response may include a suggestedFix.fixId. Pass that fixId to the relevant tool call to approve the recovery:

{ "name": "mouse_click",           "arguments": { "fixId": "fix-..." } }
{ "name": "keyboard(action='type')",         "arguments": { "fixId": "fix-...", "text": "hello" } }
{ "name": "click_element",         "arguments": { "fixId": "fix-..." } }
{ "name": "browser_click", "arguments": { "fixId": "fix-..." } }

The fix is one-shot and expires in 15 seconds. The server revalidates the target process identity before executing.


Diagnostic log

The server keeps an append-only log of events that never reach a tool response, at %USERPROFILE%\.desktop-touch-mcp\logs\diagnostic.log (one JSON object per line). It records crashes and slow calls, and — since the diagnostic log became the place to look when input lands in the wrong place — how each windowTitle was resolved and where each write went:

  • a resolve record per title lookup: how many windows matched, which one was picked, the ones that lost, and a flag when the terminal process-name fallback fired because nothing matched by title;

  • a dispatch_sink record per input dispatch — keyboard, terminal, scroll and desktop_act's background writes: which channel was used, which window it was addressed to, and which window was in the foreground at that moment;

  • a correlation id shared by all records from one tool call, so a resolution can be matched to the write it produced even when calls overlap.

If an input call ever seems to type into the wrong window, this is the file that says which window it picked and why. A record is written immediately before the write leaves the process, so a dispatch that is refused or fails first is not on record as having happened.

The log rolls over: once diagnostic.log passes 64 MiB it becomes diagnostic.log.1, and at most two rolled generations are kept. With one server running, the newest records are always in diagnostic.log, but when you are searching for something that happened a while ago, search diagnostic.log* rather than the one file. That glob also catches diagnostic.log.<pid>.rotating, which is where a server parks the live file for the moment it is being rolled. One of these left behind means a roll did not finish: the server was killed partway through, or the roll failed after the file was parked and the server could not put it back — it will not if a fresh diagnostic.log has been started in the meantime, and the move back can fail for the same reason the roll did. Nothing in it is lost. The pid in the name says which server it belonged to, and it is filed back into the numbered generations by the next roll — by that server if it is still running, and otherwise by any other server once the original has exited or, because process ids are reused, once the file has been parked for an hour — so a crashed server's log is not left sitting on disk forever.

Every record is measured against the limit before it is written, so a server left running for days rolls the file as it goes — there is no scheduled job, nothing to restart, and nothing to clean up by hand. A server sitting idle never rolls anything, because the check only runs when there is something to write.

The ceiling is a size, not an age. Three generations hold 192 MiB of records, and how far back that reaches depends entirely on how busy the machine is: on the install that prompted this limit, averaging roughly 170 MB a day, it is a little over one day. If you want to keep a particular incident, copy the file out rather than expecting to find it next week; if you would rather trade disk space for reach, raise DESKTOP_TOUCH_DIAGNOSTIC_LOG_MAX_BYTES.

Two situations go past that figure, and both are worth knowing about:

  • Several servers sharing one log. Every MCP client starts its own server, and by default they all write to the same file. Each tracks the bytes it has written itself and only re-measures the real file every few MB, so the live file can overshoot before one of them rolls it. The overshoot grows with the number of servers running, not without limit. Two servers can also roll at the same moment and step on each other's rename: that costs a generation, and can leave one server's newest record in diagnostic.log.1 instead of the live file. Grepping diagnostic.log* rather than the one file covers both.

  • A live file that cannot be renamed — held open by another program, or permission denied. Rotation then cannot happen and the log keeps growing at full speed; a parked .rotating file that is held open stops a roll the same way, and is checked before any numbered generation is touched. Nothing is lost — a roll that fails leaves the live file where it was, or at worst parked under the .rotating name above for a later roll to file — but this is the one case the limit does not cover, so it is not silent: a log_rotation_failed record is written into the log itself, once per stretch of failed rolls rather than once per line. It is the first thing to grep for if you find an oversized diagnostic.log after updating.

One record is never allowed to be larger than the file it lives in, so an event carrying an unusually large payload is written as a shortened stand-in: same kind, plus record_truncated, the original size, and a head field holding the first few KB of what it would have been.

Variable

Default

Meaning

DESKTOP_TOUCH_RESOLVE_LOG_RAW

(unset = off)

Window titles and the titles you search for are recorded as a short hash plus their length, because a title can contain a file name, a mail subject, or a browser page title. Set to 1 to also record the text in clear (the hash stays, so a log with both is still readable end to end).

DESKTOP_TOUCH_DIAGNOSTIC_LOG_DISABLE

(unset = on)

Set to 1 to stop writing the log entirely.

DESKTOP_TOUCH_DIAGNOSTIC_LOG_PATH

(per-user log dir)

Write the log somewhere else. A symbolic link works: the roll follows it, so the link keeps pointing at the live log and the rolled generations appear beside the real file rather than beside the link.

DESKTOP_TOUCH_DIAGNOSTIC_LOG_MAX_BYTES

67108864 (64 MiB)

Roll the live log to diagnostic.log.1 once it passes this size. Two rolled generations are kept, so the log directory ordinarily holds about three times this value — see above for the two situations that go past it. A value below 1 MiB is raised to 1 MiB and one above 1 GiB is lowered to 1 GiB, and anything that is not a positive whole number falls back to the default — a typo here cannot switch rotation off in either direction, whether you mean bytes and write MiB or the other way round. To stop logging entirely, use DESKTOP_TOUCH_DIAGNOSTIC_LOG_DISABLE.


Advanced response options

browser_eval Structured Mode

Pass withPerception: true to receive a structured JSON response with post.perception instead of raw text:

{ "name": "browser_eval", "arguments": { "expression": "document.title", "withPerception": true } }

Returns { ok: true, result: "...", post: { perception: { status: "ok", ... } } }.

mouse_drag Cross-Window Guard

mouse_drag now guards both start and end coordinates. Drags that cross window boundaries (or reach the desktop wallpaper) are blocked by default. To allow intentional cross-window or range-selection drags:

{ "name": "mouse_drag", "arguments": { "startX": 100, "startY": 100, "endX": 900, "endY": 900, "allowCrossWindowDrag": true } }

Performance (v0.15 — Rust Native Engine)

The Rust native engine (@harusame64/desktop-touch-engine) replaces PowerShell process spawning with direct COM calls over a persistent MTA thread. It loads automatically as a .node addon — no configuration needed.

UIA Benchmark (vs PowerShell baseline)

Function

Rust Native

PowerShell

Speedup

getFocusedElement

2.2 ms

366 ms

163.9×

getUiElements (Explorer, ~60 elements)

106.5 ms

346 ms

3.3×

Weighted average

~82×

Image Diff Benchmark (SSE2 SIMD)

Function

Rust (SSE2)

TypeScript

Speedup

computeChangeFraction (1920×1080)

0.26 ms

3.8 ms

~15×

dHash (perceptual hash)

0.09 ms

1.2 ms

~13×

Architecture

Claude CLI / MCP Client
    │  stdio or HTTP (MCP protocol)
    ▼
desktop-touch-mcp (TypeScript)
    │
    ├── Rust Native Engine (.node addon)          ← NEW in v0.15
    │   ├── UIA: 13 functions via napi-rs + windows-rs 0.62
    │   │   └── Dedicated COM thread (MTA) + batch BFS algorithm
    │   └── Image: SSE2 SIMD pixel diff + perceptual hashing
    │
    └── PowerShell Fallback (automatic)
        └── Activates transparently if .node is unavailable

Why getUiElements is 3.3× (not 160×)

The 160× speedup on getFocusedElement comes from eliminating PowerShell process startup (~200 ms) and .NET assembly loading. For getUiElements, the bottleneck shifts to the UIA provider inside the target application (e.g., Explorer) — it must enumerate its UI tree regardless of who asks. The Rust engine uses a batch BFS algorithm (FindAllBuildCache + TreeScope_Children) that minimizes cross-process RPC calls and supports maxElements early exit, making it dramatically faster on large trees (VS Code, browsers with 1000+ elements).


UI Operating Layer (V2)

Status: Default ON since v0.17. desktop_discover and desktop_act are available out of the box.

V2 introduces two new tools that replace coordinate-based clicking with entity-based interaction:

Tool

Description

desktop_discover

Observe a window or browser tab. Returns interactive entities with leases — no raw screen coordinates. Supports UIA (native), CDP (browser), terminal, and visual GPU lanes.

desktop_act

Interact with an entity returned by desktop_discover. Validates the lease before executing. Returns a semantic diff (entity_disappeared, modal_appeared, focus_shifted, …). When diffUnchecked is present, the diff did not look for the kinds it lists, so their absence from the diff does not mean they did not happen. On visual-only targets a successful act can bundle a roiCapture (a PNG crop of the changed region + a lease-less next-target preview) so you confirm the result and find the next target in one call — controlled by returnCapture (on-change, the default on a visible change; never to suppress; always to force).

Clicking — priority order

When multiple tools could perform the same click, prefer them in this order:

  1. browser_click(selector) — Chrome / Edge over CDP (stable across repaints)

  2. desktop_act(lease) — native windows, dialogs, visual-only targets (entity-based; use after desktop_discover)

  3. click_element(name | automationId) — native UIA fallback when desktop_act returns ok:false

  4. mouse_click(x, y) — pixel-level last resort (origin + scale from dotByDot screenshots only)

Disabling V2 (kill switch)

To hide desktop_discover / desktop_act from the tool catalog, add the disable flag and restart:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_DISABLE_FUKUWARAI_V2": "1"
      }
    }
  }
}

All V1 tools continue to work without interruption — no reinstall required. Remove the env entry and restart to re-enable.

Flag semantics (exact-match: only the literal string "1" counts):

DISABLE_FUKUWARAI_V2

V2 state

unset / not "1"

ON (default)

"1"

OFF (kill switch)

Removed: DESKTOP_TOUCH_ENABLE_FUKUWARAI_V2

This was the opt-in switch in v0.16.x. V2 is on by default since v0.17, so the flag no longer has any effect and is safe to delete from your config. To turn V2 off, set DESKTOP_TOUCH_DISABLE_FUKUWARAI_V2=1.

Recovery when V2 fails

If desktop_act returns ok: false, read reason and follow the built-in recovery hints in the tool description. Common paths:

  • lease_expired / *_mismatch / entity_not_found → re-call desktop_discover

  • modal_blocking → response.blockingElement (when present) carries { name, role, automationId?, hwnd? }. With role: "dialog" the blocker is a separate dialog window that has disabled the target's window, and hwnd is it — desktop_discover with target.hwnd = blockingElement.hwnd, answer it, then retry. Any other role: a window the discover snapshot holds, which the OS could not confirm — desktop_discover with target.hwnd = blockingElement.hwnd when hwnd is present, otherwise click_element(name=blockingElement.name); then re-discover the original target and act on the new lease (the same lease is refused again)

  • entity_outside_viewport → the element moved off screen: scroll / scroll(action='to_element') when it scrolled out of its own window, or re-call desktop_discover when the window itself moved or closed

  • origin_window_not_visible → focus_window(windowTitle) to restore the minimised / hidden window, then re-call desktop_discover

  • coordinate_outside_reachable_bounds → the target is on no connected monitor — usually stale coordinates: re-run desktop_discover. (Without the built-in Windows input module, only the primary monitor is reachable; the message says so.)

  • cursor_placement_blocked → the pointer could not be placed there (an app is holding the cursor, or the session is not interactive), so nothing was clicked: free the cursor or reconnect the session, or use click_element (UIA invoke, cursor-free)

  • keyboard_target_unsafe → nothing was typed: the characters would have gone to a different control or window, or to a read-only control, or the field you named (or its window) is disabled (if_unexpected.detail says which; for disabled, wait out whatever disabled it and re-discover). Put the focus on the field you named, then type again — not through a foreground keyboard type. By title: desktop_act with action='click'. By handle: re-discover by title first — except for a common dialog, which a title resolves to by handle as well, where nothing here can focus its text field yet. For another window: focus_window first

  • executor_failed → fall back to click_element / mouse_click / browser_click

For desktop_discover warnings (visual_provider_unavailable, visual_provider_warming, cdp_provider_failed, …), the coordinate-based tools (screenshot(detail='text'), click_element, mouse_click, terminal, …) remain available as an escape hatch.


Known limitations

Limitation

Detail

Workaround

Games / video players may return black or hang in PrintWindow capture

DirectX fullscreen apps may not redraw under PW_RENDERFULLCONTENT. Window-targeted screenshot(detail='image') already falls back to BitBlt automatically when PrintWindow returns no data or an all-black + zero-variance frame, but DirectX surfaces that hang the call don't surface as fallback.

Retry with screenshot({mode:'background', fullContent:false}) to switch to the legacy PrintWindow flag; if still black, the BitBlt fallback path (default mode='normal') will at least return the on-screen rect — hints.captureFallbackReason will say printwindow-all-black

UIA call overhead

~2 ms (focus) / ~100 ms (tree) via Rust native engine; ~300 ms via PowerShell fallback

Rust engine loads automatically; workspace_snapshot uses a 2 s timeout internally

Chrome / WinUI3 UIA elements are empty

Chromium exposes only limited UIA

screenshot(detail='text') auto-detects Chromium and falls back to Windows OCR (hints.chromiumGuard=true). For richer DOM access use browser_open + browser_locate

Chromium title-regex misses when sites rewrite document.title

Guard relies on the - Google Chrome suffix being present; some sites push it off the end of a long title

Title is treated as plain Chrome (UIA runs). OCR path is still reachable via ocrFallback='always' or when UIA returns <5 elements (uiaSparse)

browser_* CDP tools need Chrome launched with --remote-debugging-port

If Chrome is already running on the default profile without the flag, browser_open fails. The CDP E2E suite (tests/e2e/browser-cdp.test.ts) will also fail in that state

Close Chrome first, then browser_open({launch:{}}) will relaunch it in debug mode, or start Chrome manually with --remote-debugging-port=9222 --user-data-dir=C:\tmp\cdp

Layer buffer TTL

Buffer auto-clears after 90s of inactivity → next diffMode becomes an I-frame

After long waits, call workspace_snapshot to explicitly reset the buffer

keyboard(action='type') / keyboard(action='press') follow focus

When window_dock(action='dock')(pin=true) keeps another window on top (e.g. Claude CLI), keystrokes may be absorbed by that window

Call focus_window(title=...) first and verify isActive=true via screenshot(detail='meta') before sending keys

keyboard(action='type') em-dash / smart quotes in Chrome/Edge

Non-ASCII punctuation (em-dash —, en-dash –, smart quotes "" '') can be intercepted as keyboard accelerators, shifting focus to the address bar

Always use use_clipboard=true when the text contains such characters

browser_eval(action='js') on React / Vue / Svelte inputs

Setting element.value = ... or dispatching synthetic events does not update the framework's internal state

Use browser_fill(selector, value) — it uses native prototype setter + InputEvent which does update React/Vue/Svelte state


Token cost reference

Mode

Tokens

Use case

screenshot (768px PNG)

~443 tok

General visual check

screenshot(dotByDot=true) window

~800 tok

Precise clicking (no coordinate math)

screenshot(diffMode=true)

~160 tok

Post-action diff

screenshot(detail="text")

~100–300 tok

UI interaction (no image)

workspace_snapshot

~2000 tok

Full session orientation


Thanks

Huge thanks to everyone who tried a desktop-automation MCP server, filed issues, opened PRs, and shared what broke. Every bug report made the next release better. Thank you for building with me!


License

MIT

Available Tools

30 tools
browser_clickA

Click a DOM element in Chrome/Edge. Two ways to target: (1) selector — a CSS selector (combines browser_locate + mouse_click; stable across repaints); or (2) by-axis (semantic) — by:'text'|'regex'|'role'|'ariaLabel' + pattern, so you do not have to build a CSS selector for dynamic-class SPAs. by-axis resolves to a SINGLE actionable element (climbing to a clickable ancestor up to 3 levels, hit-testing for occlusion) and STOPS with code:'BrowserAmbiguousTarget' (candidates[] + next[] hints) when 2+ actionable elements match, or code:'BrowserNoActionableTarget' when matches exist but none is clickable — it never guesses. If the target is behind a modal dialog blocking the page, BOTH targeting modes STOP with code:'BrowserModalBlocking' (context.blockingElement {name, role}) instead of clicking through to the backdrop — dismiss the dialog (its close button or Escape) and retry; a plain navigation drawer does not count as blocking. Optionally add role to filter (by:'text',pattern:'Save',role:'button') and scope to narrow the search. Provide EITHER selector OR by+pattern (not both). Pass tabId+port so the server auto-guards (verifies tab readyState and identity) and returns post.perception.status. lensId is optional for advanced pinned-tab workflows. Caveats: selector mode fails if the element is outside the visible viewport — scroll it into view with browser_eval("document.querySelector('sel').scrollIntoView()") first (by-axis only resolves in-viewport actionable targets). hints.verifyDelivery:{status:'delivered'|'unverifiable', reason, observedSignals:{mutationCount,urlChanged,activeElementChanged}} reports the post-click observation in 2 values: 'delivered' fires only when mutationCount>0 OR urlChanged (activeElementChanged is recorded in observedSignals but intentionally NOT a delivery signal — plain clicks on focusable controls always update focus, treating that as 'delivered' would mask silent-fail regressions); 'unverifiable' reason ∈ {'iframe_context_mismatch','no_dom_mutation','probe_install_failed','probe_read_failed'}. CDP emits 2 values only (focus_only is a UIA-path concept, N/A here). BrowserClickNotDelivered is reserved-only (false-positive risk too high to emit) — degradation reads from 'unverifiable' status.

ParametersJSON Schema
NameRequiredDescriptionDefault
byNoSemantic axis to target by INSTEAD of a CSS selector: 'text' (visible text), 'regex', 'role' (ARIA/implicit role), 'ariaLabel'. Pair with pattern. Resolves to a SINGLE actionable element and STOPS with candidates when ambiguous.
portNoChrome/Edge CDP remote debugging port.
roleNoOptional ARIA/implicit-role filter AND-combined with by (e.g. by:'text', pattern:'Save', role:'button').
fixIdNoApprove a pending suggestedFix (one-shot, 15s TTL). Selector mode only.
scopeNoOptional CSS selector to limit the by-axis search scope (disambiguation).
tabIdNoTab ID from browser_open. Omit to use the first page tab.
lensIdNoOptional perception lens ID. Guards (target.identityStable) are evaluated before clicking, and a perception envelope is attached to post.perception on success.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
narrateNoNarration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on.minimal
patternNoValue matched against the chosen by axis (required when by is set).
selectorNoCSS selector for the target element (e.g. '#submit', '.btn'). Provide EITHER selector OR by+pattern.
caseSensitiveNoCase-sensitive matching for by:'text'/'regex' (default false).
scrollIntoViewNoWhen true, if the target is outside the viewport, scroll it into view (centered) before clicking, instead of failing with ElementNotInViewport. Default false preserves the explicit scrollIntoView-then-retry workflow. Selector mode only (by-axis resolves only in-viewport actionable targets).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to rely on, the description carries full behavioral burden and succeeds richly. It discloses auto-guards (tab readyState/identity), 'never guesses' ambiguity behavior, modal-blocking stop codes, delivery verification semantics, and the reserved-only BrowserClickNotDelivered status. These are genuine behavioral traits well beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, with major concepts front-loaded (targeting modes, ambiguity handling, modal blocking) and caveats/verification details following. Every sentence serves a purpose, though the delivery-verification section is quite dense and could overwhelm an agent scanning for quick facts. Still, it is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description covers what the agent needs to handle outcomes: error codes (BrowserAmbiguousTarget, BrowserNoActionableTarget, BrowserModalBlocking), verification status values, reasons, and observed signals. It also handles edge cases like iframe_context_mismatch and modal-vs-navigation-drawer distinction. Nothing critical is left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial meaning beyond parameter names: selector vs. by-axis exclusivity, role as an AND-filter, scope narrowing, tabId+port auto-guards, lensId purpose, scrollIntoView workflow implications, and include/narrate envelope behavior. This is far more than a restatement of the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Click a DOM element in Chrome/Edge' and immediately defines two distinct targeting modes (selector and by-axis), explicitly referencing siblings browser_locate and mouse_click. This gives a specific verb, resource, and clear differentiation from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use guidance: selector for stable CSS-based targeting, by-axis for dynamic-class SPAs. It also gives actionable exclusions and workarounds, such as using browser_eval to scroll elements into view, and instructs to retry after dismissing a modal dialog. It clearly states 'Provide EITHER selector OR by+pattern (not both).'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_evalA

Purpose: Inspect or operate on a browser tab via 3 actions: 'js' (evaluate JS), 'dom' (get HTML), 'appState' (extract SSR-injected SPA state). Details: action='js' — Run a JS expression. withPerception:true wraps in {ok, result, post}. action='dom' — Return outerHTML of selector (or document.body), truncated to maxLength. action='appState' — Scan Next/Nuxt/Remix/Apollo/GitHub/Redux SSR injected JSON; pass selectors to override defaults. Prefer: Use action='appState' BEFORE 'dom' or 'js' on SPAs where rendered HTML is sparse — single CDP call. Use 'dom' when 'appState' is empty and you need page structure. Use 'js' as the escape hatch for arbitrary scripting. Caveats: DOM nodes cannot be returned from action='js' directly (circular refs are serialized safely). React/Vue/Svelte controlled inputs cannot be set via element.value — use keyboard(action='type') / browser_fill instead. readyState is strictly checked; guard blocks if page is still loading. Typed errors: code:'BrowserNotConnected' on CDP disconnect (re-attach via browser_open); code:'AutoGuardBlocked' when the auto-guard refuses (e.g. page still loading) — the error message preserves the guard's 1-sentence recommended next step (most often wait_until({condition:'ready_state'}) or browser_eval readyState polling, then retry). Examples: browser_eval({action:'js', expression:'document.title'}) → page title browser_eval({action:'dom', selector:'#main', maxLength:5000}) → outerHTML browser_eval({action:'appState'}) → default SPA state probes

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoChrome/Edge CDP remote debugging port.
tabIdNoTab ID from browser_open. Omit to use the first page tab.
actionYesAction selector — one of: js, dom, appState. Per-action required fields are enforced at call time (see the tool description); this flat schema lists every action's fields as optional.
lensIdNoOptional perception lens ID. Guards (target.identityStable) are evaluated before eval.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
maxBytesNoMax bytes per individual payload (default 4000). Larger payloads are truncated.
selectorNoCSS selector for root element. Omit for document.body.
maxLengthNoMax characters of HTML to return (default 10000).
selectorsNoCustom probe selectors. Omit to use the default SPA framework set (__NEXT_DATA__ / __NUXT_DATA__ / __REMIX_CONTEXT__ / __APOLLO_STATE__ / window:__INITIAL_STATE__ etc.). Window globals must be prefixed with 'window:'.
expressionNoJavaScript expression to evaluate. The server automatically wraps snippets in an async IIFE to avoid repeated const/let collisions. For multi-statement snippets, use an explicit final return value. Declarations (const/let/var) are scoped per snippet — use window.* / globalThis.* for persistence. A single eval is bounded by the CDP per-command timeout (~15s): do NOT write in-page polling loops here — use wait_until (element_matches / url_matches / ready_state) to wait for conditions instead.
includeContextNoWhen true, append activeTab and readyState context to the response.
withPerceptionNoWhen true, return structured JSON {ok, result, post} with post.perception attached. Default false preserves raw-text return.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description fully discloses behaviors: actions, truncation, withPerception wrapping, readyState checks, error codes (BrowserNotConnected, AutoGuardBlocked), serialization limitations, and execution timeout. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Purpose, Details, Prefer, Caveats, Examples) and front-loads the purpose. It is slightly verbose but every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (12 parameters, 3 actions), the description covers all aspects, including return structures, error types, and recommended usage patterns. Examples illustrate typical calls. No output schema exists, but the description adequately hints at return shapes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds significant context beyond the schema, such as the purpose of each action, the effect of withPerception, and the use of selectors for appState. It provides examples that clarify parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Inspect or operate on a browser tab via 3 actions'. It specifies each action (js, dom, appState) and differentiates from sibling tools by focusing on evaluation/scripting rather than navigation, clicks, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Prefer:' section explicitly tells when to use each action, e.g., 'Use action='appState' BEFORE 'dom' or 'js' on SPAs'. The 'Caveats' section specifies when not to use this tool (e.g., for controlled inputs) and provides error handling guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_fillA

Fill a form input with a value via CDP — works on React/Vue/Svelte controlled inputs that reject browser_eval value assignment. Two ways to target: (1) selector — a CSS selector (use browser_overview / browser_locate to find one); or (2) by-axis (semantic) — by:'text'|'regex'|'role'|'ariaLabel' + pattern (e.g. by:'ariaLabel', pattern:'Email address', or by:'role', pattern:'textbox'), so you do not have to build a CSS selector. by-axis resolves to a SINGLE fillable element and STOPS with code:'BrowserAmbiguousTarget' (candidates[] + next[] hints) when 2+ match, or code:'BrowserNoActionableTarget' when the match is not a fillable input/textarea/contenteditable — it never guesses. Optionally add role to filter and scope to narrow. Provide EITHER selector OR by+pattern (not both). Use this over browser_eval when setting a controlled input's value via JS does not update framework state. Caveats: Requires browser_open (CDP active). actual in the response shows the element's value after fill; verify it matches the intended value (for a field the page masks, such as a password, neither the value you passed nor the one read back is echoed — valueWithheld:'masked' — and the match is still checked, in the page). Typed errors: code:'BrowserFillNotDelivered' on post-fill value mismatch — note the false-positive case where a React controlled input's onChange transforms the value (delivery actually succeeded; hints.verifyDelivery.subReason:'controlled_input_transform' for that case; the actual value is authoritative).

ParametersJSON Schema
NameRequiredDescriptionDefault
byNoSemantic axis to target by INSTEAD of a CSS selector: 'text' (visible text), 'regex', 'role' (ARIA/implicit role), 'ariaLabel'. Pair with pattern. Resolves to a SINGLE actionable element and STOPS with candidates when ambiguous.
portNoChrome/Edge CDP remote debugging port.
roleNoOptional ARIA/implicit-role filter AND-combined with by (e.g. by:'text', pattern:'Save', role:'button').
scopeNoOptional CSS selector to limit the by-axis search scope (disambiguation).
tabIdNoTab ID from browser_open. Omit to use the first page tab.
valueYesText to fill into the input element
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
patternNoValue matched against the chosen by axis (required when by is set).
selectorNoCSS selector for the input element. Provide EITHER selector OR by+pattern.
caseSensitiveNoCase-sensitive matching for by:'text'/'regex' (default false).
includeContextNoWhen true, append activeTab and readyState context to the response.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly. It discloses CDP dependency, that it stops with specific error codes on ambiguous or non-actionable targets, that it never guesses, that actual value is verified and authoritative, that masked fields suppress echoing, and even the false-positive controlled-input transform case with subReason hints. No contradiction with structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

While long, every sentence earns its place. The structure is logical: purpose and use-case first, then targeting methods, then caveats and error handling. It front-loads the core intent and flows from selection to verification. No redundancy; the density is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 11 parameters, no output schema, and no annotations, the description is remarkably complete. It covers targeting options, disambiguation, error codes and hints, verification via 'actual' and 'valueWithheld', and the CDP prerequisite. The only minor omission is a full enumeration of all response fields, but the description highlights the critical ones. This is sufficient for an agent to invoke it correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial semantic depth beyond the schema: it explains the behavior of the 'by' axis resolution, the semantics of 'pattern' paired with 'by', the role and scope filters, and the error scenarios tied to parameters. For instance, it clarifies that 'by' must be paired with 'pattern' and that 'selector' is an exclusive alternative. This enriches understanding beyond field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the exact action ('Fill a form input with a value via CDP') and explicitly differentiates from siblings by noting it works on controlled inputs that reject browser_eval assignment, and names browser_eval and browser_locate as alternatives. It also enumerates two distinct targeting methods, making its purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance on when to use this tool over browser_eval ('Use this over browser_eval when setting a controlled input's value via JS does not update framework state'), and when to choose selector vs. by-axis ('by-axis resolves to a SINGLE fillable element...'), with a direct constraint: 'Provide EITHER selector OR by+pattern (not both).' It also warns about the CDP prerequisite (browser_open).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_formA

Inspect all form fields (input, select, textarea, button) within a CSS-selector-specified container and return their name, type, id, current value (withheld for a field the page masks, such as a password: value is null, valueWithheld is 'masked', and hasValue says whether it holds one), hint text, disabled/readOnly state, and associated label text (resolved via for[id], ancestor LABEL, aria-labelledby, aria-label in that order). Use this before browser_fill to discover exact field selectors and avoid accidentally targeting the wrong input (e.g. a global search bar). Caveats: Requires browser_open (CDP active). Hidden inputs (type=hidden) are excluded by default — set includeHidden:true if needed. Value text is truncated at 200 chars.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoChrome/Edge CDP remote debugging port.
tabIdNoTab ID from browser_open. Omit to use the first page tab.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
selectorYesCSS selector for the form or container element to inspect (e.g. '#login-form', '.search-bar'). All input, select, textarea, and button descendants are returned.
maxResultsNoMaximum number of form fields to return (default 100).
includeHiddenNoWhen true, include hidden inputs (type=hidden). Default false to avoid CSRF-token / serialized-state clutter.
includeContextNoWhen true, append activeTab and readyState context to the response.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for behavioral disclosure. It discloses the CDP prerequisite, the default exclusion of hidden inputs, the 200-char truncation, and the masking behavior (value null, valueWithheld 'masked', hasValue) with the exact resolution order for labels. This is thorough and leaves no critical behavior unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than minimal but every sentence adds essential context. It front-loads the core purpose and output details, then covers caveats. The structure flows logically from what the tool does, to how to use it, to behavioral notes. It could be slightly tightened (e.g., combining some caveats), but it is efficient and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, no output schema), the description provides a complete picture: it explains the returned data structure (including edge cases like masked fields), the prerequisite, the default behavior for hidden inputs, and the label resolution order. An agent has everything needed to invoke it correctly and interpret results. No output schema exists, so the description's coverage of the response is essential and fully supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining the semantics of includeHidden (why default false: to avoid CSRF-token clutter), the output shape (masking details), and the container concept for selector. However, it does not elaborate on port, tabId, or include's envelope behavior, though those are already well described in the schema. The added context justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb ('Inspect'), a specific resource ('all form fields within a CSS-selector-specified container'), and enumerates the exact attributes returned. It also names the sibling tool it supports ('browser_fill') and the risk it mitigates (targeting the wrong input), making its purpose unambiguous and distinct from other browser tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use this before browser_fill' and explains why (to discover exact field selectors and avoid accidental targeting). It also provides caveats: requires browser_open (CDP active), hidden inputs are excluded by default, and when to set includeHidden:true. This gives clear when-to-use guidance and a precondition, though it does not name alternatives beyond browser_fill.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_locateA

Find a DOM element by CSS selector and return its physical screen coordinates — compatible directly with mouse_click. Prefer browser_click to find+click in one step. Prefer browser_overview to discover selectors. Caveats: Coordinates are captured at call time; if the page reflows before mouse_click, coords may be stale. Typed errors: code:'BrowserNotConnected' (call browser_open first), code:'ElementNotFound' (selector did not match — re-discover via browser_overview / browser_search).

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoChrome/Edge CDP remote debugging port.
tabIdNoTab ID from browser_open. Omit to use the first page tab.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
selectorYesCSS selector for the target element (e.g. '#submit', '.btn', 'button[type=submit]').
includeContextNoWhen true, append activeTab and readyState context to the response.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses capture-time coordinates, staleness risk, and typed errors. Could mention more about response shape, but sufficient for safe use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose, zero wasted words. Caveats and error patterns are clearly separated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but description states return of physical screen coordinates. More detail on return format (e.g., object with x,y) would improve completeness, but error types and compatibility hints are good.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, baseline 3. Description adds little beyond schema for parameters, but mentions coordinate compatibility with mouse_click, which is helpful context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states action (find DOM element) and output (physical screen coordinates), explicitly distinguishing from sibling tools browser_click and browser_overview.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly suggests using browser_click for find+click in one step and browser_overview to discover selectors. Also includes caveat about stale coordinates after page reflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_navigateA

Navigate a browser tab to a URL via CDP Page.navigate — more reliable than clicking the address bar. Pass tabId+port so the server auto-guards (verifies tab readyState) and returns post.perception.status. lensId is optional for advanced pinned-tab workflows. Caveats: Does not block until page load completes — the Page.navigate ack confirms only that the navigation request was accepted (frameStoppedLoading / loaderId observation is internal). Follow with wait_until({condition:'ready_state' or 'element_matches'}) or repeated browser_eval polling for slow pages. Typed errors: code:'NavigateFailed' (Page.navigate rejected — DNS failure, malformed URL, network unreachable; check URL + connectivity), code:'BrowserNotConnected' (CDP disconnect — re-attach via browser_open), code:'AutoGuardBlocked' when the auto-guard refuses (e.g. tab still loading) — the error message preserves the guard's 1-sentence recommended next step (most often wait_until({condition:'ready_state'}) then retry).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to navigate to
portNoChrome/Edge CDP remote debugging port.
tabIdNoTab ID from browser_open. Omit to use the first page tab.
lensIdNoOptional perception lens ID. Guards (target.identityStable) are evaluated before navigating, and a perception envelope is attached to post.perception on success.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
narrateNoNarration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on.minimal
waitForLoadNoWhen true (default), wait for document.readyState === 'complete' before returning. Use waitForLoad:false for the legacy behavior (return immediately after Page.navigate). Accepts the strings "true"/"false".
loadTimeoutMsNoMax milliseconds to wait for page load when waitForLoad=true (default 15000). On timeout, returns ok:true with readyState set to current state and hints.warnings=['NavigateTimeout'].

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses several valuable non-obvious traits: ack semantics, auto-guard behavior, post.perception.status, and three typed errors with remediation. However, it asserts 'Does not block until page load completes' while the schema's waitForLoad defaults to true and waits for readyState 'complete', a conflict that can mislead agents about default behavior. Since no annotations exist, the description carries the full burden, making this inconsistency more damaging.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded and information-dense, with a clear caveat label and a structured error taxonomy. It is slightly overstuffed—repeated wait_until advice and internal CDP jargon (frameStoppedLoading/loaderId) could be trimmed—but overall every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema or annotations, the description covers return perception status, error codes, auto-guard refusal, and recommended follow-ups for an 8-parameter CDP tool. The only serious omission is the waitForLoad/loadTimeoutMs behavior, which is contradicted by the caveat and leaves slow-page guidance incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds useful context for tabId/port (auto-guard) and lensId (pinned-tab workflows), but it does not compensate for the waitForLoad inconsistency nor add semantics for include/narrate beyond what the schema already describes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb (navigate), resource (browser tab to URL), and mechanism (CDP Page.navigate), and explicitly contrasts with clicking the address bar, which separates it from sibling browser tools. The lensId and perception-status details further clarify its scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly directs follow-up with wait_until or browser_eval polling for slow pages, and names browser_open for reconnection after BrowserNotConnected. It also frames the tool as more reliable than the address-bar alternative, giving an agent enough context to choose among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_openA

Connect to Chrome/Edge running with --remote-debugging-port and return open tab IDs — required before all other browser_* tools. Pass launch:{} (or with overrides) to auto-spawn a debug-mode browser when no CDP endpoint is live (idempotent: an already-running endpoint is preferred). Returns tabs[] with id, url, title, active — pass tabId to browser_* tools to target a specific tab. Caveats: CDP connection is per-process; if Chrome restarts, call browser_open again to get fresh tab IDs. A Chrome session started without --remote-debugging-port cannot be taken over — close it first or use a separate userDataDir. If the CDP endpoint is unreachable and launch is omitted, returns ok:false (typically code:'BrowserNotConnected' when the fetch surfaces ECONNREFUSED, otherwise code:'ToolError' with error 'Cannot reach Chrome/Edge CDP...'); re-call with launch:{} (idempotent) to auto-spawn or start Chrome manually with --remote-debugging-port=9222.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoChrome/Edge CDP remote debugging port.
launchNoIf set, spawn a debug-mode browser when no CDP endpoint is live on the target port (idempotent: an already-running endpoint is preferred and the spawn step is skipped). Pass {} to use defaults (chrome, C:\tmp\cdp, no initial URL). Omit to perform pure connect.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description fully discloses behavior: returns tabs array with specific fields, idempotent launch, error codes, killExisting warning data loss, and CDP connection per-process. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single paragraph but well-structured with clear logical flow. Slightly redundant on 'idempotent' but overall concise given the information density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all necessary aspects: prerequisite, input/output, errors, edge cases (browser restart, existing session, unreachable endpoint). References sibling tools. Complete for a setup tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already describes all parameters, but description adds context: explains launch parameter purpose, default behavior, example usage, and caveats for killExisting. Adds value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool connects to a debug-mode browser and returns open tab IDs, and it distinguishes from siblings by being the prerequisite for all other browser_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit instructions on when to use (required before other browser tools), when to use launch:{} vs pure connect, and caveats (Chrome restart, cannot take over existing session without debug port). Also gives error handling guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_overviewA

List all interactive elements (links, buttons, inputs, ARIA controls) on the current page with CSS selectors, visible text (an input's name — never what is typed in it), and viewport status — use before browser_click to discover stable selectors, and prefer this over screenshot when verifying button/toggle state after submission (no image tokens, structured output). scope limits to a CSS subsection (e.g. '.sidebar'). Returns state (checked/pressed/selected/expanded) for ARIA custom controls. Also returns a modal: section — whether a true modal dialog is blocking the page (isModal + blocker {name, role} + the signals it was judged on); it is ALWAYS present (isModal:false when no modal), and a navigation drawer is NOT reported as a modal (only an aria-modal / alertdialog / native showModal dialog, or a backdrop-backed dialog that locks the page, is treated as modal). Caveats: Selectors are CDP-generated snapshots — re-call after page navigates or re-renders. An input's text is its name — from aria-labelledby, aria-label, its , title, then its hint text — never what is typed in it: read values with browser_form, which withholds the value of a field the page masks (a password). Typed errors: code:'BrowserNotConnected' (CDP not attached — call browser_open or browser_open({launch:{}})). Note: a non-matching scope CSS selector silently falls back to the full document (does not raise an error) — verify the selector via browser_eval if scoped enumeration is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoChrome/Edge CDP remote debugging port.
scopeNoCSS selector to limit the search scope (e.g. '.s-main-slot', '#nav-search-form'). Omit to scan the full page.
tabIdNoTab ID from browser_open. Omit to use the first page tab.
typesNoElement types to include. Default 'all' returns links, buttons, and inputs.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
maxResultsNoMaximum number of elements to return (default 50).
inViewportOnlyNoWhen true, only return elements currently visible in the viewport.
includeContextNoWhen true, append activeTab and readyState context to the response.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure, and it does so richly. It discloses that selectors are CDP-generated snapshots requiring re-calls after navigation or re-render, that input text is always the name and never the typed value, that browser_form withholds masked-password values, that a non-matching scope silently falls back to the full document, and it names the BrowserNotConnected error condition. No annotation contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place: main purpose is front-loaded, usage guidance follows, and caveats are grouped and clearly labeled. Though lengthy, the complexity of the tool (8 parameters, modal behavior, error conditions, scope fallback) justifies the length. There is no filler or redundant restating of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description does a strong job of explaining what the agent will receive: ARIA state values, the always-present modal section with isModal and blocker details, and structured output. It also covers error conditions, scope fallback, and masked input behavior. The definition is complete enough for correct invocation and expectation-setting.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds materially beyond the schema by explaining the non-matching scope selector behavior (silent fallback to the full document) and by clarifying input-text semantics as they relate to output. Most other parameter details are already well-documented in the schema, so the added value is meaningful but not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'List all interactive elements (links, buttons, inputs, ARIA controls) on the current page' and enumerates the returned data (CSS selectors, visible text, viewport status). It also explicitly distinguishes the tool from browser_click and screenshot, so an agent can immediately tell what it does and why it exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: 'use before browser_click to discover stable selectors' and 'prefer this over screenshot when verifying button/toggle state after submission'. It also routes value reading to browser_form and suggests browser_eval for verifying a scope selector, providing clear context and alternatives without leaving inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

click_elementA

Invoke a UI element by name or automationId via UIA InvokePattern — no screen coordinates needed. The server auto-guards using windowTitle (verifies identity, foreground, modal) and returns post.perception.status. Prefer over mouse_click for buttons, menu items, and links in native Windows apps. Use desktop_discover first to discover automationIds. Pass fixId from a suggestedFix to re-target after window identity drift. lensId is optional for advanced pinned-lens use. Caveats: Typed errors: code:'InvokePatternNotSupported' — the control does not expose InvokePattern, fall back to mouse_click; code:'ElementDisabled' — the element is in a disabled state, re-check preconditions before retry; code:'GuardFailed' — read the perception envelope (attention / guard fields) and choose recovery (re-focus, wait, or pass the suggestedFix.fixId on the next call). Some custom controls do not expose InvokePattern at all; fall back to mouse_click for those.

ParametersJSON Schema
NameRequiredDescriptionDefault
hwndNoDirect window handle ID (takes precedence over windowTitle). String to avoid 64-bit precision issues.
nameNoElement name/label (partial match, case-insensitive)
fixIdNoApprove a pending suggestedFix (one-shot, 15s TTL).
lensIdNoOptional perception lens ID. Guards (safe.keyboardTarget, target.identityStable) are evaluated before clicking, and a perception envelope is attached to post.perception on success.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
narrateNoNarration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on.minimal
controlTypeNoControl type filter, e.g. 'Button', 'MenuItem'
windowTitleYesPartial window title of the target window. Use '@active' for the current foreground window.
automationIdNoExact AutomationId of the element

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears full behavioral disclosure burden. It goes beyond basics: explains server-side guarding using windowTitle, the return of post.perception.status, and enumerates typed error codes with precise recovery actions (InvokePatternNotSupported, ElementDisabled, GuardFailed). It also discloses limitations of custom controls. This is exceptional transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but tightly structured: purpose, usage guidance, then error handling caveats. Every sentence earns its place—no filler or redundancy. Despite its length, it remains front-loaded with the core purpose and immediately actionable guidance, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 9 parameters, no annotations, and no output schema, the description covers all essential contextual aspects: purpose, prerequisites (desktop_discover), selection criteria, error recovery, and limitations. It omits nothing an agent needs to call it correctly. The schema handles parameter details, and the description handles decision-making and error handling comprehensively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond schema by explaining the intent of fixId ('re-target after window identity drift') and lensId ('advanced pinned-lens use'), and mentions that include and narrate affect response shape. While the schema already documents each parameter, these contextual clarifications improve agent comprehension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Invoke a UI element by name or automationId via UIA InvokePattern — no screen coordinates needed.' It clearly differentiates from siblings like mouse_click and browser_click by identifying the use case (native Windows UI controls) and explicitly naming alternative tools. This is a precise, actionable purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'Prefer over mouse_click for buttons, menu items, and links in native Windows apps.' It also instructs to 'Use desktop_discover first to discover automationIds' and explains when to pass fixId. It even describes fallback cases, e.g., 'fall back to mouse_click' when InvokePattern is unsupported. No ambiguity remains about selecting this tool vs alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clipboardA

Read or write the Windows clipboard. action='read' returns current text content (empty string if non-text). action='write' replaces clipboard with given text and verifies delivery by reading the clipboard back and comparing the bytes (UTF-16LE) for exact equality. Caveats: Non-text clipboard payloads (images, files) return empty string on read. Reading text larger than about 8MB (roughly 4 million characters, far above anything this tool can write) is refused immediately rather than copied. Calls give up after 4s (read) / 5s (write) — the usual cause is that the application owning the clipboard has stopped responding, which blocks the clipboard for every application on the machine, so retrying does not help until it recovers or is closed. A write that gives up leaves the clipboard in an indeterminate state (it may still complete once that application recovers), unlike code:'ClipboardWriteNotDelivered' where the write is known not to have landed — re-read before relying on the contents. Overwrites existing clipboard content on write. action='write' delivery-verification failure returns code:'ClipboardWriteNotDelivered' — typical causes: a third-party clipboard manager intercepts SetClipboardData, DLP / endpoint protection blocks the payload, RDP / Citrix clipboard transcoding strips the text, or another process clears the clipboard between Set and the read-back. Recovery: retry the write, or fall back to keyboard(action='type', use_clipboard=false) for short text. On builds without the native addon (backend:'powershell') writes are additionally capped at about 12000 characters and return code:'ClipboardWriteTooLargeForFallback' above it. Diagnostics: every response reports backend:'native'|'powershell' — the implementation that served the call. A successful write adds postCloseChecked: whether the read that catches a clipboard manager swapping the payload actually ran (native uses a separate second read; powershell's single read-back is that read), plus postCloseSkipReason when it did not run. On backend:'native' only, a successful write also reports sequenceAfterWrite (a Windows clipboard sequence number, for diagnosis only — the delivery verdict is always the byte comparison) and, when the post-close read alone confirmed the write, inSessionReadable:false. Examples: clipboard({action:'write', text:'hello'}) → write+verify; clipboard({action:'read'}) → returns current text.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoText to place on the clipboard
actionYesAction selector — one of: read, write. Per-action required fields are enforced at call time (see the tool description); this flat schema lists every action's fields as optional.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must reveal all behavioral traits. It discloses timeouts (4s/5s), failure codes (ClipboardWriteNotDelivered), indeterminate state after a timed-out write, backend differences (native vs powershell), delivery verification, and diagnostic fields. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Despite its length, the description is front-loaded with the core purpose and uses clear sections (Caveats, Recovery, Diagnostics) to structure dense information. Every sentence adds actionable detail, so the length is justified rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description thoroughly covers return behavior (empty string for non-text), error conditions, environment variations, and diagnostics. It anticipates edge cases like clipboard managers and RDP transcoding, and offers recovery steps. No significant gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds critical semantics beyond it: the exact meaning of each action, the 8MB read limit, 12000-char fallback limit, and the delivery-verification process. It also explains the 'include' parameter's effect indirectly through response shape context, though not named.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Read or write the Windows clipboard,' clearly specifying the verb and resource. It further distinguishes the two actions ('read' returns text; 'write' replaces and verifies) and avoids confusion with sibling tools like keyboard by focusing on clipboard interaction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit alternative guidance for failed writes: 'fall back to keyboard(action="type", use_clipboard=false) for short text.' It also warns when retrying is futile (unresponsive clipboard owner) and when reading non-text payloads yields empty strings, effectively defining when-not-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_stateA

Purpose: Read-only observation of the current desktop state. Returns focused window/element, modal flag, attention signal from Auto Perception. Phase 4 absorbs former get_active_window / get_cursor_position / get_screen_info / get_document_state via include* flags. Details: Always returns: focusedWindow (title, hwnd, processName), focusedElement (name, type, value, automationId), cursorPos {x,y}, cursorOverElement (name, type), cursorOverWindow, hasModal (boolean), pageState ('ready'|'loading'|'dialog'), attention, visibleWindows count. Optional fields (default off): includeCursor:true → cursor {x,y,monitorId} (richer than cursorPos). includeScreen:true → screen {virtualScreen, displays[], displayCount, primaryIndex}. includeDocument:true → document {url, title, readyState, selection, scroll, viewport} via CDP (silently omitted on non-Chromium foreground). includeSessionContext:true (or include:['sessionContext']) → sessionContext {origin, consoleSessionId, sessionLabel, sessionState, ownWinStation} for Terminal Services session classification (ADR-017, observability-only). Chromium: cursorOverElement is null (UIA sparse); focusedElement may fall back to CDP document.activeElement; hints.focusedElementSource reports which path produced the row ('view' = engine-perception latest_focus, 'uia' = direct UIA query, 'cdp' = document.activeElement). Does NOT enumerate descendants — use desktop_discover for actionable entity list and window list. Prefer: Use after each action to confirm state. Cheapest observation tool — cheaper than any screenshot. attention='ok' means safe to proceed; other values require recovery (see suggest[]). Set include* flags only when you need the extra data (each adds one syscall or CDP round-trip). Caveats: Cannot detect non-UIA elements (custom-drawn UIs, game overlays). hasModal only detects modal dialogs exposed via UIA — browser alert/confirm dialogs may not appear here. includeDocument requires browser_open (CDP active); silently omitted otherwise with hints.documentUnavailable. focusedElement.value is the focused field's current text, so a plain credential field's value comes back like any other. Masked fields are withheld on the CDP road by rule; on the UIA road the value is whatever the provider serves and nothing here checks for a masked control — a masked field can arrive as mask characters, one per character of the secret, which nothing here distinguishes from its real value. It is also not always present, and hints.focusedElementValueAbsent says why when it is not: 'view_road_has_no_value' (the perception view is preferred and carries no values at all) or 'masked_on_this_road' (the CDP read dropped a masked field). No hint means nothing was dropped on the road that answered — on the UIA road an absent value can still be a field whose provider will not serve one.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoCDP port for includeDocument (default 9222).
tabIdNoOptional CDP tab id for includeDocument; omit for the focused tab.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
includeCursorNoWhen true, add a richer `cursor` field with monitor index alongside the lightweight `cursorPos`. Phase 4: absorbs former get_cursor_position. Default false.
includeScreenNoWhen true, add a `screen` field with all connected display info (resolution, position, DPI, scale). Phase 4: absorbs former get_screen_info. Default false. Use the displayId values returned here in screenshot / window_dock(action='dock').
includeDocumentNoWhen true, add a `document` field with the focused Chrome tab's url, title, readyState, selection, and scroll position via CDP. Phase 4: absorbs former get_document_state. Default false. Requires browser_open (CDP active); silently omitted on non-Chromium foreground.
includeSessionContextNoWhen true, add a `sessionContext` field with the Terminal Services session classification (origin, consoleSessionId, sessionLabel: 'console'|'rdp'|'other', sessionState: 'active'|'connected'|'disconnected'|'locked'|'unknown', ownWinStation). Default false. Equivalent to `include: ['sessionContext']`. Per ADR-017: observability-only — does not gate input. `sessionState: 'locked'` is a heuristic (active + foreground=null + previous sample within 60s saw a non-null foreground); treat it as a generic input-pause signal — it can also fire on secure-desktop transitions (UAC prompt, Credential UI), where the user-visible state is not strictly 'locked' but input is equally unavailable to this session.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and exceeds expectations. It discloses extensive caveats: non-UIA element detection limits, hasModal only seeing UIA dialogs, CDP fallback behavior, masked-field handling with road-specific details, and the 'attention' signal interpretation. It also explains why focusedElement.value may be absent and the meaning of hints. This level of transparency is exemplary and far beyond typical tool descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with Purpose, Details, Prefer, and Caveats sections. Every sentence carries substantive information, and the critical guidance (attention='ok' and prefer-after-action) is front-loaded. It is appropriately detailed for a complex tool with 7 optional parameters, though it could be tightened slightly without losing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 optional params, no required, no output schema, many edge cases), the description is remarkably complete. It covers all return fields, optional inclusions, failure modes (browser_open requirement, non-Chromium, masked fields), and even hints for troubleshooting. An agent has everything needed to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value beyond the schema: includeCursor is described as 'richer than cursorPos', includeDocument explains silent omission on non-Chromium, and includeSessionContext includes ADR-017 context and a heuristic explanation for 'locked' state. These are meaningful additions that help an agent decide when to enable flags.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise purpose: 'Read-only observation of the current desktop state.' It lists exactly what is returned (focused window, element, cursor, modal flag, etc.) and clearly distinguishes itself from desktop_discover by noting it does not enumerate descendants. The verb 'observe' and resource 'desktop state' are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance says 'Use after each action to confirm state' and notes it is the cheapest observation tool compared to screenshots. It explains when to enable each include* flag and warns that includeDocument requires browser_open. It also compares itself to desktop_discover for actionable entity lists, giving clear when-to-use vs when-not-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

excelA

Purpose: Author and run VBA macros against Excel via COM late binding (ADR-015). Headline differentiator against Claude for Excel which writes formulas but cannot run VBA. Details: action='run_vba' authors a Sub in a fresh workbook, saves into the managed Trusted Location (%LOCALAPPDATA%\desktop-touch-mcp\trusted-vba), and Application.Run the macro. Requires HKCU AccessVBOM=1 + VBAWarnings=1 + a registered Trusted Location (all configured by node scripts/enable-access-vbom.mjs). Trust setup: Excel must restart after the CLI runs (values cached at process start). action='check_access_vbom' is a read-only preflight returning {trusted, lockedByPolicy, scope}. Prefer: Run check_access_vbom first when a workflow depends on macro execution; the remediation hint pre-empts an opaque HRESULT 0x800a03ec failure inside run_vba. Caveats: macroName MUST appear as Sub <name>(...) in code (else VbaMacroNotFound). VBA Editor UI is structurally bypassed — no UIA tree walk needed. Excel COM is STA: each call serialises through the bridge's worker thread, so long-running macros block subsequent excel() calls on the same MCP server. Examples: excel({action:'check_access_vbom'}) → {trusted:true, scope:'hkcu'} excel({action:'run_vba', code:'Sub DesktopTouchAdHoc()\n Range("A1").Value = "Hello"\nEnd Sub'}) → {ok:true, workbookPath:'...\trusted-vba\dt_vba_.xlsm'} excel({action:'run_vba', code:'Sub Demo()\n MsgBox "hi"\nEnd Sub', macroName:'Demo', visible:true}) → demo recording path

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behavioral traits: required registry keys (AccessVBOM=1, VBAWarnings=1), Trusted Location setup, need for Excel restart, COM STA serialization causing blocking, and the requirement that macroName must match the Sub name in code. Failure modes (HRESULT 0x800a03ec) are also noted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Purpose, Details, Prefer, Caveats, Examples). It is comprehensive but each sentence adds necessary information, avoiding redundancy. Front-loading the purpose and distinction aids quick understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (VBA execution, COM, permissions), the description covers all essential aspects: purpose, setup requirements, prevalidation, caveats about naming and blocking, and complete examples with return values. No gaps remain despite the lack of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema is empty with additionalProperties: true, so it provides no parameter definitions. The description compensates by defining the key parameters (action, code, macroName, visible) and their meanings, including constraints like macroName must match Sub name. Examples show expected parameter combinations and their effects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the purpose: author and run VBA macros against Excel via COM late binding. It distinguishes from 'Claude for Excel' which writes formulas but cannot run VBA. The two actions run_vba and check_access_vbom are explicitly defined, making the tool's functionality unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides direct guidance: 'Prefer: Run check_access_vbom first when a workflow depends on macro execution'. It explains the preflight check and remediation for failures. Examples illustrate correct usage for both actions, clarifying when to use each.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

focus_windowA

Bring a window to the foreground by partial title match (case-insensitive). Use when a tool does not accept a windowTitle param, or when you need to switch focus before a sequence of actions. Use chromeTabUrlContains to activate a specific Chrome/Edge tab by URL substring before focusing — only the active tab's title appears in the windows list. If CDP is unavailable, chromeTabUrlContains is silently skipped — check response.hints.warnings. Returns WindowNotFound if no match exists; call desktop_discover to see available titles. Caveats: On some apps focus may be immediately stolen back (modal dialogs, UAC prompts) — verify with desktop_state after focusing. Win11 foreground refusal (UIPI cross-elevation / admin-only target / call from a background process or service) returns code:'ForegroundRestricted' ok:false instead of silently failing — recover by switching to a tool that does not require foreground transfer: desktop_act / click_element use UIA InvokePattern (no foreground needed); keyboard BG path bypasses foreground for terminal-class targets only (Windows Terminal / cmd / PowerShell — keyboard with windowTitle on non-terminal apps still hits the same ForegroundRestricted refusal). browser_* tools target by tabId/selector, not windowTitle.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesPartial window title to search for (case-insensitive)
cdpPortNoCDP port for chromeTabUrlContains (default 9222)
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
forceFocusNoWhen set, use AttachThreadInput-based foreground escalation on the first attempt. When omitted (default), focus_window first tries the standard SetForegroundWindow path and auto-escalates to force-focus only if Win11 refused the default attempt (issue #197). Override env: DESKTOP_TOUCH_FORCE_FOCUS=1 sets the implicit default to true. If both default and force paths fail, focus_window now returns ok:false code:'ForegroundRestricted' instead of the previous silent ok:true with windowChanged:false.
chromeTabUrlContainsNoWhen set, activate the Chrome/Edge tab whose URL contains this substring before focusing the window. Requires Chrome/Edge running with --remote-debugging-port (default 9222). Use this when the target is a Chrome tab that is not currently active — the active tab title is the only one visible in the window title list.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses case-insensitivity, partial match, CDP availability (silently skipped), focus stealing, Win11 foreground refusal with specific error code, and recovery options. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is relatively long but every sentence adds value. Front-loaded with core purpose. Well-structured, no redundancy. Concisely covers all necessary aspects.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given complexity (5 params, no output schema), description covers usage, alternatives, failure modes, caveats, and error recovery. Complete for an AI agent to select and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%. Description adds context beyond schema: explains when to use chromeTabUrlContains ('when the target is a Chrome tab that is not currently active'), details forceFocus behavior (auto-escalation), and include parameter options. Adds meaning beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states 'Bring a window to the foreground by partial title match (case-insensitive)', which is a specific verb and resource. It distinguishes from siblings by mentioning when to use focus_window vs chromeTabUrlContains and other tools like desktop_act.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance: 'Use when a tool does not accept a windowTitle param, or when you need to switch focus before a sequence of actions.' Also details when to use chromeTabUrlContains and provides alternatives for recovery from ForegroundRestricted (desktop_act, click_element, keyboard).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

keyboardA

Purpose: Send keyboard input to a window: 'type' for text, 'press' for key combos, 'sequence' for atomic multi-step chords. Details: action='type' inserts text (auto-clipboard for non-ASCII, bypassing IME conversion). action='press' sends key combos like 'ctrl+c'/'alt+tab'. action='sequence' runs ordered steps in one keyboard lock — use for Alt+letter, letter mnemonic chains where intermediate tool calls would close the menu. windowTitle or hwnd is REQUIRED (blank/whitespace counts as neither) — the server focuses and auto-guards that window (identity, foreground, modal) first, and a call with neither stops with DestinationRequired before any key is sent. Use windowTitle:'@active' to aim at the foreground window on purpose; an hwnd naming a titleless window works only while that window is already foreground. DESKTOP_TOUCH_REQUIRE_DESTINATION=0 downgrades the stop to a warning. Prefer: Set lensId for perception guards. Use desktop_act({action:'setValue'}) for UIA ValuePattern text fields. Caveats: win+r/win+x/win+s/win+l blocked. action='type' does not handle CJK IME composition — use use_clipboard=true or desktop_act({action:'setValue'}); neither lands while an IME composition is pending — commit or cancel it first. hints.clipboard reports the backend and whether the clipboard was restored. Non-ASCII text (CJK / emoji / diacritics / smart-quote-class punctuation) auto-clipboards to prevent silent-drop and Chrome accelerator hijack; pass forceKeystrokes:true to disable. Background (PostMessage/WM_CHAR) auto-engages for terminal-class windows (Windows Terminal / cmd / PowerShell); DTM_BG_AUTO=1 enables globally. Foreground non-terminal type runs a per-chunk leash; user focus-steal mid-stream aborts with FocusLostDuringType + context.typed/remaining; pass abortOnFocusLoss:false to disable. BG type verifies WM_CHAR via UIA TextPattern read-back; mismatch returns BackgroundInputNotDelivered (see SUGGESTS for false-positive notes). BG press read-back is scoped to terminal-class + enter/tab/arrow; other combos return verifyDelivery:'unverifiable', failure returns BackgroundKeyNotDelivered. action='sequence' is FG-only (BG/foreground_flash schema-rejected); emits verifyDelivery:'focus_only'; mid-loop focus theft returns MenuFocusLostMidSequence + context.remaining: Step[]. Win11 FG refusal returns ForegroundRestricted — terminal-class targets auto-engage BG (except a combo with ctrl/shift/alt and type with replaceAll, which go through FG and can return ForegroundRestricted too); non-terminal switch to desktop_act / click_element. Examples: keyboard({action:'type', text:'hello', windowTitle:'Untitled - Notepad'}) → text injected (guarded) keyboard({action:'type', text:'hello', windowTitle:'@active'}) → typed into the foreground window keyboard({action:'press', keys:'ctrl+c', windowTitle:'Untitled - Notepad'}) → copy keyboard({action:'press', keys:'escape', windowTitle:'Dialog'}) → dismiss dialog keyboard({action:'sequence', steps:[{keys:'alt+i', gapMs:100},{keys:'m'}], windowTitle:'Microsoft Visual Basic'}) → Insert > Module (atomic)

ParametersJSON Schema
NameRequiredDescriptionDefault
hwndNoDirect window handle ID (takes precedence over windowTitle). Obtain from get_windows response (hwnd field). String type to avoid 64-bit precision issues.
keysNoKey combo string, e.g. 'ctrl+c', 'alt+tab', 'enter', 'ctrl+shift+s'. Note: win+r, win+x, win+s, win+l are blocked for security.
textNoThe text to type (max 10,000 characters)
fixIdNoApprove a pending suggestedFix (one-shot, 15s TTL). Pass the fixId returned by a previous failed keyboard(action='type') to re-attempt with guard-validated args.
stepsNoOrdered list of key-press steps. Min 1, max 16. Total duration must not exceed 5000ms (excludes settleMs and focus acquisition). N=1 is allowed but inherits the sequence verification contract (hints.verifyDelivery.status='focus_only'); if you want the stricter keyboard:press contract, call keyboard({action:'press', keys}) directly (issue #278, matrix doc §3.1).
actionYesAction selector — one of: type, press, sequence. Per-action required fields are enforced at call time (see the tool description); this flat schema lists every action's fields as optional.
lensIdNoOptional perception lens ID. Guards (safe.keyboardTarget) are evaluated before typing, and a perception envelope is attached to post.perception on success.
methodNoInput method. background = WM_CHAR PostMessage (no focus change); foreground = SendInput (current default); auto = pick automatically.auto
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
narrateNoNarration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on.minimal
settleMsNoMilliseconds to wait before checking post-action state.
forceFocusNoBypass Windows foreground-stealing protection before focusing.
replaceAllNoWhen true, send Ctrl+A to select all existing text before typing. Equivalent to Ctrl+A → keyboard(action='type') in one call (requires field already focused). Default false.
trackFocusNoDetect if focus was stolen after the action.
forceImeOffNoIssue #245 系統②: when true, query the target window's IME open-status via Imm32 before typing; if ON, switch OFF for the duration of this call and restore the prior state in `finally`. Prevents silent romaji conversion when the user's Japanese IME is active but the LLM is typing ASCII commands. Requires `windowTitle` or `hwnd` (otherwise no target to query). Default false — existing use_clipboard auto-promotion still handles non-ASCII symbols transparently. No-op when the addon predates the IMM bridge (call proceeds with whatever IME state is in effect).
windowTitleNoPartial title of the window that should receive keyboard input.
use_clipboardNoIf true, copy text to clipboard and paste with Ctrl+V instead of simulating keystrokes. Use this when typing URLs, paths, or ASCII text into apps with Japanese IME active — pasted text is not run through IME conversion. Note this does not help while an IME composition is already in progress: the paste keystroke is consumed by the IME and nothing is inserted, so commit or cancel the composition first. Your clipboard is replaced for the duration of the call and put back afterwards; hints.clipboard reports which backend served the paste and whether the restore ran. On builds without the native addon this path is capped at about 12000 characters and fails with code:'ClipboardWriteTooLargeForFallback' above it. Default false.
forceKeystrokesNoWhen true, always use keystroke mode even if text contains non-ASCII content (CJK, emoji, diacritics, em-dash, smart quotes, etc.) that would normally trigger auto-clipboard. Default false — auto-clipboard is enabled.
abortOnFocusLossNoFocus Leash Phase B: when true, the foreground keystroke send is split into chunks (default 8 chars; override via DTM_LEASH_CHUNK_SIZE env) and the target window's foreground state is verified between chunks. If the user grabs focus mid-stream, the call aborts and returns FocusLostDuringType with context.typed (chars delivered to target) and context.remaining (unsent tail) so the caller can re-focus and retry the unsent portion. Default: true when windowTitle is provided, false otherwise. Has no effect on the clipboard path (atomic Ctrl+V) or the BG (WM_CHAR) path (HWND-targeted, foreground-independent).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries full weight and delivers: auto-clipboard behavior for non-ASCII, IME composition failure mode, focus-leash chunking with FocusLostDuringType, terminal-class BG auto-engagement, WM_CHAR verification via UIA read-back, sequence atomicity, blocked Windows key combos, and IME-bridge forceImeOff. It even exposes failure contract differences between sequence and press. This is far beyond a minimal safety note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and front-loaded with a clear Purpose line, then Details, Prefer, Caveats, and Examples. However, it is very long and mixes many distinct concerns (blocked combos, return contracts, env vars, issue references, read-back details) into a single prose block. The Examples section helps, but the Caveats section in particular is a wall of text that would benefit from segmentation; every sentence earns its place, but not every sentence earns its position.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 19 params, no output schema, and no annotations, the description covers the missing pieces: return hints (hints.clipboard, verifyDelivery status, context.typed/remaining), failure modes (FocusLostDuringType, BackgroundInputNotDelivered, ForegroundRestricted, MenuFocusLostMidSequence), and per-action verification contracts. It loses the 5th point because it omits a few edge states like the exact shape of the raw response and the interaction between include:['envelope'] and the narrated/rich post state, though the schema's include/narrate descriptions partially cover that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema descriptions are already rich (e.g., steps, use_clipboard, forceImeOff, abortOnFocusLoss cover their semantics thoroughly). The tool description adds cross-cutting semantics like action-specific validation, clipboard backend reporting, and the relationship between sequence and focus_only, but the per-parameter meaning is already fully carried by the schema. Base is 3; the description's additions are contextual rather than per-field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb–resource pair ('Send keyboard input to a window') and enumerates the three action modes ('type', 'press', 'sequence') with their distinct purposes. It differentiates from sibling tools by naming the intended alternatives (desktop_act setValue, click_element) rather than overlapping with browser/keyboard-free siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance is given for when to use each action, when to prefer setValue for UIA fields, and when to switch to desktop_act/click_element for non-terminal windows. It also states the exact destination requirements (windowTitle or hwnd REQUIRED, @active for foreground targeting) and the env-var downgrade path. The Prefer section and Caveats cover both positive and negative selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

key_lockerA

Purpose: Manage credentials the terminal autofills for you (SSH key passphrases, sudo / login passwords). Secrets are entered once into the locker's own secure dialog and stored encrypted on this machine (Windows DPAPI, current user); these stored secrets are NEVER shown to the assistant or sent to this tool — this covers what the locker holds, not every password on the machine. Details: action='save' pre-seeds a credential for a binding URI (ssh://user@host:22, sudo://host/root, https-cred://host:443, sshkey:SHA256:…): it opens the locker's secure entry dialog and stores the secret; the first save also shows a one-time enable confirmation. action='list' shows saved bindings (metadata only, never secrets). action='forget' deletes a binding and its secret. action='set_policy' toggles per-binding autofill confirmation. action='status' reports whether the locker is enabled (consent) and how many bindings exist. action='launch_console' opens (or reuses) an autofill-capable anchored pane and returns {paneId, windowTitle} — by default a new tab in the user's current Windows Terminal window (host:'classic' opens a dedicated classic console window instead). Prefer: Autofill is AUTOMATIC when a bound command triggers a credential prompt in the terminal — there is no manual fill action. But autofill ONLY fires in a pane opened by launch_console (a pre-existing terminal is never autofilled): to autofill, first launch_console, then run the ssh / sudo command with terminal({action:'run'|'send', paneId}) — pass the paneId field, not the windowTitle. Keep the returned paneId; there is no pane-listing action, but launch_console with fresh:false reuses the most-recent pane and returns its paneId again. Use save to enroll, list/status to inspect. Caveats: Windows-only. The anchored pane defaults to a Windows Terminal tab (autofill and terminal reads operate while that tab is the ACTIVE tab — switching away pauses them safely); host:'classic' opens a dedicated classic console window instead, and is the retry when Windows Terminal is not installed (KeyLockerWtUnavailable). The human can also see and type into the pane. Enabling the locker (first save or launch_console) grants BOTH credential autofill AND the ability for the assistant to launch a locker-owned pane. Disable the whole feature with DESKTOP_TOUCH_DISABLE_KEY_LOCKER=1. An ssh save needs the host key already in known_hosts (connect once first). API-token / env-var credentials are not supported yet. Examples: key_locker({action:'status'}) → {consentAccepted:false, disabled:false, bindingCount:0} key_locker({action:'save', uri:'sudo://buildbox/root'}) → opens the secure dialog → {captured:true} key_locker({action:'list'}) → {bindings:[{displayUri:'sudo://buildbox/root', scheme:'sudo', …}]} key_locker({action:'launch_console'}) → {paneId:'wt:31264:13322426700123', windowTitle:'dtm-locker-console-…'} → then terminal({action:'send', paneId:'wt:31264:13322426700123', input:'ssh user@host'}) key_locker({action:'launch_console', host:'classic'}) → {paneId:'12345678', windowTitle:'dtm-locker-console-…'} (dedicated classic console window)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It discloses that secrets are never shown to the assistant, that the first save shows a one-time enable confirmation, that autofill only works in an active tab, that enabling grants both credential autofill and pane-launch ability, that Windows-only, and that API-token/env-var credentials are not supported. It also explains the retry path for KeyLockerWtUnavailable. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: Purpose, Details, Prefer, Caveats, and Examples. It is front-loaded with the core purpose and action list, then workflow guidance, then caveats, then concrete examples. It could be slightly tightened (e.g., the parenthetical about host:'classic' appears twice), but the structure is logical and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 actions, credential security, terminal integration, Windows-specific behavior), the description is remarkably complete. It covers prerequisites (known_hosts for ssh saves), failure modes (KeyLockerWtUnavailable), security boundaries (secrets never shown), and the exact workflow with terminal. The examples show realistic input/output pairs. Nothing an agent needs to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema is empty with additionalProperties: true, so the description must document the action parameter and its values, which it does extensively. It explains each action (save, list, forget, set_policy, status, launch_console) and the uri format for save (ssh://user@host:22, sudo://host/root, https-cred://host:443, sshkey:SHA256:…). It also documents host:'classic' and fresh:false. The only minor gap is that it doesn't enumerate every possible property in the schema, but with 0 params and 100% schema coverage, the description compensates well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Manage credentials the terminal autofills for you (SSH key passphrases, sudo / login passwords).' It clearly distinguishes this from sibling tools like terminal, clipboard, and run_macro by focusing on credential storage and autofill. The action list (save, list, forget, set_policy, status, launch_console) further disambiguates the tool's scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Prefer' section explicitly states when to use this tool vs alternatives: autofill is automatic, but only fires in a pane opened by launch_console, and the agent must pass paneId, not windowTitle. It also gives a concrete workflow: 'to autofill, first launch_console, then run the ssh / sudo command with terminal({action:'run'|'send', paneId}).' This is explicit when-to-use guidance with no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mouse_clickA

Click at screen coordinates. Normally pass windowTitle so the server auto-guards the click (verifies target identity, foreground, coordinate is inside the target rect) and returns post.perception without a confirmation screenshot. origin+scale from dotByDot=true screenshots are converted to screen coords before guarding. doubleClick:true for double-click; tripleClick:true for triple-click (selects a full line of text). Prefer click_element (UIA) for native apps, prefer browser_click for Chrome. Examples: mouse_click({windowTitle:'Notepad', x:200, y:150}) // guarded — post.perception.status='ok'. mouse_click({x:100, y:100}) // unguarded — post.perception.status='unguarded'. If a guard failure returns a suggestedFix, pass its fixId to approve the fix: mouse_click({fixId:'fix-...'}) // one-shot, expires in 15s. lensId is optional and only for advanced pinned-target workflows; omit it for normal use. Caveats: origin+scale are meaningful ONLY with dotByDot=true screenshot responses. hints.verifyDelivery:{status:'delivered'|'focus_only'|'unverifiable', reason} reports the post-click observation; the status is decided from three signals — focused-element shift, the element under the cursor changing between two readable reads, and window-foreground change — or from none of them firing. Win11 foreground refusal during the homing path (UIPI cross-elevation / admin-only target / call from a background process or service) returns code:'ForegroundRestricted' ok:false rather than landing the click on the wrong window — recover by switching to a tool that accepts windowTitle directly (click_element / desktop_act) — browser_* tools target by tabId/selector, not windowTitle. MouseClickNotDelivered is reserved-only (false-positive risk is too high to emit a typed code), so degradation is expressed via the 'unverifiable' status, not a separate error.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesX coordinate. Screen-absolute by default. When 'origin' is provided, treated as image-local (pixel position within the screenshot).
yYesY coordinate. Screen-absolute by default. When 'origin' is provided, treated as image-local.
hwndNoDirect window handle ID (takes precedence over windowTitle). Obtain from get_windows response (hwnd field). String type to avoid 64-bit precision issues.
fixIdNoOne-shot fix approval ID. If a previous mouse_click returned a suggestedFix, pass that fixId here to approve it. The server revalidates the fix and executes with corrected args. fixId expires in 15 seconds and can only be used once.
scaleNoScale factor from screenshot response (only when dotByDotMaxDimension caused a resize). Omit if the screenshot was 1:1. Only used when 'origin' is also provided.
speedNoCursor movement speed in px/sec. 0 = instant.
buttonNoMouse button to clickleft
homingNoEnable homing correction if the target window moved.
lensIdNoOptional perception lens ID for advanced pinned-target workflows. When provided, guards are evaluated before clicking (safe.clickCoordinates, target.identityStable) and a perception envelope is attached to post.perception in the response. For normal use, omit lensId and pass windowTitle directly — Auto Perception handles tracking.
originNoWhen set, (x,y) are image-local coords from a screenshot. Server converts to screen coords: screen_x = origin.x + x / (scale ?? 1), screen_y = origin.y + y / (scale ?? 1). Copy origin values directly from the screenshot response text. This eliminates manual coord math and prevents out-of-window clicks.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
narrateNoNarration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on.minimal
settleMsNoMilliseconds to wait before checking post-action state.
elementIdNoAutomationId of the UI element.
forceFocusNoBypass Windows foreground-stealing protection before focusing.
trackFocusNoDetect if focus was stolen after the action.
doubleClickNoWhether to double-click
elementNameNoName or label of the UI element.
tripleClickNoWhether to triple-click (select a line of text). Takes precedence over doubleClick when both are true.
windowTitleNoPartial title of the target window.
verifyDeliveryYesParameter 'verifyDeliveryParam' from the Windows server schema.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and excels. It discloses auto-guard behavior, the post.perception status, the conversion of origin+scale, double/triple-click precedence, the suggestedFix mechanism with 15s expiry, the verifyDelivery status signals, Win11 foreground refusal handling with a specific error code, and the reserved-only MouseClickNotDelivered. Every notable behavioral trait is explicitly documented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but each sentence carries unique information. It is front-loaded with the core action and examples, then adds caveats and error recovery. The structure is logical, though the density could be slightly intimidating; however, there is no redundancy or filler, making it appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 21 parameters, nested objects, and no output schema, the description is remarkably complete. It explains return shapes (post.perception.status, hints.verifyDelivery), error recovery (fixId, ForegroundRestricted), parameter interactions, and even the rationale behind reserved error codes. An agent has everything needed to call this tool correctly and handle failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds real value by explaining the interplay of origin+scale with conversion math, the semantics of verifyDelivery statuses ('delivered'|'focus_only'|'unverifiable' and the three signals), the one-shot fixId expiry, and the precedence of tripleClick over doubleClick. It does not rehash every parameter but focuses on the non-obvious ones, elevating it above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Click at screen coordinates.' It immediately distinguishes from siblings by naming alternatives ('Prefer click_element (UIA) for native apps, prefer browser_click for Chrome') and clarifies scope. This gives an agent a clear, unambiguous purpose and differentiates it from similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool vs alternatives ('Prefer click_element (UIA)... prefer browser_click...'), when to pass windowTitle for guarded clicks, and when lensId is appropriate ('only for advanced pinned-target workflows'). It also explains the unguarded fallback and how to recover from guard failures via fixId. Usage context is fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mouse_dragA

Click and drag from (startX, startY) to (endX, endY) holding the left mouse button — for sliders, drag-and-drop, canvas drawing, and window resizing. Pass windowTitle so the server auto-guards the start coordinate and returns post.perception. Examples: mouse_drag({windowTitle:'Notepad', startX:50, startY:50, endX:200, endY:200}). lensId is optional and only for advanced pinned-target workflows. Caveats: Left button only. Both start and endpoint are guarded. Cross-window and desktop drags are blocked by default — pass allowCrossWindowDrag:true to confirm intent; that refusal is code:'CrossWindowDragBlocked'. A drag starting in a tabbed application's tab strip returns code:'TabDragBlocked' — pass allowTabDrag:true when tearing off or rearranging a tab is intended. hints.verifyDelivery:{status:'delivered'|'focus_only'|'unverifiable', reason} reports the post-drop observation in the same 3-value shape as mouse_click. MouseDragNotDelivered is SUGGESTS-registered but reserved-only (not emitted) — degradation is expressed via the 'unverifiable' status rather than a typed code. Win11 foreground refusal (UIPI cross-elevation / admin-only target / call from a background process or service) returns code:'ForegroundRestricted' ok:false from the homing path.

ParametersJSON Schema
NameRequiredDescriptionDefault
endXYes
endYYes
hwndNoDirect window handle ID (takes precedence over windowTitle). Obtain from get_windows response (hwnd field). String type to avoid 64-bit precision issues.
speedNoCursor movement speed in px/sec. 0 = instant.
homingNoEnable homing correction if the target window moved.
lensIdNoOptional perception lens ID. Guards and envelope same as mouse_click.
startXYes
startYYes
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
narrateNoNarration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on.minimal
windowTitleNoPartial title of the target window.
allowTabDragNoWhen true, allow drags that start in the title-bar / tab-strip area of a tabbed app (Notepad, Terminal, Edge, Chrome, etc.). Default false — such drags are blocked because they detach the tab into a new window rather than moving the window. Pass true only when you intentionally want to rearrange or detach a tab. Note: active only when auto-guard is enabled (same scope as allowCrossWindowDrag).
verifyDeliveryYesParameter 'verifyDeliveryParam' from the Windows server schema.
allowCrossWindowDragNoWhen true, allow dragging the endpoint into a different window or the desktop background. Default false — cross-window drags (including desktop/wallpaper) are blocked to prevent accidents. Pass true to confirm intent for deliberate cross-window or desktop-area drags.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behavioral traits: left-button only, both start and endpoint guarded, cross-window and tab drags blocked by default with specific error codes, verification status reporting, homing correction, and Win11 foreground restrictions. It also explains that MouseDragNotDelivered is reserved-only, showing deep transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and information-rich, front-loaded with purpose and examples, followed by caveats. It is longer than ideal but every sentence adds critical operational detail. The structure flows from core action to edge cases, making it usable despite the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 14 parameters, no output schema, and no annotations, the description covers all essential behavioral aspects: error codes, verification, guards, exceptions, and even server-specific refusal conditions. It leaves no critical operational gap for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, so the description compensates for missing schema descriptions on startX/startY/endX/endY by illustrating them in the example. It adds meaning to verifyDelivery (3-value shape) and clarifies the purpose of windowTitle, allowCrossWindowDrag, and allowTabDrag. Minor gap: no explicit description of the coordinate semantics beyond the example, but it's sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (click and drag), the resource (mouse coordinates), and specific use cases (sliders, drag-and-drop, canvas drawing, window resizing). It differentiates from mouse_click by specifying the drag behavior and includes an explicit example with parameter names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use the tool (e.g., for drag operations) and when to use specific flags (allowCrossWindowDrag, allowTabDrag) to override default blocks. It also mentions advanced lensId workflows, giving clear context for parameter choices.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notification_showA

Show a Windows system tray balloon notification to alert the user. Use at the end of a long-running task so the user knows it finished without watching the screen. Caveats: toast の user reach は原理的に観測不能 (matrix §3.1 line 158 規範整合)。Focus Assist (Do Not Disturb) / Notifications-off setting / consent UI sink いずれも tool 側からは判別不能のため、successful response は常に hints.verifyDelivery を含む (status="unverifiable", reason="user_visible_side_effect_uninspectable", channel="win32_balloon_tip" — 全 double-quoted JSON literal)。caller は user 側の post-notification behavior (例: wait_until(focus_changes)) で間接観測することが望ましい。Uses System.Windows.Forms — no external modules needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesNotification body text
titleYesNotification title
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the successful response always includes a hints.verifyDelivery object with status 'unverifiable' and explains the reasons (Focus Assist, DND, consent UI). It also mentions the underlying technology (System.Windows.Forms) and that no external modules are needed. However, it does not clarify if the tool is blocking or asynchronous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately concise but includes Japanese text and a JSON literal that may hinder readability. The first sentence is clear, followed by technical caveats that are all relevant but could be more structured. Every sentence serves a purpose, but the mix of languages and technical jargon reduces conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has only 3 parameters, no output schema, and no nested objects, the description covers the key aspects: purpose, usage context, delivery verification behavior, and a suggested indirect observation method. It is complete enough for an agent to use correctly, though it does not cover behavior for multiple rapid calls or notification click handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description does not add significant new meaning beyond what the schema already provides for the title, body, and include parameters. The description focuses on behavioral aspects rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Show a Windows system tray balloon notification to alert the user', providing a specific verb and resource. It distinguishes from sibling tools which are focused on browser interactions, desktop actions, and data processing, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use at the end of a long-running task so the user knows it finished without watching the screen', providing clear context for when to use. It also includes caveats about delivery verification and suggests indirect observation via wait_until, but does not mention alternative tools or explicitly state when not to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_macroA

Purpose: Execute multiple tools sequentially in one MCP call — eliminates round-trip latency for predictable multi-step workflows. Details: steps[] is an array of {tool, params} objects. Not every tool is a valid step: excel, key_locker, server_status, screenshot_query, screenshot_gc and run_macro itself never are, and whether desktop_discover/desktop_act or get_windows/get_ui_elements/set_element_value are depends on this server's configuration — each pair is refused in the other one. Everything else in the tool list is a step. The tool field's own description names the exact set THIS server dispatches. Plus a special sleep pseudo-step: {tool:"sleep", params:{ms:N}} (max 10000ms per step). stop_on_error=true (default) halts on first failure. Max 50 steps. The LLM cannot inspect intermediate results during execution — all steps run to completion (or first error) before any output is returned. Prefer: Use for predictable fixed sequences (focus → sleep → type → screenshot). Do not use for conditional logic — return to the LLM between branches so it can inspect intermediate state. Caveats: If any step may fail conditionally (e.g. a dialog that may or may not appear), split the macro at that point. Each screenshot step within a macro incurs the same token cost as a standalone call. Examples: [{tool:'focus_window',params:{windowTitle:'Notepad'}},{tool:'sleep',params:{ms:300}},{tool:'keyboard',params:{action:'type',text:'Hello'}},{tool:'screenshot',params:{detail:'text',windowTitle:'Notepad'}}] [{tool:'browser_navigate',params:{url:'https://example.com'}},{tool:'wait_until',params:{condition:'element_matches',target:{by:'text',pattern:'Example Domain'}}}]

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsNoOrdered list of tool calls to execute sequentially (max 50 steps).
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
stop_on_errorNoStop execution on the first error (default true). Set false to collect all results.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so excellently. It discloses which tools are never valid steps, that sleep is a pseudo-step, the max 50-step limit, stop_on_error behavior, that intermediate results cannot be inspected, and the token cost of screenshots. This is rich behavioral context well beyond what the schema provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section earns its place: purpose, details, usage preference, caveats, and two concrete examples. It is well-structured with clear headers and front-loads the core purpose. The length is justified by the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is highly complete for a complex tool: it explains step validity, execution flow, failure behavior, and examples. The only minor gap is that it doesn't explicitly describe the shape of the returned results (e.g., an array of per-step outputs), though it does note that output is only returned after all steps complete or an error occurs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds substantial meaning beyond the schema: it defines the structure of steps[] with {tool, params}, explains the special sleep pseudo-step, enumerates which tools are invalid as steps, and clarifies the include/stop_on_error parameters. This goes well beyond the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's specific purpose: 'Execute multiple tools sequentially in one MCP call.' It distinguishes itself from sibling tools by emphasizing batching and latency elimination, making it easy for an agent to understand this is a meta-orchestration tool rather than a single-action tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance is provided: 'Prefer: Use for predictable fixed sequences... Do not use for conditional logic — return to the LLM between branches.' It also gives concrete caveats about splitting macros when steps may fail conditionally, which tells the agent exactly when this tool is and isn't appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotA

Purpose: Capture desktop, window, or region across detail levels (meta / text / image / som / ocr) and capture modes (normal / background). Details: detail='meta' (default) returns window titles+positions only (~20 tok/window, no image). detail='text' returns UIA actionable elements with clickAt coords, no image (~100-300 tok). detail='som' returns OCR-detected elements with IDs plus a Set-of-Marks annotated image delivered by-ref by default (bypasses UIA entirely). detail='ocr' returns Windows OCR words with screen-pixel clickAt coords (Phase 4: absorbs former screenshot_ocr — use when UIA is sparse and you want to force OCR unconditionally). detail='image' and detail='som' both return a cheap by-ref resource_link by default (no inline base64); pass confirmImage=true to also embed the inline image (the annotated bitmap for som). mode='background' captures hidden/minimised/occluded windows via PrintWindow (Phase 4: absorbs former screenshot_background) — pair with windowTitle/hwnd. dotByDot=true returns 1:1 pixel WebP; compute screen coords: screen_x = origin_x + image_x (or screen_x = origin_x + image_x / scale when dotByDotMaxDimension is set — scale printed in response). diffMode=true returns only changed windows after the first call (~160 tok). region={x,y,width,height} captures a sub-rectangle (Phase 4: absorbs former scope_element when paired with windowTitle/hwnd — discover element bounds via desktop_discover, then pass region here). Data reduction: grayscale=true (−50%), dotByDotMaxDimension=1280 (caps longest edge), windowTitle+region (sub-crop to exclude browser chrome — e.g. region={x:0, y:120, width:1920, height:900}). Prefer: Use meta to orient, text before clicking, dotByDot only when precise pixel coords are needed. Use detail='som' for native apps or games that do not expose UIA elements (UIA-Blind). Use detail='ocr' for OCR-only (skip UIA entirely). Use mode='background' when the target window must stay hidden or cannot be brought to foreground. Prefer browser_* tools for Chrome. Use diffMode after actions to confirm state changed. Only use image+confirmImage when text returned 0 actionable elements and visual inspection is genuinely required. Caveats: Default mode scales to maxDimension=768 — image pixels ≠ screen pixels; apply the scale formula before passing to mouse_click. Foreground detail='image' returns a by-ref resource_link by default; pass confirmImage=true to also receive inline pixels. diffMode requires a prior full-capture baseline (non-diff call or workspace_snapshot) — calling diffMode cold returns a full frame, not a diff. mode='background' requires windowTitle or hwnd, and only composes with detail in {'image','meta'} — detail='text'/'som'/'ocr' run only against foreground capture (the dispatcher rejects the conflicting combination). Passing mode='background' is itself the acknowledgement that image pixels are wanted, so confirmImage is NOT required for it (matches the former screenshot_background contract). fullContent=false enables legacy mode (faster but GPU windows may be black). detail='ocr' requires windowTitle or hwnd; first call may take ~1s (WinRT cold-start) and the matching OCR language pack must be installed. Examples: screenshot() → meta orientation of all windows screenshot({detail:'text', windowTitle:'Notepad'}) → clickable elements with coords screenshot({detail:'ocr', windowTitle:'PDF', ocrLanguage:'ja'}) → OCR words with screen-pixel coords screenshot({mode:'background', windowTitle:'Chrome', dotByDot:true, dotByDotMaxDimension:1280, grayscale:true}) → background-capture pixel-accurate Chrome screenshot({windowTitle:'Notepad', region:{x:0,y:120,width:600,height:400}}) → cropped sub-region (zoom into element after desktop_discover)

ParametersJSON Schema
NameRequiredDescriptionDefault
hwndNoDirect window handle ID (takes precedence over windowTitle). Obtain from desktop_discover (windows[].hwnd). String type to avoid 64-bit precision issues.
modeNoCapture mode. 'normal' — default. Window-targeted captures (windowTitle / hwnd) use Win32 PrintWindow with automatic BitBlt fallback when PrintWindow returns no data or an all-black frame; the route used is reported in hints.captureSource. Fullscreen / displayId captures use BitBlt. 'background' — explicit Win32 PrintWindow capture, retained for back-compat and explicit selection. Requires windowTitle (or hwnd). Pair with fullContent for GPU-rendered apps.normal
detailNoResponse detail level (omit to let the server pick a smart default): omitted — auto: 'image' when dotByDot/region/displayId is specified, else 'meta' 'meta' — window title + screen region only (~20 tok/window, cheapest) 'text' — UIA element tree as JSON with text values (~100-300 tok/window, no image) 'image' — actual screenshot pixels. Returns a cheap by-ref resource_link by default (no inline base64); pass confirmImage=true to ALSO embed the inline image. 'som' — Set-of-Marks elements + annotated image (bypasses UIA entirely). Returns the OCR elements[] plus a cheap by-ref resource_link by default (no inline base64); pass confirmImage=true to ALSO embed the annotated bitmap. 'ocr' — Windows OCR words with screen-pixel clickAt coords (Phase 4: absorbs former screenshot_ocr). Use when UIA returns no actionable elements (WinUI3 custom-drawn UIs, game overlays, PDF viewers). Note: detail='text' auto-falls back to OCR via ocrFallback='auto'; choose detail='ocr' only when forcing OCR unconditionally.
regionNoCapture only this sub-region. Without windowTitle: virtual screen coordinates. With windowTitle: window-local coordinates — useful to exclude browser chrome (tabs/address bar). Example: windowTitle='Chrome', region={x:0, y:120, width:1920, height:900} skips the 120px browser chrome.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
diffModeNoLayer diff mode — compares each window against the buffered previous frame. First call = full I-frame (all windows). Subsequent calls = only changed windows (P-frame). Implicitly enables dotByDot. Best used with windowTitle=undefined to snapshot all windows.
dotByDotNo1:1 pixel mode — no scaling, WebP compression. Window captures include 'origin: (x,y)' so you can compute screen position: screen_x = origin_x + image_x. When dotByDotMaxDimension is also set, scale factor is included: screen_x = origin_x + image_x / scale.
displayIdNoCapture a specific monitor (0 = primary). Use desktop_state({includeScreen:true}) to list displays.
grayscaleNoConvert to grayscale before encoding. Reduces file size ~50% for text-heavy content (e.g. AWS console, code editors). Avoid when color is meaningful (charts, status indicators).
fullContentNoWhen mode='background', use PW_RENDERFULLCONTENT to capture GPU-rendered windows (Chrome, Electron, WinUI3). Default true. Set false for legacy mode (faster but GPU windows may appear black). Ignored unless mode='background'.
ocrFallbackNoOCR fallback behaviour when detail='text'. 'auto' (default): fire Windows OCR if UIA returns 0 actionable elements OR hints.uiaSparse=true (UIA returned <5 elements, typical for Chrome). 'always': always augment actionable[] with OCR words. 'never': disable OCR entirely.auto
ocrLanguageNoBCP-47 language tag for the OCR engine (e.g. 'ja', 'en-US'). Auto-detects from system locale when omitted. Used when detail='text' (OCR fallback) or detail='ocr' (direct OCR).
webpQualityNoWebP quality when dotByDot=true or diffMode=true. 40=layout only, 60=general (default), 80=fine text.
windowTitleNoCapture only the window whose title contains this string. Use '@active' for the current foreground window. Prefer over full-screen when target window is known.
confirmImageNoEmbed inline image pixels in the response. detail='image' now returns a cheap by-ref resource_link WITHOUT this flag (it is no longer blocked); confirmImage=true ADDITIONALLY embeds the inline image for immediate vision. detail='som' likewise returns its elements[] + a by-ref resource_link by default; confirmImage=true ADDITIONALLY inlines the annotated SoM bitmap. Prefer detail='text' / diffMode=true / dotByDot=true first — set confirmImage=true only when inline visual inspection is genuinely required.
maxDimensionNoMax width or height in pixels (default 768). Use 1280 to read small text, code, or fine UI details. Ignored when dotByDot=true.
preprocessPolicyNoOCR preprocessing scale policy for detail='som' and OCR fallback paths. 'auto' (default): clamp scale to 1 on OOM (>8MP) or high-DPI (≥150%). 'aggressive': relaxes DPI clamp to 175%, preserving upscale on 150%-DPI monitors (e.g. Outlook PWA). Also auto-enables adaptive binarization. 'minimal': always scale=1 regardless of DPI/resolution.auto
preprocessAdaptiveNoWhen true, apply Sauvola adaptive binarization after contrast stretch. Improves recognition of thin text on low-contrast or gradient backgrounds. Automatically enabled when preprocessPolicy='aggressive'. Requires Rust native engine; silently skipped otherwise.
dotByDotMaxDimensionNoCap the longest edge (pixels) when dotByDot=true. Reduces payload while preserving coordinate math. Example: 1280 on a 1920×1080 screen → scale≈0.667. Response includes scale factor: screen_x = origin_x + image_x / scale. Recommended for Chrome: dotByDot=true, dotByDotMaxDimension=1280, grayscale=true.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description thoroughly discloses behavioral traits: default scaling, diffMode cold-start behavior, background capture constraints, OCR fallback logic, and coordinate computation. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but well-structured with sections for purpose, details, prefer, caveats, and examples. Front-loaded with key info. Could be slightly more concise, but justified by tool complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but description covers return types (resource_link, inline image, elements) and important notes like the need for a baseline for diffMode. Addresses all key aspects for a complex tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value with examples, inter-parameter dependencies (e.g., region with windowTitle), and business logic (e.g., confirmImage behavior).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool captures desktop, window, or region with multiple detail levels and capture modes. It distinguishes from sibling tools like browser_* and screenshot_gc, and provides clear usage boundaries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Extensive guidance on when to use each detail level, mode, and combination. Includes a 'Prefer' section with explicit recommendations and a 'Caveats' section detailing restrictions and prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshot_gcA

Reclaim disk space from cached screenshots by retention policy. By DEFAULT this is a dry run: it returns the captures that WOULD be deleted (candidates) plus a count/size of leftover orphan files, and deletes nothing. To actually delete, pass BOTH dryRun:false AND confirm:true. Retention caps (all optional): maxCount (keep newest N), maxTotalBytes (keep newest under a byte budget), maxAgeMs (delete older than). When you pass none, the cache's env defaults apply (newest 200 / 256 MiB). Scope to a single tag with tag (other tags are never touched); includeOrphans (default true) also reclaims leftover on-disk files with no index entry. The newest capture is always kept by the count/byte caps. Only ever touches files inside the screenshot cache — never any other path.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoLimit deletion to captures under this tag (case-insensitive). Other tags are never touched.
dryRunNoDefault true: only LIST what would be deleted, delete nothing. Set false (with confirm:true) to actually delete.
confirmNoSafety gate: deletion happens ONLY when dryRun:false AND confirm:true. Otherwise the call is forced to a dry run.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
maxAgeMsNoDelete captures older than this many milliseconds (opt-in; can clear even the newest).
maxCountNoKeep only the newest N captures; delete the rest. The single newest is always kept.
maxTotalBytesNoKeep the newest captures under this total byte budget; delete older ones beyond it. The newest is always kept.
includeOrphansNoDefault true: also reclaim leftover on-disk image files that are not tracked in the cache index (e.g. files left behind by a crash).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description fully bears burden. Discloses dry-run default, two-flag safety gate, retention cap behaviors (always keeps newest), and scope limitation to screenshot cache only. Thoroughly covers behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single dense paragraph, front-loaded with key behavior and safety. Every sentence provides value; however, could benefit from clearer structure (e.g., bullet points for retention caps). Still highly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all 8 parameters, defaults, safety, scope, and return value in dry-run (candidates + orphan stats). No output schema, but description adequately explains what the call returns. Complete for a cleanup tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description adds value: explains default retention values (newest 200 / 256 MiB), safety interplay of dryRun/confirm, purpose of include (response shape), and includeOrphans default. Goes beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Reclaim disk space from cached screenshots by retention policy', specifying the action (reclaim) and resource (cached screenshots). It distinguishes from sibling tools like screenshot (capture) and screenshot_query (query) by focusing on garbage collection and cache cleanup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains dry-run default and safety condition (dryRun:false + confirm:true for actual deletion). Provides context on when to use (disk space reclamation) and scope options (tag, includeOrphans, retention caps). Implicitly distinguishes from capture/query tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshot_queryA

List screenshots already saved in the disk-cache WITHOUT re-reading any pixels. The screenshot tools return each capture as a cheap by-ref link (screenshot://by-ref/{captureId}); this lists what is in the cache — captureId + by-ref uri, dimensions, size in bytes, timestamp, and tag/window — so you can find and re-open a specific earlier capture, or check how much the cache holds before reclaiming space with screenshot_gc. The response also carries whole-cache totals (totalCaptures / totalBytes). Reading a capture's bytes still costs tokens, so open a by-ref link only when you actually need to inspect the pixels. Filter by tag (case-insensitive) / windowUuid / since / until; page with limit (default 50) and offset. Results are newest-first and never include a filesystem path.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoFilter to captures stored under this tag (case-insensitive). Omit to list all.
limitNoMaximum rows to return, newest first (default 50, max 500).
sinceNoOnly captures taken at/after this time (epoch milliseconds, inclusive).
untilNoOnly captures taken at/before this time (epoch milliseconds, inclusive).
offsetNoRows to skip from the newest end, for paging (default 0).
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
windowUuidNoFilter to captures of a specific window (the window's stable id).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behavior: non-destructive, lists returned fields (captureId, by-ref uri, dimensions, size, timestamp, tag/window, totals), notes ordering (newest-first), and warns about token costs for reading pixels. Also explains response shape options via 'include' parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five well-structured sentences: first states core purpose, then elaborates on return content, cost warning, filter/paging options, and ordering. No fluff, front-loaded with the main action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters (0 required) and no output schema, the description covers the tool's purpose, return shape (fields and totals), filtering, pagination, ordering, and token-cost warning. It adequately equips an agent to use the tool correctly without needing additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. Description adds extra context: defaults for limit (50), offsets, case-insensitivity for tag, and ordering (newest-first). While some info repeats schema, the description organizes and clarifies usage for pagination and filtering, adding meaningful value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'List' and resource 'screenshots saved in the disk-cache', emphasizing it is non-destructive and fast. It distinguishes from sibling tools like 'screenshot' and 'screenshot_gc' by describing its specific function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit use cases: 'so you can find and re-open a specific earlier capture, or check how much the cache holds before reclaiming space with screenshot_gc.' Also warns against unnecessary token cost when opening by-ref links. Differentiates from siblings and advises on when to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrollA

Purpose: Scroll a window or page. 5 strategies via action: 'raw' (wheel notches), 'to_element' (UIA name/automationId or CSS selector), 'smart' (auto-detect target with multi-strategy fallback), 'capture' (full-page stitched image), 'read' (scroll+OCR+dedupe → stitched text). Details: action='raw': send raw mouse-wheel notches at (x,y) or current cursor, optional window focus. Scroll scale — one amount unit is one wheel notch on every dispatch path, and ≈10 notches move one screenful: UIA-capable apps (Notepad, Explorer, WPF) step ≈1/10 of the visible area per notch, browsers and WebView apps (Chrome, Electron, Tauri) ≈100 px. amount:3 (default) ≈ a third of a screen. action='to_element': scroll a named element into viewport (UIA or CDP). action='smart': handles nested scroll layers, virtualised lists, sticky-header occlusion. action='capture': stitches full-page images (caps at ~700KB raw); sizeReduced=true means downscaled. action='read': scrolls page-by-page, OCRs each viewport, deduplicates overlapping lines, returns stitched text; language auto-detected from OS locale if omitted. Prefer: Use action='to_element' or action='smart' for click target out-of-viewport recovery (entity_outside_viewport) — scrolling only helps when the target scrolled out of its own window; if desktop_act reported origin_window_not_visible, restore the window with focus_window and re-run desktop_discover instead. Use action='capture' for reading long pages as images. Use action='read' for extracting text from long native-app documents (PDF readers, text editors, terminals) where copy-paste is unavailable. For simple scroll without target, use action='raw'. Caveats: action='capture' returns stitched image — pixels do NOT match screen coords when sizeReduced=true, use for reading only, not mouse_click. action='smart' CDP path requires browser_open. action='to_element' native path requires element to implement UIA ScrollItemPattern. action='read' uses OCR (imperfect accuracy) and requires the window to be visible; for browser pages prefer browser_eval or browser_overview for accurate DOM text. action='raw' typed errors: code:'ScrollNotDelivered' on silent drop (overlay / non-scrollable / UIPI low-IL); already-at-boundary is success via pre/post-percent disambiguation. hints.verifyDelivery.{channel,reason} per ADR-018 §2.6 — channel is the transport used ('uia'/'cdp'/'postmessage'/'wheel_send_input'); reason='pixel_delta_observed' = pixel evidence only (weak). action='smart' typed errors: code:'OverflowHiddenAncestor' (retry with expandHidden:true), code:'VirtualScrollExhausted' (provide virtualIndex). Examples: scroll({action:'raw', direction:'down', amount:5, windowTitle:'Chrome'}) scroll({action:'to_element', name:'OK', windowTitle:'Dialog'}) scroll({action:'smart', target:'#create-release-btn'}) scroll({action:'capture', windowTitle:'Chrome', maxScrolls:10}) scroll({action:'read', windowTitle:'Acrobat', maxPages:15}) // OCR + dedupe long PDF

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoX coordinate to scroll at (moves cursor there first)
yNoY coordinate to scroll at
hintNoScroll direction hint for binary-search (image path). Seeds lo/hi bounds to reduce attempts.
hwndNoDirect window handle ID (takes precedence over windowTitle).
nameNoPartial name/label of the element (UIA name match). Use for native app elements. At least one of name or selector must be provided.
portNoCDP port for Chrome path (default 9222)
blockNoVertical alignment after scroll — start/center/end/nearest (Chrome path only, default: center)center
speedNoCursor movement speed in px/sec (0=teleport, omit=default)
tabIdNoTab ID (Chrome path only). Omit for first page tab.
actionYesAction selector — one of: raw, to_element, smart, capture, read. Per-action required fields are enforced at call time (see the tool description); this flat schema lists every action's fields as optional.
amountNoNumber of scroll notches (default 3, max 1000). One notch = one physical wheel detent on every dispatch path. Count screenfuls, not lines: ≈10 notches move one screenful, so amount:3 (default) is about a third of a screen. UIA-capable apps (Notepad, Explorer, WPF) step ≈1/10 of the window's visible area per notch; browsers and WebView-based apps (Chrome, Electron, Tauri) move ≈100 px per notch. Exact distance depends on the app, the window size and the OS wheel-speed setting. The 1000-notch ceiling exists because each notch is dispatched as real wheel input; to reach a specific place in a long document use action='to_element' or action='smart' instead of a huge amount.
homingNoApply window-movement homing correction to (x,y) before scrolling. Default true.
inlineNoVertical alignment after scroll (CDP path). Default: center.center
targetNoCSS selector (Chrome/Edge) or partial UIA name (native apps). For CDP path, must be a valid CSS selector (starts with #, ., tag, or [ ). For UIA path, a partial name match against element Name property.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
languageNoOCR language code (e.g. 'ja', 'en', 'zh'). Omit to auto-detect from Windows system locale via Intl.DateTimeFormat().resolvedOptions().locale. Default: auto.
maxDepthNoMax number of ancestor scroll containers to walk. Default 3.
maxPagesNoMaximum number of scroll steps / OCR pages (default 20, max 50).
maxWidthNoMax size of the short edge of the final image (default 1280). For 'down': caps the image width; height is unconstrained. For 'right': caps the image height; width is unconstrained.
selectorNoCSS selector for the element (Chrome/Edge only). At least one of name or selector must be provided.
strategyNoauto (default): try CDP → UIA → image in order. cdp: Chrome/Edge only. uia: native Windows UIA. image: image + Win32 binary-search.auto
directionNoScroll direction
scrollKeyNoKey sent to scroll one page. PageDown (default): full-page scroll for most apps. Space: web/PDF readers. ArrowDown: line-by-line slow scroll.PageDown
maxScrollsNoMaximum scroll iterations before stopping (default 10, max 30)
retryCountNoMax scroll attempts (image path binary-search). Default 3, cap 4.
windowTitleNoPartial window title. When provided, the server focuses this window first.
expandHiddenNoTemporarily set overflow:hidden ancestors to overflow:auto to unlock scroll. Mutates live CSS.
virtualIndexNoTarget row index in a virtualised list (0-based). Enables direct TanStack/data-index seeking.
virtualTotalNoTotal row count in a virtualised list. Required when virtualIndex is set.
scrollDelayMsNoMilliseconds to wait after each scroll for rendering to settle (default 400). Increase for slow/animated pages.
verifyWithHashNoVerify scroll effectiveness via perceptual hash comparison. Automatically enabled for image path.
stopWhenNoChangeNoStop automatically when two consecutive pages yield no new lines after deduplication (page-end detection). Default true.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and delivers extensively: it discloses pixel-vs-screen-coordinate mismatch for capture, the mutating CSS side effect of expandHidden, OCR accuracy limits, the requirement for window visibility in read mode, typed error codes (ScrollNotDelivered, OverflowHiddenAncestor, VirtualScrollExhausted), and the weak-evidence caveat for verifyDelivery pixel_delta_observed. It also explains the wheel-notch semantics and the 1000-notch ceiling rationale.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely packed with high-value information, organized into Purpose, Details, Prefer, Caveats, and Examples sections. It front-loads the five strategies and their one-line definitions before diving into per-action details. The length is justified by the tool's complexity (5 modes, 32 parameters), though a few details (e.g., ADR-018 §2.6 reference) are terse to the point of being cryptic for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 32-parameter, 5-mode tool with no output schema and no annotations, the description is remarkably complete. It covers per-action semantics, error codes, prerequisites (browser_open for CDP, UIA ScrollItemPattern for native), fallback ordering, and even provides five usage examples covering different actions. The only minor gap is that the return shape for each action is not explicitly described, but the description's examples and caveats largely compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds substantial meaning beyond the schema: it explains the relationship between amount and screenfuls, clarifies that one notch equals one physical wheel detent, gives per-app step sizes, and explains per-action required fields are enforced at call time. It also adds context for parameters like sizeReduced (mentioned in description though not in schema) and virtualIndex/virtualTotal via the VirtualScrollExhausted error. Minor deduction because the description is not exhaustive for every parameter (e.g., homing, scrollKey, verifyWithHash are only in schema).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and resource ('Scroll a window or page') and immediately enumerates the five action strategies with one-line definitions. It distinguishes itself from siblings like mouse_click, browser_eval, and screenshot by naming its own sub-modes (raw, to_element, smart, capture, read) and their outputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Prefer' section explicitly states when to use each action and when not to: use to_element/smart for entity_outside_viewport recovery, use capture for reading long pages as images, use read for OCR text extraction, and use raw for simple scrolling. It also names alternatives (focus_window, desktop_discover, browser_eval, browser_overview) and gives exclusion conditions like 'scrolling only helps when the target scrolled out of its own window'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

server_statusA

Return MCP server status. engine: native engine availability — uia: 'native' = Rust UIA addon (~2 ms focus / ~100 ms tree); 'powershell' = PS fallback (~366 ms focus). nativeUia says why: 'native', 'disabled' (DESKTOP_TOUCH_DISABLE_NATIVE_UIA=1 takes the native UIA engine out and keeps the rest of the addon: UIA calls with a PowerShell version use it, the rest are skipped), or 'unavailable' (the addon has no UIA engine). nativeUiaEvidence says whether the native UIA engine actually ran in this process, as the engine and the OS answer rather than the switch: comThreadStarts and tasksSent (the engine's own counts of attempts), tasksDone (tasks the engine finished; far below tasksSent means native UIA is stuck and every UIA call is waiting out its timeout), focusRegistration ('pending' for long means focus events are off because a desktop-wide registration is not returning), uiaCoreLoaded (whether UIAutomationCore.dll is loaded, supporting evidence only); null = the addon cannot say. It changes per call, so read it after the act it should cover. imageDiff: 'native' = Rust SSE2 SIMD (0.26 ms @ 1080p); 'typescript' = TS fallback (~3.8 ms). health: process diagnostic snapshot (issue #365) — pid (which process these numbers are from; the evidence counts restart from zero in a new one), uptimeSec, memory.{rssBytes,heapUsedBytes,heapTotalBytes}, cpu.{userUs,systemUs} (cumulative since startup), shutdown.{pending,graceMs,inflightCount} (pending=true means stdin EOF received and grace timer is running), lastRpc.{receivedAt(ISO),method} (last JSON-RPC request observed on stdio transport; HTTP transport is not tracked). Diagnostic metadata — do not surface unless the user asks about performance/troubleshooting. engine values other than nativeUiaEvidence are stable for the process lifetime; nativeUiaEvidence and health change per call.

ParametersJSON Schema
NameRequiredDescriptionDefault
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that results change per call, that health is a process diagnostic snapshot, that nativeUiaEvidence reflects actual engine runs rather than the switch, that shutdown.pending indicates stdin EOF, and that lastRpc only tracks stdio transport. These are exactly the kind of behavioral caveats an agent needs, and they are clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but deliberately structured into field groups (engine, nativeUia, imageDiff, health) with terse explanations. The opening sentence states the core purpose, and each section earns its place by defining non-obvious semantics and failure modes. It could be trimmed slightly, but the density is justified by the tool's diagnostic complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description thoroughly documents every return value, sub-field, and their dynamic behavior. It explains limitations (HTTP transport not tracked), how to interpret counts (tasksDone vs tasksSent), and why a value might be null. An agent can call this tool and correctly interpret the response with no further information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the include parameter is fully documented in the input schema. The description adds nothing about the include parameter beyond what the schema says (envelope/raw), so the baseline of 3 applies. It does not worsen or improve the semantic gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Return MCP server status', a specific verb+resource pair. It then enumerates the distinct field groups (engine, nativeUia, imageDiff, health, lastRpc) with precise meanings, making it unmistakably different from sibling tools like desktop_state or workspace_snapshot. The resource and scope are fully stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: 'Diagnostic metadata — do not surface unless the user asks about performance/troubleshooting' signals when the tool is appropriate. It also advises 'read it after the act it should cover' because values change per call. It does not name explicit alternatives or exclude any sibling, but it orients the agent well.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terminalA

Purpose: Interact with a terminal window: read output, send input, or run+wait+read in one call. action='read' / action='send' absorb the formerly-standalone read/send tools (Phase 4). Details: action='run' is the recommended high-level workflow: send command → wait until quiet/pattern/timeout → read output. The command text is passed as input (the legacy parameter name command is also accepted as a deprecated alias — see issue #245). Returns completion={reason, elapsedMs} first-class plus outputIntegrity:'ok'|'baseline_lost' so callers can detect when scrollback could not be anchored to the pre-send buffer. action='read' reads current text via UIA TextPattern (falls back to OCR); use sinceMarker for incremental diff. action='send' sends a command with focus management. Prefer: action='run' for command execution + result. For long-running commands (test runners, builds, deploys) use until:{mode:'pattern', pattern:''} — the default quiet mode is tuned for short interactive commands and may complete prematurely on multi-second silent gaps mid-run. Use action='read'/'send' for fine-grained control or when you need to interleave other actions. read/send/run also accept a launch_console paneId (pass the paneId field, NOT windowTitle) to keep targeting a pane after its title changes. Caveats: Do not screenshot the terminal — action='read' is cheaper and structured. action='run' supports completion reasons: quiet | pattern_matched | exited | timeout | window_closed | window_not_found | send_failed (send rejected on a live window — see warnings for the underlying error code). until:{mode:'exit', shell:'bash'|'powershell'} (issue #386) returns completion.exitCode + reason:'exited' via an echo-immune sentinel that works for multiline input that pattern mode cannot anchor; pass shell explicitly (auto fails as ExitModeShellAmbiguous on WT/conhost/SSH), cmd is unsupported (ExitModeShellUnsupported), open-construct input is rejected (ExitModeUnsafeInput). When outputIntegrity:'baseline_lost' is returned, output is forced to '' and readError.code='BaselineMarkerLost' is set: rerun with until:{mode:'pattern',...} or longer timeoutMs. action='run' may also emit warnings prefixed FileLockCollision: when output reveals an EBUSY/Windows-lock/EAGAIN-EDEADLK file collision (e.g. shell '>' redirect colliding with the script's own writer — issue #236). Default quietMs=1500 (issue #196); long silences require pattern mode. preferClipboard=true (send default) replaces the clipboard. Hidden-input prompts emit verifyDelivery.unverifiable (reason:'hidden_input_prompt') — use method:'foreground'. action='read' typed errors: TerminalWindowNotFound, TerminalTextPatternUnavailable (force source:'ocr'); stale sinceMarker → hints.terminalMarker.previousMatched:false on ok:true (omit sinceMarker). FG-path Win11 foreground refusal returns code:'ForegroundRestricted' — switch to method:'background' or DTM_BG_AUTO=1. BG path auto-engages only when the target class is ConsoleWindowClass (conhost: cmd / PowerShell / pwsh) OR env DTM_BG_AUTO=1. Windows Terminal (CASCADIA_HOSTING_WINDOW_CLASS) is EXCLUDED (issue #173): WT runs on WinUI/XAML and silently drops WM_CHAR, so the FG path is default — pass sendOptions:{method:'background'} only if your WT build accepts BG input. Examples: terminal({action:'run', windowTitle:'PowerShell', input:'npm test', until:{mode:'pattern', pattern:'Test Files'}}) → recommended for test runners; matches when vitest summary appears terminal({action:'run', windowTitle:'pwsh', input:'ls'}) → quiet 1500ms wait, returns output (short interactive) terminal({action:'run', windowTitle:'pwsh', command:'ls'}) → identical to the above; command is a deprecated alias of input (issue #245) terminal({action:'read', windowTitle:'PowerShell', sinceMarker:'...'}) → incremental diff using the read action terminal({action:'send', windowTitle:'PowerShell', input:'echo hello'}) → sends text + Enter using the send action

ParametersJSON Schema
NameRequiredDescriptionDefault
inputNoCommand to send (Enter is appended automatically). Either `input` or its deprecated alias `command` is required.
untilNo
actionYesAction selector — one of: read, send, run. Per-action required fields are enforced at call time (see the tool description); this flat schema lists every action's fields as optional.
paneIdNo
commandNo[Deprecated alias of `input`] Accepted for callers that mis-remember the parameter name; new code should use `input`. If both are set, `input` wins.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
timeoutMsNoHard timeout in ms (default 30s)
readOptionsNoExtra options forwarded to terminal read (lines, source, ocrLanguage, etc.)
sendOptionsNoExtra options forwarded to terminal send (method, chunkSize, etc.)
windowTitleNoPartial title of the terminal window (e.g. 'PowerShell', 'pwsh', 'WindowsTerminal'). Provide windowTitle OR paneId (paneId takes precedence).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full burden. It richly discloses behavioral traits: completion reasons (quiet, pattern_matched, exited, timeout, etc.), outputIntegrity values, baseline_lost behavior, file-lock warnings, default quietMs, platform-specific nuances (Windows Terminal vs conhost, FG/BG paths), and hidden-input prompt handling. This transparency far exceeds typical tool descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely long but well-structured with clear sections (Purpose, Details, Prefer, Caveats, Examples) and front-loaded purpose. While every sentence is information-dense, some redundancy exists (e.g., command alias mentioned multiple times, multiple issue references). Still, for a tool with this complexity, the length is largely justified, though a trim would improve conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description thoroughly explains return values (completion={reason, elapsedMs}, outputIntegrity, error codes), behavioral details, and edge cases. Examples cover run/read/send with various options. The description is complete enough for an agent to select and invoke the tool correctly across diverse scenarios, including Windows Terminal limitations and error handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, but the description adds substantial meaning: it explains the deprecated `command` alias, clarifies paneId vs windowTitle precedence, details until modes (pattern/exit/quiet) with examples, and describes readOptions/sendOptions forwarding. It also gives concrete input examples that map parameters to usage scenarios, going well beyond the schema's field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: "Interact with a terminal window: read output, send input, or run+wait+read in one call." It clearly distinguishes three actions (read/send/run) and explicitly notes that this tool absorbs the formerly-standalone read/send tools, making its scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance: "Prefer: action='run' for command execution + result" and "Use action='read'/'send' for fine-grained control or when you need to interleave other actions." It also instructs users to avoid screenshots of the terminal, directing them to read instead, and explains when to use pattern mode for long-running commands.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_untilA

Purpose: Server-side poll for an observable condition — eliminates screenshot-polling loops when waiting for state changes. Details: condition selects what to watch: window_appears/window_disappears (target.windowTitle required), focus_changes (optional target.fromHwnd), element_appears/value_changes (target.windowTitle + target.elementName required, UIA; min 500ms interval), ready_state (target.windowTitle; visible + not minimized), terminal_output_contains (target.windowTitle + target.pattern required [+target.regex:true], needs terminal tools loaded), element_matches (target.by + target.pattern required, needs browser tools loaded), url_matches (target.pattern required [+target.regex:true]; matches the active tab's location.href via CDP — use for SPA route changes, redirects, OAuth flows). Returns {ok:true, elapsedMs, observed} on success, or WaitTimeout error with suggest hints. timeoutMs default 5000 (max 60000). Prefer: Use instead of run_macro({sleep:N}) + screenshot loops. Use terminal_output_contains to detect CLI command completion. Use element_matches for browser DOM readiness after navigation. Use url_matches when the URL is the most reliable signal (SPA routing / redirect cascades). Caveats: terminal_output_contains/element_matches/url_matches need a browser CDP connection (open --remote-debugging-port=9222 first). element_appears/value_changes spawn a UIA process per poll (interval floor 500ms). On timeout: {ok:false, code:'WaitTimeout', error, suggest:[...]}; suggest[] may open with a line earned by the last poll — read it, do not index. Those two also return context.lastLook: did the element resolve, why not, and for value_changes baseline/latest — the field's VALUE, so a masked credential arrives as mask characters. Other codes: 'ToolError' (validation / missing hook — read the message), 'BrowserNotConnected' (re-attach via browser_open), and 'WindowExcluded' for a target this server may not act through, answered at once rather than polled. Branch on code. Examples: wait_until({condition:'window_appears', target:{windowTitle:'Save As'}, timeoutMs:10000}) wait_until({condition:'terminal_output_contains', target:{windowTitle:'Terminal', pattern:'$ '}, timeoutMs:30000}) wait_until({condition:'element_matches', target:{by:'text', pattern:'Submit', scope:'#checkout-form'}}) wait_until({condition:'url_matches', target:{pattern:'/dashboard'}, timeoutMs:15000}) wait_until({condition:'url_matches', target:{pattern:'^https://app\\.example\\.com/orders/[0-9]+$', regex:true}})

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoTarget descriptor — fields used depend on condition. Accepts an object literal or a JSON-stringified object.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
conditionYesCondition to wait for. See per-condition target requirements.
timeoutMsNoMaximum time to wait (default 5000ms)
intervalMsNoPoll interval (default 200ms — terminal_output_contains uses 500 internally)

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly: it discloses success/error response shapes, all relevant error codes, browser/CDP prerequisites, per-poll UIA process spawning, the 500ms interval floor, and context.lastLook behavior including masked credentials. This goes far beyond what the schema or annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with Purpose, Details, Prefer, Caveats, and Examples sections, front-loading the most decision-relevant information. Each sentence earns its place given the nine-condition matrix, and the examples cover representative invocation patterns without becoming redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description fully covers return values, error codes, prerequisites, condition-specific target requirements, and example calls. An agent can select and invoke this tool correctly for any of the nine conditions without needing additional documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds critical meaning: per-condition target field requirements, optional regex flags, the internal 500ms interval for terminal_output_contains, the minimum interval for UIA conditions, and URL/SPA matching semantics. The target property in the schema is only described as 'depends on condition', so this text is essential for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Server-side poll for an observable condition', and immediately ties the purpose to eliminating screenshot-polling loops. It also differentiates itself from sibling tools such as run_macro and screenshot by explicitly contrasting those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Prefer' section explicitly states when to use this tool instead of run_macro({sleep:N}) + screenshot loops, and maps each condition to its intended use case (terminal completion, DOM readiness, SPA routing, redirects). The caveats further clarify when CDP prerequisites or UIA constraints apply.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

window_dockA

Purpose: Decorate a window: pin (always-on-top), unpin, or dock (move + resize + optional pin). Details: action='pin' makes window always-on-top until unpin/duration_ms. action='unpin' removes always-on-top. action='dock' positions to corner with width/height (default 480×360 bottom-right) and optionally pins. Minimized windows are automatically restored before docking. Prefer: Use action='dock' for terminal/CLI window auto-positioning at session start. Use action='pin' alone when you only need always-on-top without moving or resizing. Caveats: Pin survives minimize/restore; explicit action='unpin' needed to release. Dock fails on elevated processes. Dock overrides any existing Win+Arrow snap arrangement. Examples: window_dock({action:'dock', title:'PowerShell', corner:'bottom-right', width:480, height:360}) window_dock({action:'pin', title:'Settings', duration_ms:5000}) window_dock({action:'unpin', title:'Settings'})

ParametersJSON Schema
NameRequiredDescriptionDefault
pinNoIf true, set always-on-top so the docked window stays visible on top of other windows. Use window_dock(action='unpin') to remove the topmost flag later. Default true.
titleNoPartial window title (case-insensitive)
widthNoWindow width in pixels after docking. Default 480.
actionYesAction selector — one of: pin, unpin, dock. Per-action required fields are enforced at call time (see the tool description); this flat schema lists every action's fields as optional.
cornerNoScreen corner to snap the window to. Default 'bottom-right'.bottom-right
heightNoWindow height in pixels after docking. Default 360.
marginNoPixel padding between the window and the screen edge. Default 8.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
monitorIdNoMonitor to dock on (from desktop_state({includeScreen:true})). Omit for primary monitor.
duration_msNoAuto-unpin after this many ms (0–60000). Omit to pin indefinitely.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully bears the burden of behavioral disclosure. It explains pin survival across minimize/restore, need for explicit unpin, dock failure on elevated processes, override of Windows snap, and auto-restoration of minimized windows before docking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with labeled sections (Purpose, Details, Prefer, Caveats, Examples), front-loads the purpose, and every sentence adds value. It is concise yet comprehensive, fitting all necessary information into a manageable length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 parameters, 3 actions) and lack of output schema, the description provides thorough coverage of usage and behavior, including examples. However, it does not describe the return value or error handling, which would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds value beyond the schema by explaining defaults (480×360 bottom-right), the relationship between action and required fields, and the interplay between pin and duration_ms. However, the schema itself is already clear, so the description only slightly enhances understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs (decorate, pin, unpin, dock) and resources (window), distinguishing it from sibling tools like focus_window. The three actions are explicitly listed and explained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Prefer' section advising when to use action='dock' versus action='pin' alone, and includes caveats about elevated processes and overriding snap arrangements, giving clear guidance on appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_launchA

Purpose: Launch an application and wait for its new window to appear, returning title, HWND, and PID. Details: Runs the command via ShellExecute, snapshots the window list before launch, then polls until a new HWND appears (compared by HWND, not title). Returns {windowTitle, hwnd, pid, elapsedMs}. Works for localized window titles (e.g. '電卓' for calc.exe) because detection is HWND-based, not title-based. timeoutMs default 10000. detach=true fires without waiting and returns no window info. Prefer: Use instead of run_macro({exec, sleep, desktop_discover}) combos. Follow with focus_window(windowTitle) to interact with the launched app. Caveats: Single-instance apps that reuse an existing window will not register as a new HWND — call desktop_discover first to check if the window is already open. detach=true returns immediately with no window title or hwnd. Examples: workspace_launch({command:'notepad.exe'}) → {windowTitle:'', hwnd:'...', pid:...} workspace_launch({command:'calc.exe', timeoutMs:15000})

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoCommand-line arguments (max 20). Shell metacharacters (; & | ` $() ${}) are not allowed.
waitMsNoMilliseconds to wait for the window to appear (default 2000)
commandYesExecutable name or full path (e.g. 'notepad.exe', 'calc.exe'). Shell interpreters (cmd.exe, powershell.exe, etc.) are blocked.
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and richly discloses behavior: HWND-based detection, polling mechanism, localized title handling, and detach effects. It explains the snapshot-then-poll logic and default timeout, leaving no ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Purpose, Details, Prefer, Caveats, Examples). It is front-loaded with the main purpose and every sentence adds value without redundancy. Efficient use of space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description fully explains the return shape and edge cases. It covers localization, single-instance apps, and detach behavior, making it complete for an AI agent to understand invocation and outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage, but the description adds valuable context: explains command and args restrictions, default timeout (though slightly inconsistent with schema's waitMs), and includes examples. However, the minor discrepancy in default value (timeoutMs vs waitMs) prevents a perfect score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Launch an application and wait for its new window to appear, returning title, HWND, and PID.' It specifies the verb (launch), resource (application), and return values, distinguishing from sibling tools like run_macro.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly recommends using this tool instead of run_macro combos and provides follow-up actions like focus_window. It also includes caveats for single-instance apps and detach behavior, offering clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_snapshotA

Purpose: Orient fully in one call — returns display layouts, all window thumbnails (WebP), and per-window actionable element lists with clickAt coords. Details: uiSummary.actionable[] per window includes: action ('click'|'type'|'expand'|'select'), clickAt {x,y} (pass directly to mouse_click), value (current text for editable fields). Runs parallel internally; latency ≈ max(single screenshot), not N×screenshots. Also resets the diffMode buffer so subsequent screenshot(diffMode=true) returns only changes (P-frame). Prefer: Use at session start or after major workspace changes. Use screenshot(detail='meta') for cheap re-orientation within a session. Use screenshot(detail='text', windowTitle=X) for a single-window update. Caveats: Thumbnails are scaled, not 1:1 — use screenshot(dotByDot=true, windowTitle=X) for pixel-accurate coords on a specific window after snapshot. Also: this call resets the screenshot diff baseline (I-frame) and identity tracker as a side effect, so subsequent screenshot(diffMode=true) starts fresh from this snapshot. The reset is not currently exposed in causal/working memory — record an explicit 'workspace_snapshot' step if you need to track the reset point in your causal trail (ADR-010 §11 OQ carry-over for full visibility).

ParametersJSON Schema
NameRequiredDescriptionDefault
includeNoOptional response-shape opt-in. `['envelope']` returns the self-documenting envelope (`_version` / `data` / `as_of` / `confidence`). `['raw']` forces raw shape (overrides DESKTOP_TOUCH_ENVELOPE=1 server default). Default behaviour is raw shape (compat with existing clients).
includeUiSummaryNoWhether to include UI element summaries for each window
thumbnailMaxDimensionNoMax size of per-window thumbnail images (default 400px)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses important behavioral traits: parallel execution, resetting the diffMode buffer, side effects on diff baseline and identity tracker. It also warns about thumbnail scaling and recommends tracking the reset point. This is thorough and transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with headings (Purpose, Details, Prefer, Caveats) and each section adds value. However, it is somewhat verbose, especially the caveat referencing ADR-010. It could be slightly more concise without losing information, but it remains effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters, no output schema, and no annotations, the description is remarkably complete. It explains what is returned, side effects, usage patterns, and caveats. It fully equips the agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (all three parameters have descriptions in the input schema). The description does not add significant new meaning beyond the schema; it references 'includeUiSummary' indirectly and mentions thumbnail dimensions, but the schema already covers these. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Orient fully in one call' and enumerates what it returns (display layouts, thumbnails, actionable element lists). It is a specific verb+resource combination and distinguishes from sibling tools like screenshot and screenshot variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Use at session start or after major workspace changes.' It also offers alternatives: 'Use screenshot(detail='meta') for cheap re-orientation' and 'screenshot(detail='text', windowTitle=X) for a single-window update.' This helps the agent decide when to use this tool versus others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv2.0.0
    • Changedbrowser_click1 field changed
      • changedInput schema / properties / narrate / description
        Previous value: -"Narration level. rich includes UIA or browser state diff when supported."New value: +"Narration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on."
    • Changedbrowser_navigate1 field changed
      • changedInput schema / properties / narrate / description
        Previous value: -"Narration level. rich includes UIA or browser state diff when supported."New value: +"Narration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on."
    • Changedclick_element1 field changed
      • changedInput schema / properties / narrate / description
        Previous value: -"Narration level. rich includes UIA or browser state diff when supported."New value: +"Narration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on."
    • Changedkeyboard1 field changed
      • changedInput schema / properties / narrate / description
        Previous value: -"Narration level. rich includes UIA or browser state diff when supported."New value: +"Narration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on."
    • Changedmouse_click1 field changed
      • changedInput schema / properties / narrate / description
        Previous value: -"Narration level. rich includes UIA or browser state diff when supported."New value: +"Narration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on."
    • Changedmouse_drag1 field changed
      • changedInput schema / properties / narrate / description
        Previous value: -"Narration level. rich includes UIA or browser state diff when supported."New value: +"Narration level. rich includes UIA or browser state diff when supported, and is withheld with post.rich.diffDegraded when the diff cannot be shown to describe the window that was acted on."
    • Changedscroll4 fields changed
      • changedInput schema / properties / amount / description
        Previous value: -"Number of scroll notches (default 3). UIA-capable apps (Notepad, Explorer, WPF — Tier 1): empirically ≈1 text line per notch; amount:3 (default) ≈ 3 lines (small nudge), amount:10 ≈ 10 lines (~½ visible area). Legacy apps (SendInput path): each amount unit sends 3 wheel ticks; at Windows default 3 lines/tick that is ≈9 text lines per unit — distance varies by app/OS wheel-speed settings."New value: +"Number of scroll notches (default 3, max 1000). One notch = one physical wheel detent on every dispatch path. Count screenfuls, not lines: ≈10 notches move one screenful, so amount:3 (default) is about a third of a screen. UIA-capable apps (Notepad, Explorer, WPF) step ≈1/10 of the window's visible area per notch; browsers and WebView-based apps (Chrome, Electron, Tauri) move ≈100 px per notch. Exact distance depends on the app, the window size and the OS wheel-speed setting. The 1000-notch ceiling exists because each notch is dispatched as real wheel input; to reach a specific place in a long document use action='to_element' or action='smart' instead of a huge amount."
      • addedInput schema / properties / amount / maximum
        Added value: +1000
      • addedInput schema / properties / amount / minimum
        Added value: +1
      • addedInput schema / properties / target / minLength
        Added value: +1
  2. 1 tool updatev1.14.3
    • Changedkeyboard1 field changed
      • changedInput schema / properties / use_clipboard / description
        Previous value: -"If true, copy text to clipboard and paste with Ctrl+V instead of simulating keystrokes. Use this when typing URLs, paths, or ASCII text into apps with Japanese IME active — prevents IME from converting characters. Default false."New value: +"If true, copy text to clipboard and paste with Ctrl+V instead of simulating keystrokes. Use this when typing URLs, paths, or ASCII text into apps with Japanese IME active — pasted text is not run through IME conversion. Note this does not help while an IME composition is already in progress: the paste keystroke is consumed by the IME and nothing is inserted, so commit or cancel the composition first. Your clipboard is replaced for the duration of the call and put back afterwards; hints.clipboard reports which backend served the paste and whether the restore ran. On builds without the native addon this path is capped at about 12000 characters and fails with code:'ClipboardWriteTooLargeForFallback' above it. Default false."
  3. 1 tool updatev1.13.0
    • Changedterminal2 fields changed
      • addedInput schema / properties / paneId
        Added value: +{
        +  "type": "string"
        +}
      • changedInput schema / properties / windowTitle / description
        Previous value: -"Partial title of the terminal window (e.g. 'PowerShell', 'pwsh', 'WindowsTerminal')."New value: +"Partial title of the terminal window (e.g. 'PowerShell', 'pwsh', 'WindowsTerminal'). Provide windowTitle OR paneId (paneId takes precedence)."
  4. 20 tool updatesv1.12.0
    • Addedbrowser_click
    • Addedbrowser_eval
    • Addedbrowser_fill
    • Addedbrowser_form
    • Addedbrowser_locate
    • Addedbrowser_navigate
    • Addedbrowser_open
    • Addedbrowser_overview
    • Addedbrowser_search
    • Addedclick_element
    • Addedclipboard
    • Addeddesktop_state
    • Addedexcel
    • Addedfocus_window
    • Addedkey_locker
    • Addedkeyboard
    • Addedmouse_click
    • Changedscreenshot2 fields changed
      • changedInput schema / properties / confirmImage / description
        Previous value: -"Must be true to receive image pixels when detail='image'. Without this flag, detail='image' is blocked and a guidance message is returned instead. Prefer detail='text' / diffMode=true / dotByDot=true first — only set confirmImage=true when visual inspection is genuinely required."New value: +"Embed inline image pixels in the response. detail='image' now returns a cheap by-ref resource_link WITHOUT this flag (it is no longer blocked); confirmImage=true ADDITIONALLY embeds the inline image for immediate vision. detail='som' likewise returns its elements[] + a by-ref resource_link by default; confirmImage=true ADDITIONALLY inlines the annotated SoM bitmap. Prefer detail='text' / diffMode=true / dotByDot=true first — set confirmImage=true only when inline visual inspection is genuinely required."
      • changedInput schema / properties / detail / description
        Previous value: -"Response detail level (omit to let the server pick a smart default):\n  omitted — auto: 'image' when dotByDot/region/displayId is specified, else 'meta'\n  'meta'  — window title + screen region only (~20 tok/window, cheapest)\n  'text'  — UIA element tree as JSON with text values (~100-300 tok/window, no image)\n  'image' — actual screenshot pixels. BLOCKED unless confirmImage=true is also passed.\n  'som'   — Set-of-Marks image + OCR elements (bypasses UIA entirely). BLOCKED unless confirmImage=true is also passed.\n  'ocr'   — Windows OCR words with screen-pixel clickAt coords (Phase 4: absorbs former screenshot_ocr). Use when UIA returns no actionable elements (WinUI3 custom-drawn UIs, game overlays, PDF viewers). Note: detail='text' auto-falls back to OCR via ocrFallback='auto'; choose detail='ocr' only when forcing OCR unconditionally."New value: +"Response detail level (omit to let the server pick a smart default):\n  omitted — auto: 'image' when dotByDot/region/displayId is specified, else 'meta'\n  'meta'  — window title + screen region only (~20 tok/window, cheapest)\n  'text'  — UIA element tree as JSON with text values (~100-300 tok/window, no image)\n  'image' — actual screenshot pixels. Returns a cheap by-ref resource_link by default (no inline base64); pass confirmImage=true to ALSO embed the inline image.\n  'som'   — Set-of-Marks elements + annotated image (bypasses UIA entirely). Returns the OCR elements[] plus a cheap by-ref resource_link by default (no inline base64); pass confirmImage=true to ALSO embed the annotated bitmap.\n  'ocr'   — Windows OCR words with screen-pixel clickAt coords (Phase 4: absorbs former screenshot_ocr). Use when UIA returns no actionable elements (WinUI3 custom-drawn UIs, game overlays, PDF viewers). Note: detail='text' auto-falls back to OCR via ocrFallback='auto'; choose detail='ocr' only when forcing OCR unconditionally."
    • Addedscreenshot_gc
    • Addedscreenshot_query
  5. 16 tool updatesv1.10.4
    • Removedbrowser_click
    • Removedbrowser_eval
    • Removedbrowser_fill
    • Removedbrowser_form
    • Removedbrowser_locate
    • Removedbrowser_navigate
    • Removedbrowser_open
    • Removedbrowser_overview
    • Removedbrowser_search
    • Removedclick_element
    • Removedclipboard
    • Removeddesktop_state
    • Removedexcel
    • Removedfocus_window
    • Removedkeyboard
    • Removedmouse_click
  6. 1 tool updatev1.10.3
    • Changedscreenshot2 fields changed
      • removedInput schema / properties / ocrLanguage / default
        Removed value: -"ja"
      • changedInput schema / properties / ocrLanguage / description
        Previous value: -"BCP-47 language tag for the OCR engine (e.g. 'ja', 'en-US'). Used when detail='text' (OCR fallback) or detail='ocr' (direct OCR)."New value: +"BCP-47 language tag for the OCR engine (e.g. 'ja', 'en-US'). Auto-detects from system locale when omitted. Used when detail='text' (OCR fallback) or detail='ocr' (direct OCR)."
  7. 27 tool updatesv1.9.2
    • Addedbrowser_click
    • Addedbrowser_eval
    • Addedbrowser_fill
    • Addedbrowser_form
    • Addedbrowser_locate
    • Addedbrowser_navigate
    • Addedbrowser_open
    • Addedbrowser_overview
    • Addedbrowser_search
    • Addedclick_element
    • Addedclipboard
    • Addeddesktop_state
    • Addedexcel
    • Addedfocus_window
    • Addedkeyboard
    • Addedmouse_click
    • Addedmouse_drag
    • Addednotification_show
    • Addedrun_macro
    • Addedscreenshot
    • Addedscroll
    • Addedserver_status
    • Addedterminal
    • Addedwait_until
    • Addedwindow_dock
    • Addedworkspace_launch
    • Addedworkspace_snapshot
  8. 27 tool updatesv1.8.0
    • Removedbrowser_click
    • Removedbrowser_eval
    • Removedbrowser_fill
    • Removedbrowser_form
    • Removedbrowser_locate
    • Removedbrowser_navigate
    • Removedbrowser_open
    • Removedbrowser_overview
    • Removedbrowser_search
    • Removedclick_element
    • Removedclipboard
    • Removeddesktop_state
    • Removedexcel
    • Removedfocus_window
    • Removedkeyboard
    • Removedmouse_click
    • Removedmouse_drag
    • Removednotification_show
    • Removedrun_macro
    • Removedscreenshot
    • Removedscroll
    • Removedserver_status
    • Removedterminal
    • Removedwait_until
    • Removedwindow_dock
    • Removedworkspace_launch
    • Removedworkspace_snapshot
  9. 27 tool updatesv1.6.0
    • Addedbrowser_click
    • Addedbrowser_eval
    • Addedbrowser_fill
    • Addedbrowser_form
    • Addedbrowser_locate
    • Addedbrowser_navigate
    • Addedbrowser_open
    • Addedbrowser_overview
    • Addedbrowser_search
    • Addedclick_element
    • Addedclipboard
    • Addeddesktop_state
    • Addedexcel
    • Addedfocus_window
    • Addedkeyboard
    • Addedmouse_click
    • Addedmouse_drag
    • Addednotification_show
    • Addedrun_macro
    • Addedscreenshot
    • Addedscroll
    • Addedserver_status
    • Addedterminal
    • Addedwait_until
    • Addedwindow_dock
    • Addedworkspace_launch
    • Addedworkspace_snapshot

TDQS

A4.1/5.0

Scored across 30 tools

Disambiguation4/5

Each tool has a well-defined primary purpose with detailed 'Prefer' guidance, but there is some overlap between screenshot (text/som detail), desktop_state, and workspace_snapshot as observation tools, and the three click-related tools (browser_click, click_element, mouse_click) require careful reading of descriptions to distinguish. Overall, the boundaries are mostly clear.

Naming Consistency3/5

The browser_* and screenshot_* prefixes are consistent, but the overall convention is mixed: bare nouns (keyboard, terminal, scroll, excel, clipboard) coexist with verb_noun (focus_window, click_element, wait_until) and noun_verb (mouse_click, mouse_drag) patterns. This inconsistency makes predicting tool names by analogy difficult.

Tool Count3/5

30 tools is on the heavy side, but the server covers multiple subdomains (browser automation, native UI control, terminal interaction, clipboard, Excel macros, window management, screenshot caching), so each tool roughly earns its place. Still, the count exceeds the comfortable 3-15 range and approaches the limit where discoverability suffers.

Completeness3/5

The surface covers most core operations for desktop automation, but several tool descriptions reference helpers that are not present in this set (desktop_discover, desktop_act), which would cause dead ends when agents follow the documented 'Prefer' guidance. Basic browser workflows (navigation, click, fill, wait) are well-covered, and workarounds exist for missing operations like refresh (browser_eval).

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    GUI automation MCP server that enables AI agents to see and control the Windows desktop using a local Vision LLM (Ollama), supporting screenshot analysis, mouse/keyboard actions, and autonomous task execution.
    4
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A Windows computer use agent — FastMCP server that gives AI assistants hands on the real desktop: windows, UI elements, mouse, keyboard, screenshots, OCR, shortcuts, dialogs, and outcome verification.
    40
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Allow AI agents to see and control a real Windows PC you own: observe (UIA + screenshots), click/type/drag/scroll, launch apps, owner Live View. BYOH — your machine, your key.
    16
    Apache 2.0
  • F
    license
    A
    quality
    A
    maintenance
    Cross-platform desktop automation MCP server that lets AI agents capture screenshots, run OCR with UI-element classification, control mouse/keyboard, and launch programs on Linux, macOS, and Windows.
    20
    1
    -