Skip to main content
Glama

๐Ÿ–ฅ๏ธ Screen Control

A local remote-control system for AI agents: watch your computer's screen live and send mouse/keyboard commands to it. Everything runs on your own machine โ€” no data ever leaves it, no cloud middleman.


Why Screen Control?

AI agents today can write code and call APIs โ€” but they can't see or touch your desktop. Screen Control gives any agent general-purpose computer use over a clean, safety-gated HTTP/MCP interface:

  • Perceive โ€” OCR for text, single frames or a live MJPEG stream for vision-capable models, and a text-only diff endpoint for models that can't consume images at all.

  • Act โ€” absolute and relative mouse, Unicode-safe keyboard, window management, background (focus-free) control, virtual desktops.

  • Stay safe โ€” token auth, blocked deadly shortcuts, focus guard, a stuck-input watchdog and an emergency failsafe are all enforced server-side, no matter how confused the agent gets.

One process, zero configuration, works with any language that can speak HTTP โ€” or natively through MCP in Claude Desktop, Cursor, VS Code and cloud agents.

Performance Is Agent-Bound

Screen Control is the perception and actuation layer โ€” the eyes and hands. The effective speed and capability of any agent using it are bounded by that agent itself and by the environment it runs in:

  • Thinking speed โ€” one action per agent "turn": the perceive โ†’ plan โ†’ act โ†’ verify loop lives in the agent, so model inference latency and reasoning depth directly set the pace. The API itself adds only milliseconds per call.

  • Context capacity โ€” screen readings (OCR text, frames, diffs) consume the agent's context window; a larger window means more situational awareness before verification degrades.

  • Runtime environment โ€” network latency, MCP/HTTP round-trip overhead, tool-call limits and hosting constraints all stack on top of the loop.

In practice this means: the same repo makes a fast reasoning model fast and capable, and makes a slow model slow โ€” the toolchain is not the bottleneck. Real-time or action-heavy tasks need an agent with fast inference and tight tool-loop latency; slower agents should prefer deliberate, verification-heavy tasks.


Related MCP server: pov

Table of Contents


Features

Feature

Description

๐Ÿ–ผ๏ธ Live screen feed

Continuously refreshing screenshot in the browser

๐Ÿ–ฑ๏ธ Mouse control

Click, right-click, double-click, scroll, drag & drop via live screenshot

โŒจ๏ธ Keyboard control

Text typing (Unicode/Turkish included, layout-independent), keys and shortcuts (Ctrl+C, Alt+Tabโ€ฆ)

๐Ÿ‘๏ธ OCR

Converts on-screen text to machine-readable format

๐Ÿ“ท Vision access

Raw-pixel paths for image-capable models: single frames, MJPEG stream, text-based motion detection

๐ŸชŸ Window management

List, focus, safe close (WM_CLOSE), kill (task-manager style)

๐Ÿ–ฅ๏ธ Focus-free control

Read/write background windows via PostMessage without stealing focus

๐ŸŽฎ Game mode

Camera look via relative mouse movement, hold-to-move keys

๐Ÿ” Token auth

Every request requires X-Auth-Token (CSRF protection)

๐Ÿฆบ Stuck-input watchdog

Auto-releases held keys after 30 s of inactivity

๐Ÿ›Ÿ Failsafe

Cursor to top-left corner aborts all commands (disabled in game mode)


Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Browser (Web UI)                      โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚
โ”‚  โ”‚  Live     โ”‚  โ”‚  Control  โ”‚  โ”‚  Windows / Game Mode   โ”‚ โ”‚
โ”‚  โ”‚  View     โ”‚  โ”‚  Panel    โ”‚  โ”‚  Panel                 โ”‚ โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚
โ”‚       โ”‚              โ”‚                     โ”‚              โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
        โ”‚              โ”‚                     โ”‚
        โ–ผ              โ–ผ                     โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                  HTTP API (Flask)                        โ”‚
โ”‚                  127.0.0.1:8745                           โ”‚
โ”‚                                                         โ”‚
โ”‚  /api/screenshot    /api/mouse     /api/key              โ”‚
โ”‚  /api/vision/*      /api/ocr       /api/window           โ”‚
โ”‚  /api/game          /api/held      /api/release_all      โ”‚
โ”‚  /api/windows       /api/desktops  /api/desktop          โ”‚
โ”‚                                                         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
โ”‚  โ”‚ Auth Layer  โ”‚  โ”‚  Watchdog    โ”‚  โ”‚  OCR Engine  โ”‚   โ”‚
โ”‚  โ”‚ (token)     โ”‚  โ”‚  (30s auto)  โ”‚  โ”‚  (RapidOCR)  โ”‚   โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
        โ”‚              โ”‚                     โ”‚
        โ–ผ              โ–ผ                     โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                  control.py (Core)                        โ”‚
โ”‚                                                         โ”‚
โ”‚  Screen:  mss (fast capture), PIL (processing)          โ”‚
โ”‚  Mouse:   pyautogui (absolute), SendInput (relative)    โ”‚
โ”‚  Keyboard: pyautogui + SendInput+KEYEVENTF_UNICODE      โ”‚
โ”‚  Windows: Win32 API (EnumWindows, SetForegroundWindow)   โ”‚
โ”‚  Background: PrintWindow (capture), PostMessage (input)  โ”‚
โ”‚  Virtual Desktops: pyvda                                 โ”‚
โ”‚  Game Mode: ClipCursor + MOUSE_MOVE_RELATIVE             โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚              backends/ (pluggable)           โ”‚
โ”‚  WindowsBackend  โ”‚ LinuxBackend โ”‚ MacOSBackendโ”‚
โ”‚      (full)      โ”‚   (stub)     โ”‚   (stub)   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Coordinates & Concurrency

Per-Monitor DPI awareness. control.py calls SetProcessDpiAwarenessContext(PER_MONITOR_AWARE_V2) at import time โ€” before the pyautogui import, because pyautogui touches coordinate APIs during import and would otherwise lock the process to the interpreter manifest's default (system-aware). With PMv2 active, every coordinate in the system is a physical pixel end to end: mss capture, OCR bounding boxes, pyautogui/SendInput clicks, ClipCursor. On High-DPI displays (125%/150% scaling) nothing drifts between what OCR reports and where the mouse clicks.

Lock architecture. The server uses two independent locks instead of one global lock:

Lock

Protects

Endpoints

_input_lock

mouse, keyboard, game mode, window ops

/api/mouse, /api/key, /api/game, /api/window/post, ...

_read_lock

capture, OCR, vision, enumeration

/api/screenshot, /api/ocr, /api/vision/*, /api/windows, ...

A slow OCR (3โ€“5 s on a busy screen) no longer freezes concurrent screenshot or vision reads โ€” reads queue behind reads, inputs behind inputs.

Live-Loop Working Principle

This system is designed for a live perceive-act loop, not pre-written command chains:

  1. READ โ€” OCR or vision reads the screen before and after every action

  2. ONE ACTION โ€” each round sends a single command

  3. VERIFY โ€” acceptance is "it appeared on screen", not "I sent it"

  4. ADAPT โ€” if verification fails, the next step changes based on what is actually seen

This is enforced by the expect_hwnd guard: typing is refused (409) if the foreground window doesn't match the target.


Installation

cd screen-control
pip install -r requirements.txt

Requirements

Package

Purpose

Required?

mss

Fast screen capture

โœ… Yes

pyautogui

Mouse/keyboard control

โœ… Yes

pyvda

Virtual desktop management

โœ… Yes

flask

HTTP server

โœ… Yes

Pillow

Image processing

โœ… Yes

rapidocr-onnxruntime

OCR (screen text reading)

โš ๏ธ Optional

Note: The OCR package is large and may take a while to install. If it fails, everything else still works โ€” only the OCR feature is unavailable.

System Requirements

  • OS: Windows 10/11 (x64)

  • Python: 3.10+

  • Display: Any resolution; the system adapts automatically


Quick Start

# 1. Start the server
cd screen-control
python server.py

# 2. Open in browser
#    http://127.0.0.1:8745

# 3. Or control via API
TOKEN=$(cat .token)
curl -H "X-Auth-Token: $TOKEN" http://127.0.0.1:8745/api/screenshot -o screen.jpg

API Reference

Capabilities

Returns the active backend name and what it can do. Agents should call this first (see ROADMAP.md for the multi-platform plan).

GET /api/capabilities
โ†’ {"ok": true, "backend": "windows",
   "capabilities": {"screen_capture": true, "game_mode": true, ...}}

Feature values: true (supported), false (absent), null (unknown โ€” stub backend), "optional" (depends on an optional dependency).

Platform Support Matrix

Capability

Windows

Linux X11

Linux Wayland

macOS

Screen capture

Full

Full

Portal-dependent

Permission required

OCR

Full/optional

Full/optional

Full/optional

Full/optional

Mouse control

Full

Full

Restricted

Accessibility permission

Keyboard control

Full

Full

Restricted

Accessibility permission

Window enumeration

Full

WM-dependent

Limited

Accessibility/API-dependent

Background input

Strong

WM/app-dependent

Usually unavailable

Limited

Virtual desktops

Supported

DE/WM-dependent

DE/WM-dependent

Spaces-specific

Game mode

Supported

Experimental

Limited

Experimental

Linux and macOS backends are currently fail-closed stubs: every operation returns BACKEND_UNAVAILABLE (501) until implemented (ROADMAP Phases 5โ€“7). Windows is the reference backend.

Authentication

Every request must include the X-Auth-Token header. The token is generated on each server start and written to .token.

TOKEN=$(cat .token)

Code

Meaning

401

Missing or invalid token

415

POST without Content-Type: application/json

Token bootstrap (for the bundled web UI):

GET /token
โ†’ {"ok": true, "token": "abc123..."}

The /token endpoint is safe: Same-Origin Policy prevents foreign pages from reading it.


Screen Capture

GET /api/screenshot

Returns a JPEG screenshot.

Parameter

Type

Default

Description

monitor

int

1

Monitor index

region

string

โ€”

x,y,w,h sub-region

curl -H "X-Auth-Token: $TOKEN" -o screen.jpg http://127.0.0.1:8745/api/screenshot
curl -H "X-Auth-Token: $TOKEN" "http://127.0.0.1:8745/api/screenshot?region=0,0,800,600"

GET /api/info

Returns screen dimensions and system state.

{"ok": true, "width": 1920, "height": 1080, "ocr_available": true,
 "failsafe": true, "game_mode": false}

Mouse Control

POST /api/mouse

action

Required params

Optional params

Description

move

x, y

duration (default 0.15)

Move cursor to absolute position

click

x, y

button (left/right), clicks (default 1)

Click at position

scroll

clicks

x, y

Scroll wheel (positive=up)

drag

x1, y1, x, y

duration, button

Drag between two points

down

button (default "left")

โ€”

Press and hold mouse button

up

button (default "left")

โ€”

Release held mouse button

# Click at center of screen
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"action":"click","x":960,"y":540,"button":"left"}'

# Right-click
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"action":"click","button":"right","x":960,"y":540}'

# Scroll down
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"action":"scroll","clicks":-3}'

Keyboard Control

POST /api/key

action

Required params

Description

press

key

Press and release a key

down

key

Hold a key down (tracked for watchdog)

up

key

Release a held key

hotkey

keys (array)

Key combination (e.g. ["ctrl","c"])

type

text

Type text (Unicode, layout-independent)

Optional param

Default

Description

expect_hwnd

โ€”

Window handle to verify focus (409 if mismatch)

interval

0.03

Delay between characters for type

# Press Enter
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"press","key":"enter"}'

# Ctrl+C
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"hotkey","keys":["ctrl","c"]}'

# Type text (Turkish characters supported)
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"type","text":"Merhaba dรผnya"}'

# Hold W key down (for walking in games)
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"down","key":"w"}'

OCR (Screen Reading)

POST /api/ocr

Converts on-screen text to machine-readable format.

Param

Type

Default

Description

region

array

โ€”

[x, y, w, h] sub-region (faster)

{
  "ok": true,
  "text": "Hello World\nFile Edit View",
  "lines": ["Hello World", "File Edit View"],
  "items": [
    {"text": "Hello World", "x": 960, "y": 40},
    {"text": "File Edit View", "x": 100, "y": 15}
  ]
}
# Full screen OCR
curl -X POST http://127.0.0.1:8745/api/ocr -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{}'

# Region-only (faster, ~10x for small regions)
curl -X POST http://127.0.0.1:8745/api/ocr -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"region":[0,0,800,100]}'

Vision Access (Image Models)

Three endpoints for models that can consume images:

Endpoint

Description

GET /api/vision/frame

Single JPEG frame (raw or base64)

GET /api/stream

MJPEG live stream

POST /api/vision/diff

Text-based motion detection (no vision needed)

GET /api/vision/frame

Param

Default

Description

scale

1.0

Downscale factor (0.5 = half size)

gray

0

1 for greyscale

quality

80

JPEG quality (20-95)

format

โ€”

base64 for JSON response

region

โ€”

x,y,w,h sub-region

# Half-size greyscale frame as base64 (for text-only models)
curl "http://127.0.0.1:8745/api/vision/frame?scale=0.5&gray=1&format=base64" \
  -H "X-Auth-Token: $TOKEN"

GET /api/stream

MJPEG live stream. Drop into <img src> or consume frame-by-frame.

Param

Default

Description

fps

10

Frames per second (1-30)

quality

70

JPEG quality

scale

1.0

Downscale factor

region

โ€”

x,y,w,h sub-region

POST /api/vision/diff

Text-based motion detection โ€” no vision model required.

Body

Description

{}

Compare against last stored frame

{"grab":"gray"}

Store current frame for next comparison

{"b64_prev":"..."}

Compare against provided previous frame

{
  "ok": true,
  "changed": true,
  "changed_pct": 12.5,
  "bbox": [100, 200, 400, 350],
  "tiles": [
    {"row": 2, "col": 4, "pct": 35.2, "center": [1000, 390]}
  ]
}

Window Management

GET /api/windows

List all visible windows.

{
  "ok": true,
  "windows": [
    {
      "hwnd": 123456,
      "title": "My Application",
      "process": "app.exe",
      "pid": 7890,
      "focused": true,
      "rect": [0, 0, 1920, 1080],
      "desktop": 1
    }
  ]
}

POST /api/window

action

Required

Optional

Description

focus

hwnd

โ€”

Bring window to foreground

close

hwnd

expect_title, expect_process

Safe close via WM_CLOSE

kill

hwnd, pid

โ€”

Force kill (task-manager style)

topmost

hwnd

โ€”

Set always-on-top

untopmost

hwnd

โ€”

Remove always-on-top

maximize

hwnd

โ€”

Maximise window

# Focus a window
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"hwnd":12345,"action":"focus"}'

# Safe close (with title verification)
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"action":"close","hwnd":12345,"expect_title":"Notepad"}'

# Kill process
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"kill","hwnd":12345,"pid":7890}'

Focus-Free (Background) Control

Read and control windows without stealing focus โ€” the user keeps working on their main desktop.

GET /api/window/capture

Capture a window via PrintWindow (works even on another virtual desktop).

Param

Description

hwnd (required)

Window handle

client

1 = client area only

ocr

1 = return OCR text instead of image

# Capture window as PNG
curl "http://127.0.0.1:8745/api/window/capture?hwnd=12345" \
  -H "X-Auth-Token: $TOKEN" -o window.png

# Capture + OCR in one call
curl "http://127.0.0.1:8745/api/window/capture?hwnd=12345&ocr=1" \
  -H "X-Auth-Token: $TOKEN"

POST /api/window/post

Send input to a window, choosing the delivery path automatically.

action

Description

type

Type text (Unicode-safe)

key

Send a key press

hotkey

Send a key combination

click

Click at client coordinates

scroll

Scroll the window

drag

Drag inside the window

Optional mode parameter controls routing:

mode

Behavior

auto (default)

Decided by input-mode probe (see below)

background

Force PostMessage path (window keeps focus/z-order)

focused

Force focus + SendInput path

# Type into a background Notepad
curl -X POST http://127.0.0.1:8745/api/window/post -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"hwnd":12345,"action":"type","text":"Hello from background!"}'

Routing rules (mode=auto):

  • postmessage โ€” classic Win32 app: background PostMessage, no focus change.

  • uia โ€” WinUI/UWP/XAML surface (single DirectX canvas, no Win32 child controls): posted messages are silently swallowed, so the window is focused and the action is replayed through SendInput (client coords converted to screen). This is the documented fallback for modern apps.

  • focused โ€” window is already foreground: focused SendInput path.

  • invalid โ€” HTTP 409; not a reachable top-level window.

GET /api/window/input-mode

Classify how a window receives input before posting to it. Returns one of focused | postmessage | uia | invalid.

curl "http://127.0.0.1:8745/api/window/input-mode?hwnd=12345" \
  -H "X-Auth-Token: $TOKEN"

WinUI note: New Notepad (and other XAML-hosted apps) has no classic child Edit control to post to โ€” the whole UI is one DirectX surface. input-mode reports uia for these; /api/window/post then automatically uses the focused SendInput path. /api/window/children remains useful for classic apps with real child controls.

GET /api/window/children

List child controls of a window (class name + title + hwnd).

curl "http://127.0.0.1:8745/api/window/children?hwnd=12345" -H "X-Auth-Token: $TOKEN"

Virtual Desktops

GET /api/desktops

List all virtual desktops.

POST /api/desktop

action

Params

Description

switch

number

Switch to desktop N

create

โ€”

Create a new desktop

curl http://127.0.0.1:8745/api/desktops -H "X-Auth-Token: $TOKEN"

curl -X POST http://127.0.0.1:8745/api/desktop -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"switch","number":2}'

Game Mode

action

Params

Description

start

sensitivity (default 12)

Lock cursor to center, enable game input

move

dx, dy, sensitivity

Rotate camera (relative mouse)

stop

โ€”

Release cursor + all held input

heartbeat

โ€”

Keep-alive for long holds

# Start game mode
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"start","sensitivity":12}'

# Look right
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"move","dx":50,"dy":0}'

# Hold W to walk forward
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"down","key":"w"}'

# ... later ...
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"up","key":"w"}'

# Stop game mode
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"stop"}'

Safety Endpoints

GET /api/held

Returns currently held keys/buttons and watchdog status.

{
  "ok": true,
  "keys": ["w", "shift"],
  "buttons": ["left"],
  "game_mode": true,
  "idle_seconds": 5.2,
  "watchdog_count": 0,
  "last_watchdog": null
}

POST /api/release_all

Emergency: release everything (held keys, mouse buttons, game-mode cursor lock).

curl -X POST http://127.0.0.1:8745/api/release_all -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{}'


๐Ÿค– For AI Agents

A dedicated, comprehensive guide for AI agents (LLMs, vision models, automation frameworks) is available in AGENT_GUIDE.md.

It covers:

  • Perceive-act loop (read โ†’ plan โ†’ act โ†’ verify)

  • Focus guard (expect_hwnd) to prevent wrong-window accidents

  • App automation and game control workflows

  • Vision access for image-capable models

  • Text-based motion detection

  • Bandwidth optimization

  • Complete curl examples


MCP Support (One-Click Cloud Agents)

Model Context Protocol (MCP) turns this project into a plug-and-play toolbox for any MCP-capable agent: Claude Desktop, Claude Code, Cursor, VS Code Copilot Agent mode, custom cloud agents โ€” no custom glue code, no curl scripts. The agent discovers and calls the tools natively.

How it works

MCP agent (cloud or desktop)
        โ”‚  MCP protocol (stdio or streamable-HTTP)
        โ–ผ
  mcp_server.py   โ† thin wrapper: tools โ†’ HTTP calls, token auto-read
        โ”‚  REST + X-Auth-Token (localhost only)
        โ–ผ
  server.py       โ† the single source of truth:
                    auth, locks, watchdog, focus guard, all safety rules

mcp_server.py adds no new powers โ€” every safety mechanism (auth token, input/read locks, watchdog, Alt+F4 block, focus guard, failsafe) stays enforced by server.py.

Setup

pip install mcp            # optional dependency (see requirements.txt)
python server.py           # start the REST server first (it writes .token)

The MCP server auto-reads the token from .token (or the SCREEN_CONTROL_TOKEN env var) โ€” zero configuration.

Desktop agents (stdio transport)

Claude Desktop โ€” claude_desktop_config.json:

{
  "mcpServers": {
    "screen-control": {
      "command": "python",
      "args": ["C:/path/to/screen-control/mcp_server.py"]
    }
  }
}

Claude Code: claude mcp add screen-control -- python C:/path/to/screen-control/mcp_server.py

Cursor / VS Code: add the same entry to their MCP config files.

Remote / cloud agents (streamable-HTTP transport)

python mcp_server.py --http --port 8751
# MCP endpoint: http://127.0.0.1:8751/mcp

The HTTP transport is token-protected: every request must carry the X-Auth-Token header (same token as the REST server). Query-string tokens (?token=...) are rejected by design โ€” URLs leak into proxy/tunnel logs, browser history and shared links, and this token grants full desktop control. Clients that cannot send custom headers should run a local stdio mcp_server.py instead. Only GET /health is open, for liveness probes. DNS-rebinding protection is disabled on this transport deliberately โ€” tunneled requests arrive with a foreign Host header, and the rebinding threat is already covered by the token guard.

For a cloud agent, expose it through a tunnel:

cloudflared tunnel --url http://127.0.0.1:8751
# โ†’ prints a https://<random>.trycloudflare.com URL

Then configure the agent's MCP connection with <tunnel-url>/mcp plus the token from .token as a header (X-Auth-Token).

Headerless connectors (scoped keys)

For clients that cannot send custom headers (e.g. web connectors that only take an endpoint URL), create a scoped API key โ€” a persistent, optionally time-limited credential โ€” and embed it in the URL path:

# Create a 24-hour scoped key (requires the server to be running)
curl -X POST http://127.0.0.1:8745/api/keys -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"action":"create","name":"spark","expires_in_hours":24}'

# Connector endpoint becomes:
#   https://<tunnel-url>/mcp/<scoped-key>

Design guarantees (SC-06):

  • The master session token is refused in URLs (403) โ€” only scoped keys may travel there

  • Scoped keys expire automatically; expired keys authenticate nothing

  • Scoped keys are revocable by name at any moment via POST /api/keys ({"action":"revoke","name":"spark"}) โ€” revocation takes effect immediately on every endpoint

  • The .apikeys file stores only SHA-256 hashes, never raw keys

โš ๏ธ A tunnel exposes PC control to the internet. Keep the token secret, prefer short-lived tunnels and scoped keys for headerless connectors, and stop the server when not in use.

One-Command Startup (launcher + auto-tunnel)

start-server.bat automates the whole cloud setup and prints everything your cloud agent needs, ready to paste:

  1. Downloads cloudflared.exe if missing (portable, no admin required)

  2. Stops leftover instances from a previous run

  3. Starts the REST server (port 8745) and the MCP HTTP server (port 8751)

  4. Waits until both are healthy (/token and /health probes)

  5. Starts a cloudflared quick tunnel, extracts its public URL from tunnel.log, and prints the summary:

 ============================================================
  ALL SYSTEMS RUNNING
 ============================================================
  Local REST API  : http://127.0.0.1:8745
  Local MCP       : http://127.0.0.1:8751/mcp
  Public MCP URL  : https://<random>.trycloudflare.com/mcp

  ------------------------------------------------------------
  PASTE INTO YOUR CLOUD AGENT  (MCP connector settings)
  ------------------------------------------------------------
  Endpoint : https://<random>.trycloudflare.com/mcp
  Header   : X-Auth-Token: <token>
  URL form : https://<random>.trycloudflare.com/mcp?token=<token>
             (only if the connector cannot send headers)
  ------------------------------------------------------------

stop-server.bat stops all three (REST, MCP, tunnel) in one go.

Available tools (16)

Category

Tools

Perception

get_info, ocr_screen, screenshot (real image block for vision models), motion_diff

Mouse / keyboard

mouse, keyboard (with expect_hwnd), get_held, release_all

Windows

list_windows, focus_window, window_children, window_input_mode, window_post, window_capture_ocr, close_window

Game mode

game (start / move / stop / heartbeat)

Which transport for whom

Consumer

Transport

Command

Claude Desktop / Cursor / VS Code (local)

stdio

python mcp_server.py

Claude Code

stdio

claude mcp add ... (above)

Cloud / remote agents

streamable-HTTP

start-server.bat (recommended) or python mcp_server.py --http --port 8751 + cloudflared tunnel --url http://127.0.0.1:8751

Note: This project targets MCP Python SDK 2.x (MCPServer API). With SDK 1.x, replace the import with from mcp.server.fastmcp import FastMCP, Image and MCPServer with FastMCP.


Security Model

Threat: Malicious Web Pages (CSRF)

Even bound to 127.0.0.1, a malicious page in the browser can trigger non-preflighted requests (text/plain fetch, HTML form POST) to localhost. The browser blocks the response but not the request โ€” the server would still execute the command.

Mitigation: Every request requires X-Auth-Token. A foreign page cannot read this token (Same-Origin Policy), so it cannot authenticate.

Additional layers:

  • POST requests must use Content-Type: application/json (415 otherwise)

  • This blocks form-encoded and text-plain POSTs even if the token leaked

  • Host header trust (DNS rebinding): when bound to loopback, requests carrying a non-loopback Host header are refused with 421 โ€” a rebinding page that resolves its domain to 127.0.0.1 cannot read /token or call the API

  • /token and / responses carry Cache-Control: no-store so the credential is never persisted by browsers or proxies

Threat: Stuck Keys / Game Mode Lock

In game mode, ClipCursor pins the cursor to a 2ร—2 box โ€” the classic pyautogui failsafe (cursor to top-left) does not work.

Mitigations:

  1. Physical Esc / Alt+Tab โ€” real hardware input; this API cannot block it, and it always works

  2. POST /api/release_all โ€” instant release of everything

  3. Watchdog (automatic) โ€” 30 s of server-side inactivity with held input triggers automatic release

Threat: Wrong Window Typing

Mitigations:

  • expect_hwnd guard on /api/key โ€” if the foreground window doesn't match, typing is refused with 409

  • The focused path of /api/window/post verifies the focus after the focus switch and before any synthetic input (409 on mismatch) โ€” input is never replayed into whatever window happens to be foreground

  • focus_window() raises on failure instead of silently returning

Threat: Dangerous Key Combos

Mitigation: Blocked at the API level (403) on every delivery path โ€” the direct /api/key route, the background /api/window/post route (PostMessage), and the focused fallback route share one safety policy (control._assert_allowed):

  • Alt+F4 โ€” the only banned Alt combo (Alt+Tab, Alt+menu are legitimate)

  • Win key โ€” prevents Start menu, task switching

  • Ctrl+Alt+Del โ€” system security screen

  • Shift+Delete style โ€” prevents permanent deletion

Threat: Killing System Processes

Mitigations:

  • The process name is resolved from the PID directly (Win32 toolhelp snapshot), not from the visible-window inventory โ€” windowless/background system processes get the same protection as visible ones

  • Critical system processes are blacklisted (default-deny for unknown PIDs): winlogon.exe, csrss.exe, smss.exe, services.exe, lsass.exe, svchost.exe, system, registry, dwm.exe

  • Optional expect_process confirmation: a mismatch aborts the kill with 409 โ€” protects against killing a newly-reused PID

Threat: Resource Exhaustion (rogue agent / DoS)

A token-holding but misbehaving client should not be able to exhaust memory or starve the input lock.

Mitigations:

  • MAX_CONTENT_LENGTH = 1 MB โ€” oversized request bodies are rejected (413)

  • region width/height/area and scale are bounded (400 otherwise)

  • text payloads are capped at 10,000 characters per input call

  • MJPEG streams are capped at 10 concurrent clients (429 beyond that)

  • Pillow decompression-bomb limit is set for client-supplied images

Network Access

The server binds to 127.0.0.1 by default. To expose it to the network:

python server.py --host 0.0.0.0  # โš ๏ธ anyone on the network can control this machine

Game Mode Guide

Setup

# 1. Focus the game window
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"hwnd":GAME_HWND,"action":"focus"}'

# 2. Start game mode
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"start","sensitivity":12}'

Camera Look

# Look right
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"move","dx":50,"dy":0}'

# Look down
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"move","dx":0,"dy":30}'

Movement

# Walk forward (hold W)
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"down","key":"w"}'

# ... walk for a while ...

# Release W
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"up","key":"w"}'

Minecraft-Specific

# Place block (right-click)
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"action":"click","button":"right","x":960,"y":540}'

# Break block (hold left-click)
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"down","button":"left"}'

# ... after breaking ...

curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"up","button":"left"}'

# Select hotbar slot
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"press","key":"1"}'

# Open inventory
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
  -H "Content-Type: application/json" -d '{"action":"press","key":"e"}'

Suitability

Game Type

Suitable?

Notes

Minecraft (building)

โœ… Yes

Place blocks, walk, mine

Minecraft (PvP)

โŒ No

Too slow for fast combat

Turn-based games

โœ… Yes

Ample time for readโ†’actโ†’verify

RPG / adventure

โœ… Yes

Inventory, dialogue, exploration

Fast FPS

โŒ No

Reaction time insufficient

Puzzle games

โœ… Yes

Click-based, read-heavy


Vision Access Guide

For Image-Capable Models

If the consuming model can process images, use the vision endpoints directly:

GET /api/vision/frame?scale=0.5&gray=1&quality=70

This returns a single JPEG that the model can analyze for:

  • Game HUD elements (health, mana, inventory)

  • On-screen text (menus, chat, tooltips)

  • Visual scene understanding (blocks, entities, terrain)

For Text-Only Models

Use the diff endpoint for motion detection without vision:

POST /api/vision/diff {"grab":"gray"}   โ†’ first call: stores frame
POST /api/vision/diff                    โ†’ subsequent calls: returns diff

The response tells you where things changed (tile coordinates) and how much (percentage), which is sufficient for:

  • Detecting that an action had an effect

  • Locating moving elements on screen

  • Tracking animation state changes

Bandwidth Optimization

Approach

Payload

Use Case

scale=1.0, gray=0

~500 KB

Full detail

scale=0.5, gray=1

~50 KB

Good for most vision models

scale=0.25, gray=1

~10 KB

Maximum compression

diff (text)

~1 KB

Text-only agents

region=...

Variable

Focus on specific area


Troubleshooting

"OCR engine not installed"

pip install rapidocr-onnxruntime

Server won't start (port in use)

# Find the process using port 8745
netstat -ano | findstr ":8745"

# Kill it
taskkill /PID <pid> /F

"Focus mismatch" (409) when typing

The foreground window changed between the focus call and the type call. Solution: always pass expect_hwnd and verify focus before typing.

Window not found

The window may have been closed or may be a system window that EnumWindows doesn't expose. Try:

curl http://127.0.0.1:8745/api/windows -H "X-Auth-Token: $TOKEN"

Game mode cursor stuck

Use POST /api/release_all or press Esc / Alt+Tab physically.

High OCR latency

OCR on a full 1920ร—1080 screen can take from a few seconds up to ~30 s depending on your CPU and on-screen complexity. Use a region โ€” small crops are typically 10ร— faster:

{"region": [0, 0, 800, 100]}

Turkish characters not appearing

The system uses SendInput + KEYEVENTF_UNICODE which is layout-independent. If characters still don't appear, the target app may not support Unicode input โ€” try POST /api/window/post with action: "type" instead.


Project Structure

screen-control/
โ”œโ”€โ”€ server.py            # Flask HTTP server + all API endpoints
โ”œโ”€โ”€ core/                # PlatformBackend interface + standardized errors
โ”‚   โ”œโ”€โ”€ backends.py      # Abstract backend + lazy discovery
โ”‚   โ””โ”€โ”€ errors.py        # ApiError envelope + error codes
โ”œโ”€โ”€ backends/            # OS implementations behind PlatformBackend
โ”‚   โ”œโ”€โ”€ windows.py       # Reference backend (moved from control.py)
โ”‚   โ”œโ”€โ”€ linux.py         # Fail-closed stub (ROADMAP Phase 5)
โ”‚   โ”œโ”€โ”€ macos.py         # Fail-closed stub (ROADMAP Phase 7)
โ”‚   โ”œโ”€โ”€ fake.py          # In-memory backend for tests
โ”‚   โ””โ”€โ”€ forbidden.py     # Shared blocked-key policy
โ”œโ”€โ”€ control.py           # Compatibility shim re-exporting backends.windows
โ”œโ”€โ”€ tests/unit/          # Offline unit + integration tests (no real input)
โ”œโ”€โ”€ mcp_server.py        # MCP server (stdio + streamable-HTTP) โ€” thin wrapper over the API
โ”œโ”€โ”€ sdk/
โ”‚   โ””โ”€โ”€ screen_control.py  # Python SDK client (pip-installable style)
โ”œโ”€โ”€ index.html           # Bundled web UI (live view + control panels)
โ”œโ”€โ”€ docs/images/         # README assets (demo GIF captured by the API itself)
โ”œโ”€โ”€ requirements.txt     # Python dependencies
โ”œโ”€โ”€ start-server.bat     # One command: REST + MCP + cloud tunnel (Windows)
โ”œโ”€โ”€ stop-server.bat      # Stop all three processes
โ”œโ”€โ”€ test-security.py     # Security + game-mode test suite (34 checks)
โ”œโ”€โ”€ test-game.py         # Live game-mechanics test (app launch โ†’ draw โ†’ safe close)
โ”œโ”€โ”€ test-endtoend.py     # End-to-end test: open Notepad โ†’ type โ†’ save โ†’ verify
โ”œโ”€โ”€ .github/workflows/   # CI: runs the security suite on every push
โ”œโ”€โ”€ AGENT_GUIDE.md       # AI agent integration guide (separate from this file)
โ”œโ”€โ”€ README.md            # This file
โ””โ”€โ”€ .token               # Auto-generated auth token (gitignored)

Testing

Prerequisites

The server must be running:

cd screen-control
python server.py

Security test suite

Tests authentication, blocked key combos, window management, safe close, critical process protection, and game mode โ€” all non-destructive.

cd screen-control
python test-security.py

Expected output:

== Token Authentication ==
โœ“ Missing token -> 401
โœ“ Wrong token -> 401
โœ“ Correct token -> 200
โœ“ Non-JSON POST -> 415

== Blocked Key Combos ==
โœ“ Alt+F4 blocked (403)
โœ“ Win key blocked (403)
โœ“ Win+D blocked (403)
โœ“ Delete blocked (403)

== Window List ==
โœ“ Windows list requires GET
โœ“ Window list is non-empty  โ€” 8 windows
โœ“ Exactly one focused window

== Safe Close Verification ==
โœ“ Wrong title aborts close

== Critical Process Protection ==
โœ“ System process (pid 4) rejected (403)
โœ“ pid 0 rejected (403)

== Game Mode ==
โœ“ Game mode started
โœ“ Relative camera look
โœ“ Game mode stopped

== Watchdog (dry run) ==
โœ“ Held state returns ok
โœ“ Watchdog count reported

========================================
RESULT: 19 passed, 0 failed

Live game-mechanics test

Launches a real application (mspaint or notepad), performs hold-to-draw game mechanics, verifies via pixel analysis, then safely closes with "Don't Save" dialog handling.

cd screen-control
python test-game.py

Note: This test launches a real application. It handles cleanup automatically (sends WM_CLOSE and clicks "Don't Save" if a dialog appears).


Contributing

  1. Fork the repository

  2. Create a feature branch

  3. Make your changes

  4. Test on a Windows machine

  5. Submit a pull request

Code Style

  • Python: PEP 8, type hints, docstrings on all public functions

  • Docstrings: English, Google style

  • Error messages: English, descriptive

  • Comments: English, explain why not what


License

MIT License. See LICENSE for details.


Built with โค๏ธ for local automation and AI agent research.

Available Tools

16 tools
close_windowA

Close a window safely (WM_CLOSE after verification โ€” never blind Alt+F4). Pass expect_title/expect_process to refuse (409) if the window changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
hwndYes
expect_titleNo
expect_processNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does a good job: it discloses the safe close mechanism, the verification step, and the 409 refusal condition. It could further clarify what 'verification' checks exactly and what success/failure responses look like, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded, and every sentence earns its place: the first sentence states the action and mechanism, the second explains the safety parameters. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter tool with an output schema, the description is nearly complete: it covers the action, the safe method, and the guard parameters. Minor gaps include not stating what happens when no expect_* params are passed or how success is indicated, but these are acceptable given the low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the purpose of expect_title and expect_process (safety guards that trigger a 409 on window change), adding real meaning beyond the schema. hwnd is not described, but its role as a window handle is reasonably inferable from the tool name and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Close a window safely') and provides the exact mechanism (WM_CLOSE after verification, never blind Alt+F4), which clearly distinguishes it from sibling tools like window_post or keyboard-based closing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives actionable usage guidance: pass expect_title/expect_process to refuse if the window changed, and explicitly warns against blind Alt+F4. It does not name alternative tools or state when not to use this tool, but the context is clear enough for an agent to recognize when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

focus_windowB

Bring a window to the foreground (focus it).

ParametersJSON Schema
NameRequiredDescriptionDefault
hwndYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the intended effect without mentioning potential side effects, failure conditions (e.g., if the window is minimized or unresponsive), or whether it requires special permissions. This leaves the agent without knowledge of how the operation might behave beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that immediately states the purpose with no redundant words. It is appropriately concise for a simple operation and front-loads the key action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has a required parameter (hwnd) that is undocumented, and no behavioral details are provided, the description is incomplete. The existence of an output schema (per context) is noted, but the description does not clarify what the tool returns (e.g., success/failure), leaving gaps in the agent's understanding of how to call and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'hwnd' has no description in the schema (coverage 0%), and the tool description does not explain it at all. An agent cannot infer that 'hwnd' is a window handle or how to obtain it, making parameter usage opaque. The description fails to compensate for the schema's lack of parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Bring a window to the foreground (focus it)' clearly states the action (bring/focus) on a specific resource (a window), distinguishing it from siblings like close_window or window_children. It is unambiguous and immediately tells an agent what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when you want to focus a window, but it provides no explicit guidance on when to prefer it over alternatives or any exclusions. There is no mention of conditions like 'use only when the window is not already focused' or when to avoid it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gameB

Game mode for FPS-style control. action: start (clip cursor to center) | move (relative camera look by dx,dy pixels * sensitivity) | stop | heartbeat (keep-alive during long holds). Combine with keyboard() down/up for WASD and mouse(action='down'/'up') for shooting/building.

ParametersJSON Schema
NameRequiredDescriptionDefault
dxNo
dyNo
actionYes
sensitivityNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses that 'start' clips cursor to center, 'move' does relative camera look by dx,dy pixels * sensitivity, and 'heartbeat' is a keep-alive during long holds. However, it doesn't disclose side effects like cursor visibility changes, whether stop restores the cursor, or what happens on errors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the tool's purpose. The action list is dense but efficient. The pipe-separated format is a bit terse but saves space. It earns its place by covering purpose, actions, and usage in two sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema (though not shown in detail) and 4 parameters. The description covers the core actions and how to combine with keyboard/mouse, but lacks details like what the output looks like, whether sensitivity applies only to move, and edge cases like what happens if start is called twice. For a game-control tool, this is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains dx/dy as relative camera look pixels and sensitivity as a multiplier, but it doesn't explain the 'action' parameter's allowed values beyond listing them in the description (start, move, stop, heartbeat) โ€” which is helpful but not mapped to the schema. The 'action' parameter has no enum in the schema, so the description's list is the only guidance, but it's not structured as parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a game mode for FPS-style control and enumerates the specific actions (start, move, stop, heartbeat). It distinguishes itself from siblings like mouse and keyboard by framing it as a mode that combines with them, though it doesn't explicitly name a sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool: for FPS-style control, and it instructs combining with keyboard() down/up for WASD and mouse(action='down'/'up') for shooting/building. It implies this is the alternative to direct mouse/keyboard for game contexts, but doesn't explicitly state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_heldA

List currently held keys/buttons and watchdog status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It states the tool lists items, implying a non-destructive read operation, but offers no details on side effects, error behavior, or what 'watchdog status' entails. For a simple getter, this is minimally transparent but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, direct sentence with no padding. It efficiently conveys the core function without irrelevant detail, achieving maximum conciseness for the tool's simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has zero parameters and an output schema (assumed to document return structure), the description is sufficient for an agent to call it. The mention of 'watchdog status' adds context beyond the name, though it could elaborate on what that status means. Overall, nothing critical is missing for a straightforward query tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema already fully describes the interface. The description adds no parameter semantics because none exist; a baseline of 4 is appropriate when there are no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('List') and a specific resource ('currently held keys/buttons and watchdog status'). It distinguishes this from sibling tools like 'release_all' and 'keyboard' by implying a read-only state query, though it doesn't explicitly name alternatives. The purpose is clear and not tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus siblings. It doesn't mention scenarios where 'get_held' is preferred, nor does it state any exclusions. The tool name itself is suggestive, but the description leaves all usage context to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_infoA

Server/PC status: screen size, monitors, OCR availability, failsafe and game-mode state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses what data is reported, which is the core behavioral trait for a status query, but it does not explicitly state that the tool is non-destructive or that it does not alter state beyond reading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence that front-loads the purpose ('Server/PC status:') and then packs the returned field list compactly. Every word earns its place; there is no filler or repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateless zero-parameter query tool with a provided output schema, the description fully covers what the agent needs to know. The schema handles return-value structure, and the description enumerates the reported categories, leaving nothing missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters, the baseline is 4. There is nothing for the description to clarify about arguments; instead, it properly explains what the tool's output means, which is the only relevant semantic information given no schema to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific implicit verb (retrieve/query) over a concrete resource (server/PC status) and enumerates the exact data fields returned: screen size, monitors, OCR availability, failsafe, and game-mode state. This clearly distinguishes it from siblings like screenshot and ocr_screen, which capture images or extract text, and from get_held, which reports keyboard state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The listed fields (OCR availability, failsafe, game-mode) imply a pre-flight environment-status use case, so usage is plausibly inferred. However, the description never explicitly says when to call this tool versus alternatives, nor states any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

keyboardA

Control the keyboard. action: press | hotkey | type | down | up. press: key name (enter, f11, a...); hotkey: keys=["ctrl","s"]; type: Unicode-safe text. expect_hwnd refuses (409) the action if the foreground window differs โ€” use it to avoid typing into the wrong app.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNo
keysNo
textNo
actionNopress
intervalNo
expect_hwndNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a key behavioral trait: expect_hwnd causes a 409 refusal if the foreground window differs, and notes that type is Unicode-safe. Since no annotations are provided, the description carries the full burden, but it omits other behaviors such as whether actions are synchronous, error handling beyond 409, or any permission requirements. It covers some but not all behavioral aspects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, two sentences long, and front-loads the core purpose and actions. It includes practical examples and a caution note without any fluff, making it highly readable and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 6 parameters and no annotations, the description covers the main actions and key parameters but omits details about the interval parameter and the return value (though an output schema exists). For a simple keyboard tool, this may be sufficient for basic use, but it leaves some aspects undocumented, making it not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates well by explaining the action values and giving examples for key (press), keys (hotkey), and text (type), and clarifies the purpose of expect_hwnd. It does not explain interval or the exact format of key names beyond the examples, but it provides enough to use the primary parameters effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it controls the keyboard and enumerates the actions (press, hotkey, type, down, up) with concrete examples for each, making its purpose unmistakable and distinguishing it from mouse and other input tools. It is a specific verb+resource definition without any tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a usage guideline for the expect_hwnd parameter, advising to use it to avoid typing into the wrong app. However, it does not explicitly discuss when to use this tool versus sibling tools like mouse or get_held, or any exclusions. The guidance is present but not comprehensive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_windowsA

List visible top-level windows: hwnd, title, process name, pid.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the burden of explaining side effects and filtering. It clearly states the tool lists only visible top-level windows (no hidden or child windows), and the word 'List' implies a non-mutating read operation. It does not mention ordering or enumeration caveats, but these are minor for a zero-parameter read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence conveys the action, scope, and output fields with no filler. The key scoping terms ('visible', 'top-level') appear up front.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that no parameters are required and an output schema exists, the description supplies the essential identifying fields and the central filter. An agent can call this tool correctly without needing additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so there is nothing to document beyond the schema's empty properties object. The description adds value by specifying what the returned rows contain (hwnd, title, process name, pid), which the schema cannot show.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('List') and resource ('visible top-level windows'), then enumerates the exact fields returned. It distinguishes itself from sibling tools like window_children by narrowing scope to top-level windows and from screenshot by being an enumeration rather than a capture.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to use this tool instead of related window tools such as window_children, focus_window, or close_window. There are no exclusions, prerequisites, or alternative references, leaving the agent to infer the intended use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

motion_diffA

Text-only change detection between two snapshots: which tiles of the screen changed and their clickable centers. Use to verify an action had an effect or to spot movement without OCR/vision.

ParametersJSON Schema
NameRequiredDescriptionDefault
regionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It mentions 'text-only' and 'without OCR/vision' but does not disclose how snapshots are obtained (e.g., whether it captures them internally), potential side effects, or performance implications. This ambiguity leaves the agent uncertain about the tool's execution model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise, front-loaded sentences pack the core purpose and usage into minimal space. No redundancy; every clause adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the what and when but leaves a critical gap: the mechanism of 'two snapshots' is vague. Since the tool only accepts an optional region, it's unclear whether the agent must supply snapshots or if the tool manages them internally. The output schema exists, so return format is covered, but this ambiguity undermines completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0% and the description does not mention the 'region' parameter at all. The parameter is left completely unexplained, forcing the agent to guess whether it restricts detection to a screen area or serves another purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (detect), the resource (screen changes between two snapshots), and the output (changed tiles and clickable centers). It explicitly contrasts with OCR/vision, distinguishing it from sibling tools like ocr_screen.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete use cases ('verify an action had an effect' or 'spot movement') and explicitly frames it as an alternative to OCR/vision. It doesn't name specific sibling tools, but the context is sufficient for an agent to decide when to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mouseB

Control the mouse. action: move | click | scroll | drag | down | up. click uses x,y (optional = current pos); drag goes (x1,y1)->(x,y); scroll uses 'clicks' (negative = down). Physical-pixel coordinates.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
x1No
y1No
actionNomove
buttonNoleft
clicksNo
durationNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses meaningful traits: physical-pixel coordinates, optional current-position click, and negative clicks for downward scroll. However, it does not mention side effects (e.g., moving the actual system cursor, affecting other applications), permissions, or reversibility. Some behavioral info is present, but gaps remain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only three short sentences, front-loading the core purpose and then compactly enumerating action semantics. Every phrase earns its place, using terse notation that is easy to scan. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 8 parameters, no annotations, and an output schema that likely documents return values. The description covers the most important parameter semantics but omits button and duration, and does not explain the effect of each action beyond coordinate mechanics. It is adequate for basic use but not complete for all parameters and possible edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the meaning of action values, x/y as click coordinates (optional), x1/y1 for drag start, and clicks for scroll amount. It does not describe the button or duration parameters, leaving parts of the parameter space undocumented. Still, it adds substantial meaning beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Control the mouse' and lists the supported actions (move, click, scroll, drag, down, up), giving a clear verb+resource statement. It doesn't explicitly distinguish from sibling tools, but the action list makes it self-evident compared to keyboard, screenshot, or window controls.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use the mouse tool versus alternatives like keyboard or screenshot. It explains how to perform specific actions (e.g., 'click uses x,y', 'scroll uses clicks'), but not when this tool should be preferred or avoided. There is no mention of context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ocr_screenA

Read the screen as text with coordinates (OCR). Best for menus, dialogs, editors, chat โ€” anything text-heavy. Returns lines plus per-item centers for clicking. 'region' is an optional "x,y,w,h" crop.

ParametersJSON Schema
NameRequiredDescriptionDefault
regionNo
scan_focusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It usefully discloses the return shape (lines plus per-item centers for clicking) and the optional crop, and 'Read' implies a non-destructive operation. But it leaves scan_focus behavior unexplained and does not mention permission requirements or performance characteristics (OCR latency), which are material for an agent deciding whether to call it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero waste: purpose is front-loaded, usage guidance follows, and behavior plus parameter semantics are packed into one efficient sentence. Every clause adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value detail is already covered, and the description's mention of 'lines plus per-item centers' aligns with it. The main gaps are the unexplained scan_focus parameter and lack of explicit differentiation from window_capture_ocr among a large sibling list โ€” both are things an agent needs to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully explains 'region' (format 'x,y,w,h' and optional crop semantics), but scan_focus is entirely undocumented in both the schema and the description โ€” the agent cannot know what focusing behavior it toggles. With half the parameters unexplained, a 2 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb+resource ('Read the screen as text with coordinates (OCR)') and differentiates from siblings: 'best for menus, dialogs, editors, chat โ€” anything text-heavy' distinguishes it from screenshot (visual capture) and the sibling window_capture_ocr (window-scoped OCR). An agent can tell what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear when-to-use context ('text-heavy' interfaces such as menus, dialogs, editors, chat). However, it does not name the alternative window_capture_ocr or state when to prefer that sibling over whole-screen OCR, so exclusions are only implied, not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

release_allA

Release every held key and button (emergency reset). Safe to call anytime.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of disclosing behavior. It does state that the operation is an emergency reset and safe to call anytime, which is useful context, but it does not explain potential side effects, whether it affects the OS-level key state or only the tool's internal state, or what the response will be.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short, front-loaded sentences with no filler. The core action and emergency-reset intent come first, and the safety reassurance is appended efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema, the description is complete: it names the action, the scope, the emergency-reset intent, and the safety profile. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema is empty and there are zero parameters, so the description does not need to explain parameter semantics. The phrase 'every held key and button' also conveys the scope of the operation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Release') and a precise resource ('every held key and button'), and adds the clarifying parenthetical 'emergency reset'. This makes the tool's purpose unmistakable and distinguishes it from siblings like get_held, which queries held state rather than clearing it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'emergency reset' implies the intended use case, and 'Safe to call anytime' provides clear, unconditional context for when it can be used. However, it does not explicitly name alternatives or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotA

Capture the screen as an image (for vision-capable models). 'region' = "x,y,w,h"; 'scale' 0-1 shrinks to save bandwidth; 'quality' 1-100.

ParametersJSON Schema
NameRequiredDescriptionDefault
scaleNo
regionNo
monitorNo
qualityNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It does explain parameter behavior (scale shrinks to save bandwidth, quality range, region format), which is useful. However, it does not explicitly state that the operation is read-only, nor does it disclose output format, permission requirements, or multi-monitor behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler. The main purpose is front-loaded, and the parameter hints are compact and actionable. Nothing extraneous is included.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and 0% schema description coverage, the description should handle all context. It misses the meaning of the 'monitor' parameter and fails to state the output format (e.g., base64, URL, path), making it incomplete for an agent to call correctly in multi-monitor or output-sensitive scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clearly defines region as 'x,y,w,h', scale as 0-1, and quality as 1-100, adding meaning to three of four parameters. It entirely omits the 'monitor' parameter, leaving it unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Capture the screen as an image'. 'For vision-capable models' further clarifies the intended use case, and the plain-image capture distinguishes it from OCR siblings like ocr_screen and window_capture_ocr.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for vision-capable models' implies a use case, but there is no explicit guidance on when to use this tool versus alternatives like ocr_screen or window_capture_ocr. No exclusions or when-not-to-use conditions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

window_capture_ocrA

Capture a background window WITHOUT focusing it; with ocr=1 (default) returns its text directly โ€” read other apps' content while the user works.

ParametersJSON Schema
NameRequiredDescriptionDefault
ocrNo
hwndYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a key behavioral trait: it does NOT focus the window, which is a significant side-effect avoidance. It also explains the default ocr behavior. With no annotations provided, this is valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the most important behavior (no focus), and the ocr default is explained inline. Zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are covered. The description explains the core behavior and default. It could mention what happens when ocr=0, but the output schema likely covers that. Overall adequate for a 2-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the ocr parameter's default and effect (returns text directly), but doesn't explain hwnd beyond the schema's integer type. The description adds some meaning but not full parameter coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool captures a background window without focusing it, and with ocr=1 (default) returns its text directly. This distinguishes it from siblings like screenshot, ocr_screen, and focus_window.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when you need to read another app's content without disrupting the user's work. It doesn't explicitly name alternatives or exclusions, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

window_childrenA

List child controls of a window (class name, title, hwnd). Useful for classic Win32 apps with real child controls (e.g. an Edit box) โ€” target the child hwnd in window_post. Modern WinUI/UWP apps have no classic children; use window_input_mode to detect them.

ParametersJSON Schema
NameRequiredDescriptionDefault
hwndYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden. It clearly signals a non-mutating list operation and adds a useful platform caveat about WinUI/UWP apps having no classic children. It does not discuss edge cases like direct-versus-recursive enumeration, but for a read-only inspection tool the core behavior is sufficiently disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence states the core action and output fields; the second adds usage context and an alternative. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description is complete enough. It explains the target scenario, the practical follow-up action (window_post), and the main exception (modern apps). No critical calling information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description only mentions hwnd incidentally as the child handle to pass to window_post. It does not explicitly state that the input hwnd is the parent window handle or where it should come from, though the tool name makes this mostly inferable. The description partially compensates for the missing schema detail but could be more explicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and a clear resource ('child controls of a window'), and it names the returned attributes (class name, title, hwnd). It also distinguishes the tool from siblings by framing it as the Win32 child-enumeration tool versus window_input_mode for modern apps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when this tool is useful ('classic Win32 apps with real child controls'), what to do with the results ('target the child hwnd in window_post'), and which alternative to use for modern apps ('use window_input_mode to detect them'). This is clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

window_input_modeA

Classify how a window receives input BEFORE posting to it: postmessage (classic Win32) | uia (WinUI/UWP: focus+SendInput fallback) | focused (already foreground) | invalid (dead window).

ParametersJSON Schema
NameRequiredDescriptionDefault
hwndYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the output categories, including the 'invalid (dead window)' case, and implies a read-only operation ('Classify') with no side effects. However, it does not explicitly state that it performs no mutation or that it requires a valid handle, though the 'invalid' output suggests handling of dead windows. This is adequate but not rich in behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured line with line breaks separating the classification options. It contains zero filler and front-loads the purpose and timing. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, an output schema presumably describing the classification results), the description covers the essential context: when to use, what it returns, and edge cases (invalid). It could mention that it is read-only or requires a valid handle, but these are minor gaps for a classification tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It implies that 'hwnd' refers to the window being classified, but does not explicitly define the parameter or its format. Since the tool name and description make the role of hwnd obvious, a 3 is appropriate, but explicit clarification would improve it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Classify') with a clear resource ('how a window receives input') and explicitly lists the four possible classification values. This makes the tool's purpose unambiguous and distinguishes it from siblings like window_post (which posts input) and focus_window (which focuses).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'BEFORE posting to it' provides clear temporal context, indicating this tool should be called before window_post. The classification values also imply decision logic, but no explicit exclusions or alternatives are mentioned. Still, the guidance is sufficiently clear for an agent to know when to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

window_postB

Send input to a window in the background. action: type | key | hotkey | click | scroll. mode: auto (recommended โ€” routes WinUI/UWP apps through the focused SendInput path automatically) | background | focused. Note: 'uia' routing focuses the window (unavoidable for modern apps).

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
keyNo
hwndYes
keysNo
modeNoauto
textNo
actionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does disclose that 'uia' routing focuses the window, which is a key side effect. It also implies background operation, but it does not explain what 'background' means in terms of window visibility or input capture, nor does it describe side effects for actions like click or scroll. The behavior is partially transparent but lacks detail on edge cases or requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact paragraph that front-loads the core purpose and then lists the key enums. It uses a clear 'action: ...' and 'mode: ...' format, and the note about uia routing is placed at the end as an important caveat. The structure is logical and efficient, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 8 parameters, 2 required, and an output schema exists, but the description is insufficient for an agent to use it correctly. It omits parameter meanings, prerequisites (e.g., whether the window must be visible or minimized), and concrete examples. The note about uia routing is helpful but not enough. While the output schema may describe return values, the input side is under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate for the 8 parameters. It explains the 'action' and 'mode' enums, but leaves x, y, key, keys, text, and hwnd entirely unexplained. For instance, x and y are likely coordinates for click/scroll, and key/keys/text are input payloads, but the description provides no such meaning. This is a significant gap that forces the agent to guess or inspect the schema further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Send input to a window in the background.' It lists the supported action types (type, key, hotkey, click, scroll), giving a concrete sense of what the tool does. However, it does not explicitly differentiate itself from sibling tools like mouse, keyboard, or focus_window, so an agent must infer that this is specifically for targeting a window handle (hwnd) rather than global input.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides some usage guidance by explaining mode options and recommending 'auto' for WinUI/UWP apps, and notes that 'uia' routing forces focus. However, it does not explicitly state when to use this tool versus alternatives like mouse or keyboard, nor does it clarify the trade-offs between background and focused modes beyond the note. The guidance is partial, leaving the agent to infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 16 tool updatesv0.1.0
    • First observedclose_window
    • First observedfocus_window
    • First observedgame
    • First observedget_held
    • First observedget_info
    • First observedkeyboard
    • First observedlist_windows
    • First observedmotion_diff
    • First observedmouse
    • First observedocr_screen
    • First observedrelease_all
    • First observedscreenshot
    • First observedwindow_capture_ocr
    • First observedwindow_children
    • First observedwindow_input_mode
    • First observedwindow_post

TDQS

A3.7/5.0

Scored across 16 tools

Disambiguation5/5

Each tool targets a distinct mechanismโ€”screen capture, OCR, motion detection, global input, window-specific input, window introspection, and game mode. Potentially similar tools like ocr_screen and window_capture_ocr are cleanly separated by foreground screen versus background window. No two tools appear to do the same job.

Naming Consistency4/5

Most tools follow a readable snake_case convention with verb-first names like list_windows, focus_window, close_window, and release_all. A few noun-style names like mouse, keyboard, and game, plus window_post and window_capture_ocr, are minor deviations but still clear and predictable.

Tool Count4/5

Sixteen tools is on the higher end but each earns its place across capture, OCR, input, window management, and game mode. The count feels slightly heavy but remains well-scoped for a desktop automation server.

Completeness4/5

Core workflows are well covered: screen capture, OCR, change detection, global and background input, window discovery, focusing, posting, and closing. Minor gaps like querying the current cursor position or window geometry are absent but most operations have workarounds.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers