Screen Control
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Screen Controltake a screenshot and summarize what's on my screen"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ฅ๏ธ Screen Control
A local remote-control system for AI agents: watch your computer's screen live and send mouse/keyboard commands to it. Everything runs on your own machine โ no data ever leaves it, no cloud middleman.
Why Screen Control?
AI agents today can write code and call APIs โ but they can't see or touch your desktop. Screen Control gives any agent general-purpose computer use over a clean, safety-gated HTTP/MCP interface:
Perceive โ OCR for text, single frames or a live MJPEG stream for vision-capable models, and a text-only diff endpoint for models that can't consume images at all.
Act โ absolute and relative mouse, Unicode-safe keyboard, window management, background (focus-free) control, virtual desktops.
Stay safe โ token auth, blocked deadly shortcuts, focus guard, a stuck-input watchdog and an emergency failsafe are all enforced server-side, no matter how confused the agent gets.
One process, zero configuration, works with any language that can speak HTTP โ or natively through MCP in Claude Desktop, Cursor, VS Code and cloud agents.
Performance Is Agent-Bound
Screen Control is the perception and actuation layer โ the eyes and hands. The effective speed and capability of any agent using it are bounded by that agent itself and by the environment it runs in:
Thinking speed โ one action per agent "turn": the perceive โ plan โ act โ verify loop lives in the agent, so model inference latency and reasoning depth directly set the pace. The API itself adds only milliseconds per call.
Context capacity โ screen readings (OCR text, frames, diffs) consume the agent's context window; a larger window means more situational awareness before verification degrades.
Runtime environment โ network latency, MCP/HTTP round-trip overhead, tool-call limits and hosting constraints all stack on top of the loop.
In practice this means: the same repo makes a fast reasoning model fast and capable, and makes a slow model slow โ the toolchain is not the bottleneck. Real-time or action-heavy tasks need an agent with fast inference and tight tool-loop latency; slower agents should prefer deliberate, verification-heavy tasks.
Related MCP server: pov
Table of Contents
Features
Feature | Description |
๐ผ๏ธ Live screen feed | Continuously refreshing screenshot in the browser |
๐ฑ๏ธ Mouse control | Click, right-click, double-click, scroll, drag & drop via live screenshot |
โจ๏ธ Keyboard control | Text typing (Unicode/Turkish included, layout-independent), keys and shortcuts (Ctrl+C, Alt+Tabโฆ) |
๐๏ธ OCR | Converts on-screen text to machine-readable format |
๐ท Vision access | Raw-pixel paths for image-capable models: single frames, MJPEG stream, text-based motion detection |
๐ช Window management | List, focus, safe close (WM_CLOSE), kill (task-manager style) |
๐ฅ๏ธ Focus-free control | Read/write background windows via PostMessage without stealing focus |
๐ฎ Game mode | Camera look via relative mouse movement, hold-to-move keys |
๐ Token auth | Every request requires |
๐ฆบ Stuck-input watchdog | Auto-releases held keys after 30 s of inactivity |
๐ Failsafe | Cursor to top-left corner aborts all commands (disabled in game mode) |
Architecture
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Browser (Web UI) โ
โ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ Live โ โ Control โ โ Windows / Game Mode โ โ
โ โ View โ โ Panel โ โ Panel โ โ
โ โโโโโโฌโโโโโโ โโโโโโฌโโโโโโ โโโโโโโโโโโโโฌโโโโโโโโโโโโโ โ
โ โ โ โ โ
โโโโโโโโโผโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโ
โ โ โ
โผ โผ โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ HTTP API (Flask) โ
โ 127.0.0.1:8745 โ
โ โ
โ /api/screenshot /api/mouse /api/key โ
โ /api/vision/* /api/ocr /api/window โ
โ /api/game /api/held /api/release_all โ
โ /api/windows /api/desktops /api/desktop โ
โ โ
โ โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โ
โ โ Auth Layer โ โ Watchdog โ โ OCR Engine โ โ
โ โ (token) โ โ (30s auto) โ โ (RapidOCR) โ โ
โ โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ โ
โผ โผ โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ control.py (Core) โ
โ โ
โ Screen: mss (fast capture), PIL (processing) โ
โ Mouse: pyautogui (absolute), SendInput (relative) โ
โ Keyboard: pyautogui + SendInput+KEYEVENTF_UNICODE โ
โ Windows: Win32 API (EnumWindows, SetForegroundWindow) โ
โ Background: PrintWindow (capture), PostMessage (input) โ
โ Virtual Desktops: pyvda โ
โ Game Mode: ClipCursor + MOUSE_MOVE_RELATIVE โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ backends/ (pluggable) โ
โ WindowsBackend โ LinuxBackend โ MacOSBackendโ
โ (full) โ (stub) โ (stub) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโCoordinates & Concurrency
Per-Monitor DPI awareness. control.py calls
SetProcessDpiAwarenessContext(PER_MONITOR_AWARE_V2) at import time โ
before the pyautogui import, because pyautogui touches coordinate APIs
during import and would otherwise lock the process to the interpreter
manifest's default (system-aware). With PMv2 active, every coordinate in the
system is a physical pixel end to end: mss capture, OCR bounding boxes,
pyautogui/SendInput clicks, ClipCursor. On High-DPI displays (125%/150%
scaling) nothing drifts between what OCR reports and where the mouse clicks.
Lock architecture. The server uses two independent locks instead of one global lock:
Lock | Protects | Endpoints |
| mouse, keyboard, game mode, window ops |
|
| capture, OCR, vision, enumeration |
|
A slow OCR (3โ5 s on a busy screen) no longer freezes concurrent screenshot or vision reads โ reads queue behind reads, inputs behind inputs.
Live-Loop Working Principle
This system is designed for a live perceive-act loop, not pre-written command chains:
READ โ OCR or vision reads the screen before and after every action
ONE ACTION โ each round sends a single command
VERIFY โ acceptance is "it appeared on screen", not "I sent it"
ADAPT โ if verification fails, the next step changes based on what is actually seen
This is enforced by the expect_hwnd guard: typing is refused (409) if
the foreground window doesn't match the target.
Installation
cd screen-control
pip install -r requirements.txtRequirements
Package | Purpose | Required? |
| Fast screen capture | โ Yes |
| Mouse/keyboard control | โ Yes |
| Virtual desktop management | โ Yes |
| HTTP server | โ Yes |
| Image processing | โ Yes |
| OCR (screen text reading) | โ ๏ธ Optional |
Note: The OCR package is large and may take a while to install. If it fails, everything else still works โ only the OCR feature is unavailable.
System Requirements
OS: Windows 10/11 (x64)
Python: 3.10+
Display: Any resolution; the system adapts automatically
Quick Start
# 1. Start the server
cd screen-control
python server.py
# 2. Open in browser
# http://127.0.0.1:8745
# 3. Or control via API
TOKEN=$(cat .token)
curl -H "X-Auth-Token: $TOKEN" http://127.0.0.1:8745/api/screenshot -o screen.jpgAPI Reference
Capabilities
Returns the active backend name and what it can do. Agents should call this first (see ROADMAP.md for the multi-platform plan).
GET /api/capabilities
โ {"ok": true, "backend": "windows",
"capabilities": {"screen_capture": true, "game_mode": true, ...}}Feature values: true (supported), false (absent), null (unknown โ
stub backend), "optional" (depends on an optional dependency).
Platform Support Matrix
Capability | Windows | Linux X11 | Linux Wayland | macOS |
Screen capture | Full | Full | Portal-dependent | Permission required |
OCR | Full/optional | Full/optional | Full/optional | Full/optional |
Mouse control | Full | Full | Restricted | Accessibility permission |
Keyboard control | Full | Full | Restricted | Accessibility permission |
Window enumeration | Full | WM-dependent | Limited | Accessibility/API-dependent |
Background input | Strong | WM/app-dependent | Usually unavailable | Limited |
Virtual desktops | Supported | DE/WM-dependent | DE/WM-dependent | Spaces-specific |
Game mode | Supported | Experimental | Limited | Experimental |
Linux and macOS backends are currently fail-closed stubs: every operation returns
BACKEND_UNAVAILABLE(501) until implemented (ROADMAP Phases 5โ7). Windows is the reference backend.
Authentication
Every request must include the X-Auth-Token header. The token is
generated on each server start and written to .token.
TOKEN=$(cat .token)Code | Meaning |
401 | Missing or invalid token |
415 | POST without |
Token bootstrap (for the bundled web UI):
GET /token
โ {"ok": true, "token": "abc123..."}The
/tokenendpoint is safe: Same-Origin Policy prevents foreign pages from reading it.
Screen Capture
GET /api/screenshot
Returns a JPEG screenshot.
Parameter | Type | Default | Description |
| int | 1 | Monitor index |
| string | โ |
|
curl -H "X-Auth-Token: $TOKEN" -o screen.jpg http://127.0.0.1:8745/api/screenshot
curl -H "X-Auth-Token: $TOKEN" "http://127.0.0.1:8745/api/screenshot?region=0,0,800,600"GET /api/info
Returns screen dimensions and system state.
{"ok": true, "width": 1920, "height": 1080, "ocr_available": true,
"failsafe": true, "game_mode": false}Mouse Control
POST /api/mouse
action | Required params | Optional params | Description |
|
|
| Move cursor to absolute position |
|
|
| Click at position |
|
|
| Scroll wheel (positive=up) |
|
|
| Drag between two points |
|
| โ | Press and hold mouse button |
|
| โ | Release held mouse button |
# Click at center of screen
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" \
-d '{"action":"click","x":960,"y":540,"button":"left"}'
# Right-click
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" \
-d '{"action":"click","button":"right","x":960,"y":540}'
# Scroll down
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" \
-d '{"action":"scroll","clicks":-3}'Keyboard Control
POST /api/key
action | Required params | Description |
|
| Press and release a key |
|
| Hold a key down (tracked for watchdog) |
|
| Release a held key |
|
| Key combination (e.g. |
|
| Type text (Unicode, layout-independent) |
Optional param | Default | Description |
| โ | Window handle to verify focus (409 if mismatch) |
| 0.03 | Delay between characters for |
# Press Enter
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"press","key":"enter"}'
# Ctrl+C
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"hotkey","keys":["ctrl","c"]}'
# Type text (Turkish characters supported)
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"type","text":"Merhaba dรผnya"}'
# Hold W key down (for walking in games)
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"down","key":"w"}'OCR (Screen Reading)
POST /api/ocr
Converts on-screen text to machine-readable format.
Param | Type | Default | Description |
| array | โ |
|
{
"ok": true,
"text": "Hello World\nFile Edit View",
"lines": ["Hello World", "File Edit View"],
"items": [
{"text": "Hello World", "x": 960, "y": 40},
{"text": "File Edit View", "x": 100, "y": 15}
]
}# Full screen OCR
curl -X POST http://127.0.0.1:8745/api/ocr -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{}'
# Region-only (faster, ~10x for small regions)
curl -X POST http://127.0.0.1:8745/api/ocr -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"region":[0,0,800,100]}'Vision Access (Image Models)
Three endpoints for models that can consume images:
Endpoint | Description |
| Single JPEG frame (raw or base64) |
| MJPEG live stream |
| Text-based motion detection (no vision needed) |
GET /api/vision/frame
Param | Default | Description |
| 1.0 | Downscale factor (0.5 = half size) |
| 0 | 1 for greyscale |
| 80 | JPEG quality (20-95) |
| โ |
|
| โ |
|
# Half-size greyscale frame as base64 (for text-only models)
curl "http://127.0.0.1:8745/api/vision/frame?scale=0.5&gray=1&format=base64" \
-H "X-Auth-Token: $TOKEN"GET /api/stream
MJPEG live stream. Drop into <img src> or consume frame-by-frame.
Param | Default | Description |
| 10 | Frames per second (1-30) |
| 70 | JPEG quality |
| 1.0 | Downscale factor |
| โ |
|
POST /api/vision/diff
Text-based motion detection โ no vision model required.
Body | Description |
| Compare against last stored frame |
| Store current frame for next comparison |
| Compare against provided previous frame |
{
"ok": true,
"changed": true,
"changed_pct": 12.5,
"bbox": [100, 200, 400, 350],
"tiles": [
{"row": 2, "col": 4, "pct": 35.2, "center": [1000, 390]}
]
}Window Management
GET /api/windows
List all visible windows.
{
"ok": true,
"windows": [
{
"hwnd": 123456,
"title": "My Application",
"process": "app.exe",
"pid": 7890,
"focused": true,
"rect": [0, 0, 1920, 1080],
"desktop": 1
}
]
}POST /api/window
action | Required | Optional | Description |
|
| โ | Bring window to foreground |
|
|
| Safe close via WM_CLOSE |
|
| โ | Force kill (task-manager style) |
|
| โ | Set always-on-top |
|
| โ | Remove always-on-top |
|
| โ | Maximise window |
# Focus a window
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"hwnd":12345,"action":"focus"}'
# Safe close (with title verification)
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" \
-d '{"action":"close","hwnd":12345,"expect_title":"Notepad"}'
# Kill process
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"kill","hwnd":12345,"pid":7890}'Focus-Free (Background) Control
Read and control windows without stealing focus โ the user keeps working on their main desktop.
GET /api/window/capture
Capture a window via PrintWindow (works even on another virtual desktop).
Param | Description |
| Window handle |
| 1 = client area only |
| 1 = return OCR text instead of image |
# Capture window as PNG
curl "http://127.0.0.1:8745/api/window/capture?hwnd=12345" \
-H "X-Auth-Token: $TOKEN" -o window.png
# Capture + OCR in one call
curl "http://127.0.0.1:8745/api/window/capture?hwnd=12345&ocr=1" \
-H "X-Auth-Token: $TOKEN"POST /api/window/post
Send input to a window, choosing the delivery path automatically.
action | Description |
| Type text (Unicode-safe) |
| Send a key press |
| Send a key combination |
| Click at client coordinates |
| Scroll the window |
| Drag inside the window |
Optional mode parameter controls routing:
mode | Behavior |
| Decided by input-mode probe (see below) |
| Force PostMessage path (window keeps focus/z-order) |
| Force focus + SendInput path |
# Type into a background Notepad
curl -X POST http://127.0.0.1:8745/api/window/post -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" \
-d '{"hwnd":12345,"action":"type","text":"Hello from background!"}'Routing rules (mode=auto):
postmessageโ classic Win32 app: background PostMessage, no focus change.uiaโ WinUI/UWP/XAML surface (single DirectX canvas, no Win32 child controls): posted messages are silently swallowed, so the window is focused and the action is replayed through SendInput (client coords converted to screen). This is the documented fallback for modern apps.focusedโ window is already foreground: focused SendInput path.invalidโ HTTP 409; not a reachable top-level window.
GET /api/window/input-mode
Classify how a window receives input before posting to it. Returns one of
focused | postmessage | uia | invalid.
curl "http://127.0.0.1:8745/api/window/input-mode?hwnd=12345" \
-H "X-Auth-Token: $TOKEN"WinUI note: New Notepad (and other XAML-hosted apps) has no classic child Edit control to post to โ the whole UI is one DirectX surface.
input-modereportsuiafor these;/api/window/postthen automatically uses the focused SendInput path./api/window/childrenremains useful for classic apps with real child controls.
GET /api/window/children
List child controls of a window (class name + title + hwnd).
curl "http://127.0.0.1:8745/api/window/children?hwnd=12345" -H "X-Auth-Token: $TOKEN"Virtual Desktops
GET /api/desktops
List all virtual desktops.
POST /api/desktop
action | Params | Description |
|
| Switch to desktop N |
| โ | Create a new desktop |
curl http://127.0.0.1:8745/api/desktops -H "X-Auth-Token: $TOKEN"
curl -X POST http://127.0.0.1:8745/api/desktop -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"switch","number":2}'Game Mode
action | Params | Description |
|
| Lock cursor to center, enable game input |
|
| Rotate camera (relative mouse) |
| โ | Release cursor + all held input |
| โ | Keep-alive for long holds |
# Start game mode
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"start","sensitivity":12}'
# Look right
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"move","dx":50,"dy":0}'
# Hold W to walk forward
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"down","key":"w"}'
# ... later ...
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"up","key":"w"}'
# Stop game mode
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"stop"}'Safety Endpoints
GET /api/held
Returns currently held keys/buttons and watchdog status.
{
"ok": true,
"keys": ["w", "shift"],
"buttons": ["left"],
"game_mode": true,
"idle_seconds": 5.2,
"watchdog_count": 0,
"last_watchdog": null
}POST /api/release_all
Emergency: release everything (held keys, mouse buttons, game-mode cursor lock).
curl -X POST http://127.0.0.1:8745/api/release_all -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{}'๐ค For AI Agents
A dedicated, comprehensive guide for AI agents (LLMs, vision models, automation frameworks) is available in AGENT_GUIDE.md.
It covers:
Perceive-act loop (read โ plan โ act โ verify)
Focus guard (
expect_hwnd) to prevent wrong-window accidentsApp automation and game control workflows
Vision access for image-capable models
Text-based motion detection
Bandwidth optimization
Complete curl examples
MCP Support (One-Click Cloud Agents)
Model Context Protocol (MCP) turns this project into a plug-and-play toolbox for any MCP-capable agent: Claude Desktop, Claude Code, Cursor, VS Code Copilot Agent mode, custom cloud agents โ no custom glue code, no curl scripts. The agent discovers and calls the tools natively.
How it works
MCP agent (cloud or desktop)
โ MCP protocol (stdio or streamable-HTTP)
โผ
mcp_server.py โ thin wrapper: tools โ HTTP calls, token auto-read
โ REST + X-Auth-Token (localhost only)
โผ
server.py โ the single source of truth:
auth, locks, watchdog, focus guard, all safety rulesmcp_server.py adds no new powers โ every safety mechanism
(auth token, input/read locks, watchdog, Alt+F4 block, focus guard,
failsafe) stays enforced by server.py.
Setup
pip install mcp # optional dependency (see requirements.txt)
python server.py # start the REST server first (it writes .token)The MCP server auto-reads the token from .token (or the
SCREEN_CONTROL_TOKEN env var) โ zero configuration.
Desktop agents (stdio transport)
Claude Desktop โ claude_desktop_config.json:
{
"mcpServers": {
"screen-control": {
"command": "python",
"args": ["C:/path/to/screen-control/mcp_server.py"]
}
}
}Claude Code: claude mcp add screen-control -- python C:/path/to/screen-control/mcp_server.py
Cursor / VS Code: add the same entry to their MCP config files.
Remote / cloud agents (streamable-HTTP transport)
python mcp_server.py --http --port 8751
# MCP endpoint: http://127.0.0.1:8751/mcpThe HTTP transport is token-protected: every request must carry the
X-Auth-Token header (same token as the REST server). Query-string tokens
(?token=...) are rejected by design โ URLs leak into proxy/tunnel
logs, browser history and shared links, and this token grants full desktop
control. Clients that cannot send custom headers should run a local stdio
mcp_server.py instead. Only GET /health is open, for liveness probes.
DNS-rebinding protection is disabled on this transport deliberately โ
tunneled requests arrive with a foreign Host header, and the rebinding
threat is already covered by the token guard.
For a cloud agent, expose it through a tunnel:
cloudflared tunnel --url http://127.0.0.1:8751
# โ prints a https://<random>.trycloudflare.com URLThen configure the agent's MCP connection with <tunnel-url>/mcp plus the
token from .token as a header (X-Auth-Token).
Headerless connectors (scoped keys)
For clients that cannot send custom headers (e.g. web connectors that only take an endpoint URL), create a scoped API key โ a persistent, optionally time-limited credential โ and embed it in the URL path:
# Create a 24-hour scoped key (requires the server to be running)
curl -X POST http://127.0.0.1:8745/api/keys -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" \
-d '{"action":"create","name":"spark","expires_in_hours":24}'
# Connector endpoint becomes:
# https://<tunnel-url>/mcp/<scoped-key>Design guarantees (SC-06):
The master session token is refused in URLs (403) โ only scoped keys may travel there
Scoped keys expire automatically; expired keys authenticate nothing
Scoped keys are revocable by name at any moment via
POST /api/keys({"action":"revoke","name":"spark"}) โ revocation takes effect immediately on every endpointThe
.apikeysfile stores only SHA-256 hashes, never raw keys
โ ๏ธ A tunnel exposes PC control to the internet. Keep the token secret, prefer short-lived tunnels and scoped keys for headerless connectors, and stop the server when not in use.
One-Command Startup (launcher + auto-tunnel)
start-server.bat automates the whole cloud setup and prints everything
your cloud agent needs, ready to paste:
Downloads
cloudflared.exeif missing (portable, no admin required)Stops leftover instances from a previous run
Starts the REST server (port 8745) and the MCP HTTP server (port 8751)
Waits until both are healthy (
/tokenand/healthprobes)Starts a cloudflared quick tunnel, extracts its public URL from
tunnel.log, and prints the summary:
============================================================
ALL SYSTEMS RUNNING
============================================================
Local REST API : http://127.0.0.1:8745
Local MCP : http://127.0.0.1:8751/mcp
Public MCP URL : https://<random>.trycloudflare.com/mcp
------------------------------------------------------------
PASTE INTO YOUR CLOUD AGENT (MCP connector settings)
------------------------------------------------------------
Endpoint : https://<random>.trycloudflare.com/mcp
Header : X-Auth-Token: <token>
URL form : https://<random>.trycloudflare.com/mcp?token=<token>
(only if the connector cannot send headers)
------------------------------------------------------------stop-server.bat stops all three (REST, MCP, tunnel) in one go.
Available tools (16)
Category | Tools |
Perception |
|
Mouse / keyboard |
|
Windows |
|
Game mode |
|
Which transport for whom
Consumer | Transport | Command |
Claude Desktop / Cursor / VS Code (local) | stdio |
|
Claude Code | stdio |
|
Cloud / remote agents | streamable-HTTP |
|
Note: This project targets MCP Python SDK 2.x (
MCPServerAPI). With SDK 1.x, replace the import withfrom mcp.server.fastmcp import FastMCP, ImageandMCPServerwithFastMCP.
Security Model
Threat: Malicious Web Pages (CSRF)
Even bound to 127.0.0.1, a malicious page in the browser can trigger
non-preflighted requests (text/plain fetch, HTML form POST) to localhost.
The browser blocks the response but not the request โ the server
would still execute the command.
Mitigation: Every request requires X-Auth-Token. A foreign page
cannot read this token (Same-Origin Policy), so it cannot authenticate.
Additional layers:
POST requests must use
Content-Type: application/json(415 otherwise)This blocks form-encoded and text-plain POSTs even if the token leaked
Host header trust (DNS rebinding): when bound to loopback, requests carrying a non-loopback
Hostheader are refused with421โ a rebinding page that resolves its domain to127.0.0.1cannot read/tokenor call the API/tokenand/responses carryCache-Control: no-storeso the credential is never persisted by browsers or proxies
Threat: Stuck Keys / Game Mode Lock
In game mode, ClipCursor pins the cursor to a 2ร2 box โ the classic
pyautogui failsafe (cursor to top-left) does not work.
Mitigations:
Physical
Esc/Alt+Tabโ real hardware input; this API cannot block it, and it always worksPOST /api/release_allโ instant release of everythingWatchdog (automatic) โ 30 s of server-side inactivity with held input triggers automatic release
Threat: Wrong Window Typing
Mitigations:
expect_hwndguard on/api/keyโ if the foreground window doesn't match, typing is refused with 409The focused path of
/api/window/postverifies the focus after the focus switch and before any synthetic input (409 on mismatch) โ input is never replayed into whatever window happens to be foregroundfocus_window()raises on failure instead of silently returning
Threat: Dangerous Key Combos
Mitigation: Blocked at the API level (403) on every delivery path โ
the direct /api/key route, the background /api/window/post route
(PostMessage), and the focused fallback route share one safety policy
(control._assert_allowed):
Alt+F4โ the only banned Alt combo (Alt+Tab, Alt+menu are legitimate)Win key โ prevents Start menu, task switching
Ctrl+Alt+Delโ system security screenShift+Deletestyle โ prevents permanent deletion
Threat: Killing System Processes
Mitigations:
The process name is resolved from the PID directly (Win32 toolhelp snapshot), not from the visible-window inventory โ windowless/background system processes get the same protection as visible ones
Critical system processes are blacklisted (default-deny for unknown PIDs):
winlogon.exe,csrss.exe,smss.exe,services.exe,lsass.exe,svchost.exe,system,registry,dwm.exeOptional
expect_processconfirmation: a mismatch aborts the kill with409โ protects against killing a newly-reused PID
Threat: Resource Exhaustion (rogue agent / DoS)
A token-holding but misbehaving client should not be able to exhaust memory or starve the input lock.
Mitigations:
MAX_CONTENT_LENGTH= 1 MB โ oversized request bodies are rejected (413)regionwidth/height/area andscaleare bounded (400 otherwise)textpayloads are capped at 10,000 characters per input callMJPEG streams are capped at 10 concurrent clients (429 beyond that)
Pillow decompression-bomb limit is set for client-supplied images
Network Access
The server binds to 127.0.0.1 by default. To expose it to the network:
python server.py --host 0.0.0.0 # โ ๏ธ anyone on the network can control this machineGame Mode Guide
Setup
# 1. Focus the game window
curl -X POST http://127.0.0.1:8745/api/window -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"hwnd":GAME_HWND,"action":"focus"}'
# 2. Start game mode
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"start","sensitivity":12}'Camera Look
# Look right
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"move","dx":50,"dy":0}'
# Look down
curl -X POST http://127.0.0.1:8745/api/game -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"move","dx":0,"dy":30}'Movement
# Walk forward (hold W)
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"down","key":"w"}'
# ... walk for a while ...
# Release W
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"up","key":"w"}'Minecraft-Specific
# Place block (right-click)
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" \
-d '{"action":"click","button":"right","x":960,"y":540}'
# Break block (hold left-click)
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"down","button":"left"}'
# ... after breaking ...
curl -X POST http://127.0.0.1:8745/api/mouse -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"up","button":"left"}'
# Select hotbar slot
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"press","key":"1"}'
# Open inventory
curl -X POST http://127.0.0.1:8745/api/key -H "X-Auth-Token: $TOKEN" \
-H "Content-Type: application/json" -d '{"action":"press","key":"e"}'Suitability
Game Type | Suitable? | Notes |
Minecraft (building) | โ Yes | Place blocks, walk, mine |
Minecraft (PvP) | โ No | Too slow for fast combat |
Turn-based games | โ Yes | Ample time for readโactโverify |
RPG / adventure | โ Yes | Inventory, dialogue, exploration |
Fast FPS | โ No | Reaction time insufficient |
Puzzle games | โ Yes | Click-based, read-heavy |
Vision Access Guide
For Image-Capable Models
If the consuming model can process images, use the vision endpoints directly:
GET /api/vision/frame?scale=0.5&gray=1&quality=70This returns a single JPEG that the model can analyze for:
Game HUD elements (health, mana, inventory)
On-screen text (menus, chat, tooltips)
Visual scene understanding (blocks, entities, terrain)
For Text-Only Models
Use the diff endpoint for motion detection without vision:
POST /api/vision/diff {"grab":"gray"} โ first call: stores frame
POST /api/vision/diff โ subsequent calls: returns diffThe response tells you where things changed (tile coordinates) and how much (percentage), which is sufficient for:
Detecting that an action had an effect
Locating moving elements on screen
Tracking animation state changes
Bandwidth Optimization
Approach | Payload | Use Case |
| ~500 KB | Full detail |
| ~50 KB | Good for most vision models |
| ~10 KB | Maximum compression |
| ~1 KB | Text-only agents |
| Variable | Focus on specific area |
Troubleshooting
"OCR engine not installed"
pip install rapidocr-onnxruntimeServer won't start (port in use)
# Find the process using port 8745
netstat -ano | findstr ":8745"
# Kill it
taskkill /PID <pid> /F"Focus mismatch" (409) when typing
The foreground window changed between the focus call and the type call.
Solution: always pass expect_hwnd and verify focus before typing.
Window not found
The window may have been closed or may be a system window that
EnumWindows doesn't expose. Try:
curl http://127.0.0.1:8745/api/windows -H "X-Auth-Token: $TOKEN"Game mode cursor stuck
Use POST /api/release_all or press Esc / Alt+Tab physically.
High OCR latency
OCR on a full 1920ร1080 screen can take from a few seconds up to ~30 s depending on your CPU and on-screen complexity. Use a region โ small crops are typically 10ร faster:
{"region": [0, 0, 800, 100]}Turkish characters not appearing
The system uses SendInput + KEYEVENTF_UNICODE which is layout-independent.
If characters still don't appear, the target app may not support Unicode
input โ try POST /api/window/post with action: "type" instead.
Project Structure
screen-control/
โโโ server.py # Flask HTTP server + all API endpoints
โโโ core/ # PlatformBackend interface + standardized errors
โ โโโ backends.py # Abstract backend + lazy discovery
โ โโโ errors.py # ApiError envelope + error codes
โโโ backends/ # OS implementations behind PlatformBackend
โ โโโ windows.py # Reference backend (moved from control.py)
โ โโโ linux.py # Fail-closed stub (ROADMAP Phase 5)
โ โโโ macos.py # Fail-closed stub (ROADMAP Phase 7)
โ โโโ fake.py # In-memory backend for tests
โ โโโ forbidden.py # Shared blocked-key policy
โโโ control.py # Compatibility shim re-exporting backends.windows
โโโ tests/unit/ # Offline unit + integration tests (no real input)
โโโ mcp_server.py # MCP server (stdio + streamable-HTTP) โ thin wrapper over the API
โโโ sdk/
โ โโโ screen_control.py # Python SDK client (pip-installable style)
โโโ index.html # Bundled web UI (live view + control panels)
โโโ docs/images/ # README assets (demo GIF captured by the API itself)
โโโ requirements.txt # Python dependencies
โโโ start-server.bat # One command: REST + MCP + cloud tunnel (Windows)
โโโ stop-server.bat # Stop all three processes
โโโ test-security.py # Security + game-mode test suite (34 checks)
โโโ test-game.py # Live game-mechanics test (app launch โ draw โ safe close)
โโโ test-endtoend.py # End-to-end test: open Notepad โ type โ save โ verify
โโโ .github/workflows/ # CI: runs the security suite on every push
โโโ AGENT_GUIDE.md # AI agent integration guide (separate from this file)
โโโ README.md # This file
โโโ .token # Auto-generated auth token (gitignored)Testing
Prerequisites
The server must be running:
cd screen-control
python server.pySecurity test suite
Tests authentication, blocked key combos, window management, safe close, critical process protection, and game mode โ all non-destructive.
cd screen-control
python test-security.pyExpected output:
== Token Authentication ==
โ Missing token -> 401
โ Wrong token -> 401
โ Correct token -> 200
โ Non-JSON POST -> 415
== Blocked Key Combos ==
โ Alt+F4 blocked (403)
โ Win key blocked (403)
โ Win+D blocked (403)
โ Delete blocked (403)
== Window List ==
โ Windows list requires GET
โ Window list is non-empty โ 8 windows
โ Exactly one focused window
== Safe Close Verification ==
โ Wrong title aborts close
== Critical Process Protection ==
โ System process (pid 4) rejected (403)
โ pid 0 rejected (403)
== Game Mode ==
โ Game mode started
โ Relative camera look
โ Game mode stopped
== Watchdog (dry run) ==
โ Held state returns ok
โ Watchdog count reported
========================================
RESULT: 19 passed, 0 failedLive game-mechanics test
Launches a real application (mspaint or notepad), performs hold-to-draw game mechanics, verifies via pixel analysis, then safely closes with "Don't Save" dialog handling.
cd screen-control
python test-game.pyNote: This test launches a real application. It handles cleanup automatically (sends WM_CLOSE and clicks "Don't Save" if a dialog appears).
Contributing
Fork the repository
Create a feature branch
Make your changes
Test on a Windows machine
Submit a pull request
Code Style
Python: PEP 8, type hints, docstrings on all public functions
Docstrings: English, Google style
Error messages: English, descriptive
Comments: English, explain why not what
License
MIT License. See LICENSE for details.
Built with โค๏ธ for local automation and AI agent research.
Available Tools
16 toolsclose_windowA
Close a window safely (WM_CLOSE after verification โ never blind Alt+F4). Pass expect_title/expect_process to refuse (409) if the window changed.
| Name | Required | Description | Default |
|---|---|---|---|
| hwnd | Yes | ||
| expect_title | No | ||
| expect_process | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does a good job: it discloses the safe close mechanism, the verification step, and the 409 refusal condition. It could further clarify what 'verification' checks exactly and what success/failure responses look like, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded, and every sentence earns its place: the first sentence states the action and mechanism, the second explains the safety parameters. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter tool with an output schema, the description is nearly complete: it covers the action, the safe method, and the guard parameters. Minor gaps include not stating what happens when no expect_* params are passed or how success is indicated, but these are acceptable given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the purpose of expect_title and expect_process (safety guards that trigger a 409 on window change), adding real meaning beyond the schema. hwnd is not described, but its role as a window handle is reasonably inferable from the tool name and schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Close a window safely') and provides the exact mechanism (WM_CLOSE after verification, never blind Alt+F4), which clearly distinguishes it from sibling tools like window_post or keyboard-based closing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives actionable usage guidance: pass expect_title/expect_process to refuse if the window changed, and explicitly warns against blind Alt+F4. It does not name alternative tools or state when not to use this tool, but the context is clear enough for an agent to recognize when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
focus_windowB
Bring a window to the foreground (focus it).
| Name | Required | Description | Default |
|---|---|---|---|
| hwnd | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the intended effect without mentioning potential side effects, failure conditions (e.g., if the window is minimized or unresponsive), or whether it requires special permissions. This leaves the agent without knowledge of how the operation might behave beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that immediately states the purpose with no redundant words. It is appropriately concise for a simple operation and front-loads the key action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that the tool has a required parameter (hwnd) that is undocumented, and no behavioral details are provided, the description is incomplete. The existence of an output schema (per context) is noted, but the description does not clarify what the tool returns (e.g., success/failure), leaving gaps in the agent's understanding of how to call and interpret the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The sole parameter 'hwnd' has no description in the schema (coverage 0%), and the tool description does not explain it at all. An agent cannot infer that 'hwnd' is a window handle or how to obtain it, making parameter usage opaque. The description fails to compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Bring a window to the foreground (focus it)' clearly states the action (bring/focus) on a specific resource (a window), distinguishing it from siblings like close_window or window_children. It is unambiguous and immediately tells an agent what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used when you want to focus a window, but it provides no explicit guidance on when to prefer it over alternatives or any exclusions. There is no mention of conditions like 'use only when the window is not already focused' or when to avoid it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gameB
Game mode for FPS-style control. action: start (clip cursor to center) | move (relative camera look by dx,dy pixels * sensitivity) | stop | heartbeat (keep-alive during long holds). Combine with keyboard() down/up for WASD and mouse(action='down'/'up') for shooting/building.
| Name | Required | Description | Default |
|---|---|---|---|
| dx | No | ||
| dy | No | ||
| action | Yes | ||
| sensitivity | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that 'start' clips cursor to center, 'move' does relative camera look by dx,dy pixels * sensitivity, and 'heartbeat' is a keep-alive during long holds. However, it doesn't disclose side effects like cursor visibility changes, whether stop restores the cursor, or what happens on errors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the tool's purpose. The action list is dense but efficient. The pipe-separated format is a bit terse but saves space. It earns its place by covering purpose, actions, and usage in two sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (though not shown in detail) and 4 parameters. The description covers the core actions and how to combine with keyboard/mouse, but lacks details like what the output looks like, whether sensitivity applies only to move, and edge cases like what happens if start is called twice. For a game-control tool, this is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains dx/dy as relative camera look pixels and sensitivity as a multiplier, but it doesn't explain the 'action' parameter's allowed values beyond listing them in the description (start, move, stop, heartbeat) โ which is helpful but not mapped to the schema. The 'action' parameter has no enum in the schema, so the description's list is the only guidance, but it's not structured as parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a game mode for FPS-style control and enumerates the specific actions (start, move, stop, heartbeat). It distinguishes itself from siblings like mouse and keyboard by framing it as a mode that combines with them, though it doesn't explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool: for FPS-style control, and it instructs combining with keyboard() down/up for WASD and mouse(action='down'/'up') for shooting/building. It implies this is the alternative to direct mouse/keyboard for game contexts, but doesn't explicitly state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_heldA
List currently held keys/buttons and watchdog status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for behavioral disclosure. It states the tool lists items, implying a non-destructive read operation, but offers no details on side effects, error behavior, or what 'watchdog status' entails. For a simple getter, this is minimally transparent but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, direct sentence with no padding. It efficiently conveys the core function without irrelevant detail, achieving maximum conciseness for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has zero parameters and an output schema (assumed to document return structure), the description is sufficient for an agent to call it. The mention of 'watchdog status' adds context beyond the name, though it could elaborate on what that status means. Overall, nothing critical is missing for a straightforward query tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema already fully describes the interface. The description adds no parameter semantics because none exist; a baseline of 4 is appropriate when there are no parameters to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('List') and a specific resource ('currently held keys/buttons and watchdog status'). It distinguishes this from sibling tools like 'release_all' and 'keyboard' by implying a read-only state query, though it doesn't explicitly name alternatives. The purpose is clear and not tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus siblings. It doesn't mention scenarios where 'get_held' is preferred, nor does it state any exclusions. The tool name itself is suggestive, but the description leaves all usage context to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_infoA
Server/PC status: screen size, monitors, OCR availability, failsafe and game-mode state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses what data is reported, which is the core behavioral trait for a status query, but it does not explicitly state that the tool is non-destructive or that it does not alter state beyond reading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the purpose ('Server/PC status:') and then packs the returned field list compactly. Every word earns its place; there is no filler or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateless zero-parameter query tool with a provided output schema, the description fully covers what the agent needs to know. The schema handles return-value structure, and the description enumerates the reported categories, leaving nothing missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the baseline is 4. There is nothing for the description to clarify about arguments; instead, it properly explains what the tool's output means, which is the only relevant semantic information given no schema to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific implicit verb (retrieve/query) over a concrete resource (server/PC status) and enumerates the exact data fields returned: screen size, monitors, OCR availability, failsafe, and game-mode state. This clearly distinguishes it from siblings like screenshot and ocr_screen, which capture images or extract text, and from get_held, which reports keyboard state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The listed fields (OCR availability, failsafe, game-mode) imply a pre-flight environment-status use case, so usage is plausibly inferred. However, the description never explicitly says when to call this tool versus alternatives, nor states any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
keyboardA
Control the keyboard. action: press | hotkey | type | down | up. press: key name (enter, f11, a...); hotkey: keys=["ctrl","s"]; type: Unicode-safe text. expect_hwnd refuses (409) the action if the foreground window differs โ use it to avoid typing into the wrong app.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | ||
| keys | No | ||
| text | No | ||
| action | No | press | |
| interval | No | ||
| expect_hwnd | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a key behavioral trait: expect_hwnd causes a 409 refusal if the foreground window differs, and notes that type is Unicode-safe. Since no annotations are provided, the description carries the full burden, but it omits other behaviors such as whether actions are synchronous, error handling beyond 409, or any permission requirements. It covers some but not all behavioral aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, two sentences long, and front-loads the core purpose and actions. It includes practical examples and a caution note without any fluff, making it highly readable and immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 6 parameters and no annotations, the description covers the main actions and key parameters but omits details about the interval parameter and the return value (though an output schema exists). For a simple keyboard tool, this may be sufficient for basic use, but it leaves some aspects undocumented, making it not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates well by explaining the action values and giving examples for key (press), keys (hotkey), and text (type), and clarifies the purpose of expect_hwnd. It does not explain interval or the exact format of key names beyond the examples, but it provides enough to use the primary parameters effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it controls the keyboard and enumerates the actions (press, hotkey, type, down, up) with concrete examples for each, making its purpose unmistakable and distinguishing it from mouse and other input tools. It is a specific verb+resource definition without any tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a usage guideline for the expect_hwnd parameter, advising to use it to avoid typing into the wrong app. However, it does not explicitly discuss when to use this tool versus sibling tools like mouse or get_held, or any exclusions. The guidance is present but not comprehensive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_windowsA
List visible top-level windows: hwnd, title, process name, pid.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the burden of explaining side effects and filtering. It clearly states the tool lists only visible top-level windows (no hidden or child windows), and the word 'List' implies a non-mutating read operation. It does not mention ordering or enumeration caveats, but these are minor for a zero-parameter read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence conveys the action, scope, and output fields with no filler. The key scoping terms ('visible', 'top-level') appear up front.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that no parameters are required and an output schema exists, the description supplies the essential identifying fields and the central filter. An agent can call this tool correctly without needing additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so there is nothing to document beyond the schema's empty properties object. The description adds value by specifying what the returned rows contain (hwnd, title, process name, pid), which the schema cannot show.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('List') and resource ('visible top-level windows'), then enumerates the exact fields returned. It distinguishes itself from sibling tools like window_children by narrowing scope to top-level windows and from screenshot by being an enumeration rather than a capture.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool instead of related window tools such as window_children, focus_window, or close_window. There are no exclusions, prerequisites, or alternative references, leaving the agent to infer the intended use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
motion_diffA
Text-only change detection between two snapshots: which tiles of the screen changed and their clickable centers. Use to verify an action had an effect or to spot movement without OCR/vision.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It mentions 'text-only' and 'without OCR/vision' but does not disclose how snapshots are obtained (e.g., whether it captures them internally), potential side effects, or performance implications. This ambiguity leaves the agent uncertain about the tool's execution model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise, front-loaded sentences pack the core purpose and usage into minimal space. No redundancy; every clause adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the what and when but leaves a critical gap: the mechanism of 'two snapshots' is vague. Since the tool only accepts an optional region, it's unclear whether the agent must supply snapshots or if the tool manages them internally. The output schema exists, so return format is covered, but this ambiguity undermines completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0% and the description does not mention the 'region' parameter at all. The parameter is left completely unexplained, forcing the agent to guess whether it restricts detection to a screen area or serves another purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (detect), the resource (screen changes between two snapshots), and the output (changed tiles and clickable centers). It explicitly contrasts with OCR/vision, distinguishing it from sibling tools like ocr_screen.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete use cases ('verify an action had an effect' or 'spot movement') and explicitly frames it as an alternative to OCR/vision. It doesn't name specific sibling tools, but the context is sufficient for an agent to decide when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mouseB
Control the mouse. action: move | click | scroll | drag | down | up. click uses x,y (optional = current pos); drag goes (x1,y1)->(x,y); scroll uses 'clicks' (negative = down). Physical-pixel coordinates.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| x1 | No | ||
| y1 | No | ||
| action | No | move | |
| button | No | left | |
| clicks | No | ||
| duration | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses meaningful traits: physical-pixel coordinates, optional current-position click, and negative clicks for downward scroll. However, it does not mention side effects (e.g., moving the actual system cursor, affecting other applications), permissions, or reversibility. Some behavioral info is present, but gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only three short sentences, front-loading the core purpose and then compactly enumerating action semantics. Every phrase earns its place, using terse notation that is easy to scan. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 8 parameters, no annotations, and an output schema that likely documents return values. The description covers the most important parameter semantics but omits button and duration, and does not explain the effect of each action beyond coordinate mechanics. It is adequate for basic use but not complete for all parameters and possible edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of action values, x/y as click coordinates (optional), x1/y1 for drag start, and clicks for scroll amount. It does not describe the button or duration parameters, leaving parts of the parameter space undocumented. Still, it adds substantial meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Control the mouse' and lists the supported actions (move, click, scroll, drag, down, up), giving a clear verb+resource statement. It doesn't explicitly distinguish from sibling tools, but the action list makes it self-evident compared to keyboard, screenshot, or window controls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use the mouse tool versus alternatives like keyboard or screenshot. It explains how to perform specific actions (e.g., 'click uses x,y', 'scroll uses clicks'), but not when this tool should be preferred or avoided. There is no mention of context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ocr_screenA
Read the screen as text with coordinates (OCR). Best for menus, dialogs, editors, chat โ anything text-heavy. Returns lines plus per-item centers for clicking. 'region' is an optional "x,y,w,h" crop.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | ||
| scan_focus | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It usefully discloses the return shape (lines plus per-item centers for clicking) and the optional crop, and 'Read' implies a non-destructive operation. But it leaves scan_focus behavior unexplained and does not mention permission requirements or performance characteristics (OCR latency), which are material for an agent deciding whether to call it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero waste: purpose is front-loaded, usage guidance follows, and behavior plus parameter semantics are packed into one efficient sentence. Every clause adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value detail is already covered, and the description's mention of 'lines plus per-item centers' aligns with it. The main gaps are the unexplained scan_focus parameter and lack of explicit differentiation from window_capture_ocr among a large sibling list โ both are things an agent needs to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It fully explains 'region' (format 'x,y,w,h' and optional crop semantics), but scan_focus is entirely undocumented in both the schema and the description โ the agent cannot know what focusing behavior it toggles. With half the parameters unexplained, a 2 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb+resource ('Read the screen as text with coordinates (OCR)') and differentiates from siblings: 'best for menus, dialogs, editors, chat โ anything text-heavy' distinguishes it from screenshot (visual capture) and the sibling window_capture_ocr (window-scoped OCR). An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use context ('text-heavy' interfaces such as menus, dialogs, editors, chat). However, it does not name the alternative window_capture_ocr or state when to prefer that sibling over whole-screen OCR, so exclusions are only implied, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
release_allA
Release every held key and button (emergency reset). Safe to call anytime.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of disclosing behavior. It does state that the operation is an emergency reset and safe to call anytime, which is useful context, but it does not explain potential side effects, whether it affects the OS-level key state or only the tool's internal state, or what the response will be.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short, front-loaded sentences with no filler. The core action and emergency-reset intent come first, and the safety reassurance is appended efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description is complete: it names the action, the scope, the emergency-reset intent, and the safety profile. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty and there are zero parameters, so the description does not need to explain parameter semantics. The phrase 'every held key and button' also conveys the scope of the operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Release') and a precise resource ('every held key and button'), and adds the clarifying parenthetical 'emergency reset'. This makes the tool's purpose unmistakable and distinguishes it from siblings like get_held, which queries held state rather than clearing it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'emergency reset' implies the intended use case, and 'Safe to call anytime' provides clear, unconditional context for when it can be used. However, it does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotA
Capture the screen as an image (for vision-capable models). 'region' = "x,y,w,h"; 'scale' 0-1 shrinks to save bandwidth; 'quality' 1-100.
| Name | Required | Description | Default |
|---|---|---|---|
| scale | No | ||
| region | No | ||
| monitor | No | ||
| quality | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It does explain parameter behavior (scale shrinks to save bandwidth, quality range, region format), which is useful. However, it does not explicitly state that the operation is read-only, nor does it disclose output format, permission requirements, or multi-monitor behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The main purpose is front-loaded, and the parameter hints are compact and actionable. Nothing extraneous is included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, no annotations, and 0% schema description coverage, the description should handle all context. It misses the meaning of the 'monitor' parameter and fails to state the output format (e.g., base64, URL, path), making it incomplete for an agent to call correctly in multi-monitor or output-sensitive scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clearly defines region as 'x,y,w,h', scale as 0-1, and quality as 1-100, adding meaning to three of four parameters. It entirely omits the 'monitor' parameter, leaving it unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Capture the screen as an image'. 'For vision-capable models' further clarifies the intended use case, and the plain-image capture distinguishes it from OCR siblings like ocr_screen and window_capture_ocr.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for vision-capable models' implies a use case, but there is no explicit guidance on when to use this tool versus alternatives like ocr_screen or window_capture_ocr. No exclusions or when-not-to-use conditions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
window_capture_ocrA
Capture a background window WITHOUT focusing it; with ocr=1 (default) returns its text directly โ read other apps' content while the user works.
| Name | Required | Description | Default |
|---|---|---|---|
| ocr | No | ||
| hwnd | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a key behavioral trait: it does NOT focus the window, which is a significant side-effect avoidance. It also explains the default ocr behavior. With no annotations provided, this is valuable behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the most important behavior (no focus), and the ocr default is explained inline. Zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values are covered. The description explains the core behavior and default. It could mention what happens when ocr=0, but the output schema likely covers that. Overall adequate for a 2-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the ocr parameter's default and effect (returns text directly), but doesn't explain hwnd beyond the schema's integer type. The description adds some meaning but not full parameter coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool captures a background window without focusing it, and with ocr=1 (default) returns its text directly. This distinguishes it from siblings like screenshot, ocr_screen, and focus_window.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when you need to read another app's content without disrupting the user's work. It doesn't explicitly name alternatives or exclusions, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
window_childrenA
List child controls of a window (class name, title, hwnd). Useful for classic Win32 apps with real child controls (e.g. an Edit box) โ target the child hwnd in window_post. Modern WinUI/UWP apps have no classic children; use window_input_mode to detect them.
| Name | Required | Description | Default |
|---|---|---|---|
| hwnd | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden. It clearly signals a non-mutating list operation and adds a useful platform caveat about WinUI/UWP apps having no classic children. It does not discuss edge cases like direct-versus-recursive enumeration, but for a read-only inspection tool the core behavior is sufficiently disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence states the core action and output fields; the second adds usage context and an alternative. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description is complete enough. It explains the target scenario, the practical follow-up action (window_post), and the main exception (modern apps). No critical calling information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description only mentions hwnd incidentally as the child handle to pass to window_post. It does not explicitly state that the input hwnd is the parent window handle or where it should come from, though the tool name makes this mostly inferable. The description partially compensates for the missing schema detail but could be more explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a clear resource ('child controls of a window'), and it names the returned attributes (class name, title, hwnd). It also distinguishes the tool from siblings by framing it as the Win32 child-enumeration tool versus window_input_mode for modern apps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when this tool is useful ('classic Win32 apps with real child controls'), what to do with the results ('target the child hwnd in window_post'), and which alternative to use for modern apps ('use window_input_mode to detect them'). This is clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
window_input_modeA
Classify how a window receives input BEFORE posting to it: postmessage (classic Win32) | uia (WinUI/UWP: focus+SendInput fallback) | focused (already foreground) | invalid (dead window).
| Name | Required | Description | Default |
|---|---|---|---|
| hwnd | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the output categories, including the 'invalid (dead window)' case, and implies a read-only operation ('Classify') with no side effects. However, it does not explicitly state that it performs no mutation or that it requires a valid handle, though the 'invalid' output suggests handling of dead windows. This is adequate but not rich in behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured line with line breaks separating the classification options. It contains zero filler and front-loads the purpose and timing. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, an output schema presumably describing the classification results), the description covers the essential context: when to use, what it returns, and edge cases (invalid). It could mention that it is read-only or requires a valid handle, but these are minor gaps for a classification tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implies that 'hwnd' refers to the window being classified, but does not explicitly define the parameter or its format. Since the tool name and description make the role of hwnd obvious, a 3 is appropriate, but explicit clarification would improve it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Classify') with a clear resource ('how a window receives input') and explicitly lists the four possible classification values. This makes the tool's purpose unambiguous and distinguishes it from siblings like window_post (which posts input) and focus_window (which focuses).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'BEFORE posting to it' provides clear temporal context, indicating this tool should be called before window_post. The classification values also imply decision logic, but no explicit exclusions or alternatives are mentioned. Still, the guidance is sufficiently clear for an agent to know when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
window_postB
Send input to a window in the background. action: type | key | hotkey | click | scroll. mode: auto (recommended โ routes WinUI/UWP apps through the focused SendInput path automatically) | background | focused. Note: 'uia' routing focuses the window (unavoidable for modern apps).
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| key | No | ||
| hwnd | Yes | ||
| keys | No | ||
| mode | No | auto | |
| text | No | ||
| action | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does disclose that 'uia' routing focuses the window, which is a key side effect. It also implies background operation, but it does not explain what 'background' means in terms of window visibility or input capture, nor does it describe side effects for actions like click or scroll. The behavior is partially transparent but lacks detail on edge cases or requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact paragraph that front-loads the core purpose and then lists the key enums. It uses a clear 'action: ...' and 'mode: ...' format, and the note about uia routing is placed at the end as an important caveat. The structure is logical and efficient, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 8 parameters, 2 required, and an output schema exists, but the description is insufficient for an agent to use it correctly. It omits parameter meanings, prerequisites (e.g., whether the window must be visible or minimized), and concrete examples. The note about uia routing is helpful but not enough. While the output schema may describe return values, the input side is under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate for the 8 parameters. It explains the 'action' and 'mode' enums, but leaves x, y, key, keys, text, and hwnd entirely unexplained. For instance, x and y are likely coordinates for click/scroll, and key/keys/text are input payloads, but the description provides no such meaning. This is a significant gap that forces the agent to guess or inspect the schema further.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Send input to a window in the background.' It lists the supported action types (type, key, hotkey, click, scroll), giving a concrete sense of what the tool does. However, it does not explicitly differentiate itself from sibling tools like mouse, keyboard, or focus_window, so an agent must infer that this is specifically for targeting a window handle (hwnd) rather than global input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage guidance by explaining mode options and recommending 'auto' for WinUI/UWP apps, and notes that 'uia' routing forces focus. However, it does not explicitly state when to use this tool versus alternatives like mouse or keyboard, nor does it clarify the trade-offs between background and focused modes beyond the note. The guidance is partial, leaving the agent to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v0.1.0- First observed
close_window - First observed
focus_window - First observed
game - First observed
get_held - First observed
get_info - First observed
keyboard - First observed
list_windows - First observed
motion_diff - First observed
mouse - First observed
ocr_screen - First observed
release_all - First observed
screenshot - First observed
window_capture_ocr - First observed
window_children - First observed
window_input_mode - First observed
window_post
TDQS
Scored across 16 tools
Each tool targets a distinct mechanismโscreen capture, OCR, motion detection, global input, window-specific input, window introspection, and game mode. Potentially similar tools like ocr_screen and window_capture_ocr are cleanly separated by foreground screen versus background window. No two tools appear to do the same job.
Most tools follow a readable snake_case convention with verb-first names like list_windows, focus_window, close_window, and release_all. A few noun-style names like mouse, keyboard, and game, plus window_post and window_capture_ocr, are minor deviations but still clear and predictable.
Sixteen tools is on the higher end but each earns its place across capture, OCR, input, window management, and game mode. The count feels slightly heavy but remains well-scoped for a desktop automation server.
Core workflows are well covered: screen capture, OCR, change detection, global and background input, window discovery, focusing, posting, and closing. Minor gaps like querying the current cursor position or window geometry are absent but most operations have workarounds.
Maintenance
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Eyes and hands on real Windows PCs โ observe, click, type via Glasswarp API.
Human-input bridge for AI agents with voice-first answer links, MCP tools, and HTTP APIs.
Zero-setup MCP gateway securely connecting AI to your tools with authentication and workflows
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceMCP server that enables Claude to control your computer, similar to Anthropic's computer use but easy to set up locally.1,112 npm379MIT
- AlicenseAqualityDmaintenanceEnables LLM agents to capture screenshots, control mouse/keyboard, and manage windows on desktop platforms, primarily Windows, via an MCP server.161MIT
- AlicenseAqualityBmaintenanceAllows AI clients to see and control Windows 10/11 desktops via MCP, with screenshots, UI Automation, Chrome CDP, keyboard/mouse, and terminal using semantic element targeting.30548 npmMIT
- AlicenseNot gradedqualityAmaintenanceLocal MCP server for Gateway-managed computer use that exposes computer.* tools with desktop control, OCR, and user-visible safety overlays.6 npm1MIT