desktop-automation
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@desktop-automationObserve the visible Notepad window and click the OK button"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
desktop-automation-mcp
A deny-by-default, per-action-gated MCP server that lets an MCP client observe and operate specific Windows application windows, and nothing else.
It is not a general desktop-automation tool. Access is granted per application and per action class, and every effect is re-verified against the live target immediately before it happens.
Why this design
Many desktop-automation MCP servers maximize capability: they expose the whole desktop and often feed the accessibility tree to the model as text. That makes two things easy that should be hard: acting on a window nobody authorized, and on-screen content being read as instructions.
This server takes the opposite position:
Nothing is allowed until you name it. A window must match an allowed title pattern and an exact executable path, and each action class (
observe,hover,click,key,text,drag,close) is a separate permission. A new tool never inherits an existing permission.Every effect is re-checked at the moment it happens. Window handle, process ID, process start time, visibility, foreground and occlusion are re-bound right before input is sent. A stale or ambiguous target is denied, never guessed.
Screen content is never handed over as text. The client receives pixels from policy-named regions and window-relative coordinates, so there is no channel for on-screen text to be mistaken for instructions.
Content-changing and irreversible actions need a confirmation token.
text,dragandcloserequire a single-use, short-lived token bound to the exact target and action class.Failures are loud. If focus, identity or geometry cannot be verified, the tool raises an error instead of acting on a best guess.
It also supports canvas applications without UI Automation (for example Flash content in Ruffle) through versioned, hash-verified coordinate profiles.
Related MCP server: dsh-desktop-operator
Requirements
Windows (the server calls Win32 APIs directly)
Python 3.12 or newer (CI covers 3.12, 3.13 and 3.14)
Quick start
Install from a checkout (the package is not published to PyPI):
git clone https://github.com/WRG-11/desktop-automation-mcp.git cd desktop-automation-mcp py -3.12 -m pip install .Write a policy for one application with only the actions you need. This one allows looking at Notepad and nothing more:
application: executable_path: "C:\\Windows\\System32\\notepad.exe" title_patterns: - "*Notepad*" allowed_actions: - observeValidate it before use:
py -3.12 tools/validate_policy.py notepad-observe.yamlRegister the server with your MCP client:
{ "mcpServers": { "desktop-automation": { "command": "desktop-automation-mcp", "env": { "DESKTOP_AUTOMATION_POLICY_FILE": "C:\\path\\to\\notepad-observe.yaml" } } } }Restart the client, then call
health_check()andlist_windows(). If the list is empty,policy_visibility_summary()tells you whether the title or the executable path is the part that does not match.
The quickstart guide walks through the same steps with environment-variable configuration, and the security guide explains how to keep a policy narrow.
Tools
Family | Tools | Permission |
Diagnostics |
| none, or |
Discovery |
|
|
Image and profiles |
|
|
Confirmation |
| the confirmed action / |
Pointer |
|
|
Keyboard |
|
|
Protected |
|
|
Parameters, return shapes and per-tool guarantees are in the tool reference.
Configuration
A policy can come from a version-controlled file
(DESKTOP_AUTOMATION_POLICY_FILE, validated against
schema/policy.schema.json) or from environment
variables (DESKTOP_AUTOMATION_ALLOWED_TITLES,
DESKTOP_AUTOMATION_ALLOWED_PROCESS_PATHS,
DESKTOP_AUTOMATION_ALLOWED_ACTIONS). Setting both is rejected rather than
merged. Screenshots in file mode are limited to regions the policy names.
Focus mode defaults to passive: the server never brings a window to the
front, it only acts on the window you already focused. Rate limits, the
screenshot memory budget and input limits are all configurable. Everything is
listed in the configuration reference.
Known limitations
Windows only, on the local machine, for visible (not minimized) top-level windows.
Pointer tools move the shared system cursor; do not grant
clickorhoverwhile you are using the machine yourself.No OCR or text search: the server returns pixels and coordinates, not text.
Mouse-and-key chords (holding a key while dragging) are not supported.
UAC prompts, the sign-in screen and the lock screen are outside the target model; the server relies on Windows desktop isolation and never bypasses it.
The confirmation token proves that a token was requested for this target, not that a human approved it. The client must keep approval as a separate turn; see the confirmation contract.
Documentation
Development
py -3.12 -m pip install -e ".[dev]"
py -3.12 -m ruff check src tests tools
py -3.12 -m ruff format --check src tests tools
py -3.12 -m unittest discover -s tests -vThe quality workflow runs the same checks on clean Windows runners for Python 3.12, 3.13 and 3.14, then builds the wheel and completes an MCP stdio handshake against the installed package. It never drives real windows; UI smoke tests are run separately under operator supervision.
Chaos scenarios (tests/test_chaos.py) run the real target.py,
visibility.py and server.py functions against shared fake Win32 state
(tests/fake_platform.py); tests/test_server.py keeps its own mock pattern.
Contributing and security
Read CONTRIBUTING.md before opening a pull request. Report vulnerabilities privately as described in SECURITY.md, not in a public issue. For questions, see SUPPORT.md.
License
Apache License 2.0. See NOTICE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
Securely control computers you explicitly pair through files, terminals, processes, screenshots, desktop UI/input, clipboard, browser automation, diagnostics, and document tools.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Runtime permission, approval, and audit layer for AI agent tool execution.
Related MCP Servers
- FlicenseBqualityCmaintenanceEnables local Windows UI automation and screen capture, targeting windows that are difficult to automate such as games and legacy apps, with tools for window management, mouse and keyboard input, and desktop capture.8-
- AlicenseNot gradedqualityAmaintenanceEnables safe Windows desktop automation and computer use through natural language, including window observation, UI Automation, and execution of verified actions like clicking, typing, and scrolling.2MIT
- AlicenseNot gradedqualityAmaintenanceEnables agentic control of Windows via MCP tools for screenshots, mouse, keyboard, UI Automation, OCR, windows, and clipboard, while risky actions require real-human approval through a policy-gated dialog with audit logging.5MIT
- AlicenseNot gradedqualityBmaintenanceEnables guarded, local-first perception and interaction with Windows desktops through UI Automation, screenshots, and bounded mouse/keyboard actions with deterministic verification. It prioritizes semantic/read-only inspection and excludes arbitrary shell execution or destructive system operations.Apache 2.0