Universal Computer Control MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Universal Computer Control MCPtake a screenshot, find the Save button, click it, and verify the file saved"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Universal Computer Control MCP
Give your AI agent eyes and hands on any computer. Observe → understand → act → verify. No browser MCPs. No Playwright. Just real control.
A cross-platform computer control MCP server that lets an AI agent operate a Windows or Linux computer like a human: observe the screen → understand the UI → choose an action → perform it → observe again → verify → recover if needed.
The system deliberately has no dependency on Chrome MCP, browser MCP, Playwright MCP or any other external MCP server. Browsers, IDEs, office suites, terminals, file managers, Electron apps and canvas-heavy applications are all just GUI applications. Multiple backends live inside the server and the best available mechanism is chosen automatically per action.
AI Agent / LLM
| MCP (stdio)
v
+------------------------------+
| Universal Computer MCP Server |
+---------------+--------------+
v
+------------------------------+
| Computer Control Engine | <- usable WITHOUT MCP
+---------------+--------------+
| | |
v v v
Accessibility Vision/OCR Physical Input
Layer Layer Layer
| | | | | |
UIA AT-SPI OCR OpenCV Mouse Keyboard
(Windows)(Linux) Tesseract VLM PyAutoGUI OS APIsHighlights
29 stable MCP tools under the
computer.namespace — capabilities, not implementation details. Nopyautogui_click, noocr_click, noplaywright_click.Interaction hierarchy per action: Windows UIA → Linux AT-SPI → semantic element → OCR text → image template → vision-language model → coordinates → PyAutoGUI. Every fallback is automatic.
computer.observeis the core tool: one structured snapshot with screen geometry, active window, cursor, OCR'd text and normalized UI elements (stableelement_Nids within a cycle).
Related MCP server: planchette
Requirements
Mandatory | Optional (recommended) | |
Python | 3.11+ | |
Core |
| |
Input |
| |
Windows |
| |
Linux |
| |
Screenshots |
| |
OCR |
| |
Vision |
| |
VLM |
|
Installation
Windows 10/11
git clone https://github.com/adkdev200/universal-computer-control.git && cd universal-computer-control
powershell -ExecutionPolicy Bypass -File scripts\install_windows.ps1Manual steps:
Install Python 3.11+ (
winget install -e --id Python.Python.3.12), tick Add to PATH.python -m venv .venv && .venv\Scripts\pip install -e ".[input,windows,ocr,vision,screenshot,vlm]"Install Tesseract OCR:
winget install -e --id UB-Mannheim.TesseractOCR(setocr.tesseract_cmdinconfig.yamlif it is not onPATH).Run:
.venv\Scripts\universal-computer-control(alias:.venv\Scripts\ucc)
Windows UI Automation needs no extra permissions. If you run the server elevated, non-elevated windows may refuse automation (UIA integrity levels).
Linux (X11 or XWayland)
git clone https://github.com/adkdev200/universal-computer-control.git && cd universal-computer-control
bash scripts/install_linux.sh --yes # add --no-sudo to skip system packages
./.venv/bin/ucc-doctor # verify every backend, exact fix hintsThe installer is fully non-interactive with --yes: it installs system
packages (wmctrl/xdotool/xclip/scrot/tesseract/at-spi2-core), all Python
extras, and a ready ~/.universal-computer/config/config.yaml with OCR,
template matching and the VLM provider enabled. ucc-doctor then probes every
backend live and prints copy-paste fix commands for anything missing. Note:
even on machines with no wmctrl/xdotool/xclip installed, the Linux window
backend now works through a built-in python-Xlib (EWMH/ICCCM) fallback - the
system tools only make it more robust.
Manual steps:
Python 3.11+ and system packages:
sudo apt install python3 python3-venv python3-pip \ wmctrl xdotool xclip scrot \ tesseract-ocr at-spi2-corepython3 -m venv .venv && ./.venv/bin/pip install -e ".[input,linux,ocr,vision,screenshot,vlm]"Run from a graphical session:
./.venv/bin/universal-computer-control(alias:./.venv/bin/ucc)
AT-SPI requirements (accessibility): at-spi2-core must be running (it is
on desktop distros) and for GNOME/GTK apps enable
gsettings set org.gnome.desktop.interface toolkit-accessibility true.
X11 vs Wayland: X11 is fully supported. On Wayland-with-XWayland, input
and screenshots work through XWayland but foreign window management may be
restricted by the compositor. On pure Wayland (no DISPLAY), the window
backend disables itself with a warning; AT-SPI still works for cooperating
toolkits. See docs/TROUBLESHOOTING.md.
Registering the server with an MCP client
claude_desktop_config.json (see examples/mcp-config.example*.json):
{
"mcpServers": {
"universal-computer": {
"command": "/absolute/path/to/universal-computer-control/.venv/bin/python",
"args": ["-m", "universal_computer"],
"env": {}
}
}
}Any MCP client that speaks stdio works the same way. The engine is also a plain library — no MCP required:
from universal_computer import ComputerControlEngine, load_config
computer = ComputerControlEngine(load_config())
obs = await computer.observe("normal")
await computer.click("Continue")
await computer.type("hello")Tools (MCP API)
Group | Tools |
Observation |
|
Mouse |
|
Keyboard |
|
Windows |
|
Clipboard |
|
System |
|
Admin |
|
computer.observe returns a normalized snapshot:
{
"ok": true,
"id": "obs_17",
"level": "normal",
"screen": {"width": 1920, "height": 1080, "monitors": [...]},
"active_window": {"title": "Google Chrome", "application": "chrome.exe", "process_id": 1234},
"cursor": {"x": 812, "y": 530},
"elements": [
{"id": "element_1", "role": "button", "name": "Continue",
"bbox": [1050, 700, 1170, 750], "source": "uia", "confidence": 0.99}
],
"ocr": [
{"text": "Continue", "bbox": [1050, 700, 1170, 750], "confidence": 0.97}
]
}computer.backend_status reports per-backend health:
{
"pyautogui": {"available": true, "healthy": true},
"uia": {"available": true, "healthy": true},
"atspi": {"available": false, "healthy": false,
"error": "pyatspi import failed (...); install python3-pyatspi"},
"ocr": {"available": true, "healthy": true}
}Configuration
config.yaml (see config.example.yaml) plus UCC_ environment overrides
with __ nesting:
UCC_ENGINE__VERIFICATION_MODE=strict
UCC_SECURITY__MODE=standard
UCC_SECURITY__ALLOWED_COMMANDS='["ls", "git status", "xdotool"]'
UCC_LOGGING__LEVEL=DEBUG
UCC_VISION__VLM__ENABLED=trueConfigurable: observation level & caching, verification mode
(off/basic/strict), PyAutoGUI pauses/failsafe, OCR provider & language,
vision threshold, template directory, VLM endpoint (key read from the
environment variable named in vision.vlm.api_key_env), security mode,
allowlists, rate limits, timeouts, persistence retention, log levels and
per-backend priorities (backends.priorities: {pyautogui: 5}).
Security model
Mode |
|
|
| allowed (still rate-limited); | allowed (still rate-limited) |
| requires | allowed unless a non-empty |
| only explicitly allowlisted; | exact allowlist match required |
Unquoted shell operators (;, |, &, `, <, >, $) only pass the
allowlist when an entry matches the full command string — a bare
program-name entry (git, python, ...) is not enough.
Always on: rate limiting (max_actions_per_minute), command timeouts,
credential redaction in logs and command output, no clipboard/typed-text
logging, close_window/run_command/launch_application confirmation gates,
and the emergency stop — computer.emergency_stop() halts every subsequent
action (checked before each attempt) and is persisted as a stop file so it
survives restarts and can be triggered externally.
Extending
New OCR: implement
universal_computer.vision.base.OCRProvider(detect_text(image) -> list[OCRResult]) and register it inVisionServices.from_config.New vision model: implement
VisionLanguageModel.locate_element(async, description →VisionResultwith bbox). The OpenAI-compatible implementation shows the JSON contract.New platform (macOS): subclass the
backends.baseABCs (InputBackend,ScreenshotBackend,AccessibilityBackend,WindowManagementBackend,ClipboardBackend,ApplicationBackend) and register them incore/bootstrap.py. Nothing else changes — the engine only knows capabilities.Template images: put PNGs in
vision.template_dir; refer to them by name fromcomputer.find_visualor as click targets.
Development
pip install -e ".[dev]"
pytest tests # 160 unit + integration tests, no display needed
ruff check src testsscripts/e2e_live_test.py is an end-to-end harness that boots its own
Xvfb + openbox session, creates real X11 windows, and exercises every
backend live (screenshots, OCR, window ops via both the tool paths and the
python-Xlib fallback, OpenCV template matching, the VLM pipeline against a
mock OpenAI-compatible endpoint, clipboard, input, launch, emergency stop):
python scripts/e2e_live_test.py # full stack
python scripts/e2e_live_test.py --no-tools # zero X11 tools: Xlib fallbackUnit tests run on any machine (OS-specific pieces are mocked). The
integration test boots the real server over stdio. See
docs/ARCHITECTURE.md for the module map and docs/TROUBLESHOOTING.md for
platform-specific fixes.
Contributors & Acknowledgments
Z.ai (GLM) — co-developer: MCP server, cross-platform backends, vision pipeline, python-Xlib window fallback and
ucc-doctordiagnosticsClaude Code (Anthropic) — co-developer: code generation, testing and plug-and-play hardening
@adkdev200 — creator & maintainer
This project is built collaboratively with AI pair-programming: the majority of the codebase was written and reviewed by Z.ai's GLM and Anthropic's Claude Code under the direction of the maintainer.
This server cannot be deployed
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Use your Mac, Windows or Linux computer from ChatGPT, Claude or Codex: files, commands, documents.
Related MCP Servers
- AlicenseAqualityCmaintenanceAllows AI clients to see and control Windows 10/11 desktops via MCP, with screenshots, UI Automation, Chrome CDP, keyboard/mouse, and terminal using semantic element targeting.30835 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables coding agents to control other application windows by providing primitives for clicking, typing, capturing screenshots, and window management. Supports macOS, Windows, and Linux.MIT
- FlicenseAqualityAmaintenanceCross-platform desktop automation MCP server that lets AI agents capture screenshots, run OCR with UI-element classification, control mouse/keyboard, and launch programs on Linux, macOS, and Windows.201-
- AlicenseNot gradedqualityAmaintenanceLets AI agents see and control desktop applications through the accessibility layer, enabling clicking, typing, scrolling, dragging, and window/app management across macOS, Windows, and Linux entirely on the local machine.13MIT