UI-Venus MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@UI-Venus MCPopen Chrome, search for flights to Tokyo, and click the first result"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
UI-Venus MCP
A cross-platform Computer-Use MCP server for AI agents — unified GUI automation across Windows, Linux, macOS, Android, iOS and browsers. Structured APIs first (UIA / AX / AT-SPI / UIAutomator / XCUITest / DOM), vision grounding as the fallback, with UI-Venus-2-9B (W8A8 quantized build verified) as the default — and pluggable — vision provider.
中文文档:docs/README.zh-CN.md
ZCode / Claude Code / Codex / OpenAI Agents / your agent
│
┌─────────┴─────────┐
│ Text or VLM │ (any model, any size)
└─────────┬─────────┘
│ MCP (stdio / Streamable HTTP)
▼
┌──────────────────────────────────────┐
│ Computer-Use MCP Server │
│ Orchestrator · Fusion Locator · │
│ Verifier · Recorder · DSL · Guard │
└──────────────────┬───────────────────┘
│
Platform Router (per-device sessions)
┌──────────┬──────────┬────────┴───┬────────────┬───────────┐
▼ ▼ ▼ ▼ ▼ ▼
Windows Linux macOS Android iOS Browser
UIA+SendInput AT-SPI+ AX+ UIAutomator XCUITest/ Playwright/
PowerShell xdotool SystemEvents +adb WDA+simctl CDP
└──────────┴──────────┴────────────┴────────────┴───────────┘
│ vision fallback (spec §4)
▼
UI-Venus-2-9B (remote GPU, OpenAI-compatible)The product definition: give any AI agent — text-only or multimodal, small or large — unified, cross-platform, structure-first, vision-fallback computer-use capabilities. Not "an AI mouse for Windows", and not "UI-Venus wrapped in a few click APIs": UI-Venus acts as the cross-platform GUI expert (visual grounding, next-action decision, visual verification), while platform adapters execute reliably through native semantics whenever they exist.
Highlights
17 MCP tools — targets, state, inspect, screenshot, locate, action, step, execute_task, verify, recorder (start/stop/to_script), run_script, run_ui_test, get/cancel_task. Full contract in docs/api.md.
Four agent modes (spec-level):
delegate(autonomous loop for text-only agents),assist/direct(your agent stays in control; MCP provides primitives),auto(structured-first with per-step vision fallback).Structure-first execution:
click(elementRef)becomes UIA Invoke on Windows, AXPress on macOS, AT-SPI doAction on Linux, UIAutomator tap on Android, XCUITest tap on iOS, Playwright click in browsers. Raw coordinates are the last resort — and when vision produces a point, the fusion locator snaps it back onto a structured element before acting.Cross-platform coordinate system: vision-normalized [0,1000] ↔ screenshot px ↔ logical points ↔ physical pixels, with DPI/Retina/density handled and unit-tested.
Honesty by construction: missing permissions, offline devices, Wayland restrictions, unsigned WDA — reported as
permission_required/restricted/BLOCKED, never faked as success.Anti-stagnation loop guard: same-action / same-element / same-screen (perceptual hash) detection with a recovery ladder (re-observe → alternate strategy → honest failure).
Security: app/device/action/domain allowlists + sensitive-action detection (删除/支付/转账/install…) parking tasks in
WAITING_CONFIRMATIONwith confirm tokens.Recorder → portable DSL: recordings become YAML scripts with semantic locators (never
click(432,621); sleep(2)), replayable across platforms; a UI-test runtime reportsPASS / FAIL / SKIP / BLOCKED.Remote GPU architecture: the 9B vision model runs on a server (OpenAI-compatible endpoint); clients only send screenshots — phones and laptops need no VRAM.
Quick start
git clone https://github.com/q1820926174-cpu/UI-Venus-MCP.git
cd UI-Venus-MCP
pnpm install
pnpm build
export VENUS_BASE_URL=http://<gpu-host>:8300/v1 # OpenAI-compatible
export VENUS_API_KEY=<your-key>
export VENUS_MODEL=UI-Venus-2-9B-W8A8
# stdio (ZCode / Claude Code / Codex)
node dist/index.js
# Streamable HTTP (remote agents / LAN)
node dist/index.js --http --port 8765Smoke-test the vision endpoint (one calibration request):
pnpm smoke:venusZCode / Claude Code config
{
"mcpServers": {
"ui-venus-mcp": {
"command": "node",
"args": ["/absolute/path/to/UI-Venus-MCP/dist/index.js"],
"env": {
"VENUS_BASE_URL": "http://<gpu-host>:8300/v1",
"VENUS_API_KEY": "<your-key>",
"VENUS_MODEL": "UI-Venus-2-9B-W8A8"
}
}
}
}More configs (HTTP transport, Codex, OpenAI Agents): examples/mcp-config.md.
Usage patterns
Assist mode — a multimodal agent drives; the MCP executes and grounds:
computer_inspect → screenshot + UI tree (+ optional vision description)
computer_locate → "关闭按钮" → {element, point} (structured first, vision fallback)
computer_action → { type: "click", element: {...} }
computer_verify → structured assertion, else vision verdictDelegate mode — a text-only agent hands over the whole task:
{
"tool": "computer_execute_task",
"arguments": {
"target": { "type": "device", "platform": "android", "deviceId": "emulator-5554" },
"task": "打开设置,将Wi-Fi打开",
"mode": "delegate",
"maxSteps": 30
}
}Internally: Observe → Plan → Locate → Execute → Observe → Verify → Recover/Finish, with the state machine, security gates, and loop guard enforcing honest termination. Poll with computer_get_task, cancel with computer_cancel_task; sensitive actions return WAITING_CONFIRMATION + confirmToken.
Record → script → replay:
computer_record_start → (drive the app via computer_action) → computer_record_to_scriptname: disable-auto-update
steps:
- locate: { role: button, name: 设置 }
action: click
- locate: { name: 自动更新 }
action: toggle
value: false
assert:
- element: { name: 自动更新 }
property: { checked: false }Replay with computer_run_script, or run as a UI test (PASS/FAIL/SKIP/BLOCKED) with computer_run_ui_test. DSL reference: examples/dsl/toggle-autoupdate.yaml.
Platform support matrix
Capability | Windows | Linux | macOS | Android | iOS | Browser |
Screenshot | ✅ CopyFromScreen | ✅ import/scrot/grim | ✅ screencapture | ✅ screencap | ✅ simctl/WDA | ✅ Playwright |
Accessibility tree | ✅ UIA | ✅ AT-SPI (pyatspi) | ✅ AX (System Events) | ✅ UIAutomator dump | ⚠️ WDA/idb | ✅ aria snapshot |
Semantic actions | ✅ UIA patterns | ✅ doAction/setText | ✅ AXPress/AXValue | ✅ dump+tap center | ⚠️ WDA elements | ✅ role/text/testid locators |
Global input | ✅ SendInput | ✅ xdotool / wtype | ✅ CGEvent | ✅ adb input | ⚠️ WDA/idb | ✅ keyboard/mouse |
App control | ✅ | ✅ | ✅ | ✅ monkey/am | ✅ simctl/WDA | n/a |
Unicode typing | ✅ KEYEVENTF_UNICODE | ✅ xdotool type | ✅ CGEvent unicode | ⚠️ ADBKeyboard IME | ✅ WDA | ✅ |
Vision fallback | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
⚠️ = capability depends on optional tooling/signing — the server reports exactly what is missing (capabilities.notes), per the honesty rules. Platform-specific setup: docs/install/ — macOS · Windows · Linux · Android · iOS.
The UI-Venus provider (verified endpoint contract)
The default provider speaks the OpenAI chat/completions format with
chat_template_kwargs: {enable_thinking: false} and grounds natural-language
elements to [0,1000]-normalized points. Calibrated against the W8A8 build
(2026-09-28): button at pixel (1300,740) on 1920×1080 → [676, 680]; the
official grounding prompt and the live verification live in
ui-venus-service/README.md and
tests/e2e/venus-live.test.ts.
Swapping providers (UI-TARS, Qwen-GUI, any VLM): implement
ComputerVisionProvider (src/providers/types.ts)
and register it in the registry — nothing else knows which model is in use.
Development
pnpm test # full hermetic suite (unit + integration + browser E2E)
pnpm test:e2e:macos # real macOS E2E (needs permissions) — RUN_MACOS_E2E=1
pnpm test:e2e:venus # live grounding E2E against your endpoint — RUN_VENUS_LIVE=1
pnpm typecheck && pnpm buildQA status per platform (what is really verified vs mocked): docs/qa-report.md. Architecture deep-dive: docs/architecture.md.
Repository layout
src/
├── mcp/ 17 tools, server assembly
├── orchestrator/ task loop, fusion locator, verifier, loop guard, security
├── providers/ ComputerVisionProvider + UI-Venus implementation + mock
├── platforms/ adapter interface, router, macos/linux/windows/android/ios/browser
├── scripting/ YAML DSL + runner ├── recorder/ semantic recording
├── coordinate/ space transforms ├── screenshot/ pipeline (ROI/hash/JPEG)
├── core/ types, actions, state machine, errors
ui-venus-service/ endpoint contract & serving reference
tests/ unit · integration (mock+browser+mcp) · e2e (opt-in real)License
MIT — see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
- mcp-serverOAuthcom.make
Give your AI agents the tools to build, manage, and run automation workflows.