screen-use
Enables automation of SAP client applications through natural language, using accessibility tree and optional vision models to locate and interact with UI elements.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@screen-useOpen Calculator, compute 123+456, then paste result into Notepad"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🖥️ screen-use
browser-use, but for the entire desktop.
Give any AI Agent eyes 👀 and hands 🖐️ on Windows — let Claude, Kimi, Cursor or your own agent see the screen, find UI elements, and operate any desktop app through natural language. No selectors. No scripts that break when the UI changes.

👆 An AI agent computing 123 + 456 in Calculator, then pasting the result into Notepad — two apps, zero hardcoded selectors, fully autonomous.

Why
Traditional RPA records selectors — and breaks the moment a page changes. browser-use (32k⭐) solved this for browsers. screen-use brings the same idea to the entire desktop: Excel, SAP clients, ERP software, even legacy Win32 programs.
Traditional RPA | screen-use | |
Locating elements | Recorded selectors, break easily | Understands UI via Accessibility tree + Vision models |
Scope | Browser or specific apps only | Any desktop app |
Authoring | Professional developers | Natural language |
Cost | Expensive enterprise software | Open source, local-model friendly |
How it works
Your Agent (Claude / Kimi / Cursor / custom) ← does the planning
│ MCP or Python SDK
▼
┌─────────────────────────────────────────────┐
│ screen-use │
│ Visual Loop ──► observe→think→act→verify │
│ Introspection──► difficulty playbook │
│ Meta-learning──► experience & vocab memory │
│ Perception ──► UIA tree + screenshots (SoM)│
│ Locating ──► strategy chain: │
│ ⓪ learned vocab mapping │
│ ① UIA text match (0 cost) │
│ ② Set-of-Mark + VLM │
│ Action ──► mouse / keyboard │
└─────────────────────────────────────────────┘VLM is optional, not required. The locating strategy chain hits most targets with pure Accessibility-tree text matching — zero model calls, millisecond latency. A vision model (cloud or local via Ollama) only kicks in for UIA-blind UIs.
Quickstart
git clone https://github.com/tongriyaotxt/screen-use.git
cd screen-use
pip install -r requirements.txtAs an MCP Server (recommended)
Add to claude_desktop_config.json (or any MCP-compatible agent's config):
{
"mcpServers": {
"screen-use": {
"command": "python",
"args": ["-m", "screen_use.mcp_server"],
"cwd": "path/to/screen-use"
}
}
}Then just tell your agent: "Open Calculator and compute 123 × 456."
As a Python SDK
from screen_use import ScreenUse
tools = ScreenUse()
tools.click_element("Save") # locate + click, one call
tools.type_text("Hello, 你好") # Unicode-safe (clipboard paste)
tools.hotkey("ctrl", "s")
# Atomic tools for vision-capable agents:
elements = tools.list_ui_elements() # id, name, type, bbox — no model needed
shot = tools.screenshot(annotate=True) # Set-of-Mark annotated screenshot
tools.click(500, 300)Autonomous task loop
One call, full autonomy — the agent sees, decides, acts and self-corrects:
tools.run_task("打开计算器,算 25 乘以 4") # observe → think → act → verifyIntrospection (困难分类反思): when the loop gets stuck, it classifies the difficulty — no effect / repeat loop / consecutive failures / missing elements / unexpected popup — and reflects with a targeted prompt playbook, then adjusts strategy.
Meta-learning (元学习): successful runs are remembered. Similar past tasks are recalled as experience hints, and learned vocabulary mappings (e.g. "乘号" → Multiply by) become the strategy chain's new first level. It literally gets better the more you use it. Memory lives in ~/.screen_use/.
Tools (14)
Atomic (zero model dependency): screenshot · list_ui_elements · click · double_click · right_click · click_element_id · type_text · hotkey · press · scroll
High-level: find_element (strategy-chain locating) · click_element (locate + click) · read_screen (VLM screen Q&A) · run_task (autonomous visual loop)
Vision model (optional)
Only needed when your agent has no vision AND the target app is UIA-blind. Copy .env.example to .env:
Preset | Config | Models |
Local (free, private) |
| qwen3-vl, qwen2.5vl, llama3.2-vision |
OpenAI |
| gpt-4o |
Qwen |
| qwen-vl-max |
Without any VLM configured, atomic tools and UIA matching still work fully.
Safety
🚨 Failsafe: slam your mouse to the top-left corner to abort instantly
✅
confirm_callbackhook to approve every action (SDK)🧪
ScreenUse(dry_run=True)records actions without executing
Roadmap
UIA + SoM locating strategy chain
MCP Server (14 tools)
Local VLM support (Ollama)
Autonomous visual loop (
run_task)Introspection playbook & meta-learning memory
wait_for_element/ auto-verification primitivesDrag & drop
VLM raw-coordinate fallback + OpenCV template matching (UIA-blind apps)
macOS (Accessibility API) & Linux support
PyPI release
Contributions welcome — see issues for good first tasks.
Development
pytest tests -q # 61 unit tests, no desktop/VLM needed
python examples/demo_calculator.py # end-to-end demo (real clicks!)
python examples/mcp_client_demo.py # MCP handshake + tool listLicense
MIT
screen-use = 桌面版 browser-use:让任何 AI Agent 获得看屏幕、操作桌面应用的能力。
不是传统 RPA:不录制 selector,通过无障碍树 + 视觉模型理解 UI,界面变了也不怕
跨一切桌面应用:Excel、SAP、ERP 客户端、老旧 Win32 程序
自然语言驱动:
click_element("保存按钮")一句话搞定VLM 可选:策略链第一级是纯 UIA 文本匹配(零模型、毫秒级),视觉模型只在盲区兜底,支持本地 Ollama 保护隐私
接入方式、工具列表、安全配置与上文英文版一致。
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Turns any agent into a full agentic application — branded, interactive screens generated at runtime.
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
Let ChatGPT, Claude & Cursor use your Mac: email, calendar, iMessage, Teams, files. Local, free.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/tongriyaotxt/screen-use'
If you have feedback or need assistance with the MCP directory API, please join our Discord server