open-compute-mcp
Enables screen capture and computer-use control of Blender on Windows, including GPU-composited Blender windows via the Windows.Graphics.Capture fallback.
Enables screen capture and computer-use control of Roblox Studio on Windows, including GPU-composited Roblox Studio windows via the Windows.Graphics.Capture fallback.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@open-compute-mcpcapture screen"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
open-compute-mcp
npm launcher for the open-compute MCP server β model-agnostic computer-use tools exposed over the Model Context Protocol (MCP).
EN | DE
π¦ View on npm β β’ π Security Policy β’ βοΈ Licenses β’ π€ LLM Context (llms.txt)
Quick Navigation
AI Assistant / Agent Integration: This repository contains an llms.txt file providing structured, machine-readable specifications of tools, safety modes (OC_SAFETY_MODE), and client configuration examples for RAG crawlers and autonomous agent frameworks.
The MCP client is the reasoner (no API key, model-agnostic): it calls capture
to see the screen, then acts with do / click_name / invoke. This is the keyless
Mode-A loop of open-compute, but as native tool-calls.
Key Capabilities
State-bound Perception & Window Targeting: Captures/trees return one-shot observation IDs; window enumeration returns stable window/process IDs and issued tokens. WGC remains the GPU-window fallback.
Fail-closed Action Execution: Coordinates consume one observation and exact window binding; UIA names resolve exact-first; text is segmented with focus checks and character-count postconditions.
Leased Signal Overlay & Abort Control: The glowing border/cursor signal has owner/session metadata, a bounded TTL, turn-end cleanup, and immediate human abort.
Multimodal Collaboration & Voice Notes: Push-to-talk voice recording (
talk), screen chat messaging (chat), directory monitoring (watch_dir), and macro replay (rec_replay).
Related MCP server: ScreenPilot
Architecture
graph TD
A["AI Reasoner<br/>(Claude / Antigravity / Cursor)"] -- "MCP stdio (JSON-RPC)" --> B["npx open-compute-mcp<br/>(Node.js Launcher)"]
B -- "Spawns via uvx" --> C["open-compute Python Engine<br/>(GitHub @ main)"]
C -- "Screenshots / WGC" --> D["Windows Display"]
C -- "UIA / Mouse / Keys" --> E["Windows Desktop Apps"]
C -- "Glowing Border & Cursor" --> F["Signal Overlay UI"]
subgraph Safety Gate
C -. "OC_SAFETY_MODE<br/>(confirm / read_only / allow_all)" .-> C
C -. "OC_DENY<br/>(hard action blacklist)" .-> C
endThis package is a thin launcher. It contains no server logic β it spawns the Python open-compute server (pulled from GitHub) and pipes MCP stdio through. Real screen capture and input require the interactive Windows desktop session.
Requirements
Python 3.10+ and uv on the host. The default launch uses
uvxto fetch open-compute (with themcpextra) from GitHub on first run β themcpextra tracks the GitHub repo, so this works regardless of PyPI release timing.Windows for real capture/input (mss + UIA). Other platforms import the tools but cannot drive a desktop.
Tools
Tool | Purpose |
| Return one-shot observation metadata plus an image (optionally one exact window). |
| Execute a safety-gated action; coordinates require |
| Return UIA elements and a one-shot observation ID for their coordinates. |
| Exact-first, ambiguity-safe click in a required issued window, with score/alternatives. |
| Exact-first, click-free UIA activation in a required issued window. |
| List stable window/process IDs, exact titles, issued tokens, rects and centers. |
| Virtual-desktop geometry + per-monitor breakdown (read-only). |
| Watch directories for file-system changes. |
| Feed-manager status (read-only). |
| Replay a |
| Show a configurable pre-action color/text countdown, then the mode-colored overlay, with owner/session lease and bounded TTL. |
| Hide the signal overlay. |
| Owner/session/mode/visible/expires_at + pending abort message. |
| Ask the human for a short abort reason; the message is returned for the model. |
| Humanβmodel message about screen content, optionally with screenshot. |
| Push-to-talk voice note β WAV path (hold key, speak, release; STT/TTS model-side). |
All coordinates are normalized 0..1 relative to the virtual desktop. Tool
descriptions are localized in six languages (de/en/es/ja/ru/zh) via OC_LANGUAGE.
do also accepts the hold primitives mouse_down / mouse_up / key_down /
key_up for press-and-hold sequences (rubber-band selection, modifier-held
clicking, game input); anything still held is released when the server stops.
capture(window=...) falls back to Windows.Graphics.Capture when a plain grab of
a hardware-composited window (Roblox Studio, Blender, a GPU-accelerated browser)
comes back all-black β install the wgc extra for that.
Safe Interaction & Signal Lifecycle
The v0.8 Python engine enforces observe β one action β automatic refresh. Keep
the full descriptor or window_token from list_windows, then pass it as
expected_window together with the latest observation_id from capture or
tree. click_name/invoke require that issued window too. Reuse, changed
state, focus mismatch, covered windows, and ambiguous UIA targets are rejected
before input. type returns requested/sent character
counts and complete/partial status without echoing the text. Signals have a
hard TTL and are removed at action turn end unless keep_signal=true.
An explicit signal_show starts the engine's configured pre-action grace
period. The static grace color is distinct from the mode color and the visible
text counts down Start in N Sekunden once per second. At zero, both phase and
color switch once to active. signal_status exposes the same phase, remaining
seconds, current color, and screenreader label. Duration, grace color, and text
template come from OC_SIGNAL_GRACE_SECONDS / OC_SIGNAL_CONFIG; 0 skips the
countdown. The design uses no flashing, pulsing, or motion animation.
sequenceDiagram
autonumber
actor Reasoner as AI Reasoner (Claude / AGY)
participant Launcher as Node.js Launcher (open-compute-mcp)
participant Engine as Python Engine (open-compute)
participant UI as Windows Desktop / UIA
actor Operator as Human Operator
Note over Reasoner,Operator: Phase 1: Visual Perception & State Inspection
Reasoner->>Launcher: capture(window?) / tree()
Launcher->>Engine: Forward stdio JSON-RPC
Engine->>UI: Grab Screen (mss/WGC) or Read UIA Tree
UI-->>Engine: Frame Image / Semantic Element Tree
Engine-->>Launcher: Observation ID + normalized response/image
Launcher-->>Reasoner: State-bound visual observation
Note over Reasoner,Operator: Phase 2: Signal Overlay Activation
Reasoner->>Launcher: signal_show(mode="control")
Launcher->>Engine: Invoke Signal Overlay
Engine->>UI: Render static grace color + Start in N seconds
UI-->>Operator: Text countdown + accessible window name
Engine->>UI: At zero, switch once to the mode color
Note over Reasoner,Operator: Phase 3: Action Request & Safety Gate
Reasoner->>Launcher: do(one action, window token, observation_id) / click_name(target)
Launcher->>Engine: Process Action Payload
alt OC_SAFETY_MODE == "confirm" (Default)
Engine-->>Launcher: Status "needs_confirmation" (Report Only)
Launcher-->>Reasoner: Human confirmation needed
else OC_SAFETY_MODE == "allow_all" (Isolated VM)
Engine->>UI: Execute Mouse/Keyboard / Hold Primitives
UI-->>Engine: Action Completed
Engine-->>Launcher: Post-observation + window/modal/text postconditions
Launcher-->>Reasoner: Action completed; old observation invalid
end
Note over Reasoner,Operator: Phase 4: Emergency Abort or Completion
opt Operator Triggers Emergency Abort
Operator->>Engine: Hotkey Pressed (Abort Signal)
Engine->>UI: Auto-release all held keys/mouse buttons
Engine-->>Reasoner: signal_abort message returned
end
Engine->>UI: Remove overlay on turn end/error/abort (unless keep_signal=true)Use with an MCP client
Via this npm launcher (npx):
{
"mcpServers": {
"open-compute": {
"command": "npx",
"args": ["-y", "open-compute-mcp"]
}
}
}Directly via Python (uvx), no npm:
{
"mcpServers": {
"open-compute": {
"command": "uvx",
"args": ["--from", "open-compute[mcp,local,uia] @ git+https://github.com/ellmos-ai/open-compute.git", "open-compute-mcp"]
}
}
}Configuration (environment variables)
Variable | Effect |
| Path to a |
| Full command override (whitespace-split), e.g. |
| Git ref (branch/tag/sha) to pin for the uvx launch (default: the repo's default branch). |
| Extras for the default |
| Language of the tool descriptions: |
|
|
| Comma-separated action types always denied (e.g. |
| Resize factor for every capture, |
| Cap the longest edge in pixels (default off). Setting it suppresses the scale default, so the two never shrink twice. |
|
|
| Hard overlay lease limit in seconds (default 120). |
| Additional idle timeout for explicitly kept auto-signals (default 60). |
| Pre-action countdown duration (default 20; |
| Signal JSON containing |
Capture size β why this launcher halves it by default
A vision model is billed per pixel, and every frame stays in the conversation, so a full-HD grab is charged again on each following request. The cost of a session therefore grows with the square of the number of screenshots, not linearly.
Because open-compute's coordinates are normalized 0..1, shrinking the image costs
nothing in click accuracy β do works in fractions of the image either way. Only
legibility drops, and at 0.5 buttons and field borders stay clearly identifiable; small
body text is what gets hard to read.
Setting | 1920Γ1080 grab | Cost |
| full resolution | ~1600 tokens |
| 960Γ540 | ~690 tokens |
| 768Γ432 | ~440 tokens |
The Python library itself defaults to full resolution β its callers are not necessarily paying per pixel. Only this launcher, which exists to serve agents, opts into the smaller frame and prints a one-line notice when it does.
What saves more than any scale factor: prefer tree where
the accessibility model carries the content β note that in browsers it usually exposes only
the browser chrome, not the page; and use capture(window=β¦) rather than the full desktop.
Coordinate actions deliberately follow observe β one action β automatic refresh;
do not batch multiple coordinate steps against one stale frame.
Safety
Computer-use is powerful. OC_SAFETY_MODE is an operator ceiling (confirm
default Β· read_only Β· allow_all); a per-call mode can only tighten it, never
loosen it. Because MCP stdio has no serverβclient confirm callback, confirm /
read_only report an action without performing it. For interactive use, run in
an isolated VM/session, set OC_SAFETY_MODE=allow_all, and let your client's
tool-approval dialog be the human-in-the-loop. OC_DENY (comma-separated action
types) is a hard deny list. Treat on-screen content as untrusted (prompt-injection
risk).
Troubleshooting: do/click_name only ever return needs_confirmation and never
act. That is the confirm ceiling working as designed under stdio MCP. Fix for
interactive use: set "env": {"OC_SAFETY_MODE": "allow_all"} in the server
registration and let the client's tool-approval dialog gate each action (do not
auto-allow do/click_name/invoke there). The env change only takes effect when
the server process (re)starts β an already-connected client keeps the old ceiling
until it reconnects.
License
MIT β see LICENSE. Part of the open-compute project.
ellmos-ai Ecosystem
This MCP server is part of the ellmos-ai ecosystem β AI infrastructure, MCP servers, and intelligent tools.
MCP Server Family
Server | Tools | Focus | npm |
46 | Filesystem, process management, interactive sessions, cloud-lock-safe operations | ||
22 | Code analysis, JSON repair, imports, diffs, regex | ||
12 | File repair, format conversion, batch operations | ||
18 | n8n workflow management via AI assistants | ||
20 | MCP stack discovery, profile management, control plane | ||
45 | Local-first LLM memory, knowledge, state, routing, swarm orchestration |
| |
8 | Server operations: health checks, log analysis, deploy dry-runs, mail diagnostics |
| |
3 | Headless Blender asset QA and FBX reimport verification |
| |
16 | Model-agnostic computer use: capture, safety-gated actions, Windows UIA, signal overlay & voice/chat |
|
AI Infrastructure & Sibling Tooling
Project | Description |
Local-first text-based OS for LLM agents β 113+ handlers, 550+ tools, SQLite memory | |
Model-agnostic computer-use core powering Open Compute MCP | |
Provider-neutral LLM orchestration with auto-routing and budget tracking | |
Lightweight agent memory, connectors, and automation infrastructure | |
Self-hosted AI research stack (Ollama + n8n + Rinnsal + KnowledgeDigest) | |
Autonomous agent chain framework for Claude Code | |
Minimalist database-driven LLM OS prototype (4 functions, 1 table) | |
Testing framework for LLM operating systems (7 dimensions) | |
Safe, redacted, HMAC-verified SQLite snapshot synchronizer | |
Hierarchical policy & delegation authority engine |
Open Bricks Umbrella
Our partner organization open-bricks bundles AI-native desktop applications β a modern, open-source software suite built for the age of AI. Sibling suites include DevCenter, CodeBox, MethodenAnalyser, CleanMarkdown, and PDFtoPDFocr.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
No tool schema history has been recorded yet.
This server cannot be installed
Maintenance
Related MCP Connectors
- WauldoOAuthcom.wauldo
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page β 5 MCP tools for AI agents.
Free public MCP for AI agents β 193 tools, 44 workflows. No API key.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceThe local MCP server that gives any AI agent safe desktop control. Provides 6 compact tools (computer, accessibility, window, system, browser, task) for cross-platform GUI automation with ground-truth verification.55400MIT
- FlicenseNot gradedqualityCmaintenanceEnables LLMs to take full control of your device by providing screen automation tools for capturing, clicking, typing, and scrolling, ideal for automation and education.54-
- AlicenseAqualityDmaintenanceEnables LLM agents to capture screenshots, control mouse/keyboard, and manage windows on desktop platforms, primarily Windows, via an MCP server.161MIT
- AlicenseNot gradedqualityCmaintenanceA framework-agnostic computer-use MCP server that exposes core desktop operations (screen capture, mouse, keyboard, and file access) as standard MCP tools, enabling any MCP-compatible agent to drive a computer.327MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ellmos-ai/open-compute-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server