cvision-mcp
Enables background window enumeration, capture, and synthetic mouse/keyboard input on KDE Plasma using KWin scripting over D-Bus, without stealing focus or moving the host cursor.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cvision-mcpCapture the calculator window and click the equals button"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
cvision-mcp
Cross-platform Model Context Protocol (MCP) server for zero-interference background window capture and input automation across Windows, Linux (X11), and Linux (Wayland).
Core Philosophy: Zero-Interference Automation
Standard desktop automation agents (such as Anthropic Computer Use, RobotJS, and PyAutoGUI-based tools) suffer from a critical architectural limitation: they hijack the physical mouse cursor and steal active window focus. When an agent runs using those tools, your cursor jerks across the screen, your active typing session is interrupted, and full-screen video or gaming is broken.
cvision-mcp is engineered from the ground up for silent, background execution:
No Cursor Drift: Physical hardware pointer coordinates never change.
No Focus Stealing: Target windows receive synthetic events without changing the active foreground window (
GetForegroundWindow/ active Wayland seat).Background Capture: Screenshots are captured from offscreen or occluded buffers without forcing windows to the front.
Full Multitasking: Watch 4K video, type, or play games on your primary desktop while the agent works in the background.
Related MCP server: lowlevel-computer-use-mcp
Operating System Architecture
Feature | Windows | Linux (X11 / Xwayland) | Linux (Wayland / KDE) |
Window Enumeration |
|
| KWin D-Bus Scripting ( |
Screenshot Capture |
|
| Headless Spectacle + Bounding Box Crop |
Mouse Click Injection |
|
| XSendEvent (Xwayland) / Gamescope |
Keyboard Injection |
|
| XSendEvent (Xwayland) / Gamescope |
Isolated Session Mode | Dedicated Desktop Object | Virtual Display ( | Headless |
1. Windows (backend/windows.py)
Implemented entirely in standard Python
ctypes(user32.dll,gdi32.dll,kernel32.dll) with zero external C compilation.Dispatches mouse and keyboard events directly into the target window's thread message queue via
PostMessageW.Captures occluded or minimized windows directly to a memory device context (DIBSection) using
PrintWindow.
2. Linux X11 & Xwayland (backend/linux_x11.py)
Employs
python-xlibto construct synthetic protocol events (ButtonPress,ButtonRelease,KeyPress,KeyRelease).Directs events specifically to the target
WindowXID withsame_screen=1.Verified to produce zero displacement of the X11 root pointer.
3. Linux Wayland (backend/linux_wayland.py)
Wayland's security model isolates client surfaces from inter-client surveillance and global event injection.
Integrates with KDE Plasma 6 KWin scripting via D-Bus (
org.kde.KWin /Scripting) to query window geometry, PIDs, and internal IDs in under 150ms without authorization prompts.Performs background window screenshots using headless
spectacle -b -n -o <temp>and crops the target bounding box via PIL.Provides
launch_isolated_apputilizinggamescope(--headless --expose-wayland) to execute apps in a dedicated nested display where input and capture are fully accessible with 0% risk of host desktop interference.
Exposed MCP Tools
list_windows
Lists all accessible desktop windows with dimensions, coordinates, process names, PIDs, and IDs.
Parameters:
filter_title(string, optional): Substring to filter window titles.
Returns: Array of
WindowInfoobjects.
capture_window
Captures an unoccluded screenshot of the specified window and saves it to disk as a PNG image.
Parameters:
window_id(string, required): Target window ID (HWND, XID hex0x..., internal ID, or title substring).output_path(string, optional): Destination path for PNG file.
Returns: Image metadata including saved path, width, and height.
window_click
Sends a background mouse click to specific relative coordinates (x, y) inside the target window without moving the host cursor.
Parameters:
window_id(string, required): Target window ID or title.x(integer, required): X coordinate relative to the window.y(integer, required): Y coordinate relative to the window.button(string, optional): Mouse button (left,right,middle). Default:left.clicks(integer, optional): Number of clicks (1 for single, 2 for double). Default: 1.
window_type
Types text directly into the target window's message queue without shifting keyboard focus.
Parameters:
window_id(string, required): Target window ID or title.text(string, required): Text string to type.
window_send_key
Sends key press and release sequences (e.g. Enter, Escape, Backspace, Tab, Arrow keys) with optional modifier keys (Ctrl, Shift, Alt).
Parameters:
window_id(string, required): Target window ID or title.key(string, required): Key name (Enter,Escape,Tab,BackSpace,Space,Left,Right, etc.).modifiers(array of strings, optional): List of modifiers (ctrl,shift,alt).
launch_isolated_app
Launches an application in an isolated nested display session (gamescope headless mode on Linux).
Parameters:
command(string, required): Command line to execute.width(integer, optional): Virtual screen width. Default: 1920.height(integer, optional): Virtual screen height. Default: 1080.
Installation & Setup
Requirements
Python 3.8+
Pillow (
pip install Pillow)On Linux:
python-xlib(pip install python-xlib)
git clone https://github.com/sorinqu-org/cvision-mcp.git
cd cvision-mcp
pip install -r requirements.txtMCP Client Configuration
Google Antigravity / Claude Desktop Configuration
Add the server to your mcp_config.json (or claude_desktop_config.json):
{
"mcpServers": {
"cvision-mcp": {
"command": "python3",
"args": [
"/path/to/cvision-mcp/server.py"
]
}
}
}Windows Configuration
{
"mcpServers": {
"cvision-mcp": {
"command": "python",
"args": [
"C:\\path\\to\\cvision-mcp\\server.py"
]
}
}
}Protocol Verification
You can verify stdio JSON-RPC 2.0 communication directly from the terminal:
python3 -c "
import subprocess, json
proc = subprocess.Popen(['python3', 'server.py'], stdin=subprocess.PIPE, stdout=subprocess.PIPE, text=True)
proc.stdin.write(json.dumps({'jsonrpc': '2.0', 'id': 1, 'method': 'initialize'}) + '\n')
proc.stdin.flush()
print(proc.stdout.readline())
"Expected output:
{"jsonrpc": "2.0", "id": 1, "result": {"protocolVersion": "2024-11-05", "capabilities": {"tools": {}}, "serverInfo": {"name": "cvision-mcp", "version": "1.0.0"}}}License
Apache-2.0 License. See LICENSE for details.
Available Tools
6 toolscapture_windowA
Capture a screenshot of a specific window without bringing it to the foreground or stealing focus from the user. Works while the user watches video or works.
| Name | Required | Description | Default |
|---|---|---|---|
| window_id | Yes | Window ID (HWND, XID hex 0x..., or title substring). | |
| output_path | No | Optional absolute path where the screenshot PNG will be saved. Defaults to an artifact path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It does disclose the most important behavior: capturing without foregrounding or stealing focus. However, it does not mention output behavior, failure modes, or what happens with minimized or obscured windows, leaving some ambiguity about the capture result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. The primary action is front-loaded, and the key non-intrusive behavior is stated immediately, making it easy to scan and understand.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a relatively simple tool with two parameters, both fully described in the schema. The description provides the essential operational context about focus preservation. There is no output schema, so return-value details are not specified, but an agent has enough information to invoke it correctly for the intended non-disruptive screenshot use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both window_id and output_path are already well documented in the schema. The description adds no significant parameter-level meaning beyond the schema, which establishes the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Capture a screenshot') and a specific resource ('a specific window'), and adds a distinguishing behavioral detail: it does not bring the window to the foreground or steal focus. This differentiates it from all listed siblings, which concern typing, clicking, launching, listing, or window type actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when a screenshot is needed without disrupting the user, such as 'while the user watches video or works.' It does not explicitly name alternatives or exclusion conditions, but the use case is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
launch_isolated_appA
Launch an application inside an isolated background session (gamescope / virtual display). Guarantees 100% background automation, capture, and input with zero host mouse movement or focus interference.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Virtual resolution width. | |
| height | No | Virtual resolution height. | |
| command | Yes | Command line of the application to run. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses the isolated session mechanism (gamescope/virtual display), the guarantee of background automation/capture/input, and the key side-effect-free trait of zero host mouse movement or focus interference. It does not cover lifecycle details or errors, but it substantially informs an agent about runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence with no filler. Every clause adds value: isolation mechanism, automation guarantees, and the absence of host interference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple launch tool with one required parameter and only two optional sizing parameters, the description covers the essential purpose and behavioral guarantees. It does not mention return values or how to target the launched app with sibling tools, but the absence of an output schema and the tool's straightforward nature keep this from being a major gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: width, height, and command are fully documented in the input schema. The description adds general context about the isolated background session but not parameter-specific meaning beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Launch'), a clear resource ('an application'), and defines the distinguishing scope ('inside an isolated background session (gamescope / virtual display)'). It immediately differentiates this from sibling window-control tools like window_click or capture_window, which operate on existing windows.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when this tool is appropriate: when background automation, capture, and input are needed without host desktop interference. It does not explicitly name alternatives or say 'use this instead of X,' but the isolated-session framing provides enough context to select it over the sibling window interaction tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_windowsA
List all accessible desktop windows across Windows, Linux X11, and Linux Wayland with window IDs, coordinates, geometry, and process info.
| Name | Required | Description | Default |
|---|---|---|---|
| filter_title | No | Optional substring to filter window titles. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It does convey the read-only nature through the verb 'List' and mentions platform scope and returned data, but it does not explicitly say it has no side effects, nor does it disclose any permissions, accessibility limitations, or Wayland-specific caveats beyond the word 'accessible'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the action and resource, then packs in the key context: platform coverage and returned fields. There is no redundant or filler wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one optional parameter and no output schema, the description provides a solid high-level account of the return payload. It could be more complete around error cases, empty results, or output formatting, but it supplies enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes filter_title as an optional substring filter with 100% coverage, so the baseline applies. The description does not add parameter-level detail, but none is required because the schema fully documents the only parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('desktop windows'), and clearly states the scope ('all accessible desktop windows across Windows, Linux X11, and Linux Wayland') plus the returned data (window IDs, coordinates, geometry, process info). This distinguishes it from sibling tools like window_click or window_send_key, which perform actions rather than enumeration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not state when to use this tool versus alternatives, nor does it mention when not to use it. Although siblings like capture_window and window_type are clearly different in purpose, the description provides no explicit routing guidance or conditions such as 'use this when you need to enumerate windows before targeting one'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
window_clickA
Send a background mouse click at relative coordinates (x, y) inside the target window without moving the user's physical mouse cursor or stealing focus.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate relative to the window. | |
| y | Yes | Y coordinate relative to the window. | |
| button | No | Mouse button to click. | left |
| clicks | No | Number of clicks (1 for single, 2 for double). | |
| window_id | Yes | Target window ID or title. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description takes on the full burden of behavioral disclosure. It explicitly states the click is background, does not move the physical cursor, and does not steal focus. These are the key behavioral traits needed to use the tool safely, though it does not mention error cases or return values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence that front-loads the core action, then immediately states the key constraints. There is no redundant or vague filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple click action, the description plus a fully documented schema provides enough to invoke the tool correctly. It lacks an explicit return/error behavior description, which is a minor gap given the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all parameters. The description adds a small clarifying notion that x/y are 'relative coordinates inside the target window,' but otherwise does not go beyond the schema for button, clicks, or window_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Send a background mouse click') targeted at a specific resource ('inside the target window'), with unique coordinates. This immediately distinguishes it from sibling keyboard tools like window_type, window_send_key, and capture_window.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'background... without moving the user's physical mouse cursor or stealing focus' gives clear context for when this tool is appropriate: interaction that should not disturb the user. It does not explicitly name alternative tools or exclusions, but the usage context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
window_send_keyA
Send special key events (Enter, Escape, Tab, Backspace, Arrow keys, etc.) with optional modifiers to the target window.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Key name (e.g. Enter, Escape, Tab, BackSpace, Space, Left, Right). | |
| modifiers | No | Optional list of modifiers (ctrl, shift, alt). | |
| window_id | Yes | Target window ID or title. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It accurately states that it sends special key events and supports modifiers, but it does not disclose whether the window must be focused, whether keys are pressed and released, or what side effects or return values to expect. This is acceptable but not deeply transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the main action, lists useful examples, and mentions optional modifiers without any filler. Every part of the sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description and schema together cover the core invocation details, but the tool has no annotations or output schema. The absence of any guidance about focus requirements, side effects, or expected responses leaves a moderate gap for an agent deciding whether this tool will actually perform the desired action safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% description coverage for all three parameters. The description adds a few key examples and the concept of 'special' keys, but it does not provide meaningful extra semantics beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Send') with a clear resource ('special key events... to the target window') and concrete examples that distinguish it from text-oriented tools like window_type. It is immediately obvious what the tool does and how it differs from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: for special/control keys such as Enter, Escape, and arrow keys with optional modifiers. It does not explicitly name alternatives or state when not to use it, but the context is clear enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
window_typeA
Send background keyboard text typing directly into the target window without changing the active foreground window.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to type into the window. | |
| window_id | Yes | Target window ID or title. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It usefully reveals that typing happens in the background and leaves the foreground window unchanged, which is valuable and non-obvious. However, it does not mention possible limitations, error cases, or system-level requirements for background input to work.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that immediately conveys the core purpose and the key distinguishing behavior. Every word earns its place, and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, two well-documented parameters, and clear core behavior, the description is mostly complete for an agent to select and invoke it correctly. The main gaps are lack of guidance on failure behavior and any return value, but these are minor for a simple background typing operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (window_id and text) are already fully described in the input schema. The description adds no additional semantic detail about the parameters, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (send background keyboard text typing) and a specific resource (the target window). It also distinguishes itself from normal foreground typing by noting it does not change the active foreground window, which is a clear differentiation from sibling tools like window_send_key and window_click.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it should be used for typing text into a background window, but it does not explicitly explain when to choose this tool over siblings such as window_send_key or window_click, nor does it mention prerequisites like focus requirements or accessibility permissions. The usage context is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.0- First observed
capture_window - First observed
launch_isolated_app - First observed
list_windows - First observed
window_click - First observed
window_send_key - First observed
window_type
TDQS
Scored across 6 tools
Each tool targets a distinct action: listing, capturing, clicking, typing text, sending special keys, or launching an isolated app. Even the two keyboard tools are clearly separated by text input versus key events.
The naming is readable but mixes conventions: list_windows, capture_window, and launch_isolated_app start with verbs, while window_type, window_send_key, and window_click use a noun-first prefix. The window_ prefix is recognizable but not applied consistently across all window-targeting tools.
Six tools is a well-scoped set for background desktop automation. Each tool serves a clear purpose without unnecessary bloat or redundancy.
The surface covers the core workflow of discovering windows, capturing them, sending input, and launching isolated apps. Minor gaps exist, such as window management actions like moving, resizing, or closing windows, but these are not essential for the primary automation use case.
Maintenance
Related MCP Connectors
Eyes and hands on real Windows PCs — observe, click, type via Glasswarp API.
Securely control computers you explicitly pair through files, terminals, processes, screenshots, desktop UI/input, clipboard, browser automation, diagnostics, and document tools.
Build agents to automate any background task. Works with your ChatGPT/Claude subscription.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides native macOS desktop automation for AI agents, enabling screen capture, mouse/keyboard control, window management, and iOS/Android simulator control in both foreground and background modes without focus stealing.3MIT
- AlicenseAqualityAmaintenanceEnables MCP agents to automate real GUI applications on headless desktops, providing background mouse/keyboard control, window/process management, screenshots, and safe human handoff without disturbing the user's desktop.584MIT
- FlicenseBqualityCmaintenanceEnables local Windows UI automation and screen capture, targeting windows that are difficult to automate such as games and legacy apps, with tools for window management, mouse and keyboard input, and desktop capture.8-
- AlicenseNot gradedqualityAmaintenanceEnables AI assistants to autonomously control Windows 11 and 10 desktops via sub-10ms screen capture, native UI Automation element inspection, and zero-lag keyboard and mouse input. Combines a visual plane with a semantic plane and stall detection so agents can operate real applications reliably without vision-only guessing.223 npm4MIT