kasm-use
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@kasm-useOpen Chrome and search for today's Bitcoin price on CoinMarketCap."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
kasm-use
Let an AI agent use a Kasm Workspaces desktop: start a session, look at the screen, click, type, press keys, scroll, and stop it — in plain language, from any MCP client (Claude Code, Claude Desktop, OpenCode, Cursor, …) or as a native Hermes Agent plugin.
Other Kasm MCP servers manage sessions (start, list, stop). kasm-use actually uses them: the agent sees the desktop and operates it, like computer-use or browser-use, inside your Kasm.
Works with stock Kasm images. Nothing is baked into the workspace; it drives the desktop through Kasm's own API.
You can watch and take over. Sessions are created as your Kasm user, so they show up in your Kasm dashboard. Open one to watch the agent work, or to handle a login or CAPTCHA yourself.
Your Kasm, your key. You run the server next to your own Kasm. Nothing goes through a third party.
Tools
Tool | What it does |
| Workspaces you can start, and your running sessions |
| Start a workspace (e.g. |
| Screenshot of the session, returned as an image |
| Click at x,y in the latest screenshot's coordinates |
| Type text into the focused field |
| Keys/shortcuts, e.g. |
| Scroll, optionally at a point |
| Destroy the session |
Related MCP server: AgentViewport
Requirements
Kasm Workspaces 1.19 or newer (the screenshot API needs 1.19 images).
A Kasm API key: Admin → Settings → Developers → API Keys → Add. Grant it only the session permissions (create/list/destroy sessions, screenshot, exec). It does not need user or admin rights.
Your Kasm user ID, so sessions belong to you: Admin → Access Management → Users → open your user; the ID is in the page URL.
A workspace image based on Debian/Ubuntu (all official
kasmweb/*images are).
Configuration
Variable | Required | Meaning |
| yes | e.g. |
| yes | the API key pair |
| yes | the Kasm user sessions are created for |
| no |
|
| HTTP only | bearer token clients must send |
Run it
Locally (stdio) — the usual way
Your MCP client starts the server itself. With uv installed:
{
"mcpServers": {
"kasm": {
"command": "uvx",
"args": ["kasm-use"],
"env": {
"KASM_API_URL": "https://kasm.example.com",
"KASM_API_KEY": "…",
"KASM_API_KEY_SECRET": "…",
"KASM_USER_ID": "…"
}
}
}
}Claude Code:
claude mcp add kasm -e KASM_API_URL=https://kasm.example.com -e KASM_API_KEY=… -e KASM_API_KEY_SECRET=… -e KASM_USER_ID=… -- uvx kasm-useAs a service (streamable HTTP)
docker run -d -p 8000:8000 \
-e KASM_API_URL=https://kasm.example.com -e KASM_API_KEY=… -e KASM_API_KEY_SECRET=… \
-e KASM_USER_ID=… -e KASM_USE_TOKEN="$(openssl rand -hex 32)" \
ghcr.io/rchurro/kasm-use:latestClients connect to http://host:8000/mcp with the header Authorization: Bearer <KASM_USE_TOKEN>.
Health check: GET /healthz. Run one replica per Kasm user: the server remembers each
session's last screenshot size in memory to map click coordinates.
Keep it on a private network (LAN, VPN, Tailscale). Anyone with the token can drive your desktops.
Hermes Agent
pip install kasm-use # into Hermes's Python environment
cp -r hermes-plugin/kasm ~/.hermes/plugins/kasmEnable kasm under plugins.enabled, set the environment variables above, and start a new
Hermes session (/new) so the tools load.
Limitations
First start is slow (~90 s extra). Stock images don't include
xdotool, sokasm_startinstalls it in the session as root. Reusing a running session skips this.Screenshots can lag a few seconds behind actions; the agent is told to look again to confirm.
Kasm's exec API returns no output, so actions report "sent", not "succeeded". The next screenshot is the confirmation.
CAPTCHAs, passwords, MFA and payments are handed to you. The tool descriptions tell the agent never to type them; open the session in Kasm and take over.
It's slow compared to a local browser tool: every step is a screenshot round trip.
How it works
Eyes:
POST /api/public/get_kasm_screenshot(KasmVNC renders a JPEG).Hands:
POST /api/public/exec_command_kasmrunningxdotoolinside the session.Clicks are given in screenshot pixels and scaled to the real desktop size inside the session (
xdotool getdisplaygeometry), using the JPEG's actual dimensions — Kasm keeps the desktop's aspect ratio, so the image is often not the size requested.
License
MIT
Available Tools
8 toolskasm_clickB
Click at x,y in the coordinate space of the latest kasm_look image.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| button | No | ||
| double | No | ||
| session_id | No | Kasm session id. Omit to use your newest running session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the coordinate frame dependency but says nothing about what a click does to the remote session, whether it is destructive, how button defaults are chosen, or how double-click timing is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the most important constraint (which coordinate space) is stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 5-parameter click tool with no output schema and no annotations, the description covers the core coordinate semantics but omits button/double-click behavior and session targeting, leaving gaps an agent must guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (only session_id is documented). The description adds real meaning for x/y by tying them to the kasm_look image coordinate space, but the button enum, the double flag, and the session-omission behavior are left entirely unexplained in both description and schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Click) and the exact coordinate reference frame (the latest kasm_look image), which differentiates it from kasm_type, kasm_key and kasm_scroll. However it doesn't explicitly state it acts on a remote Kasm session, so full sibling disambiguation relies on the name prefix.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'coordinate space of the latest kasm_look image' implies the prerequisite that kasm_look must be called first, which is useful implicit guidance. But there is no statement of when to prefer this over kasm_scroll or kasm_key, and no exclusions or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kasm_keyC
Press keys/shortcuts in the Kasm session, e.g. 'ctrl+l', 'Return', 'Tab', 'Escape'.
| Name | Required | Description | Default |
|---|---|---|---|
| keys | Yes | ||
| session_id | No | Kasm session id. Omit to use your newest running session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose where the keys are sent (focused element vs. session), whether it blocks or returns a result, what happens if keys are invalid, or any failure/auth behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, no filler, with the key-format examples front-loaded where an agent needs them. Efficient, though the example list could be trimmed or the tool distinction added without cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter action tool with no output schema, the description is minimally adequate: it conveys the action and the key format. The notable gap is routing versus kasm_type, which an agent must guess at, leaving the definition just barely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%: session_id is documented in the schema (including the useful 'omit to use newest session' default), but keys is undocumented there. The description compensates by giving concrete key-string examples ('ctrl+l', 'Return', 'Tab', 'Escape'), which conveys modifier syntax that the plain 'string' type does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
A specific verb (press) plus the resource acted on (keys/shortcuts in a Kasm session), with concrete examples of the key syntax. It does not differentiate itself from the sibling kasm_type, which is the closest and most confusable tool in the list, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to use this tool versus kasm_type (typing text) or kasm_click. No prerequisites, no focus/context requirements, and no when-not-to-use guidance are given; the examples imply usage but do not state it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kasm_listB
List Kasm workspaces you can start and your running sessions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It conveys that this is a read-only listing returning two distinct sets (startable workspaces and running sessions), which is useful. It omits authentication requirements, pagination, and whether results are scoped to the user.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with no filler, front-loading the verb and the two result categories. Nothing is wasted, though it is arguably too terse to be maximally helpful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter listing tool with no output schema and no annotations, the description covers what is returned but not how the two sets relate, whether kasm_start consumes the listed items, or any auth context. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty, so there is nothing for the description to clarify. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'List' plus two resources: Kasm workspaces available to start and currently running sessions. This is meaningfully more specific than the sibling tools (kasm_start, kasm_stop), though it doesn't name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or routing to alternatives is given. An agent can infer it is a discovery call, but there is no statement of prerequisites, ordering relative to kasm_start, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kasm_lookA
Screenshot the Kasm session so you can see it. Always look before clicking, and look again after each action to confirm it worked.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Kasm session id. Omit to use your newest running session. | |
| wait_seconds | No | Pause before capturing (default 2). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the full burden. It states the core action but does not disclose output format, read-only nature, or other behavioral traits beyond the basic screenshot purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by usage advice. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple screenshot tool with two parameters, no output schema, and no annotations, the description covers what it does and when to use it. It omits return details, but the purpose is clear enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are fully documented in the schema. The description adds no parameter-specific meaning beyond what the schema provides, making 3 the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Screenshot') and resource ('the Kasm session'), and its purpose ('so you can see it'). It implicitly distinguishes itself from sibling tools like kasm_click by instructing to look before clicking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use: before clicking and after each action. No explicit when-not-to-use or named alternatives, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kasm_scrollB
Scroll the Kasm session, optionally at x,y from the latest kasm_look image.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| amount | No | Wheel clicks, 1-30 (default 5) | |
| direction | No | ||
| session_id | No | Kasm session id. Omit to use your newest running session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose the default direction, the default amount (5 wheel clicks, only in the schema), whether scrolling is relative to the current viewport, or what happens if the coordinate is outside the captured image.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence, front-loaded with the verb and resource. No wasted words, though it is arguably too terse given the gaps it leaves open.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, annotation-free tool with no output schema and only 40% schema coverage, the description is thin. Key behavioral facts an agent needs — direction default, coordinate origin/reference frame, session continuity — are left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% — x and y have bare integer types with no schema text, and the description usefully clarifies they refer to coordinates on the latest kasm_look image. It adds nothing about the direction parameter (enum with no stated default) or the amount range beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scroll) and resource (Kasm session), and the optional x,y anchor ties it to the kasm_look workflow. It does not contrast itself with siblings like kasm_click, but the purpose is immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'from the latest kasm_look image' implies the intended workflow: look first, then scroll at a coordinate. However, there is no explicit when-to-use/when-not guidance, no mention of the direction default, and no note on staying in the same session as look/click.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kasm_startA
Start a Kasm workspace session (e.g. 'Chrome', 'Terminal') as the owner, so they can watch it. Takes ~2 minutes (it installs the input tool). Then call kasm_look.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace | No | Workspace name, e.g. Chrome |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does useful work: it warns the operation takes ~2 minutes because it installs the input tool, and states the session runs under the owner's identity. It omits what happens if a session is already running or what failure modes exist, but the latency and side-effect disclosure is genuinely valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and immediately followed by the cost estimate and the required next step. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter startup tool with no output schema, the description covers what it does, how long it takes, and what to call next. It could mention behavior on repeat invocation, but an agent has enough to act correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the single 'workspace' parameter with an example. The description's inline examples ('Chrome', 'Terminal') largely duplicate that, adding no new format or constraint information. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a Kasm workspace session') with concrete examples ('Chrome', 'Terminal'). It implicitly separates itself from siblings by naming kasm_look as the follow-up, though it does not explicitly contrast with kasm_stop or kasm_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context ('as the owner, so they can watch it') and an explicit sequencing instruction: 'Then call kasm_look.' No when-not-to-use guidance or alternatives are offered, but the invocation path is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kasm_stopA
Stop (destroy) a Kasm session when the task is finished or the owner asks.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Kasm session id. Omit to use your newest running session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. Calling out 'destroy' usefully signals the operation is destructive and terminal, but it says nothing about irreversibility, what happens to session state/output, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and the usage cue trailing. No filler and nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, likely terminal tool with no output schema, the description covers what it does and when to use it. It is slightly thin on the consequences of destruction, which an agent deciding to kill a session might want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already explains the session_id default ('omit to use your newest running session'). The description adds no additional parameter meaning beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('stop') plus the resource ('Kasm session'), and the parenthetical 'destroy' removes ambiguity about what stopping entails. It is immediately distinguishable from the sibling kasm_start.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear triggering condition ('when the task is finished or the owner asks'), which is enough to know when to reach for this over the other session tools. It does not name an alternative or state when-not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kasm_typeA
Type text into the focused field of the Kasm session. NEVER type passwords, MFA codes, or payment details — stop and ask the owner to take over the session for those.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| session_id | No | Kasm session id. Omit to use your newest running session. | |
| press_enter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the sensitive-data prohibition and the required human takeover, but says nothing about side effects on the live session, whether text submission occurs, or what happens if no field is focused.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, well front-loaded: the action first, then the safety constraint. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation-style tool with no annotations and no output schema, the safety handoff is covered, but the undocumented press_enter parameter and the absence of any note on session state or failure behavior leave real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%: only session_id is documented in the schema. The description never explains text semantics or the press_enter boolean that presumably submits the input, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Type text into the focused field of the Kasm session.' This clearly separates it from siblings like kasm_key (keystrokes) and kasm_click, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit when-not rule with a fallback action: never type passwords, MFA codes, or payment details, and hand off to the owner instead. It lacks positive guidance on focusing a field first or when to prefer kasm_key, but the exclusion is clearly actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
kasm_click - First observed
kasm_key - First observed
kasm_list - First observed
kasm_look - First observed
kasm_scroll - First observed
kasm_start - First observed
kasm_stop - First observed
kasm_type
TDQS
Scored across 8 tools
Each tool maps to a distinct session action: lifecycle management, visual inspection, clicking, scrolling, text entry, and key presses. kasm_type and kasm_key both provide keyboard input, but they are clearly separated by text entry versus special keys/shortcuts, so the overlap is minimal.
All eight tools follow the same kasm_ prefix followed by a short action-oriented name. This creates a predictable and readable naming convention with no mixed styles.
Eight tools is well-scoped for a remote Kasm session controller, covering lifecycle, observation, and basic input operations. Each tool serves a clear purpose in the interaction loop without obvious bloat.
The set covers starting, listing, and stopping sessions, plus screenshotting, clicking, scrolling, typing, and key presses, which supports many basic remote-control tasks. However, common UI interactions such as mouse move/hover, drag-and-drop, right-click, and clipboard/file transfer are missing, leaving minor gaps that an agent must work around.
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Automate 1,000+ services from any MCP-compatible AI agent: build Applets, run actions and queries.
Connect any AI agent to 1,000+ apps and 27,000+ actions through one remote MCP server (OAuth).
Manage SRG+ hubs, channels, content, assets, users, and workspaces from any MCP-aware AI agent.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to manage and interact with Kasm Workspaces containerized desktop infrastructure, providing tools for session management, command execution, file operations, and user management via the Model Context Protocol.4MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to view and control the local desktop through browser-based viewport and MCP tools.MIT
- AlicenseAqualityBmaintenanceAllows AI clients to see and control Windows 10/11 desktops via MCP, with screenshots, UI Automation, Chrome CDP, keyboard/mouse, and terminal using semantic element targeting.30548 npmMIT
- AlicenseAqualityAmaintenanceEnables MCP agents to automate real GUI applications on headless desktops, providing background mouse/keyboard control, window/process management, screenshots, and safe human handoff without disturbing the user's desktop.584MIT