Skip to main content
Glama

kasm-use

Let an AI agent use a Kasm Workspaces desktop: start a session, look at the screen, click, type, press keys, scroll, and stop it — in plain language, from any MCP client (Claude Code, Claude Desktop, OpenCode, Cursor, …) or as a native Hermes Agent plugin.

Other Kasm MCP servers manage sessions (start, list, stop). kasm-use actually uses them: the agent sees the desktop and operates it, like computer-use or browser-use, inside your Kasm.

  • Works with stock Kasm images. Nothing is baked into the workspace; it drives the desktop through Kasm's own API.

  • You can watch and take over. Sessions are created as your Kasm user, so they show up in your Kasm dashboard. Open one to watch the agent work, or to handle a login or CAPTCHA yourself.

  • Your Kasm, your key. You run the server next to your own Kasm. Nothing goes through a third party.

Tools

Tool

What it does

kasm_list

Workspaces you can start, and your running sessions

kasm_start

Start a workspace (e.g. Chrome) — takes ~2 min, see Limitations

kasm_look

Screenshot of the session, returned as an image

kasm_click

Click at x,y in the latest screenshot's coordinates

kasm_type

Type text into the focused field

kasm_key

Keys/shortcuts, e.g. ctrl+l, Return, Tab

kasm_scroll

Scroll, optionally at a point

kasm_stop

Destroy the session

Related MCP server: AgentViewport

Requirements

  • Kasm Workspaces 1.19 or newer (the screenshot API needs 1.19 images).

  • A Kasm API key: Admin → Settings → Developers → API Keys → Add. Grant it only the session permissions (create/list/destroy sessions, screenshot, exec). It does not need user or admin rights.

  • Your Kasm user ID, so sessions belong to you: Admin → Access Management → Users → open your user; the ID is in the page URL.

  • A workspace image based on Debian/Ubuntu (all official kasmweb/* images are).

Configuration

Variable

Required

Meaning

KASM_API_URL

yes

e.g. https://kasm.example.com

KASM_API_KEY / KASM_API_KEY_SECRET

yes

the API key pair

KASM_USER_ID

yes

the Kasm user sessions are created for

KASM_VERIFY_TLS

no

false to accept a self-signed Kasm certificate (default true)

KASM_USE_TOKEN

HTTP only

bearer token clients must send

Run it

Locally (stdio) — the usual way

Your MCP client starts the server itself. With uv installed:

{
  "mcpServers": {
    "kasm": {
      "command": "uvx",
      "args": ["kasm-use"],
      "env": {
        "KASM_API_URL": "https://kasm.example.com",
        "KASM_API_KEY": "…",
        "KASM_API_KEY_SECRET": "…",
        "KASM_USER_ID": "…"
      }
    }
  }
}

Claude Code:

claude mcp add kasm -e KASM_API_URL=https://kasm.example.com -e KASM_API_KEY=… -e KASM_API_KEY_SECRET=… -e KASM_USER_ID=… -- uvx kasm-use

As a service (streamable HTTP)

docker run -d -p 8000:8000 \
  -e KASM_API_URL=https://kasm.example.com -e KASM_API_KEY=… -e KASM_API_KEY_SECRET=… \
  -e KASM_USER_ID=… -e KASM_USE_TOKEN="$(openssl rand -hex 32)" \
  ghcr.io/rchurro/kasm-use:latest

Clients connect to http://host:8000/mcp with the header Authorization: Bearer <KASM_USE_TOKEN>. Health check: GET /healthz. Run one replica per Kasm user: the server remembers each session's last screenshot size in memory to map click coordinates.

Keep it on a private network (LAN, VPN, Tailscale). Anyone with the token can drive your desktops.

Hermes Agent

pip install kasm-use                      # into Hermes's Python environment
cp -r hermes-plugin/kasm ~/.hermes/plugins/kasm

Enable kasm under plugins.enabled, set the environment variables above, and start a new Hermes session (/new) so the tools load.

Limitations

  • First start is slow (~90 s extra). Stock images don't include xdotool, so kasm_start installs it in the session as root. Reusing a running session skips this.

  • Screenshots can lag a few seconds behind actions; the agent is told to look again to confirm.

  • Kasm's exec API returns no output, so actions report "sent", not "succeeded". The next screenshot is the confirmation.

  • CAPTCHAs, passwords, MFA and payments are handed to you. The tool descriptions tell the agent never to type them; open the session in Kasm and take over.

  • It's slow compared to a local browser tool: every step is a screenshot round trip.

How it works

  • Eyes: POST /api/public/get_kasm_screenshot (KasmVNC renders a JPEG).

  • Hands: POST /api/public/exec_command_kasm running xdotool inside the session.

  • Clicks are given in screenshot pixels and scaled to the real desktop size inside the session (xdotool getdisplaygeometry), using the JPEG's actual dimensions — Kasm keeps the desktop's aspect ratio, so the image is often not the size requested.

License

MIT

Available Tools

8 tools
kasm_clickB

Click at x,y in the coordinate space of the latest kasm_look image.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
buttonNo
doubleNo
session_idNoKasm session id. Omit to use your newest running session.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the coordinate frame dependency but says nothing about what a click does to the remote session, whether it is destructive, how button defaults are chosen, or how double-click timing is handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the most important constraint (which coordinate space) is stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 5-parameter click tool with no output schema and no annotations, the description covers the core coordinate semantics but omits button/double-click behavior and session targeting, leaving gaps an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only session_id is documented). The description adds real meaning for x/y by tying them to the kasm_look image coordinate space, but the button enum, the double flag, and the session-omission behavior are left entirely unexplained in both description and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Click) and the exact coordinate reference frame (the latest kasm_look image), which differentiates it from kasm_type, kasm_key and kasm_scroll. However it doesn't explicitly state it acts on a remote Kasm session, so full sibling disambiguation relies on the name prefix.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'coordinate space of the latest kasm_look image' implies the prerequisite that kasm_look must be called first, which is useful implicit guidance. But there is no statement of when to prefer this over kasm_scroll or kasm_key, and no exclusions or conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kasm_keyC

Press keys/shortcuts in the Kasm session, e.g. 'ctrl+l', 'Return', 'Tab', 'Escape'.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysYes
session_idNoKasm session id. Omit to use your newest running session.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose where the keys are sent (focused element vs. session), whether it blocks or returns a result, what happens if keys are invalid, or any failure/auth behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, no filler, with the key-format examples front-loaded where an agent needs them. Efficient, though the example list could be trimmed or the tool distinction added without cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter action tool with no output schema, the description is minimally adequate: it conveys the action and the key format. The notable gap is routing versus kasm_type, which an agent must guess at, leaving the definition just barely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: session_id is documented in the schema (including the useful 'omit to use newest session' default), but keys is undocumented there. The description compensates by giving concrete key-string examples ('ctrl+l', 'Return', 'Tab', 'Escape'), which conveys modifier syntax that the plain 'string' type does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

A specific verb (press) plus the resource acted on (keys/shortcuts in a Kasm session), with concrete examples of the key syntax. It does not differentiate itself from the sibling kasm_type, which is the closest and most confusable tool in the list, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to use this tool versus kasm_type (typing text) or kasm_click. No prerequisites, no focus/context requirements, and no when-not-to-use guidance are given; the examples imply usage but do not state it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kasm_listB

List Kasm workspaces you can start and your running sessions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It conveys that this is a read-only listing returning two distinct sets (startable workspaces and running sessions), which is useful. It omits authentication requirements, pagination, and whether results are scoped to the user.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with no filler, front-loading the verb and the two result categories. Nothing is wasted, though it is arguably too terse to be maximally helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with no output schema and no annotations, the description covers what is returned but not how the two sets relate, whether kasm_start consumes the listed items, or any auth context. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty, so there is nothing for the description to clarify. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'List' plus two resources: Kasm workspaces available to start and currently running sessions. This is meaningfully more specific than the sibling tools (kasm_start, kasm_stop), though it doesn't name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or routing to alternatives is given. An agent can infer it is a discovery call, but there is no statement of prerequisites, ordering relative to kasm_start, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kasm_lookA

Screenshot the Kasm session so you can see it. Always look before clicking, and look again after each action to confirm it worked.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNoKasm session id. Omit to use your newest running session.
wait_secondsNoPause before capturing (default 2).

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the full burden. It states the core action but does not disclose output format, read-only nature, or other behavioral traits beyond the basic screenshot purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by usage advice. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple screenshot tool with two parameters, no output schema, and no annotations, the description covers what it does and when to use it. It omits return details, but the purpose is clear enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are fully documented in the schema. The description adds no parameter-specific meaning beyond what the schema provides, making 3 the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Screenshot') and resource ('the Kasm session'), and its purpose ('so you can see it'). It implicitly distinguishes itself from sibling tools like kasm_click by instructing to look before clicking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when to use: before clicking and after each action. No explicit when-not-to-use or named alternatives, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kasm_scrollB

Scroll the Kasm session, optionally at x,y from the latest kasm_look image.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
amountNoWheel clicks, 1-30 (default 5)
directionNo
session_idNoKasm session id. Omit to use your newest running session.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose the default direction, the default amount (5 wheel clicks, only in the schema), whether scrolling is relative to the current viewport, or what happens if the coordinate is outside the captured image.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence, front-loaded with the verb and resource. No wasted words, though it is arguably too terse given the gaps it leaves open.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter, annotation-free tool with no output schema and only 40% schema coverage, the description is thin. Key behavioral facts an agent needs — direction default, coordinate origin/reference frame, session continuity — are left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40% — x and y have bare integer types with no schema text, and the description usefully clarifies they refer to coordinates on the latest kasm_look image. It adds nothing about the direction parameter (enum with no stated default) or the amount range beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scroll) and resource (Kasm session), and the optional x,y anchor ties it to the kasm_look workflow. It does not contrast itself with siblings like kasm_click, but the purpose is immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'from the latest kasm_look image' implies the intended workflow: look first, then scroll at a coordinate. However, there is no explicit when-to-use/when-not guidance, no mention of the direction default, and no note on staying in the same session as look/click.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kasm_startA

Start a Kasm workspace session (e.g. 'Chrome', 'Terminal') as the owner, so they can watch it. Takes ~2 minutes (it installs the input tool). Then call kasm_look.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoWorkspace name, e.g. Chrome

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does useful work: it warns the operation takes ~2 minutes because it installs the input tool, and states the session runs under the owner's identity. It omits what happens if a session is already running or what failure modes exist, but the latency and side-effect disclosure is genuinely valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and immediately followed by the cost estimate and the required next step. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter startup tool with no output schema, the description covers what it does, how long it takes, and what to call next. It could mention behavior on repeat invocation, but an agent has enough to act correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents the single 'workspace' parameter with an example. The description's inline examples ('Chrome', 'Terminal') largely duplicate that, adding no new format or constraint information. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a Kasm workspace session') with concrete examples ('Chrome', 'Terminal'). It implicitly separates itself from siblings by naming kasm_look as the follow-up, though it does not explicitly contrast with kasm_stop or kasm_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context ('as the owner, so they can watch it') and an explicit sequencing instruction: 'Then call kasm_look.' No when-not-to-use guidance or alternatives are offered, but the invocation path is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kasm_stopA

Stop (destroy) a Kasm session when the task is finished or the owner asks.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNoKasm session id. Omit to use your newest running session.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. Calling out 'destroy' usefully signals the operation is destructive and terminal, but it says nothing about irreversibility, what happens to session state/output, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded and the usage cue trailing. No filler and nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, likely terminal tool with no output schema, the description covers what it does and when to use it. It is slightly thin on the consequences of destruction, which an agent deciding to kill a session might want.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains the session_id default ('omit to use your newest running session'). The description adds no additional parameter meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('stop') plus the resource ('Kasm session'), and the parenthetical 'destroy' removes ambiguity about what stopping entails. It is immediately distinguishable from the sibling kasm_start.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering condition ('when the task is finished or the owner asks'), which is enough to know when to reach for this over the other session tools. It does not name an alternative or state when-not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

kasm_typeA

Type text into the focused field of the Kasm session. NEVER type passwords, MFA codes, or payment details — stop and ask the owner to take over the session for those.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
session_idNoKasm session id. Omit to use your newest running session.
press_enterNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the sensitive-data prohibition and the required human takeover, but says nothing about side effects on the live session, whether text submission occurs, or what happens if no field is focused.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, well front-loaded: the action first, then the safety constraint. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation-style tool with no annotations and no output schema, the safety handoff is covered, but the undocumented press_enter parameter and the absence of any note on session state or failure behavior leave real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: only session_id is documented in the schema. The description never explains text semantics or the press_enter boolean that presumably submits the input, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Type text into the focused field of the Kasm session.' This clearly separates it from siblings like kasm_key (keystrokes) and kasm_click, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit when-not rule with a fallback action: never type passwords, MFA codes, or payment details, and hand off to the owner instead. It lacks positive guidance on focusing a field first or when to prefer kasm_key, but the exclusion is clearly actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedkasm_click
    • First observedkasm_key
    • First observedkasm_list
    • First observedkasm_look
    • First observedkasm_scroll
    • First observedkasm_start
    • First observedkasm_stop
    • First observedkasm_type

TDQS

A3.7/5.0

Scored across 8 tools

Disambiguation5/5

Each tool maps to a distinct session action: lifecycle management, visual inspection, clicking, scrolling, text entry, and key presses. kasm_type and kasm_key both provide keyboard input, but they are clearly separated by text entry versus special keys/shortcuts, so the overlap is minimal.

Naming Consistency5/5

All eight tools follow the same kasm_ prefix followed by a short action-oriented name. This creates a predictable and readable naming convention with no mixed styles.

Tool Count5/5

Eight tools is well-scoped for a remote Kasm session controller, covering lifecycle, observation, and basic input operations. Each tool serves a clear purpose in the interaction loop without obvious bloat.

Completeness4/5

The set covers starting, listing, and stopping sessions, plus screenshotting, clicking, scrolling, typing, and key presses, which supports many basic remote-control tasks. However, common UI interactions such as mouse move/hover, drag-and-drop, right-click, and clipboard/file transfer are missing, leaving minor gaps that an agent must work around.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to manage and interact with Kasm Workspaces containerized desktop infrastructure, providing tools for session management, command execution, file operations, and user management via the Model Context Protocol.
    4
    MIT