Skip to main content
Glama

Pi Desktop Bridge

CI Python 3.11+ MIT License

Give Codex and other MCP-enabled assistants eyes and hands on a Raspberry Pi's Wayland desktop. Capture real screenshots, move and click the mouse, drag, scroll, type Unicode text, and use keyboard shortcuts over an existing SSH connection.

The bridge runs on your computer and starts a small user-owned agent on the Pi. It uses the Pi's existing graphical session, with no extra hardware and no exposed VNC or web port. This controls the Pi's own desktop; controlling another computer needs a separate software connection or hardware KVM.

Features

  • Visual desktop control: eleven MCP tools for screenshots, status, mouse, keyboard, and disconnect.

  • Efficient images: smaller overviews and native-resolution regions, with screenshot coordinates mapped back to the desktop.

  • Observe after acting: input tools return a fresh screenshot; visual waits sample an area until its pixels settle.

  • Controlled sessions: one client at a time, explicit disconnect, and configurable idle release after five minutes by default.

  • Recovery checks: stale views and mismatched agent source are rejected; uncertain input requires a fresh screenshot before more input.

  • SSH transport: strict host-key checking, existing SSH credentials, and private UNIX sockets on the Pi.

Related MCP server: claude-workman

Requirements

Where

Required

Raspberry Pi

A running wlroots-compatible Wayland desktop, such as Raspberry Pi OS with labwc; python3, wayvnc, grim, and wtype

SSH connection

A trusted host key and login without a password prompt; log in as the same user who owns the graphical session

Client computer

Python 3.11+, OpenSSH, uv, and Git

Assistant application

A local stdio MCP client that supports image content and tool calls, such as Codex

No Python packages are installed on the Pi. If its desktop tools are missing, install them there:

sudo apt-get install --no-install-recommends wayvnc grim wtype

Raspberry Pi OS Lite alone has no desktop to capture. A headless Pi still needs a running graphical session with an output configured by its compositor. This MCP integration is separate from Codex's native Computer Use tool.

Quick start

1. Check SSH

The examples use the SSH alias pi-desktop. Replace it with your working alias, or create a Host pi-desktop entry in your SSH config. Confirm that it connects as the desktop user:

ssh pi-desktop

Resolve host-key trust and key-based authentication in your normal SSH client first. The bridge does not prompt for passwords or silently trust new hosts.

2. Install and deploy

Run on the computer hosting your MCP client:

git clone https://github.com/0xHayd3n/pi-desktop-bridge.git
cd pi-desktop-bridge
uv sync --frozen
uv run pi-desktop-bridge deploy --host pi-desktop
uv run pi-desktop-bridge doctor --host pi-desktop

Deployment installs the agent into ~/.local/share/pi-desktop-bridge under the SSH user and verifies its SHA-256. It needs no sudo, system service, or firewall changes. doctor reports compatibility, tools, and Wayland prerequisites without taking the desktop lease; it does not guarantee a capture or prove an off-network route.

3. Connect Codex

From the cloned repository, register its virtual-environment Python executable. The optional 960-pixel setting reduces post-action image size; explicit screenshots keep their own sizing options.

Windows / PowerShell:

$bridgePython = (Resolve-Path .\.venv\Scripts\python.exe).Path
codex mcp add pi_desktop -- $bridgePython -m pi_desktop_bridge serve --host pi-desktop --capture-max-width 960
codex mcp get pi_desktop

macOS / Linux:

codex mcp add pi_desktop -- "$PWD/.venv/bin/python" -m pi_desktop_bridge serve --host pi-desktop --capture-max-width 960
codex mcp get pi_desktop

Restart the MCP server in your client's settings, or restart Codex, to load the tools. See OpenAI's MCP configuration documentation for supported Codex clients and configuration.

Other MCP clients can use the same executable and arguments; see the client configuration example. Client support varies. ChatGPT web cannot directly launch this local stdio server.

4. Use the desktop

Start by asking your assistant:

Take a 960-pixel overview of the Pi desktop and describe what is visible. Wait for my next instruction before changing anything.

When switching windows, move the pointer into the target window and inspect the returned screenshot before clicking. A local person changing focus can affect where input goes. Call desktop_disconnect when finished to release control immediately.

Tools

Tool

Purpose

desktop_health

Check prerequisites and recovery state without taking control

desktop_status

Inspect desktop size, connection, and session policy

desktop_screenshot

Capture a full desktop, overview, or region

desktop_wait_for_stable

Observe a source area until sampled pixels settle, or return a timeout image

desktop_move

Move the pointer

desktop_click

Single or double click with the left, middle, or right button

desktop_drag

Drag between two coordinates

desktop_scroll

Scroll vertically or horizontally at a target

desktop_type

Type text into the focused application

desktop_key

Press a key or shortcut

desktop_disconnect

Release control and allow later reconnection

See the tool reference for image coordinates, view_id, crop and resize examples, stability timing, idle policy, and error recovery.

Security and data handling

  • The bridge runs as the SSH user. It adds no network listener and uses an owned private WayVNC UNIX socket. OpenSSH retains your keys and authenticates the host.

  • Source hashes detect deployment drift; they do not attest that the remote machine is trustworthy. Use only a Pi and SSH account you trust.

  • Screenshots and typed content are sent to the MCP client and may be processed by its AI service. The server does not retain a screenshot history; explicit exports and opt-in verification scripts can save images locally.

  • Input is never automatically replayed. After uncertain delivery, inspect a fresh full-desktop screenshot before deciding whether to repeat an action.

  • Reports and screenshots under _local/, the .venv/ environment, and .env files are ignored by Git. Exports saved elsewhere can be tracked: keep credentials and personal captures out of commits, issues, and pull requests.

See architecture and trust boundaries for the protocol, process lifecycle, and data flow.

Remote access and current limits

For use away from home, configure the SSH alias to a reachable VPN/mesh address or another working SSH route. A .local hostname provides LAN discovery, not worldwide access. VPN login or ping alone does not prove SSH reachability; test SSH from the network where you will run the client.

The tested setup uses Raspberry Pi 5, Raspberry Pi OS / Debian 13, labwc, and a 1920 × 1080 output. Off-network SSH was not verified in that setup. Cross-window first-click delivery can depend on compositor focus; move and observe before clicking. A visual stability result means sampled pixels matched, not that an application is ready. See the verification record for evidence and remaining limits.

Upgrade

git pull --ff-only
uv sync --frozen
uv run pi-desktop-bridge deploy --host pi-desktop
uv run pi-desktop-bridge doctor --host pi-desktop

Restart the MCP server after upgrading. The expected agent hash is pinned when the client process starts.

Optional Codex plugin

A portable plugin.json, mcp.json, and repository marketplace are included:

codex plugin marketplace add 0xHayd3n/pi-desktop-bridge

Install and enable Pi Desktop Bridge from that marketplace in a compatible Codex client. Its default alias is pi-desktop; deploy the agent and make that alias work first. The launcher uses uv and keeps its environment under the client's plugin data directory.

Direct MCP configuration is the primary verified integration. The plugin files have been parsed by Codex's loader, but this repository is not a public Plugins Directory listing. See OpenAI's plugin packaging documentation.

Development and verification

uv run python -m unittest discover -s tests -v
uv build

The unit suite covers protocol framing, image validation, bounds, concurrency, cancellation, session release, deployment, and recovery. CI runs on Windows with Python 3.14 and Ubuntu with Python 3.11.

The live tests operate a real Pi desktop through the MCP SDK:

uv run python scripts/live_smoke.py --host pi-desktop
uv run python scripts/live_views.py --host pi-desktop
uv run python scripts/live_stability.py --host pi-desktop
uv run python scripts/live_lifecycle.py --host pi-desktop

These are opt-in input tests. They open a disposable Tk window, operate it, then remove their fixture. They require tkinter and an Xwayland display at :0 on the Pi. Images and reports go into ignored _local/. Run them when nobody else is using the desktop. See the verification record for checks actually completed.

Troubleshooting and removal

  • SSH failure: verify the same alias in a normal terminal. The bridge uses BatchMode=yes and StrictHostKeyChecking=yes.

  • Missing desktop or tools: run uv run pi-desktop-bridge doctor --host pi-desktop. Check that SSH uses the graphical session's user.

  • Agent source mismatch: update, deploy again, and restart the MCP server.

  • Desktop busy: disconnect the controlling client, stop its server, or wait for its configured idle release. Health checks do not take the lease.

  • Connection lost after input: the action may already have happened. Take a fresh full-desktop screenshot and inspect it before retrying.

  • Disable: run codex mcp remove pi_desktop, or disable the server in your client. Closing the server stops its owned SSH agent and WayVNC process.

  • Remove the agent: stop clients, then remove only ~/.local/share/pi-desktop-bridge under the SSH user. Existing SSH and desktop tools remain installed.

License

MIT. Independent project; not affiliated with Raspberry Pi or OpenAI.

Available Tools

11 tools
desktop_clickB
Destructive

Click in original desktop pixels, or image pixels with the current view_id, then show a fresh screenshot. Button is left, right, or middle.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
countNo
buttonNoleft
view_idNo
capture_max_widthNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds useful side-effect context (a fresh screenshot is shown afterwards) and the coordinate mode, but says nothing about permissions, repeat-click semantics, or failure behavior beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence covering coordinate mode and the screenshot side effect with little waste. The appended button list is slightly tacked-on but not verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a mutation tool with no output schema and 6 parameters at 0% schema coverage. Crucial details such as what count does (double-click), what capture_max_width controls, and the numeric format of x/y are missing from both the schema and the description, so an agent cannot call it reliably in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all 6 parameters. The description only clarifies button values (left/right/middle) and loosely ties view_id to image pixels; x, y, count, and capture_max_width remain entirely undefined, leaving more than half the parameters unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) plus the exact resource/coordinate space, and distinguishes itself from siblings like desktop_scroll, desktop_type, and desktop_key. It also discloses the post-action screenshot, which no sibling description implies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use view_id (image pixels) versus original desktop pixels, giving an implicit routing choice. However, it never states exclusions or when to prefer a sibling tool, and the conditional is worded ambiguously.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_disconnectA
Idempotent

Release this bridge's Pi desktop session. The MCP server can reconnect on the next tool call.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is not read-only, not destructive, and idempotent; the description adds the non-obvious behavioral fact that the session is lazily re-established on the next tool call, which is the key operational consequence of calling it. It does not, however, mention any auth or state-teardown caveats beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the action front-loaded and the consequence immediately following. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless session-teardown tool with an output schema and supporting annotations, the description covers the essential purpose and the reconnect behavior. Nothing critical is missing, though a note on whether other in-flight tools are affected would have made it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. No parameter information is required or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Release') and resource ('this bridge's Pi desktop session'), which is unambiguous and clearly distinct from the sibling tools (status, health, screenshot, click, etc.), none of which tear down a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives meaningful context for choosing this tool: the disconnect is safe because 'the MCP server can reconnect on the next tool call.' This tells the agent it can drop the session without permanent loss, but it does not state when an agent should prefer disconnecting versus simply leaving the session open.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_dragB
Destructive

Drag in original desktop pixels, or image pixels with the current view_id, then show a fresh screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
end_xYes
end_yYes
start_xYes
start_yYes
view_idNo
capture_max_widthNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely useful behavior the annotations lack: coordinates may be given in original desktop pixels or image pixels relative to the current view_id, and a fresh screenshot is returned afterward. It still omits what state the drag alters or any permission/side-effect detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the coordinate-frame rule comes first and the screenshot consequence last. It is dense but slightly cryptic, so not quite maximal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, open-world UI-automation tool with no output schema, the description covers the return behavior (fresh screenshot) and the coordinate-frame choice, which is the most important decision. However, with 0% schema coverage it leaves capture_max_width and drag semantics unexplained, so it is only partially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across six parameters, so the schema supplies only bare titles. The description partially compensates by explaining that start/end coordinates are pixels and that view_id selects the image-pixel frame, but capture_max_width and the required start/end pairing semantics remain undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (drag), the two coordinate frames it accepts, and the post-action screenshot. It implicitly separates itself from desktop_move/desktop_click by describing a full drag path (start to end) rather than a single-point or continuous move. It stops short of explicitly naming the sibling it differs from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no exclusions. The agent is not told when to prefer desktop_drag over desktop_move (e.g., move-to-position vs press-drag-release), nor any prerequisite such as an active connection or view context. Only the coordinate-frame distinction offers indirect routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_healthA
Read-only

Check Pi agent and desktop prerequisites without acquiring a session. This does not reserve the desktop or guarantee a later capture.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description goes beyond them by disclosing the non-reserving, non-guaranteeing semantics of the probe, which materially affects how an agent sequences it before capture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, the primary purpose front-loaded and the caveat immediately after. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and an output schema present, the description does not need to explain return values, and it covers the key behavioral caveat for a pre-flight probe. It is close to complete, missing only an explicit pointer to when a full session is warranted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. No parameter information is missing or misleading.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check') and resource ('Pi agent and desktop prerequisites'), which is clearly distinct from the sibling desktop_status. However, it does not explicitly contrast itself against that sibling, leaving the health-vs-status distinction to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear usage condition ('without acquiring a session'), which tells the agent to use this before committing to a session. It clarifies the negative case in the second sentence but never names an alternative tool or an explicit when-not scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_keyB
Destructive

Press a key combination in the focused desktop app and show a fresh screenshot. Keys is a list such as ['Control_L','a'].

ParametersJSON Schema
NameRequiredDescriptionDefault
keysYes
capture_max_widthNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and openWorldHint=true, so the mutation and open-world nature is covered. The description adds useful context that a screenshot is returned and that input targets the focused app, but says nothing about permissions, rate limits, or what happens on invalid keys. With annotations carrying the safety profile, this is a modest addition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A compact single sentence that front-loads the action and the target. No wasted words, though the trailing screenshot clause could be structured more cleanly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description partially compensates by noting a screenshot is returned. For a destructive, open-world input tool, it is adequate but leaves the optional capture_max_width parameter and any failure/permission behavior unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It does explain the required 'keys' parameter with a concrete example (['Control_L','a']), which clarifies the expected key-name format. However, capture_max_width is left completely undocumented in both schema and description, leaving half the parameters ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Press a key combination in the focused desktop app'. The extra mention of showing a fresh screenshot describes the side effect. It is clear but does not explicitly differentiate itself from siblings like desktop_type or desktop_click.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or routing to alternatives (e.g., desktop_type for text vs desktop_key for combos). Usage is only implied by the verb. An agent must infer it should use this for key combinations, not plain typing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_moveA
Destructive

Move the pointer in original desktop pixels, or image pixels with the current view_id, then show a fresh screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
view_idNo
capture_max_widthNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=false, destructiveHint=true and openWorldHint=true, the description adds a concrete side effect not captured by structured fields: it captures and returns a fresh screenshot after the move. The coordinate-mode behavior is also disclosed. It does not mention permissions or the return format beyond the screenshot, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence conveying coordinate modes and the screenshot side effect with zero filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema is present, so the description usefully notes the fresh screenshot as the result. However, for a 4-parameter tool with 0% schema coverage it leaves capture_max_width unexplained and offers no guidance on coordinate-mode errors or prerequisites, so it is only adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load. It explains x/y units and that view_id switches to image-pixel interpretation, covering 3 of 4 parameters, but says nothing about capture_max_width, leaving one parameter's behavior completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Move the pointer') and clarifies the two coordinate systems it accepts, which distinguishes it from click/drag siblings. It stops short of explicitly naming which sibling to prefer, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies a selection condition: use desktop pixels normally, or image pixels when supplying the current view_id. That is useful routing context, but it never states when to choose move over desktop_drag or desktop_click, nor any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_screenshotA
Read-only

Capture the Pi desktop, optionally a source region in original pixels and/or a downscaled image. The returned view_id enables image-pixel mouse coordinates.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
widthNo
heightNo
max_widthNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, non-destructive, closed-world, so safety is covered. The description adds genuinely useful context beyond them: that the call yields a view_id tied to pixel-space mouse coordinates and that downscaling is optional. It omits defaults (full-screen capture when no region is given) and return format details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and no filler. Every clause adds information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so by naming view_id and the two image variants. What is missing is region/default behavior and parameter specifics, but for a zero-required-param capture tool it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the schema titles are bare (X, Y, Width, Height, Max Width), so the description must carry the load. It implies x/y/width/height form a source region and max_width drives the downscaled image, but never states units, defaults, or how the two modes interact — partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (capture the Pi desktop) plus the two output modes (source region in original pixels, downscaled image). It is clearly distinguishable from siblings like desktop_click or desktop_scroll, though it does not name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the note that the returned view_id enables image-pixel mouse coordinates hints that this should be called before coordinate-based click/move operations, but no explicit when/when-not or preconditions (e.g. needing a connected desktop) are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_scrollC
Destructive

Scroll at an explicit target in original pixels (x and y together), or image pixels with the current view_id, then show a fresh screenshot. Direction is up, down, left, or right.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
ticksNo
view_idNo
directionYes
capture_max_widthNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, destructiveHint=true, so the safety profile is covered. The description adds genuinely useful behavior beyond that: scrolling also returns a fresh screenshot, and it explains the two coordinate reference frames. It says nothing about tick magnitude, side effects on the viewport, or what happens if x/y are omitted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence set with the core action and coordinate modes front-loaded and no padding. It is slightly under-informative rather than verbose, so length itself is not the problem.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, no-output-schema mutation tool with zero schema descriptions, the description leaves major gaps: ticks semantics, capture_max_width meaning, prerequisites, and any confirmation that it mutates the observable desktop state. It is not sufficient on its own to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 6 parameters, so the description must carry the load alone. It clarifies that x and y are used together and distinguishes two coordinate frames tied to view_id, but completely ignores the 'ticks' and 'capture_max_width' parameters, leaving a third of the parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a specific verb (scroll) and the resource being scrolled (the desktop view at a target), plus the coordinate systems involved, so the agent knows precisely what happens. It does not differentiate itself from nearby siblings like desktop_move, desktop_drag, or desktop_click, which also manipulate the pointer on the desktop.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains coordinate modes ('original pixels' vs 'image pixels with the current view_id') but never says when to prefer this over desktop_drag or desktop_move, nor what prerequisites (e.g. an active connection) apply. Usage is implied by the verb only.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_statusA
Read-only

Show connection and desktop dimensions for the authorized Raspberry Pi.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds useful context about the specific data surfaced (connection state and desktop dimensions), which helps an agent decide it can call this freely to inspect state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words; the informational payload is stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not enumerate return fields, and a zero-parameter read tool with annotation coverage is nearly self-sufficient. Only the lack of any usage-routing against sibling status tools (e.g., desktop_health) keeps it short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate beyond what the empty schema shows.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Show") and specific resources ("connection and desktop dimensions") for the Raspberry Pi target. It is clearly distinguishable from action siblings like desktop_click or desktop_type, though it does not explicitly differentiate itself from the closest sibling, desktop_health.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no routing to alternatives such as desktop_health or desktop_screenshot. The read-only status-check intent is only implied by the verb "Show".

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_typeB
Destructive

Type text into the focused desktop app and show a fresh screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
capture_max_widthNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, and destructiveHint=true, so the write/destructive profile is covered. The description adds the useful fact that a screenshot is captured after typing, but says nothing about irreversibility, focus prerequisites, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the action and its observable effect are stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no output schema and fully undocumented parameters, the description does not explain the second parameter, focus prerequisites, or failure modes, leaving an agent under-informed despite the mention of the screenshot result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 2 parameters, yet the description only implicitly covers the required 'text' and never mentions 'capture_max_width' at all. The description does not compensate for the documentation gap in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear specific verb (type) plus resource (focused desktop app), and it names the side effect of returning a fresh screenshot. It is distinguishable from desktop_click or desktop_scroll, though it does not explicitly contrast with the closely related desktop_key tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'focused desktop app' implies a precondition (something must already be focused), but there is no explicit when-to-use vs. sibling guidance, no mention of needing a prior desktop_click/focus step, and no stated alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_wait_for_stableB
Read-only

Observe sampled desktop pixels until unchanged for the requested duration, or sampling times out. Returns the last image and stability status; sampled equality does not prove the app is ready. A full desktop_screenshot is still required to recover from an uncertain input.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
widthNo
heightNo
poll_msNo
max_widthNo
stable_msNo
timeout_msNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, openWorldHint=false, and destructiveHint=false, so safety is covered. The description adds valuable behavioral context: it explains the sampling mechanism, that equality does not prove the app is ready (a critical caveat), and that a desktop_screenshot is required to recover from uncertain input. This goes beyond what annotations provide, though it doesn't mention timeout behavior or potential side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence followed by two concise clauses. It is front-loaded with the core action and includes essential caveats without unnecessary verbosity. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (8 parameters, no schema descriptions, no output schema), the description is incomplete. While it covers the core behavior and a key caveat, it completely omits parameter explanations, which are critical for correct invocation. Without any output schema, the description does mention the return (last image and stability status), but the lack of parameter semantics makes it inadequate for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%—none of the 8 parameters have descriptions. The tool description does not explain the meaning of any parameter (e.g., x, y, width, height, poll_ms, stable_ms, timeout_ms, max_width). For a tool with 8 parameters and zero coverage, the description must compensate, but it provides no parameter information at all. This is a severe gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Observe sampled desktop pixels until unchanged for the requested duration, or sampling times out.' This clearly distinguishes it from siblings like desktop_screenshot (single capture) and desktop_status. It's concise and specific about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (when you need to wait for the desktop to stabilize) and mentions that 'a full desktop_screenshot is still required to recover from an uncertain input,' which provides some context about limitations. However, it does not explicitly state when to choose this tool over alternatives like desktop_screenshot or desktop_status, nor does it provide clear exclusion criteria. Usage guidance is implied but not fully developed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 11 tool updatesv0.5.0
    • First observeddesktop_click
    • First observeddesktop_disconnect
    • First observeddesktop_drag
    • First observeddesktop_health
    • First observeddesktop_key
    • First observeddesktop_move
    • First observeddesktop_screenshot
    • First observeddesktop_scroll
    • First observeddesktop_status
    • First observeddesktop_type
    • First observeddesktop_wait_for_stable

TDQS

A3.7/5.0

Scored across 11 tools

Disambiguation4/5

Most tools have clearly distinct actions: input actions (click, move, drag, scroll, type, key), capture (screenshot), and session management (status, health, disconnect). Minor overlap exists between desktop_status and desktop_health, and between desktop_screenshot and desktop_wait_for_stable, but descriptions clarify their different purposes.

Naming Consistency5/5

All tool names use a consistent snake_case pattern with the desktop_ prefix and a clear action-oriented verb or noun. The convention is predictable throughout the set.

Tool Count5/5

The 11 tools are well-scoped for a remote desktop control bridge, covering observation, input, and session management without excessive or redundant tools. Each tool appears to earn its place.

Completeness4/5

The set covers essential desktop automation operations: screenshot, wait for stability, pointer movement, clicking, dragging, scrolling, typing, key presses, and session status/health/disconnect. Minor gaps like explicit double-click or clipboard access are not critical and can be worked around with existing tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for controlling Hyprland Wayland compositor: screen capture, mouse/keyboard injection, and automation via native Hyprland IPC and wlr protocols.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP clients to control a Linux/X11 desktop like a human: see the screen, move the mouse, click UI elements via the accessibility tree, type text, and manage windows.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables MCP-speaking clients to control a real desktop via the computer_use tool, including clicking, typing, scrolling, dragging, key combos, app focus, and screen/accessibility capture. It also provides a verdict system that verifies whether input actions had their intended effect, with configurable approval modes.
    1
    MIT