Skip to main content
Glama

desktop_screenshot

Read-only

Capture the Raspberry Pi Wayland desktop or a selected region, with optional downscaling, and return a view_id that maps mouse actions to image-pixel coordinates.

Instructions

Capture the Pi desktop, optionally a source region in original pixels and/or a downscaled image. The returned view_id enables image-pixel mouse coordinates.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
widthNo
heightNo
max_widthNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.0

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, non-destructive, closed-world, so safety is covered. The description adds genuinely useful context beyond them: that the call yields a view_id tied to pixel-space mouse coordinates and that downscaling is optional. It omits defaults (full-screen capture when no region is given) and return format details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and no filler. Every clause adds information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so by naming view_id and the two image variants. What is missing is region/default behavior and parameter specifics, but for a zero-required-param capture tool it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the schema titles are bare (X, Y, Width, Height, Max Width), so the description must carry the load. It implies x/y/width/height form a source region and max_width drives the downscaled image, but never states units, defaults, or how the two modes interact — partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (capture the Pi desktop) plus the two output modes (source region in original pixels, downscaled image). It is clearly distinguishable from siblings like desktop_click or desktop_scroll, though it does not name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the note that the returned view_id enables image-pixel mouse coordinates hints that this should be called before coordinate-based click/move operations, but no explicit when/when-not or preconditions (e.g. needing a connected desktop) are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.