Skip to main content
Glama

Screenshot Capture

screenshot_capture
Read-only

Captures a single frame of a display, window, or region to a PNG. Returns {path, resolution (PIXELS), scale_factor, display_id} — scale_factor is the backing scale of the display that was ACTUALLY captured (the same value list_displays reports for that display_id, by construction: both read one function), so pixels = points x scale_factor when converting a coordinate from the image to ui_click. If the display could not be determined you get scale_factor_unknown instead of a guess; resolution is always there, so you can derive the ratio yourself. Requires Screen Recording permission; without it returns an explicit permission_required, never a blank image.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
targetYesWhat to capture. Required: {kind: "display"|"window"|"window_title"|"region"} plus display_id (see list_displays), window_id (see list_windows), title, or region {x,y,w,h} in global points.
output_pathNoWhere to write the PNG (default: temp file). Missing parent folders are created.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changed
    • addedInput schema / properties / target / description
      Added value: +"What to capture. Required: {kind: \"display\"|\"window\"|\"window_title\"|\"region\"} plus display_id (see list_displays), window_id (see list_windows), title, or region {x,y,w,h} in global points."
    • changedInput schema / properties / target / properties / kind / enum
      Previous value: -[
      -  "display",
      -  "window",
      -  "region"
      -]New value: +[
      +  "display",
      +  "window",
      +  "window_title",
      +  "region"
      +]
    • addedInput schema / properties / target / properties / title
      Added value: +{
      +  "description": "Exact window title to capture (kind:\"window_title\"). Enumerates ALL windows directly — including overlays/panels like Notification Center that list_windows hides unless include_overlays:true. If more than one window has this title, the frontmost one is captured and the response reports matched_windows > 1.",
      +  "type": "string"
      +}
  2. Changed3 schema fields changed
    • removedInput schema / properties / target / description
      Removed value: -"What to capture. Required: {kind: \"display\"|\"window\"|\"window_title\"|\"region\"} plus display_id (see list_displays), window_id (see list_windows), title, or region {x,y,w,h} in global points."
    • changedInput schema / properties / target / properties / kind / enum
      Previous value: -[
      -  "display",
      -  "window",
      -  "window_title",
      -  "region"
      -]New value: +[
      +  "display",
      +  "window",
      +  "region"
      +]
    • removedInput schema / properties / target / properties / title
      Removed value: -{
      -  "description": "Exact window title to capture (kind:\"window_title\"). Enumerates ALL windows directly — including overlays/panels like Notification Center that list_windows hides unless include_overlays:true. If more than one window has this title, the frontmost one is captured and the response reports matched_windows > 1.",
      -  "type": "string"
      -}
  3. Changed3 schema fields changed
    • changedInput schema / properties / target / description
      Previous value: -"What to capture. Required: {kind: \"display\"|\"window\"|\"region\"} plus display_id (see list_displays), window_id (see list_windows), or region {x,y,w,h} in global points."New value: +"What to capture. Required: {kind: \"display\"|\"window\"|\"window_title\"|\"region\"} plus display_id (see list_displays), window_id (see list_windows), title, or region {x,y,w,h} in global points."
    • changedInput schema / properties / target / properties / kind / enum
      Previous value: -[
      -  "display",
      -  "window",
      -  "region"
      -]New value: +[
      +  "display",
      +  "window",
      +  "window_title",
      +  "region"
      +]
    • addedInput schema / properties / target / properties / title
      Added value: +{
      +  "description": "Exact window title to capture (kind:\"window_title\"). Enumerates ALL windows directly — including overlays/panels like Notification Center that list_windows hides unless include_overlays:true. If more than one window has this title, the frontmost one is captured and the response reports matched_windows > 1.",
      +  "type": "string"
      +}
  4. Changed1 schema field changed
    • addedInput schema / properties / target / description
      Added value: +"What to capture. Required: {kind: \"display\"|\"window\"|\"region\"} plus display_id (see list_displays), window_id (see list_windows), or region {x,y,w,h} in global points."
  5. Changed1 schema field changed
    • changedInput schema / properties / output_path / description
      Previous value: -"Where to write the PNG (default: temp file)."New value: +"Where to write the PNG (default: temp file). Missing parent folders are created."
  6. Added

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnlyHint=true/destructiveHint=false annotations. It discloses the exact return shape {path, resolution, scale_factor, display_id}, explains the subtle scale_factor semantics (backing scale of the ACTUALLY captured display, same as list_displays reports), reveals the failure mode (scale_factor_unknown instead of a guessed value), and warns that Screen Recording permission is required with an explicit permission_required error and 'never a blank image' — a critical safeguard against an agent trusting a useless capture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with purpose and return shape front-loaded; every sentence earns its place — the scale_factor explanation is essential and the permission warning is non-obvious. The parenthetical 'by construction: both read one function' is slightly verbose, but it strengthens agent confidence in the equivalence claim rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, a nested target object, and four capture kinds, the description covers return shape, units (PIXELS), scale_factor failure mode, coordinate conversion, and permission behavior — an unusually complete compensation for the missing output schema. The only notable gap is error behavior for invalid or not-found targets (e.g., a window title with zero matches), which the schema's matched_windows note only partially addresses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 and the schema already documents target kinds, title matching, region format, and output_path. The description adds a small amount of parameter-adjacent value by reinforcing that IDs come from list_displays/list_windows and explaining why scale_factor matters for downstream coordinate use, but it does not materially enrich the parameter semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb-resource-output statement: 'Captures a single frame of a display, window, or region to a PNG.' The four capture kinds (display/window/window_title/region) are named, which maps directly to the schema's target.kind enum and distinguishes it from siblings like screen_record_start (multi-frame video) and web_screenshot (browser content).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for when to use the tool: it points to list_displays and list_windows as the sources for IDs, and it explicitly frames the screenshot as a precursor to ui_click via the pixels = points x scale_factor conversion. However, it never states when NOT to use it — e.g., no mention that browser content should go to web_screenshot or that video capture belongs to screen_record_start — so exclusions are left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources