Skip to main content
Glama

capture_page_screenshot

Destructive

Take a viewport, full-page, or clipped screenshot of a webpage or tab, with format and quality controls, and receive the image plus metadata for visual verification or saving.

Instructions

Capture a viewport, full-page, or clipped screenshot of a page/tab via CDP with optional JPEG/WebP quality. Returns text metadata plus an attached MCP image even when save_path is set; save_path only controls disk output. image_width/image_height are DEVICE pixels (CSS x devicePixelRatio), not the CSS pixels page_click takes, and size is the byte count. If the current model cannot consume images, it has not seen the pixels and must use scan_page, execute_js, a page-specific API, or OCR instead. Base64 is included only when return_base64=true.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
clipNo
formatNopng
tab_idNo
qualityNo
timeoutNo
full_pageNo
save_pathNo
session_idNo
return_base64No

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavior beyond the annotations: an attached MCP image is always returned, base64 is only present when requested, image dimensions are device rather than CSS pixels, and size means byte count. These are non-obvious and prevent incorrect interpretation. It does not contradict the destructiveHint annotation; the save_path write behavior is acknowledged.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, front-loaded with the core action and modifiers, then essential caveats. Every sentence adds distinct value, with no filler or repetition. The only minor jargon is 'page-specific API,' but it is acceptable in context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description covers the most critical pitfalls: output side-channel behavior, disk-writing semantics, pixel units, and model-image limitations. It leaves tab/session selection and timeout semantics implicit, but these are partially supported by sibling tools and self-explanatory defaults.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the burden for 9 parameters. It explains format/quality, full_page, clip, save_path, and return_base64, but leaves timeout, tab_id, and session_id semantics implicit. Some mentioned fields like image_width/image_height appear to describe return metadata rather than input parameters, so parameter coverage is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Capture a viewport, full-page, or clipped screenshot of a page/tab via CDP.' It clearly enumerates the tool's modes and contrasts it with page_click on pixel semantics and with scan_page/execute_js on image-consumption limitations, so an agent can distinguish it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when not to rely on the result: if the model cannot consume images, it has not seen the pixels and must use scan_page, execute_js, a page-specific API, or OCR instead. It also clarifies that save_path only controls disk output and the MCP image is still returned, preventing a likely misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.