Skip to main content
Glama

browser_screenshot

Capture visual snapshots of an agent-owned tab for canvas, charts, images, or layout that observations can't describe, and compare changes against prior or another tab.

Instructions

Take a picture of an agent-owned tab. This is a last resort, not a way to find or operate controls: browser_observe returns the refs browser_act needs, and a picture returns none. Use it only for what an observation cannot describe (canvas, charts, images, visual layout). Password and payment fields are masked. The result states the viewport size at capture; a person may resize the live window, so trust it over sizes you saw earlier. state captures a ref's :hover/:focus styling; compare_with diffs against an earlier picture or another tab and boxes what changed. To check a page you are building against an original, open both and take one with compare_with='tab:': one picture of both.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
refNoscope=element or state: a ref from the latest browser_observe of this tab.
labelNoShort caption naming the picture, e.g. 'Application form before submit'; used in the evidence file name.
scopeNoviewport (default) is what is on screen now; full_page is the scrollable document, cut at 8000 px; element is one ref.
stateNo
tab_idYesThe agent-owned tab to capture.
compare_withNo'previous' (this tab's last picture), 'tab:<tab_id>' (another tab, e.g. the original site beside your build; pages of different heights come back side by side), or an evidence path; same scope and size.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden and discloses real behavioral traits: password/payment fields are masked, the result reports the capture viewport size, and the live window may be resized so the reported size should be trusted over earlier observations. It does not cover permissions/ownership rules or how the image is delivered, but the disclosure is well beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the 'last resort' caveat before elaborating on parameters, and every sentence carries information. It is dense and a touch sprawling, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a screenshot tool with no output schema and no annotations, the description covers the key behaviors an agent needs: masking, returned viewport size, comparison semantics, full_page truncation. Only the delivery format of the image itself is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83% (baseline 3), and the description adds genuine meaning: 'state captures a ref's :hover/:focus styling', compare_with 'diffs against an earlier picture or another tab and boxes what changed', and the tab:<tab_id> side-by-side example. The 8000px full_page cut is also surfaced. It stops short of explaining ref/label futher than the schema already does.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource ('Take a picture of an agent-owned tab') and immediately positions it against siblings: browser_observe returns refs, browser_act operates, a picture returns none. An agent can distinguish this from every sibling without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says it is 'a last resort, not a way to find or operate controls' and names the exact condition for use ('only for what an observation cannot describe: canvas, charts, images, visual layout'), routing the agent to browser_observe/browser_act otherwise. It even adds the build-vs-original workflow via compare_with.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.