Skip to main content
Glama

visual_diff

Compare two web pages or HTML strings pixel-by-pixel, returning a diff image that highlights visual differences along with the changed pixel count and percentage.

Instructions

Compare two web pages (or HTML strings) pixel-by-pixel and return a diff image highlighting all visual differences. Supports full-page capture, device emulation, element selectors, and all screenshot-like options. Returns the diff image, changed pixel count, and percentage changed. Costs 1 API request.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
clipNoCrop region { x, y, width, height } in pixels
clickNoCSS selector to click before capturing on both pages
delayNoMilliseconds to wait before capture on both pages (default: 0)
url_aNoURL of the first page (required if no html_a)
url_bNoURL of the second page (required if no html_b)
widthNoViewport width in pixels (default: 1280)
heightNoViewport height in pixels (default: 720)
html_aNoRaw HTML for the first page (required if no url_a)
html_bNoRaw HTML for the second page (required if no url_b)
cookiesNoCookies to set — array of "name=value" strings or { name, value, domain? } objects
headersNoExtra HTTP headers to send with the request
blockAdsNoBlock advertisements on the page
darkModeNoEmulate dark color scheme (default: false)
fullPageNoCapture the full scrollable page for both sides (default: false)
injectJsNoCustom JavaScript to execute before capturing (max 50KB)
selectorNoCSS selector — capture only this element on both pages
timeZoneNoOverride browser timezone (e.g. "America/New_York")
bypassCSPNoBypass Content-Security-Policy on the page
injectCssNoCustom CSS to inject before capturing (max 50KB)
mediaTypeNoEmulate CSS media type
thresholdNoPixelmatch sensitivity 0–1 (default: 0.1). Lower = more sensitive to subtle differences.
userAgentNoOverride the browser User-Agent string
waitUntilNoWhen to consider navigation finished (default: networkidle2)
blockChatsNoBlock live chat widgets on the page
geolocationNoEmulate geolocation { latitude, longitude, accuracy? }
blockBannersNoHide cookie consent banners (default: false)
authorizationNoAuthorization header value (e.g. "Bearer <token>")
blockRequestsNoURL patterns to block (array of strings)
blockTrackersNoBlock tracking scripts on the page
hideSelectorsNoArray of CSS selectors to hide before capture
reducedMotionNoEmulate prefers-reduced-motion to disable animations
blockResourcesNoResource types to block (e.g. ["image", "font"])
fullPageScrollNoAuto-scroll pages before capture to trigger lazy-loaded images
viewportDeviceNoDevice preset for viewport emulation (e.g. "iphone_14_pro"). Use list_devices to see all presets.
viewportMobileNoEnable mobile meta viewport emulation
waitForSelectorNoWait for this CSS selector to appear before capturing
fullPageScrollByNoPixels to scroll per step (default: viewport height)
viewportHasTouchNoEnable touch event emulation
deviceScaleFactorNoDevice pixel ratio (default: 1)
fullPageMaxHeightNoMaximum pixel height cap for full-page captures
navigationTimeoutNoNavigation timeout in ms (default: 25000)
viewportLandscapeNoLandscape orientation
fullPageScrollDelayNoDelay between scroll steps in ms (default: 400)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.17.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose concrete behavior: exact return values (diff image, changed pixel count, percentage changed) and the cost of 1 API request, neither of which appear in the schema. It omits auth/session prerequisites and timeout/reversibility behavior, but the cost and output disclosure is genuinely beyond structured data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the purpose, then capabilities, returns, and cost. Generally efficient, though the 'and all screenshot-like options' clause is vague filler against a schema that already enumerates those options.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must explain returns and it does (diff image, changed pixel count, percentage). For a 43-param tool with full schema coverage, the only meaningful gap is operational context such as session/auth prerequisites and failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with 43 parameters, so the schema already documents every option. The description's mention of full-page capture, device emulation, and element selectors only gestures at parameter groups without adding format or semantics beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (compare), resource (two web pages or HTML strings), method (pixel-by-pixel), and output (a diff image highlighting differences). This clearly distinguishes it from the sibling take_screenshot, which captures one page, so an agent can route without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its use case (diffing two pages/HTML strings) but never states when to prefer it over take_screenshot or how it fits a session workflow. No explicit alternatives or exclusions are named, leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.