Skip to main content
Glama

Diff UI against baseline

ui_diff

Compare a page against its baseline and return only changed regions as cropped images plus a summary, within a token budget.

Instructions

Re-capture a page, pixel-diff it against the baseline for its label, and return ONLY what changed: a text summary plus cropped images of changed regions (largest first) that fit within maxTokens. Omitted regions are listed so you can fetch them with ui_region. The new capture becomes the baseline. If no baseline exists, behaves like ui_capture.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesPage URL, e.g. http://localhost:3000/settings
labelNoBaseline label. Defaults to a slug of the URL path. Keep one label per route.
waitForNoCSS selector to wait for, or ms to wait
baselineNo'previous' (default) or a captureId from ui_list
darkModeNoEmulate prefers-color-scheme: dark.
fullPageNoCapture the full scrollable page (bigger, costlier). Default false.
viewportNo
maxTokensNoToken budget for this response (text + images). Default 4000.
includeFullNoAlso return the full screenshot (downscaled to budget). Default false.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and idempotentHint=false; the description goes well beyond that by disclosing the key side effect ('The new capture becomes the baseline'), the truncation policy (only changes that 'fit within maxTokens', largest first, omitted regions enumerated), and the degraded no-baseline mode. These are exactly the traits structured fields cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences with no filler: action first, then return payload, then truncation handling, then side effect and fallback. Every clause earns its place and the most decision-relevant facts are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 9 parameters, the description still fully describes the return contract (text summary plus cropped changed-region images, largest first, budget-limited, omissions listed) and the state-mutating side effect, which is the maximum an agent needs before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 89%, so the baseline is 3, but the description adds real meaning: it ties 'label' to the baseline selection ('the baseline for its label'), explains that maxTokens gates which regions are returned at all, and clarifies the default/override relationship for 'baseline'. It does not cover the viewport or waitFor semantics, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb chain and resource ('Re-capture a page, pixel-diff it against the baseline for its label') and explicitly contrasts itself with siblings ui_capture (fallback when no baseline exists) and ui_region (fetching omitted regions). An agent can distinguish it from all six sibling tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: use this to see what changed against an existing baseline, and it names the fallback behavior ('If no baseline exists, behaves like ui_capture') plus the follow-up path via ui_region for regions dropped by the token budget. It stops short of an explicit 'prefer this over ui_capture when...' rule, but the routing is inferable and mostly spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.