Skip to main content
Glama

Compare two accessibility scans

diff_scan
Read-onlyIdempotent

Compare two scans of the same page to report fixed, new, and remaining accessibility issues, confirming whether a fix caused regressions.

Instructions

Compare two scans of the same page and report what changed: fixed[] (in the baseline, gone now), new[] (regressions — not in the baseline, present now), remaining[] (still there). Page-level complement to verify_fix (one element). Baseline is a scan_history id (baselineId, local installs) or a live scan of baselineUrl; current is url (scanned live now) or another history id (currentId). Findings are matched by issue id (rule + element), so a changed class/id on a fixed element reads as fixed AND new — check new[] before calling it a regression. Typical loop: scan_page → edit → diff_scan(baselineId=, url=) → confirm new[] is empty.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoURL to scan now as the CURRENT side (deployed, staging, or http://localhost:3000). Omit when passing currentId.
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
levelNoHighest WCAG conformance level to report. Default "AA" — Level AAA criteria are not reported. Pass "AAA" to include them (naming one AAA criterion in `wcag`, e.g. "3.2.5", also includes it).
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
formatNo"compact" prints one line per element with rule metadata once. "json" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.
viewportNoViewport for live scans (default: desktop). Use the same viewport the baseline used.
currentIdNoscan_history id to use as the CURRENT side instead of scanning `url`
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
baselineIdNoscan_history id of the BASELINE scan (local installs only)
baselineUrlNoScan this URL live as the baseline (e.g. production) — use when there is no stored baseline
rootSelectorNoCSS selector to limit live scans to (optional)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv1.8.0
    • addedInput schema / properties / level
      Added value: +{
      +  "description": "Highest WCAG conformance level to report. Default \"AA\" — Level AAA criteria are not reported. Pass \"AAA\" to include them (naming one AAA criterion in `wcag`, e.g. \"3.2.5\", also includes it).",
      +  "enum": [
      +    "A",
      +    "AA",
      +    "AAA"
      +  ],
      +  "type": "string"
      +}
  2. Addedv1.7.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely non-obvious behavior beyond that: findings are matched by issue id (rule + element), and a changed class/id on a fixed element will appear as BOTH fixed and new — a false-regression trap an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the output shape (what changed) before the plumbing, and every clause carries information. It is dense — the matching-semantics sentence and the typical-loop sentence are both long — but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden and does so by naming and defining fixed[]/new[]/remaining[]. Combined with the baseline/current input pairing and the matching caveat, an agent has enough to call it correctly without follow-up.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it pairs the optional inputs (baseline side = baselineId or baselineUrl; current side = url or currentId) and notes baselineId is local-installs-only. That encoding of mutual alternatives goes beyond the flat per-parameter schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Compare two scans of the same page') and immediately enumerates the three output buckets (fixed/new/remaining). It explicitly distinguishes itself from the sibling verify_fix by scope ('Page-level complement to verify_fix (one element)'), so an agent can route between them without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use recipe ('Typical loop: scan_page → edit → diff_scan(...) → confirm new[] is empty') plus the conditions selecting each input pair (baselineId vs baselineUrl, url vs currentId). It also warns to check new[] before declaring a regression, which is actionable guidance, not just context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.