Skip to main content
Glama

WebAbility MCP

Accessibility checks your coding agent can act on: scan a page with three engines, get a structured fix for each issue, then re-check that the fix landed.

Free. Hosted at https://mcp.webability.io/mcp. No API key for scans. MIT.

Install

Claude Code (plugin)

/plugin marketplace add snayyar00/webability-mcp
/plugin install webability-accessibility@webability

or claude mcp add --transport http webability https://mcp.webability.io/mcp

Cursor: Add to Cursor

VS Code: Install in VS Code · or run code --add-mcp '{"name":"webability","type":"http","url":"https://mcp.webability.io/mcp"}'

Claude.ai / ChatGPT / any MCP client: add a custom connector with the URL https://mcp.webability.io/mcp.

Local (the browser runs on your machine, so it scans localhost directly): npx -y -p @webability/mcp webability-mcp

Related MCP server: A11y Expert MCP

What you get back

Real output, scan_page on https://demo.vercel.store (trimmed):

Found 16 high-confidence issue(s): 0 critical, 4 serious, 8 moderate, 4 minor.
20 additional finding(s) need human review — see incomplete[]. Do NOT auto-fix these.

missing_label · serious · WCAG 1.3.1 · input.text-md.w-full.rounded-lg
  fix: { op: "add-attribute", attribute: "aria-label" }   fixability: contextual
missing_table_scope · moderate · WCAG 1.3.1 · thead > tr > th:nth-of-type(1)   (vite.dev/guide)
  fix: { op: "add-attribute", attribute: "scope", value: "col" }   fixability: mechanical

verify_fix input.text-md.w-full.rounded-lg wcag=4.1.2
  NOT RESOLVED: 1 violation still present → "verified": false
  • fix.op is one of add-attribute, set-attribute, remove-attribute, add-element, remove-element, add-text-content, suggest.

  • fixability: mechanical = apply as given. contextual = the op is known, the value (alt text, a label) needs judgment. visual = needs rendered output; propose, do not auto-apply.

  • incomplete[] holds findings that need a person (contrast over images, marketing alt text). Agents are told not to fix them.

  • source: on React ≤18 and Vue dev builds, each issue carries {file, line, column, component} from the live component tree.

  • verify_fix re-scans one element and fails closed. diff_scan reports fixed[], new[], remaining[] for a page.

Free

What you get

No account

Scan and check tools on the hosted server. Fair-use limits per IP: 30 browser scans/h, 10 AI fixes/h

Free account (OAuth prompt in your client, or npx -y @webability/cli login locally)

Adds visual_audit (vision pass), start_audit / get_audit (full report), and webability-tunnel (lets the hosted server scan your localhost)

No trial, no credits, no paid tier on the MCP.

What it does not do

It cannot judge if alt text is meaningful or if a custom widget makes sense with a screen reader. Use it to clear the automated layer in source, then test with assistive technology.

Local vs hosted

Lite — local stdio (npx / webability-mcp)

Full — hosted (https://mcp.webability.io/mcp)

Account

None

None for scan tools; free account (OAuth sign-in) for visual and full audits

Scan / fix / verify

Yes (on your machine)

Yes

find_source

Yes

No

scan_history / generate_report_pdf

Yes

No

visual_audit / start_audit / get_audit

Listed as stubs → connect Full (still free)

Yes — free with your account

PostHog / dashboard analytics

Optional env only

On by default on hosted

localhost / private addresses

Yes — the browser runs on your machine

Via a tunnel — see below

Scanning a local dev server

Use Lite (stdio). It runs the browser on your machine, so http://localhost:3000 is just localhost:

claude mcp add webability-local -- npx -y -p @webability/mcp webability-mcp

Other clients: add npx -y -p @webability/mcp webability-mcp as a stdio server.

Full (hosted) cannot reach your machine directly, by design. It runs in our cloud, so localhost there means our localhost. Every URL is checked before any fetch and loopback / private / link-local addresses are refused: without that check, anyone could point the server at internal services or a cloud metadata endpoint. That check is not relaxed for anyone.

When you need hosted: webability-tunnel

CI, a remote agent, or a dashboard-triggered scan cannot run Lite, because there is no laptop in the loop. For those, open a tunnel:

WEBABILITY_API_KEY=<your-token> npx -y -p @webability/mcp webability-tunnel --port 3000

It prints a https://tunnel.webability.io/t/<id>/ URL and a secret. Pass the URL as url and the secret as tunnel_secret:

"Scan https://tunnel.webability.io/t/abc123.../ with tunnel_secret &lt;secret&gt;"

Your machine dials out to the relay, so the URL is an ordinary public hostname and the SSRF check above still applies unchanged — nothing is weakened to make this work. Same idea as ngrok, with three differences that matter when the thing on the far side is your dev machine:

  • The URL is not a credential. Every request must carry the secret header; the URL alone returns 401. URLs leak into shell history, CI logs and screenshots.

  • Only GET and HEAD reach you, on the one port you named, and Authorization / Cookie are stripped before anything crosses in.

  • It dies when you do. 30 minutes, 5 minutes idle, or the moment you press Ctrl-C.

Anyone holding both the URL and the secret can read your dev server. Treat the pair like a password, and prefer Lite whenever there is a human at a keyboard.

Third-party tunnels (ngrok, cloudflared) also work — the hosted scanner treats their hostnames like any other public site — but they expose your dev server to anyone who learns the URL. Vite users, either way: add the tunnel hostname to server.allowedHosts, or it answers 403 Blocked request to everything.

start_audit is the exception on both transports — its pipeline runs on our servers even under Lite, so it can never reach a localhost URL.

Local options

Optional env:

  • WEBABILITY_API_URL (default https://api.webability.io) for self-hosted backends.

  • POSTHOG_PROJECT_API_KEY or POSTHOG_API_KEY to enable PostHog MCP Analytics for MCP initialize, tools/list, and tool-call usage events.

  • POSTHOG_HOST (default https://us.i.posthog.com) for EU or self-hosted PostHog ingestion.

  • WEBABILITY_POSTHOG_MCP_ANALYTICS=off to force-disable PostHog MCP Analytics even when a PostHog key is present.

Scan engines

scan_page runs three engines in parallel and deduplicates the results:

Engine

Rules

What it covers

WebAbility detectors

60+

Gradient-aware contrast, weak names, decorative icons, landmark hierarchy, ARIA correctness, link consistency, target size, keyboard traps

axe-core

104

Industry-standard WCAG 2.2 baseline

HTML_CodeSniffer

200+

Section 508 + WCAG techniques cross-reference

Three-tier output (since v1.2.1)

Every scan returns:

  • issues — high-confidence violations, safe to surface as bugs

  • incomplete — findings that need human review (contrast against gradients, marketing imagery, framer-motion pre-animation states, axe-incomplete). Never auto-fix these.

  • summary — counts by severity + an incomplete count

This mirrors axe-core's violations / incomplete / passes split and prevents agents from "fixing" false positives in destructive ways.

Structured fixes (since v1.6.0)

Every issue from scan_page, flow_scan, diff_scan and scan_html carries a machine-readable fix and a fixability tier, so an agent can act without parsing prose:

{
  "id": "wa-missing_button_type-a1b2c3",
  "fixability": "mechanical",
  "fix": { "op": "add-attribute", "attribute": "type", "value": "button", "currentValue": "", "needsManualReview": false }
}

Field

Values

fix.op

add-attribute · set-attribute · remove-attribute · add-element · remove-element · add-text-content · suggest

fixability

mechanical — value known, apply as given · contextual — op known, value needs judgment (alt text, a label) · visual — needs rendered output (contrast, focus ring, target size); propose, never auto-apply

fix.value is present only when the engine already knows it. The legacy fix.attribute / currentValue / suggestedValue / needsManualReview fields are unchanged. get_rules lists the tier for every rule so you can pick the auto-fixable set up front.

Source pointers (since v1.6.0)

On a React ≤18 or Vue dev build, scan_page, flow_scan and diff_scan read the component tree of the live page and attach the JSX call site to each finding:

{ "selector": "img#hero", "source": { "framework": "react", "file": "/app/src/Hero.tsx", "line": 12, "column": 5, "component": "Hero" } }

Open that file — no find_source round-trip. React 19 dropped _debugSource; there you still get component. Production builds have no tree; pass sourceRoot (Lite only) and issues without a pointer get sourceCandidates[] from a token grep of the selector.

Output controls (since v1.6.0)

scan_page, flow_scan, scan_html and diff_scan accept:

Param

Effect

minImpact

minor · moderate · serious · critical — drop anything below

rules[]

keep only these rule ids (missing_alt, image-alt, …)

wcag[]

keep only these criteria; a prefix like 1.4 matches 1.4.3

format: "compact"

one line per element, rule metadata printed once — a fraction of the JSON tokens

Filters apply before the 50-item cap, so minImpact: "serious" returns every serious issue on a 300-issue page. The scan_history archive keeps the unfiltered result.

In-process scan_html (since v1.6.0)

scan_html no longer launches a browser by default. It runs the WebAbility detectors and axe-core inside jsdom in the MCP process — milliseconds, no network — and returns the same three-tier issues / incomplete / summary shape as scan_page, with fix.op on every finding. Fragments are wrapped into a document automatically. jsdom has no layout, so visual-tier rules (contrast, target size, focus ring) are skipped and counted in skippedVisual; pass engine: "browser" for the previous headless-Chromium axe path.

Tools

Tool

Edition

What it does

scan_page

Lite + Full

Scan a URL for WCAG accessibility issues (3 engines). source pointers on dev builds; minImpact / rules / wcag / format output controls

flow_scan

Lite + Full

Multi-page journey scan with deduplicated issues across pages

scan_html

Lite + Full

Scan a raw HTML snippet or fragment in-process (jsdom, ms, no browser); engine: "browser" for contrast

detect_framework

Lite + Full

Detect Tailwind / MUI / Bootstrap / Next.js / WP / plain CSS

generate_ai_fix

Lite + Full

Framework-aware fix alternatives. Auto-extracts brand palette from the live URL on contrast issues.

verify_fix

Lite + Full

Re-scan a fixed element and confirm the violation is gone — verified: true/false. Closes the find → fix → verify loop.

diff_scan

Lite + Full

Baseline vs current → fixed[] / new[] / remaining[]. Page-level regression check; baseline from scan_history (Lite) or a live URL.

check_color_contrast

Lite + Full

WCAG contrast check on a color pair; pass url to get brand-aligned suggestions from the live page

check_aria

Lite + Full

Validate ARIA attributes in an HTML snippet

get_rules

Lite + Full

List axe-core + WebAbility rules with fixability and a fix op template; filter by tag, tier, or engine

find_source

Lite only

Map a CSS selector back to local source files (fallback when the page has no framework source pointer)

scan_history

Lite only

Browse prior local scans under ~/.webability/scans/

generate_report_pdf

Lite only

Turn scan findings into a branded WebAbility accessibility-report PDF saved next to the project — free, no account. Pass the issues[] from scan_page. For the full audit deliverable (Excel + evidence), use start_audit.

visual_audit

Full (free w/ account; stub on Lite)

Pixel-level audit via vision (icon contrast, focus visibility, looks-like-a-button-but-isn't)

start_audit

Full (free w/ account; stub on Lite)

Kick off the full server-side audit deliverable (report + Excel workbook). Returns an id to poll.

get_audit

Full (free w/ account; stub on Lite)

Check an audit's progress and, once complete, get the severity summary + report/workbook download URLs.

When to use this MCP

  • Building a new component and want it accessible from day one

  • Auditing a localhost / staging build before pushing

  • Triaging a Lighthouse / axe report — scan_page consolidates all three engines

  • Generating fix suggestions that match the framework you're already using

  • Checking color contrast against the user's actual brand palette (not generic suggestions)

Examples

In Cursor / Claude Code:

"Scan localhost:3000 for accessibility issues"

"Walk login → dashboard → checkout and report unique issues across the flow"

"Suggest a fix for the contrast issue on .btn-primary on https://example.com — match their brand colors"

"What does WCAG 1.4.11 check?"

"Turn the issues you just found on localhost:3000 into a branded PDF report"

Privacy, scan logs & telemetry

Every scan is logged locally to ~/.webability/scans/ — a one-line-per-scan index.jsonl ledger plus the full result of your last 500 scans. Browse them with the scan_history tool ("what did we scan earlier?") or plain jq. Set WEBABILITY_SCAN_LOG=off to disable, WEBABILITY_SCAN_LOG_DIR to relocate.

The server also reports one small telemetry event per tool call (every tool, not just scans) to the WebAbility API: tool name, a short target label (URL, selector, issue type — never page content), pass/fail, duration, issue counts, and a persistent anonymous install ID. Full scan results, HTML, and generated fix code never leave your machine via telemetry. Set WEBABILITY_SCAN_TELEMETRY=off to opt out.

If POSTHOG_PROJECT_API_KEY or POSTHOG_API_KEY is configured, the server additionally enables PostHog MCP Analytics. This captures MCP usage metadata such as initialize, tools/list, tool name, duration, client name/version, and success/failure. WebAbility strips PostHog's $mcp_parameters and $mcp_response fields before send, so raw HTML snippets, screenshots, scan responses, and generated code are not sent to PostHog by this integration.

Two tools — generate_ai_fix and visual_audit — additionally send page content (an HTML snippet or a screenshot) to WebAbility's API so it can call a third-party LLM on your behalf; WebAbility doesn't store that content, but the LLM provider sees it in transit. See PRIVACY.md for the full per-tool breakdown and WebAbility's privacy policy.

License

MIT

Available Tools

20 tools
add_siteAdd a site to the accountAInspect

Add a website to the signed-in WebAbility account. The site gets the full widget for its first 30 days and a first scan. Returns the site id and plan tier. Next: put the get_install_snippet tag on the site, and use create_upgrade_link to get a payment link for the WebAbility Pro plan. Needs a free WebAbility account (an AI agent can sign itself up with "Sign in with AgentID").

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe site to add, for example https://shop.example.com. One domain per site.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuine context beyond that: the 30-day full-widget trial, the included first scan, the returned site id and plan tier, and the authentication path (agent can self-register via 'Sign in with AgentID'). It does not explicitly state non-idempotency or duplicate-domain behavior, so it is not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action in the first sentence, followed by side effects, return values, and next steps in a compact block. Every sentence carries information, though the trailing instructions are slightly dense and could be split more cleanly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter creation tool with annotations covering the safety profile, the description supplies the missing pieces: prerequisites, side effects, and return contents (no output schema exists). Error conditions and rate limits are not covered, which keeps it below a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage, including an example URL and the 'one domain per site' constraint. The description adds no parameter-level detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Add a website to the signed-in WebAbility account'), which cleanly separates it from siblings like list_sites, scan_page, or get_install_snippet. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear workflow context by naming successors ('put the get_install_snippet tag on the site', 'use create_upgrade_link'), which effectively tells the agent where this tool sits in a sequence. It states the prerequisite (a free WebAbility account) but never says when NOT to use it, so it stops short of explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_ariaValidate ARIA in HTML or on a pageA
Read-onlyIdempotent
Inspect

Validate ARIA attribute + accessible name/role/value usage — in an HTML snippet (html) or on a live page (url), optionally limited to one element and its descendants (selector). Runs axe-core cat.aria and cat.name-role-value rules (aria-* attribute correctness, role validity, required parents/children, aria-hidden-focus, accessible names). Returns violations. A selector that matches nothing, or a page that answers an HTTP error, is an error — never "no violations". Nodes cap at 5 per rule by default — every rule reports nodesTotal + truncated; raise nodeLimit (max 50) or use scan_history(id) for the full set.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoLive page to check (pass this or `html`)
htmlNoHTML to test for ARIA correctness (pass this or `url`)
selectorNoOnly report findings on this element and its descendants (e.g. the `selector` from a scan_page issue). ARIA references outside it still resolve.
nodeLimitNoMax nodes returned per rule (default 5, max 50)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive behavior, and the description adds substantial beyond that: error semantics (a non-matching selector or an HTTP-error page is an error, never 'no violations'), the per-rule node cap of 5 with nodesTotal + truncated flags, the nodeLimit ceiling of 50, and the escape hatch via scan_history. That is exactly the extra context the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is dense and front-loaded, leading with what is validated and which inputs select which mode, then layering error and truncation semantics. The parenthetical rule list and cap details make it a long single paragraph, but nearly every clause carries information an agent needs; only slight compression would help.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description states what is returned (`violations`), how results can be truncated, and how to retrieve the full set via scan_history. Combined with the error-vs-empty distinction, an agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the html/url mutual exclusivity with its 'pass this or `html`' framing, the note that ARIA references outside the `selector` still resolve, and the default/max bounds on nodeLimit tied to truncation behavior. It does not add syntax examples or edge-case formats for the string inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (validate) plus precise resources (ARIA attribute + accessible name/role/value usage) and the two input modes (html snippet or live url). It also enumerates the underlying axe-core rule categories (cat.aria, cat.name-role-value) so an agent can distinguish it from scan_page, scan_html, and check_color_contrast without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: use with `html` for a snippet or `url` for a live page, optionally narrowed with `selector`, and it names scan_history(id) as the alternative when the truncated node set is insufficient. It does not explicitly say when to prefer this over sibling scanners like scan_html or scan_page, so it stops short of full alternative-routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_color_contrastCheck color contrastA
Read-onlyIdempotent
Inspect

Check text contrast against the WCAG AA threshold (4.5:1, or 3:1 for large text); the AAA verdict is shown only when you pass level: "AAA". Either pass a foreground / background color pair, or pass url + selector to read the element's own text color, background (composited from the nearest painted ancestors), font size and weight from the live page; explicit colors win over the page. A gradient or image background is reported as an error, never a guessed ratio. When it fails, suggests BRAND-aligned replacements — extracts the actual brand palette from the url page using our scanner (CSS vars + most-used colors), or use a provided brandColors array. No url and no brandColors = ratio + pass/fail only.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoLive page: with `selector`, the colors are read from that element; on a failing pair, the brand palette is extracted from it (uses our scanner — CSS vars + dominant colors).
levelNoConformance level to judge. Default "AA". "AAA" adds the enhanced-contrast (1.4.6) verdict and AAA palette suggestions.
isBoldNoWhether text is bold (default false, or the element's computed weight with url + selector)
fontSizeNoFont size in px (default 16, or the element's computed size with url + selector)
selectorNoCSS selector of the text element on `url` (e.g. the `selector` from a scan_page issue)
backgroundNoBackground color (hex or rgb). Optional with url + selector.
foregroundNoForeground color (hex or rgb). Optional with url + selector.
brandColorsNoPre-supplied brand palette. Skips URL extraction if provided.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: the AA threshold values, AAA verdict only when level is 'AAA', error behavior for gradient/image backgrounds, brand-aligned replacement suggestions, and the scanner-based palette extraction. These details matter for correct use and interpretation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core purpose and thresholds come first, then invocation modes, then edge cases and suggestions. Every clause adds useful information, though the single paragraph has many semicolon-separated ideas that could be slightly broken up.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema and rich behavior, the description covers thresholds, input modes, precedence rules, error handling, and output shapes (ratio + pass/fail, AAA verdict, brand suggestions). It gives the agent enough to call the tool correctly and anticipate results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes further by explaining precedence between explicit colors and live-page extraction, the compositing of background from nearest painted ancestors, and that brandColors skips URL extraction, adding meaningful behavioral semantics to the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Check text contrast against the WCAG AA threshold (4.5:1, or 3:1 for large text)'. It also distinguishes itself from siblings by referencing the selector from a scan_page issue and by scoping the tool to contrast checking rather than page-wide auditing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly describes the two main invocation modes: pass a foreground/background pair, or pass url + selector to read colors from the live page. It also clarifies precedence ('explicit colors win over the page') and the fallback behavior when neither url nor brandColors is provided. No explicit when-not-to-use guidance or named alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_frameworkDetect the site frameworkA
Read-onlyIdempotent
Inspect

Detect a page's stack as two separate fields: framework — the application framework, CMS or site builder (e.g. nextjs, nuxt, sveltekit, vitepress, astro, gatsby, wordpress, shopify, mediawiki, vue, react; "unknown" when no signal) — and cssToolkit (tailwind, bootstrap, mui, plain-css). Lists the evidence it used and the value to pass as generate_ai_fix framework. scan_page reports the same two fields from the same detection.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to inspect

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and open-world behavior, so the safety profile is covered. The description adds real behavioral context beyond that: the 'unknown' fallback when no signal exists, that evidence is listed alongside the verdict, and that detection is shared/identical with scan_page's fields. It does not discuss failure modes for unreachable URLs, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense but well-structured sentence front-loads the verb and the two outputs, then chips in the evidence/interop note. Every clause (field names, example values, 'unknown' sentinel, scan_page equivalence) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one parameter and no output schema, the description adequately substitutes by naming the two returned fields with representative values and the sentinel for no-signal cases, plus the downstream consumer. Nothing critical is missing for a low-complexity detection tool, though error/empty-result behavior is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required `url` parameter already documented at 100% schema coverage ('URL to inspect'), so the schema carries the parameter burden. The description adds no format, protocol, or normalization details for the URL, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Detect a page's stack') and enumerates the two returned fields with concrete example values, so the agent knows exactly what it obtains without opening a schema. It also distinguishes itself from scan_page by noting that sibling reports the same two fields from the same detection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear use case: obtain the `framework` value to pass into generate_ai_fix, and clarifies overlap with scan_page rather than duplicating it. It stops short of stating explicit when-not conditions (e.g. use scan_html for raw HTML, or that scan_page is preferable when a full scan is already needed), so it is clear context but not full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff_scanCompare two accessibility scansA
Read-onlyIdempotent
Inspect

Compare two scans of the same page and report what changed: fixed[] (in the baseline, gone now), new[] (regressions — not in the baseline, present now), remaining[] (still there). Page-level complement to verify_fix (one element). Baseline is a scan_history id (baselineId, local installs) or a live scan of baselineUrl; current is url (scanned live now) or another history id (currentId). Findings are matched by issue id (rule + element), so a changed class/id on a fixed element reads as fixed AND new — check new[] before calling it a regression. Typical loop: scan_page → edit → diff_scan(baselineId=, url=) → confirm new[] is empty.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoURL to scan now as the CURRENT side (deployed, staging, or http://localhost:3000). Omit when passing currentId.
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
levelNoHighest WCAG conformance level to report. Default "AA" — Level AAA criteria are not reported. Pass "AAA" to include them (naming one AAA criterion in `wcag`, e.g. "3.2.5", also includes it).
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
formatNo"compact" prints one line per element with rule metadata once. "json" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.
viewportNoViewport for live scans (default: desktop). Use the same viewport the baseline used.
currentIdNoscan_history id to use as the CURRENT side instead of scanning `url`
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
baselineIdNoscan_history id of the BASELINE scan (local installs only)
baselineUrlNoScan this URL live as the baseline (e.g. production) — use when there is no stored baseline
rootSelectorNoCSS selector to limit live scans to (optional)

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely non-obvious behavior beyond that: findings are matched by issue id (rule + element), and a changed class/id on a fixed element will appear as BOTH fixed and new — a false-regression trap an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the output shape (what changed) before the plumbing, and every clause carries information. It is dense — the matching-semantics sentence and the typical-loop sentence are both long — but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden and does so by naming and defining fixed[]/new[]/remaining[]. Combined with the baseline/current input pairing and the matching caveat, an agent has enough to call it correctly without follow-up.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it pairs the optional inputs (baseline side = baselineId or baselineUrl; current side = url or currentId) and notes baselineId is local-installs-only. That encoding of mutual alternatives goes beyond the flat per-parameter schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Compare two scans of the same page') and immediately enumerates the three output buckets (fixed/new/remaining). It explicitly distinguishes itself from the sibling verify_fix by scope ('Page-level complement to verify_fix (one element)'), so an agent can route between them without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use recipe ('Typical loop: scan_page → edit → diff_scan(...) → confirm new[] is empty') plus the conditions selecting each input pair (baselineId vs baselineUrl, url vs currentId). It also warns to check new[] before declaring a regression, which is actionable guidance, not just context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_sourceFind the source file for a selectorA
Read-onlyIdempotent
Inspect

Find source files in the local project that contain a given CSS selector. Maps DOM selectors back to source code so you can edit the right file. Searches React/Vue/Svelte/HTML/PHP/Astro files.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootDirNoProject root to search (defaults to cwd)
selectorYesCSS selector from a scan issue

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety and idempotency are covered. The description adds file-type coverage (React/Vue/Svelte/HTML/PHP/Astro) but does not disclose return format, false-positive behavior, or performance characteristics. With annotations covering the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each with a purpose: action, rationale, and file-type scope. No wasted words and the core use case is front-loaded. Could be slightly tighter but is well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only lookup tool with full schema coverage and no output schema, the description provides enough context: what it does, why (edit the right file), and supported file types. Missing return format details, but those are minor for a read-only tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema fully documents both parameters (selector and rootDir). The description reinforces 'given CSS selector' but adds no syntax or format details beyond the schema. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb ('Find') and resource ('source files') with a specific scope: files containing a given CSS selector. Distinguishes from siblings like scan_page or get_audit, but does not explicitly name an alternative tool for related tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: mapping a DOM selector from a scan issue back to code. No explicit when-not-to-use or alternative tools are mentioned, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_scanScan a multi-page user flowA
Read-onlyIdempotent
Inspect

Scan a multi-page user journey. Walks startUrl plus the required autoNavigate URLs sequentially (deterministic — one page fully rendered and scanned before the next), then returns ONE consolidated report with issues deduplicated across pages, each carrying the same fix payload as scan_page. Every requested URL gets an explicit outcome in pages[] (scanned / nav_failed / scan_failed / redirected_duplicate / duplicate_request / skipped_cap / blocked — bot-challenge, not a clean page) — a page is never silently dropped. Better than per-page scans for journeys (login → checkout etc). For a single page, use scan_page.

ParametersJSON Schema
NameRequiredDescriptionDefault
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
levelNoHighest WCAG conformance level to report. Default "AA" — Level AAA criteria are not reported. Pass "AAA" to include them (naming one AAA criterion in `wcag`, e.g. "3.2.5", also includes it).
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
formatNo"compact" prints one line per element with rule metadata once. "json" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.
maxPagesNoMax pages to scan (default 10)
startUrlYesStarting URL of the journey
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
sourceRootNoLocal project root (local installs only) for `sourceCandidates[]` on issues without a framework `source` pointer.
autoNavigateYesREQUIRED — the URLs to walk after startUrl (the MCP server is headless and cannot discover a journey interactively). For a single page, use scan_page instead.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/openWorld/idempotent annotations: it discloses deterministic sequential execution (one page fully rendered before the next), single consolidated deduplicated report with the same fix payload as scan_page, and an enumerated per-URL outcome taxonomy including blocked bot-challenges. The guarantee that a page is "never silently dropped" is exactly the kind of behavior an agent needs and cannot get from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what it is, then how it walks, then what it returns, then the routing rule. Dense but every sentence carries information; the long parenthetical status list is slightly packed but each status is decision-relevant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description still covers the return contract (one consolidated report, pages[] with explicit per-URL outcomes) and the invocation constraint on autoNavigate. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description earns an extra point by adding execution semantics to autoNavigate (deterministic, sequential, headless server cannot discover the journey interactively) that the schema's terse 'URLs to walk after startUrl' does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (scan a multi-page user journey covering startUrl plus autoNavigate) and explicitly contrasts itself with the sibling scan_page. An agent can distinguish it from scan_page without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative ("For a single page, use scan_page") and the selecting condition ("Better than per-page scans for journeys (login → checkout etc)"). No inference required to route between the two scan tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_ai_fixSuggest a framework-aware fixAInspect

Generate framework-aware fix alternatives for a specific accessibility issue. For color contrast issues, returns 3 alternatives (minimal, brand-aligned, high contrast); brand palette is auto-extracted from the live URL using our scanner if brandColors is omitted. For label/ARIA issues, returns 1-2 alternatives. Each alternative includes ready-to-paste code for the detected framework.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoPage URL — also used to auto-extract brand palette for contrast issues if `brandColors` is not provided.
htmlYesThe element's outerHTML — send at most ~600 chars
issueYesIssue object from scan_page (with selector, wcag, impact, message, fix.currentValue), or a plain-text issue description
contextNoParent element outerHTML for context (~400 chars)
frameworkYesCSS framework — use detect_framework first
brandColorsNoBrand palette for brand-aligned suggestions. If omitted on a contrast issue with a `url`, auto-extracted via the scanner.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations present, the description still adds real context beyond them: the count of alternatives returned per issue class, that brand palette is auto-extracted from the live URL when brandColors is omitted, and that output is ready-to-paste code per framework. It doesn't cover error modes or rate/scanning costs, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action and then branching by issue type; each sentence carries information. Minor redundancy between 'brand-aligned' and the brand-palette explanation, but no wasted padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden and does so well (alternative counts, ready-to-paste code, per-framework targeting). With six schema-documented params and a live-scanning dependency, it is nearly complete, though it omits what happens on a framework/issue mismatch or scanning failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reinforces the brandColors/url interaction but adds little syntax or format detail beyond what the schema already documents for url, html, issue, context, framework and brandColors.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate) and resource (framework-aware fix alternatives) scoped to 'a specific accessibility issue,' and further narrows by issue type (color contrast vs label/ARIA). This is enough to separate it from siblings like verify_fix and check_color_contrast without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear the context in which it applies (contrast vs label/ARIA issues, detected framework) and notes the brandColors/url fallback, so an agent knows when the tool is the right call. It stops short of explicitly naming alternatives or stating when not to use it, so it does not reach 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_report_pdfCreate an accessibility report PDFA
Destructive
Inspect

Turn scan findings into a branded WebAbility accessibility-report PDF, saved next to the project (local installs only). Pass the issues[] array a scan_page call returned plus the page url; the tool groups findings by functionality (Low Vision / Mobility / Navigation / Content / Cognitive), computes the WCAG score with the same penalty ladder as the platform, and renders a shareable PDF — free, no account. Use it when the user wants a deliverable to attach to an email, ticket, or compliance thread. Capped at 500 findings. For the full audit deliverable (Excel workbook + evidence screenshots), use start_audit instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe scanned page URL (used for the report header and file name)
issuesYesFindings to include — pass the `issues[]` array from scan_page (message/type/wcag/selector/html/impact/fix).
widgetDetectedNoWhether the WebAbility widget was detected on the page (shown in the report's widget status).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the write/open-world/destructive profile, and the description adds meaningful context beyond them: local installs only, capped at 500 findings, free/no account, saved next to the project, and the grouping/scoring behavior. It does not explain the destructiveHint (e.g. whether an existing PDF is overwritten) or what a successful return looks like, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then input requirements, behavior, usage condition, limit, and alternative in a logical order. It is dense but nearly every sentence earns its place; only slight trimming would improve it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a file-generating tool with no output schema, the description covers what is produced, where it is saved, and the 500-item cap, which is close to sufficient. It never explicitly states the return value (e.g. the saved file path), which is the one missing piece an agent would want.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine provenance: pass the issues[] array that a scan_page call returned, tying the parameter to a specific upstream tool. The 500-finding cap is also a relevant constraint on that parameter. widgetDetected is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (turn scan findings into) and resource (branded WebAbility accessibility-report PDF), plus the input it consumes and its output. It explicitly distinguishes itself from the sibling start_audit, so an agent can pick between them without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the when-to-use condition (user wants a deliverable to attach to an email, ticket, or compliance thread) and the when-not/alternative (for the full audit deliverable, use start_audit instead). This is explicit routing rather than implied usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_auditGet audit status and reportA
Read-onlyIdempotent
Inspect

[Full — FREE with a WebAbility account. Not runnable in Lite. Connect hosted MCP at https://mcp.webability.io/mcp (dashboard or Smithery: webability/accessibility), authenticate, then call get_audit.] Check an audit started with start_audit: returns overall status, per-step progress (scan → viewports → screenshots → agent → excel → publish), and — once complete — a severity summary plus short-lived download URLs for the report (JSON) and the Excel workbook. Poll every ~15s while status is pending/running. Only the account that started an audit can read it. Free with a WebAbility account.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe audit id returned by start_audit

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the read-only annotations: ownership restriction, recommended polling cadence, the progress pipeline, and short-lived download URLs for JSON and Excel outputs. These are operationally important behaviors an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose and polling guidance are clear and mostly front-loaded after the availability note. Minor redundancy appears with 'FREE' in the bracket and 'Free with a WebAbility account' repeated at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an async audit-status tool with no output schema, the description covers return values, progress stages, completion artifacts, polling behavior, and access restrictions. The single input parameter is fully documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'id' parameter is already documented as the audit id returned by start_audit. The description reinforces that source but adds no syntax or format details beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check') and resource ('an audit started with start_audit'), and describes exactly what it returns: overall status, per-step progress, and completion artifacts. This clearly distinguishes it from the sibling start_audit tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to check an audit started with start_audit, gives polling guidance ('Poll every ~15s while status is pending/running'), and states an access exclusion ('Only the account that started an audit can read it'). It also explains the hosted MCP connection prerequisite.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_install_snippetGet the widget install tagA
Read-onlyIdempotent
Inspect

Get the one-line script tag that installs the WebAbility accessibility widget. Put it in the or before of every page of a site added with add_site. The widget finds the site by its domain, so the same tag works on every site. No account needed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive, closed-world behavior with no parameters. The description adds genuinely useful context beyond that: no account is needed, the widget resolves the site by domain, and the same tag works across all sites. It does not address caching or whether the tag changes over time, but the extra detail is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences front-loaded with what the tag is, followed by placement and the domain/no-account facts. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no input params and no output schema, the description fully covers what an agent needs: what the tool returns, where the snippet goes, and the key constraint that it is domain-based and account-free.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so nothing to document; baseline applies. The description compensates by explaining that the tag is domain-resolved rather than parameter-driven, which tells the agent why no input is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get the one-line script tag') and names the exact artifact produced (the install tag for the WebAbility widget). It is distinguishable from siblings like add_site or list_sites without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the context of use: place it in the <head> or before </body> of a site added with add_site. This ties invocation to a workflow step, though it gives no explicit when-not-to-use or alternate tool comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_rulesList accessibility rulesA
Read-onlyIdempotent
Inspect

List accessibility rules from both engines — axe-core (104) and the WebAbility detectors (90+) — with optional filters. Every rule carries fixability (mechanical | contextual | visual) and a fix op template, so you can pick the rules worth auto-fixing before scanning. Returns ruleId, engine, description, help, helpUrl, tags/wcag, fixability, fix.

ParametersJSON Schema
NameRequiredDescriptionDefault
ruleNoOnly rules whose id contains these whole words (e.g. "link" → link-name, skip-link, not blink; "color-contrast"; case and -/_ insensitive). Also accepted: ruleId, id, query.
tagsNoaxe tag filter (e.g. ["wcag2aa"], ["best-practice"], ["cat.aria"]; a category also works without "cat."). WebAbility rules match on their WCAG criterion tag (e.g. "wcag143"). An unknown tag or category is an error that lists the valid ones.
engineNoWhich engine's rules to list (default all)
fixabilityNoOnly rules of this fixability tier

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds genuine context beyond that: rules span two engines, each carries a fixability tier and a fix op template, which tells the agent what differentiates the output. It omits ordering/pagination, keeping it at a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then the value proposition and return fields in two tight sentences. The trailing return-field enumeration is slightly list-like, but it is useful given there is no output schema, so it earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates returned fields (ruleId, engine, tags/wcag, fixability, fix) and explains both engines and the filter dimensions. Minor gaps remain around result ordering or size limits, but nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema descriptions are detailed (whole-word matching, tag formats, engine enum semantics), so the schema does the heavy lifting. The description only refers to 'optional filters' generically and adds no syntax or format meaning beyond the schema, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('List accessibility rules') and scopes it precisely to two named engines with concrete counts (axe-core 104, WebAbility 90+). An agent can immediately distinguish this rule-catalog tool from scanning siblings like scan_page or check_aria.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear use context ('pick the rules worth auto-fixing before scanning'), which implies when to call it ahead of a scan. However, it never names an alternative tool or an explicit when-not-to-use case, so routing is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sitesList the account sites and plansA
Read-onlyIdempotent
Inspect

List the sites on the signed-in WebAbility account with each site id, plan tier and the date the current plan ends. Use it to check that a payment activated WebAbility Pro on a site. Needs a free WebAbility account.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and non-destructive behavior, so the safety profile is covered. The description adds genuine context beyond them: it scopes results to the signed-in account and states the auth prerequisite (free account), which the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what is returned, the use case, and the prerequisite. Every clause carries information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description compensates by naming the fields returned (site id, plan tier, plan end date) and the auth requirement. An agent has everything needed to call it and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a 0-param tool is 4. No misleading detail is added either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (sites on the signed-in WebAbility account) and even enumerates the returned fields (site id, plan tier, plan end date). It is clearly distinguishable from the audit/scan siblings and from add_site.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit use case ("check that a payment activated WebAbility Pro on a site") and a prerequisite ("Needs a free WebAbility account"), so the agent knows when to reach for it. It does not name a when-not condition or an alternative tool, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_historyList past scansA
Read-onlyIdempotent
Inspect

Browse past scans run through this MCP server. Every scan_page / flow_scan / scan_html / visual_audit / check_aria / verify_fix call is logged locally (~/.webability/scans; local installs only — the hosted server keeps no history). Without arguments, lists recent scans (when, what target, result summary). Pass id to retrieve the FULL stored result of one past scan, or filter to match a URL/tool substring.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoA scan id from the history list — returns that scan's full stored response
limitNoMax history entries to return (default 20)
filterNoSubstring match on target URL or tool name (e.g. "webability.io" or "scan_page")

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/local-only, so the safety profile is covered. The description adds meaningful context beyond that: the on-disk storage path (~/.webability/scans) and the crucial caveat that the hosted server keeps no history, which materially affects expectations. It does not describe pagination or ordering beyond the default list behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then progressively adds scope, storage caveat, and parameter behavior. Every sentence carries information an agent needs; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by stating what the default list returns (when, target, result summary) and what `id` returns (full stored result). Combined with the storage-location caveat, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3; the description earns above baseline by emphasizing that `id` returns the FULL stored result (not just a summary row) and that `filter` matches URL or tool substring. It adds real semantic shading beyond the schema, though it doesn't restate the `limit` default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (browse past scans) and immediately scopes it by naming the exact sibling tools whose calls get logged (scan_page, flow_scan, scan_html, visual_audit, check_aria, verify_fix). An agent can distinguish this history/retrieval tool from the scan-execution siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear mode-based guidance: no arguments lists recent scans, pass `id` for a full stored result, pass `filter` for substring matching. It does not name competing siblings (e.g. diff_scan, get_audit) or state when NOT to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_htmlScan an HTML snippetA
Read-onlyIdempotent
Inspect

Scan a raw HTML snippet or component markup without serving it — IN-PROCESS by default (jsdom + WebAbility detectors + axe-core): milliseconds, no browser, no network, so it fits inside a tight edit loop. Fragments are auto-wrapped into a document. Returns scan_page's shape (issues / summary) with fix.op + fixability on every finding. jsdom has no layout, so visual-tier rules (contrast, target size, focus ring) are NOT evaluated — the dropped count is reported as skippedVisual; pass engine: "browser" to run the axe-core headless-browser path for those (slower, axe rules only, returns axe violations).

ParametersJSON Schema
NameRequiredDescriptionDefault
htmlYesHTML content to test — a full document or a fragment
tagsNoWCAG tags to check (default ["wcag2a","wcag2aa","wcag21aa","wcag22aa"])
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
levelNoHighest WCAG conformance level to report. Default "AA" — Level AAA criteria are not reported. Pass "AAA" to include them (naming one AAA criterion in `wcag`, e.g. "3.2.5", also includes it).
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
widthNoViewport width (browser engine only, default 1280)
engineNo"in-process" (default): jsdom, ms, structural rules. "browser": headless Chromium + axe-core, includes contrast.
formatNo"compact" prints one line per element with rule metadata once. "json" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.
heightNoViewport height (browser engine only, default 800)
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent annotations: it discloses the execution model (jsdom + WebAbility detectors + axe-core, no browser, no network), the fragment auto-wrapping behavior, and the critical limitation that visual-tier rules (contrast, target size, focus ring) are NOT evaluated with the dropped count surfaced as skippedVisual. This is exactly the kind of non-obvious behavioral context that prevents wrong invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose before the execution details, and each sentence carries information. The heavy use of em-dash parentheticals makes it dense, but it is not padded with filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on the burden of explaining the return shape ('scan_page's shape (issues / summary) with fix.op + fixability'), plus the compact-vs-json output behavior and skippedVisual. An agent has enough to call it and interpret the result without opening anything else.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, so the baseline is 3, but the description adds real meaning to engine (in-process vs browser tradeoff and rule coverage), format defaults, and the skippedVisual consequence of the engine choice. It reinforces and disambiguates the schema rather than merely repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scan a raw HTML snippet or component markup') plus the key scoping constraint ('without serving it'), which distinguishes it from the served-page sibling scan_page. An agent can tell immediately what it does and roughly where it sits relative to scan_page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly frames the default in-process path as fitting 'inside a tight edit loop' and states the condition for switching (pass engine:"browser" for visual-tier rules), which is explicit when-to-use guidance. It does not, however, explicitly name scan_page as the alternative for already-served URLs, so sibling routing is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_pageScan a page for accessibility issuesA
Read-onlyIdempotent
Inspect

Scan a web page for WCAG accessibility issues. Works on any URL — deployed sites, localhost, staging. Returns issues (violations the scanner stands behind; uncertain findings are judged or dropped, never listed for review) and a summary. On React ≤18 / Vue dev builds each issue carries source ({file, line, column, component}) read from the live component tree. Every issue carries a structured fix.op (add-attribute | set-attribute | remove-attribute | add-element | remove-element | add-text-content | suggest) with fix.attribute / fix.value when known, and a fixability tier (mechanical = apply as given; contextual = op known, value needs judgment; visual = needs rendered output, propose only).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to scan (e.g. https://webability.io or http://localhost:3000)
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
levelNoHighest WCAG conformance level to report. Default "AA" — Level AAA criteria are not reported. Pass "AAA" to include them (naming one AAA criterion in `wcag`, e.g. "3.2.5", also includes it).
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
formatNo"compact" prints one line per element with rule metadata once. "json" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.
viewportNoViewport size (default: desktop)
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
sourceRootNoLocal project root (local installs only). Issues without a framework `source` pointer get `sourceCandidates[]` — files whose contents match the selector's id/class/attribute tokens.
rootSelectorNoCSS selector to limit scan scope (optional)

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already covering read-only/idempotent/open-world, the description still adds substantial behavioral detail beyond them: uncertain findings are 'judged or dropped, never listed for review', source pointers come only from React C18/Vue dev builds, and each issue carries a fixability tier explaining how far the agent may act on it. This is exactly the extra context the annotations cannot supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, and the dense second sentence is justified because there is no output schema to carry the return-shape contract. It is on the long side and packs return semantics into a single block, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description covers the return shape (issues, summary, source, fix.op, fixability) unusually well and discloses how uncertain findings are handled. It omits operational facts such as whether protected/authenticated pages can be scanned, timeouts, or how results relate to scan_history, which keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with thorough per-parameter docs (wcag prefixes, level default and AAA interaction, format fallback, minImpact ordering, sourceRoot candidates), so the schema carries the burden. The description spends its words on return values rather than adding parameter meaning, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Scan a web page for WCAG accessibility issues') and scopes it crisply with 'Works on any URL — deployed sites, localhost, staging', which implicitly separates it from a static-HTML sibling. However, no sibling tool (scan_html, flow_scan, visual_audit) is named, so the differentiation is left for the agent to infer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'works on any URL' scope line implies when the tool is applicable, but there is no explicit when-to-use / when-not guidance and no routing to alternatives among the 19 sibling tools. An agent must infer that a live-URL scan is different from flow_scan or visual_audit on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_auditStart a full accessibility auditAInspect

[Full — FREE with a WebAbility account. Not runnable in Lite. Connect hosted MCP at https://mcp.webability.io/mcp (dashboard or Smithery: webability/accessibility), authenticate, then call start_audit.] Kick off a FULL accessibility audit deliverable for a URL — a persistent, timestamped artifact, not an inline scan. Runs the server-side pipeline (axe + advanced checks + mobile viewports + annotated screenshots + optional agent spot-check) and produces a downloadable report and a formatted Excel workbook (Cover / Status / Barriers / ADA context sheets) stored durably. Returns immediately with an audit id; poll get_audit for progress and, when complete, download URLs. Use this when someone needs a durable artifact to attach as evidence of testing effort for a compliance officer or legal response — for iterating on code, use scan_page + verify_fix instead. Free for everyone; needs a free WebAbility account (sign in by connecting https://mcp.webability.io/mcp/auth, or run webability login). Set includeAgent:true to add the slower agentic manual-audit pass. To audit a local dev server, open a tunnel (webability-tunnel --port 3000) and pass its URL as url with the printed secret as tunnel_secret; keep the tunnel open until get_audit reports complete (about 5 minutes) — the pipeline loads the page several times.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to audit (a public/staging URL the server can reach, or a webability-tunnel URL — not localhost)
includeAgentNoAlso run the agentic manual-audit pass (keyboard/focus/modal exploration). Slower. Default false.
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly=false, idempotent=false, openWorld=true; the description adds substantially more: async fire-and-return with an audit id, polling via get_audit, durable server-side storage, downloadable report plus Excel workbook, ~5 minute runtime, the page being loaded several times, and the requirement to keep the tunnel open until completion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Content is dense and mostly earns its place, but the bracketed setup paragraph plus a second 'Free for everyone' statement duplicate the free/account requirement, and the tunnel workflow detail is long. Front-loading of the deliverable framing is good, but there is redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully carries the return-value burden: it explains the immediate id return, the get_audit polling loop, and the eventual download URLs. For a multi-step async tool with auth and tunnel prerequisites, nothing an agent needs to invoke and follow up is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented. The description restates includeAgent and tunnel_secret behavior largely in line with the schema, adding only the workflow note that the tunnel must stay open; baseline 3 applies when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (start a FULL accessibility audit deliverable for a URL) and immediately scopes it as a persistent, timestamped artifact rather than an inline scan. It explicitly names the sibling it is not (scan_page + verify_fix for iterating on code), so an agent can differentiate without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('durable artifact to attach as evidence for a compliance officer or legal response') and when-not-to-use with the named alternative (scan_page + verify_fix). It also states prerequisites: free account, hosted MCP connection, and the tunnel workflow for local dev servers.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_fixVerify an accessibility fixA
Read-onlyIdempotent
Inspect

Re-scan a specific element after applying an accessibility fix and confirm the violation is gone — closes the loop that find-only tools leave open. After you edit the code and serve it (deployed, staging, or http://localhost:3000), call this with the URL and the selector you fixed to get a machine-checked verified: true|false (DOM engines only — visual_audit findings are out of scope). Pass the WCAG criterion (e.g. "1.1.1") or axe rule id (e.g. "color-contrast") to check just that criterion; omit it to require the element be clean of ALL violations. A blocked page (bot-challenge / HTTP error) is reported as unverified, never a pass — verification fails closed. If the selector matches no element, the result is verified: false with reason "not-found" — pass the element's current selector, or re-run scan_page if your fix removed the element. Pair with scan_page → generate_ai_fix → verify_fix for a full find-fix-verify cycle.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL now serving the fix (deployed, staging, or http://localhost:3000). Also accepted as `page` or `pageUrl`.
wcagNoOptional: WCAG criterion (e.g. "1.1.1", "1.4.3") or axe rule id (e.g. "color-contrast") to verify specifically. Omit to require the element be free of ALL violations.
selectorYesCSS selector of the element you fixed — use the `selector` from the original scan_page issue
viewportNoViewport size (default: desktop). Use the same viewport the issue was found at.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/openWorldHint, but the description adds substantial behavior beyond them: verification fails closed, a blocked page (bot-challenge/HTTP error) yields unverified rather than a pass, and a no-match selector returns verified: false with reason "not-found". These are exactly the failure-mode details an agent needs to interpret results correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the loop it closes, and every sentence carries operational content. It is denser and longer than strictly necessary — several em-dash clauses and edge cases could be tightened — but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only verification tool with no output schema, the description still conveys the return shape (verified boolean plus a reason value), the engine limitation (DOM engines only), and the recovery path for a missing selector. Nothing an agent needs to call it or interpret it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description still adds meaning: the WCAG/axe parameter narrows verification to one criterion while omitting it requires the element to be clean of ALL violations, and the selector should be the one carried over from the original scan_page issue. It also advises reusing the same viewport the issue was found at, which the schema does not state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: re-scan a specific element after a fix and confirm the violation is gone, producing a machine-checked verified: true|false. It explicitly scopes itself against siblings ("closes the loop that find-only tools leave open") and excludes visual_audit findings, so an agent can distinguish it from scan_page, visual_audit, and diff_scan without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit preconditions (edit the code and serve it — deployed, staging, or localhost), the exact moment to call it, and the workflow it belongs to (scan_page → generate_ai_fix → verify_fix). It also names what is out of scope (visual_audit findings) and what to do when the selector matches nothing (pass the current selector or re-run scan_page), which is unusually complete routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visual_auditVisual accessibility auditAInspect

[Full — FREE with a WebAbility account. Not runnable in Lite. Connect hosted MCP at https://mcp.webability.io/mcp (dashboard or Smithery: webability/accessibility), authenticate, then call visual_audit.] Pixel-level accessibility audit using Claude vision. Catches issues that DOM scanners miss: icon contrast (1.4.11), focus visibility (2.4.7), "looks like a button but isn't" (4.1.2), text rendered as images (1.4.5), visual hierarchy mismatches. Takes a URL, opens it in a headless browser, screenshots, and runs vision-based detection. Complements scan_page — run both for full coverage. Free for everyone; sign in with a free WebAbility account for vision and full audits (sign in by connecting https://mcp.webability.io/mcp/auth, or run webability login). Fair-use rate limits apply.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to audit visually
fullPageNoCapture full scrolled page instead of just viewport (default: false)
viewportNoViewport size (default: desktop)
brandColorsNoBrand hex colors for context-aware filtering

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare openWorldHint=true and non-readOnly/non-idempotent, but the description adds material context annotations cannot convey: it requires authentication, has fair-use rate limits, and drives a headless browser. It does not explain why the tool is marked non-readOnly, leaving that annotation unqualified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The technical core is strong, but the text is padded with repeated onboarding/billing content — the WebAbility account requirement, MCP URL, and sign-in instructions appear twice across bracketed and closing sentences. The first sentence is entirely setup boilerplate rather than purpose, hurting front-loading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Prerequisites, capabilities, and sibling relationship are covered well, but with no output schema the description omits what the audit actually returns (report shape, issue list, severity). For an audit tool this leaves the agent guessing about the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (url, fullPage, viewport, brandColors) are already documented in the schema. The description adds only the redundant note that it 'takes a URL' and gives no extra syntax or format guidance, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Pixel-level accessibility audit using Claude vision') and enumerates the concrete issue classes it detects with WCAG references. It distinguishes itself from DOM scanners and from the sibling scan_page, so an agent can tell exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the agent relative to an alternative: 'Complements scan_page — run both for full coverage,' and frames its niche as 'issues that DOM scanners miss.' Prerequisites (account required, not runnable in Lite) are also stated. It lacks any explicit when-not-to-use guidance, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv1.8.0
    • Addedadd_site
    • Changedcheck_color_contrast1 field changed
      • addedInput schema / properties / level
        Added value: +{
        +  "description": "Conformance level to judge. Default \"AA\". \"AAA\" adds the enhanced-contrast (1.4.6) verdict and AAA palette suggestions.",
        +  "enum": [
        +    "AA",
        +    "AAA"
        +  ],
        +  "type": "string"
        +}
    • Addedcreate_upgrade_link
    • Changeddiff_scan1 field changed
      • addedInput schema / properties / level
        Added value: +{
        +  "description": "Highest WCAG conformance level to report. Default \"AA\" — Level AAA criteria are not reported. Pass \"AAA\" to include them (naming one AAA criterion in `wcag`, e.g. \"3.2.5\", also includes it).",
        +  "enum": [
        +    "A",
        +    "AA",
        +    "AAA"
        +  ],
        +  "type": "string"
        +}
    • Changedflow_scan1 field changed
      • addedInput schema / properties / level
        Added value: +{
        +  "description": "Highest WCAG conformance level to report. Default \"AA\" — Level AAA criteria are not reported. Pass \"AAA\" to include them (naming one AAA criterion in `wcag`, e.g. \"3.2.5\", also includes it).",
        +  "enum": [
        +    "A",
        +    "AA",
        +    "AAA"
        +  ],
        +  "type": "string"
        +}
    • Changedgenerate_report_pdf1 field changed
      • changedInput schema / properties / issues / description
        Previous value: -"Findings to include — pass the `issues[]` array from scan_page (message/type/wcag/selector/html/impact/fix). Needs-review `incomplete[]` items are NOT valid input."New value: +"Findings to include — pass the `issues[]` array from scan_page (message/type/wcag/selector/html/impact/fix)."
    • Addedget_install_snippet
    • Addedlist_sites
    • Changedscan_html1 field changed
      • addedInput schema / properties / level
        Added value: +{
        +  "description": "Highest WCAG conformance level to report. Default \"AA\" — Level AAA criteria are not reported. Pass \"AAA\" to include them (naming one AAA criterion in `wcag`, e.g. \"3.2.5\", also includes it).",
        +  "enum": [
        +    "A",
        +    "AA",
        +    "AAA"
        +  ],
        +  "type": "string"
        +}
    • Changedscan_page1 field changed
      • addedInput schema / properties / level
        Added value: +{
        +  "description": "Highest WCAG conformance level to report. Default \"AA\" — Level AAA criteria are not reported. Pass \"AAA\" to include them (naming one AAA criterion in `wcag`, e.g. \"3.2.5\", also includes it).",
        +  "enum": [
        +    "A",
        +    "AA",
        +    "AAA"
        +  ],
        +  "type": "string"
        +}
  2. 12 tool updatesv1.7.0
    • Changedcheck_aria5 fields changed
      • changedInput schema / properties / html / description
        Previous value: -"HTML to test for ARIA correctness"New value: +"HTML to test for ARIA correctness (pass this or `url`)"
      • addedInput schema / properties / nodeLimit
        Added value: +{
        +  "description": "Max nodes returned per rule (default 5, max 50)",
        +  "type": "number"
        +}
      • addedInput schema / properties / selector
        Added value: +{
        +  "description": "Only report findings on this element and its descendants (e.g. the `selector` from a scan_page issue). ARIA references outside it still resolve.",
        +  "type": "string"
        +}
      • addedInput schema / properties / url
        Added value: +{
        +  "description": "Live page to check (pass this or `html`)",
        +  "type": "string"
        +}
      • removedInput schema / required
        Removed value: -[
        -  "html"
        -]
    • Changedcheck_color_contrast7 fields changed
      • changedInput schema / properties / background / description
        Previous value: -"Background color (hex or rgb)"New value: +"Background color (hex or rgb). Optional with url + selector."
      • changedInput schema / properties / fontSize / description
        Previous value: -"Font size in px (default 16)"New value: +"Font size in px (default 16, or the element's computed size with url + selector)"
      • changedInput schema / properties / foreground / description
        Previous value: -"Foreground color (hex or rgb)"New value: +"Foreground color (hex or rgb). Optional with url + selector."
      • changedInput schema / properties / isBold / description
        Previous value: -"Whether text is bold (default false)"New value: +"Whether text is bold (default false, or the element's computed weight with url + selector)"
      • addedInput schema / properties / selector
        Added value: +{
        +  "description": "CSS selector of the text element on `url` (e.g. the `selector` from a scan_page issue)",
        +  "type": "string"
        +}
      • changedInput schema / properties / url / description
        Previous value: -"Live URL to extract brand palette from (uses our scanner — CSS vars + dominant colors)."New value: +"Live page: with `selector`, the colors are read from that element; on a failing pair, the brand palette is extracted from it (uses our scanner — CSS vars + dominant colors)."
      • removedInput schema / required
        Removed value: -[
        -  "foreground",
        -  "background"
        -]
    • Addeddiff_scan
    • Changedflow_scan5 fields changed
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "\"compact\" prints one line per element with rule metadata once. \"json\" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.",
        +  "enum": [
        +    "json",
        +    "compact"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / minImpact
        Added value: +{
        +  "description": "Only findings at this severity or above (critical > serious > moderate > minor)",
        +  "enum": [
        +    "critical",
        +    "serious",
        +    "moderate",
        +    "minor"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / rules
        Added value: +{
        +  "description": "Only these rule ids (WebAbility type such as \"missing_alt\" or axe rule id such as \"image-alt\"). See get_rules.",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / sourceRoot
        Added value: +{
        +  "description": "Local project root (local installs only) for `sourceCandidates[]` on issues without a framework `source` pointer.",
        +  "type": "string"
        +}
      • addedInput schema / properties / wcag
        Added value: +{
        +  "description": "Only these WCAG criteria. A prefix selects the whole guideline (\"1.4\") or principle (\"2\").",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Changedgenerate_ai_fix2 fields changed
      • changedInput schema / properties / issue / description
        Previous value: -"Issue object from scan_page (with selector, wcag, impact, message, fix.currentValue)"New value: +"Issue object from scan_page (with selector, wcag, impact, message, fix.currentValue), or a plain-text issue description"
      • changedInput schema / properties / issue / type
        Previous value: -"object"New value: +[
        +  "object",
        +  "string"
        +]
    • Addedgenerate_report_pdf
    • Changedget_rules4 fields changed
      • addedInput schema / properties / engine
        Added value: +{
        +  "description": "Which engine's rules to list (default all)",
        +  "enum": [
        +    "all",
        +    "axe",
        +    "webability"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / fixability
        Added value: +{
        +  "description": "Only rules of this fixability tier",
        +  "enum": [
        +    "mechanical",
        +    "contextual",
        +    "visual"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / rule
        Added value: +{
        +  "description": "Only rules whose id contains these whole words (e.g. \"link\" → link-name, skip-link, not blink; \"color-contrast\"; case and -/_ insensitive). Also accepted: ruleId, id, query.",
        +  "type": "string"
        +}
      • changedInput schema / properties / tags / description
        Previous value: -"Filter by tags (e.g. [\"wcag21aa\"], [\"best-practice\"], [\"cat.aria\"])"New value: +"axe tag filter (e.g. [\"wcag2aa\"], [\"best-practice\"], [\"cat.aria\"]; a category also works without \"cat.\"). WebAbility rules match on their WCAG criterion tag (e.g. \"wcag143\"). An unknown tag or category is an error that lists the valid ones."
    • Changedscan_history1 field changed
      • changedInput schema / properties / filter / description
        Previous value: -"Substring match on target URL or tool name (e.g. \"abilyo.com\" or \"scan_page\")"New value: +"Substring match on target URL or tool name (e.g. \"webability.io\" or \"scan_page\")"
    • Changedscan_html8 fields changed
      • addedInput schema / properties / engine
        Added value: +{
        +  "description": "\"in-process\" (default): jsdom, ms, structural rules. \"browser\": headless Chromium + axe-core, includes contrast.",
        +  "enum": [
        +    "in-process",
        +    "browser"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "\"compact\" prints one line per element with rule metadata once. \"json\" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.",
        +  "enum": [
        +    "json",
        +    "compact"
        +  ],
        +  "type": "string"
        +}
      • changedInput schema / properties / height / description
        Previous value: -"Viewport height (default 800)"New value: +"Viewport height (browser engine only, default 800)"
      • changedInput schema / properties / html / description
        Previous value: -"HTML content to test"New value: +"HTML content to test — a full document or a fragment"
      • addedInput schema / properties / minImpact
        Added value: +{
        +  "description": "Only findings at this severity or above (critical > serious > moderate > minor)",
        +  "enum": [
        +    "critical",
        +    "serious",
        +    "moderate",
        +    "minor"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / rules
        Added value: +{
        +  "description": "Only these rule ids (WebAbility type such as \"missing_alt\" or axe rule id such as \"image-alt\"). See get_rules.",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / wcag
        Added value: +{
        +  "description": "Only these WCAG criteria. A prefix selects the whole guideline (\"1.4\") or principle (\"2\").",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • changedInput schema / properties / width / description
        Previous value: -"Viewport width (default 1280)"New value: +"Viewport width (browser engine only, default 1280)"
    • Changedscan_page6 fields changed
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "\"compact\" prints one line per element with rule metadata once. \"json\" returns every field (html, fix values, source). Default: JSON when it fits inline (under ~24k chars), otherwise compact plus a JSON summary of every count.",
        +  "enum": [
        +    "json",
        +    "compact"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / minImpact
        Added value: +{
        +  "description": "Only findings at this severity or above (critical > serious > moderate > minor)",
        +  "enum": [
        +    "critical",
        +    "serious",
        +    "moderate",
        +    "minor"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / rules
        Added value: +{
        +  "description": "Only these rule ids (WebAbility type such as \"missing_alt\" or axe rule id such as \"image-alt\"). See get_rules.",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / sourceRoot
        Added value: +{
        +  "description": "Local project root (local installs only). Issues without a framework `source` pointer get `sourceCandidates[]` — files whose contents match the selector's id/class/attribute tokens.",
        +  "type": "string"
        +}
      • changedInput schema / properties / url / description
        Previous value: -"URL to scan (e.g. https://example.com or http://localhost:3000)"New value: +"URL to scan (e.g. https://webability.io or http://localhost:3000)"
      • addedInput schema / properties / wcag
        Added value: +{
        +  "description": "Only these WCAG criteria. A prefix selects the whole guideline (\"1.4\") or principle (\"2\").",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Changedstart_audit3 fields changed
      • changedInput schema / properties / includeAgent / description
        Previous value: -"Also run the agentic manual-audit pass (keyboard/focus/modal exploration). Slower and paid. Default false."New value: +"Also run the agentic manual-audit pass (keyboard/focus/modal exploration). Slower. Default false."
      • addedInput schema / properties / tunnel_secret
        Added value: +{
        +  "description": "Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.",
        +  "type": "string"
        +}
      • changedInput schema / properties / url / description
        Previous value: -"URL to audit (a public/staging URL the server can reach — not localhost)"New value: +"URL to audit (a public/staging URL the server can reach, or a webability-tunnel URL — not localhost)"
    • Changedverify_fix1 field changed
      • changedInput schema / properties / url / description
        Previous value: -"URL now serving the fix (deployed, staging, or http://localhost:3000)"New value: +"URL now serving the fix (deployed, staging, or http://localhost:3000). Also accepted as `page` or `pageUrl`."
  3. 14 tool updatesv1.3.1
    • First observedcheck_aria
    • First observedcheck_color_contrast
    • First observeddetect_framework
    • First observedfind_source
    • First observedflow_scan
    • First observedgenerate_ai_fix
    • First observedget_audit
    • First observedget_rules
    • First observedscan_history
    • First observedscan_html
    • First observedscan_page
    • First observedstart_audit
    • First observedverify_fix
    • First observedvisual_audit

TDQS

A4/5.0

Scored across 20 tools

Disambiguation4/5

Six tools perform some form of scanning/checking (scan_page, scan_html, flow_scan, visual_audit, check_aria, check_color_contrast), so there is real surface overlap. However, the descriptions explicitly carve out boundaries — single page vs multi-page journey vs raw HTML snippet vs pixel/vision vs focused ARIA/contrast checks — and even cross-reference each other ('For a single page, use scan_page'). An agent can reliably pick the right tool.

Naming Consistency4/5

Nearly all tools use consistent snake_case verb_noun patterns (get_install_snippet, start_audit, detect_framework, scan_page, verify_fix, diff_scan, add_site, create_upgrade_link, check_aria, generate_ai_fix). Minor deviations exist: visual_audit uses adjective+noun with no verb, and the audit family (start_audit/get_audit) mixes terminology with the scan/flow family, but overall it is readable and predictable.

Tool Count4/5

At 20 tools this sits in the heavier band, but the domain genuinely spans several sub-areas — scanning, rule lookup, fixing, verification, diffing, reporting, deep audits, and account/site/subscription management — so most tools earn their place. It is slightly over-provisioned (multiple scan variants) but not bloated.

Completeness4/5

The surface covers a full lifecycle: framework detection, rule discovery, scanning (page/flow/HTML/visual), focused checks, AI fix generation, source mapping, verification, regression diffing, and both PDF and Excel audit deliverables, plus account/site/billing management. Coverage is strong; only niche gaps remain (e.g. no explicit bulk/CI-suite or live-region tools), which agents can work around.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

  • Security, SEO and AI-visibility scanner for web apps · free scans and focused checks via MCP.

  • Zero-config MCP security scanner for AI-generated apps. 25K+ vulnerability patterns.

  • Scan URLs or HTML for WCAG 2.2 violations. 75-rule manifest, weighted score, shareable reports.

  • Run, debug, and triage tests from your IDE using natural language, no dashboard switching, no manual data transfers. The TestMu AI (formerly LambdaTest) MCP Server is a single remote server exposing four tool suites: HyperExecute — analyze your project, generate YAML configs and test runner commands, then monitor jobs and sessions. Automation — pull a TestID's details plus command, network, and console logs into one chat for instant root-cause analysis. Includes mobile app upload. SmartUI — explain pixel, layout, DOM, and perceptual changes in a visual regression run, with context-aware React/HTML/CSS fixes. Accessibility — audit any public URL or a local React app against WCAG and get ready-to-apply remediation steps. Connects over https://mcp.lambdatest.com/mcp using OAuth 2.1 — no API keys in your config. One-click install in Cursor; works with Claude, GitHub Copilot, Cline, and any MCP client. Tests execute on the TestMu AI cloud: 3,000+ browsers and 10,000+ real devices.

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server for accessibility auditing that provides WCAG 2.2 criteria lookup, HTML remediation guidance, and automated documentation generation for UI components. It enables users to analyze code snippets for issues and generate professional accessibility audit reports.
    266 npm
    MIT
  • A
    license
    A
    quality
    F
    maintenance
    An MCP server for WCAG accessibility auditing in AI coding agents, providing tools to audit HTML, files, URLs, and diffs, plus a prompt for React component auditing.
    5
    2,025 npm
    6
    MIT