Skip to main content
Glama

webability

Server Details

Free WCAG 2.2/ADA/508 accessibility MCP: scan, AI fixes, verify, vision audit, localhost tunnel

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
snayyar00/webability-mcp
GitHub Stars
0
Server Listing
@webability/mcp

TDQS

A4.4/5.0

Scored across 13 tools

Disambiguation4/5

Most tools have clearly distinct scopes (scan_page vs flow_scan vs start_audit vs visual_audit are each well-differentiated by output type and use case). However, check_aria and scan_html overlap — both validate ARIA/accessibility in raw HTML snippets, with scan_html already running the ARIA rules check_aria targets, so an agent could reasonably misselect between them.

Naming Consistency5/5

All 13 tools use snake_case with a clear verb_noun pattern (scan_page, start_audit, get_audit, verify_fix, generate_ai_fix, check_aria). Consistent convention throughout with no stylistic mixing.

Tool Count5/5

13 tools is well within the ideal range and each earns its place, covering scanning, fixing, verification, diffing, and audit-reporting distinct functions. No redundancy bloat.

Completeness4/5

Covers a full find-fix-verify lifecycle (scan → generate_ai_fix → verify_fix), plus diff, multi-page flow, visual audit, and durable audit artifacts — strong coverage. Minor gap: descriptions repeatedly reference a scan_history(id) tool that is not in the surface, forcing agents to work around its absence.

Available Tools

13 tools
check_ariaAInspect

Validate ARIA attribute + accessible name/role/value usage in an HTML snippet. Runs axe-core cat.aria and cat.name-role-value rules (aria-* attribute correctness, role validity, required parents/children, aria-hidden-focus, accessible names). Returns violations (high-confidence) and incomplete (needs human review, e.g. dangling ARIA references — do NOT auto-fix). Nodes cap at 5 per rule by default — every rule reports nodesTotal + truncated; raise nodeLimit (max 50) or use scan_history(id) for the full set.

ParametersJSON Schema
NameRequiredDescriptionDefault
htmlYesHTML to test for ARIA correctness
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
nodeLimitNoMax nodes returned per rule (default 5, max 50)
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses the two return buckets, warns that `incomplete` items (dangling ARIA references) must NOT be auto-fixed, and explains the default 5-node cap, per-rule nodesTotal/truncated reporting, and the nodeLimit max of 50. That is unusually rich behavioral disclosure for a validation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with what is validated and which rules run, then returns, then the cap escape hatch. No filler and no repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description still tells the agent what comes back (violations vs incomplete), the confidence distinction between them, the truncation behavior, and how to get complete results. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds operational meaning beyond the schema: it explains that nodeLimit controls a per-rule cap, that truncation is reported, and how to escape the cap. The context/llm_model analytics params are left to the schema, which is fine.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (validate) and resource (ARIA attribute + accessible name/role/value usage) scoped to an HTML snippet, and names the exact axe-core rule categories it runs. An agent can distinguish it from scan_html/scan_page because the scope is ARIA/name-role-value correctness rather than a general scan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The snippet scope and the rule categories imply when this is the right tool, and it explicitly routes to scan_history(id) when the capped node set is insufficient. It does not name a sibling to avoid (e.g. scan_html) or state prerequisites for obtaining an id, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_color_contrastAInspect

Check a foreground/background color pair against WCAG contrast thresholds. When it fails, suggests BRAND-aligned replacements — extracts the actual brand palette from a live URL using our scanner (CSS vars + most-used colors), or use a provided brandColors array. No url and no brandColors = ratio + pass/fail only. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoLive URL to extract brand palette from (uses our scanner — CSS vars + dominant colors).
isBoldNoWhether text is bold (default false)
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
fontSizeNoFont size in px (default 16)
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
backgroundYesBackground color (hex or rgb)
foregroundYesForeground color (hex or rgb)
brandColorsNoPre-supplied brand palette. Skips URL extraction if provided.
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely succeeds: it discloses the cloud-hosted network restriction, the tunnel_secret requirement for tunnel URLs, and the local-run workaround. It does not cover auth, rate limits, or failure behavior beyond the private-address case, leaving some operational gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Core purpose and the primary usage fork are front-loaded in the first two sentences; the trailing NOTE and tunnel instructions are dense but each clause answers a real question. It is on the long side, yet no sentence is purely filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter tool with no output schema and no annotations, the description covers environment constraints, alternatives, and the main invocation modes well. It stops short of describing the return shape (ratio, pass/fail, suggested replacements) beyond passing mention, which is the one remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds combination semantics beyond the schema: the url-vs-brandColors precedence ('skips URL extraction if provided') and the both-empty fallback mode, which is real invocation-relevant meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: checks a foreground/background color pair against WCAG contrast thresholds, and on failure proposes brand-aligned replacements. It also carves out a distinct niche versus siblings like visual_audit, scan_page, and check_aria by focusing solely on color contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit branching rules: supply a `brandColors` array or a `url` to get brand suggestions, while 'No url and no brandColors = ratio + pass/fail only.' It also states when the hosted endpoint is unusable (localhost/private addresses) and names two concrete alternatives for local dev servers.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_frameworkAInspect

Detect which framework/stack a page uses (Tailwind, MUI, Bootstrap, WordPress, Next.js, plain CSS). Use before generate_ai_fix to get framework-appropriate code. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to inspect
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses the hosted-server constraint (localhost/private addresses refused, runs in cloud, cannot reach the user's machine), the two workarounds, and that tunnel_secret is required for tunnel URLs or the relay will refuse. This is exactly the operational context an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and the primary usage rule are front-loaded in the first two sentences, with the environment caveat and workarounds following. It is dense but every clause earns its place; the only mild cost is the length of the hosting note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations and no output schema, the description covers purpose, workflow ordering, and environment constraints well. It stops short of describing what the detection result looks like, which is the one remaining gap, but nothing blocking is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema, including the tunnel_secret and conversation_id semantics. The description reinforces the url/tunnel_secret relationship but adds little that the schema does not already state, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (detect) and resource (framework/stack of a page), and enumerates the frameworks it can identify (Tailwind, MUI, Bootstrap, WordPress, Next.js, plain CSS). This clearly distinguishes it from siblings like scan_page or generate_ai_fix.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it 'before generate_ai_fix to get framework-appropriate code', naming the sibling and the ordering condition. It also gives concrete when-to-use routing for the localhost case (run locally vs. tunnel), so an agent knows exactly which path applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diff_scanAInspect

Compare two scans of the same page and report what changed: fixed[] (in the baseline, gone now), new[] (regressions — not in the baseline, present now), remaining[] (still there). Page-level complement to verify_fix (one element). Baseline is a scan_history id (baselineId, local installs) or a live scan of baselineUrl; current is url (scanned live now) or another history id (currentId). Findings are matched by issue id (rule + element), so a changed class/id on a fixed element reads as fixed AND new — check new[] before calling it a regression. Needs-review findings are diffed separately (incompleteResolved / incompleteNew) and never counted as fixed. Typical loop: scan_page → edit → diff_scan(baselineId=, url=) → confirm new[] is empty. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoURL to scan now as the CURRENT side (deployed, staging, or http://localhost:3000). Omit when passing currentId.
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
formatNo"compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json.
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
viewportNoViewport for live scans (default: desktop). Use the same viewport the baseline used.
currentIdNoscan_history id to use as the CURRENT side instead of scanning `url`
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
baselineIdNoscan_history id of the BASELINE scan (local installs only)
baselineUrlNoScan this URL live as the baseline (e.g. production) — use when there is no stored baseline
rootSelectorNoCSS selector to limit live scans to (optional)
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: findings are matched by issue id (rule + element), so a changed class/id reads as fixed AND new — a non-obvious caveat. It further discloses that needs-review findings are diffed separately and never counted as fixed, and that the hosted server refuses localhost/private addresses, along with the two workarounds. These are behavioral traits well beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Output buckets are front-loaded and the paragraph flows logically from semantics to matching rules to environment constraints to the canonical loop. It is dense and long for a single description, but nearly every clause earns its place; the localhost/tunnel passage is the only section bordering on overload.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter, no-annotation, no-output-schema tool, the description covers the return shape (fixed/new/remaining plus the incomplete buckets), matching rules, environment limits, and the recommended invocation pattern. An agent has enough to call it correctly without opening sibling docs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter is individually documented, so the baseline is 3. The description adds relationship semantics the schema does not: which parameter combinations form the baseline vs current side, that viewport must match the baseline, and that tunnel_secret is mandatory for tunnel URLs. Solid added value, though it stops short of illustrating default resolution when both sides are ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Compare two scans of the same page and report what changed") and immediately enumerates the three output buckets fixed[]/new[]/remaining[], which pins down the semantics. It also explicitly positions itself as the "page-level complement to verify_fix (one element)," so an agent can pick between the two siblings without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit typical loop (scan_page → edit → diff_scan → confirm new[] empty), the conditional selection between baselineId vs baselineUrl and url vs currentId, and the alternative paths when the hosted server cannot reach localhost (run the MCP locally, or use a tunnel). When-to-use and when-to-substitute are both stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_scanAInspect

Scan a multi-page user journey. Walks startUrl plus the required autoNavigate URLs sequentially (deterministic — one page fully rendered and scanned before the next), then returns ONE consolidated report with issues deduplicated across pages, each carrying the same fix payload / confidence / review flags as scan_page. Every requested URL gets an explicit outcome in pages[] (scanned / nav_failed / scan_failed / redirected_duplicate / duplicate_request / skipped_cap / blocked — bot-challenge, not a clean page) — a page is never silently dropped. Better than per-page scans for journeys (login → checkout etc). For a single page, use scan_page. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
formatNo"compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json.
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
maxPagesNoMax pages to scan (default 10)
startUrlYesStarting URL of the journey
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
sourceRootNoLocal project root (local installs only) for `sourceCandidates[]` on issues without a framework `source` pointer.
autoNavigateYesREQUIRED — the URLs to walk after startUrl (the MCP server is headless and cannot discover a journey interactively). For a single page, use scan_page instead.
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it declares deterministic sequential rendering, one consolidated deduplicated report, and a closed enumeration of per-URL outcomes (scanned/nav_failed/scan_failed/redirected_duplicate/duplicate_request/skipped_cap/blocked) with the guarantee that no page is silently dropped. It also discloses the localhost/private-address refusal and the tunnel_secret requirement — significant operational constraints beyond the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core behavior and the sibling routing rule, with the hosted/local caveat correctly relegated later. It is long and repeats the 'for a single page, use scan_page' guidance twice, but nearly every sentence carries information an agent needs to call it correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description explains the return shape (pages[] with explicit per-URL outcomes, consolidated deduplicated issues each carrying fix payload/confidence/review flags), the traversal semantics, and the environment constraints. An agent has everything needed to invoke and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already documented; baseline is 3. The description usefully reinforces that autoNavigate is required and walked after startUrl, but it also refers to a `url` parameter for tunnel usage that does not exist in the schema (startUrl does), which slightly muddies parameter mapping rather than clarifying it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (scan a multi-page user journey) and immediately separates itself from the closest sibling: 'For a single page, use scan_page.' The consolidated-report + dedup behavior is described concretely, so an agent can tell exactly what this tool produces versus per-page scans.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing rules: use flow_scan for journeys (login → checkout), use scan_page for a single page, and two named alternatives for local dev servers (run the MCP locally, or open a tunnel and pass the secret). It also states the hosted-host restriction so the agent knows when the tool simply cannot be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_ai_fixAInspect

Generate framework-aware fix alternatives for a specific accessibility issue. For color contrast issues, returns 3 alternatives (minimal, brand-aligned, high contrast); brand palette is auto-extracted from the live URL using our scanner if brandColors is omitted. For label/ARIA issues, returns 1-2 alternatives. Each alternative includes ready-to-paste code for the detected framework. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoPage URL — also used to auto-extract brand palette for contrast issues if `brandColors` is not provided.
htmlYesThe element's outerHTML — send at most ~600 chars
issueYesIssue object from scan_page (with selector, wcag, impact, message, fix.currentValue), or a plain-text issue description
contextNoParent element outerHTML for context (~400 chars)
frameworkYesCSS framework — use detect_framework first
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
brandColorsNoBrand palette for brand-aligned suggestions. If omitted on a contrast issue with a `url`, auto-extracted via the scanner.
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses substantial behavior: alternative counts by issue type, ready-to-paste framework code, automatic brand-palette extraction when brandColors is omitted, and the hosted-server restriction that localhost/private addresses are refused. It does not cover rate limits, timeouts, or failure modes, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence, and the issue-type branching follows logically. The lengthy NOTE on localhost/tunnel is verbose but each sentence carries actionable constraint information, so it earns its place even if it could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and nine parameters, the description does a good job explaining return values (3 alternatives named minimal/brand-aligned/high contrast, or 1-2 for label/ARIA issues, each with ready-to-paste code) and the environment constraint. It is nearly complete for invocation, lacking only error/edge-case handling detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all nine parameters are already documented, establishing a baseline of 3. The description restates the url/brandColors auto-extraction behavior but adds essentially no syntax or format detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Generate) and resource (framework-aware fix alternatives for an accessibility issue), and it is immediately distinguishable from siblings like verify_fix and scan_page. The description also scopes output by issue type (color contrast vs label/ARIA), so an agent knows exactly what it produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for use (a specific accessibility issue) and practical routing guidance: use detect_framework first for the framework value, and two concrete workarounds when scanning localhost. It stops short of naming verify_fix or other siblings as the follow-up/alternative step, so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_auditAInspect

Check an audit started with start_audit: returns overall status, per-step progress (scan → viewports → screenshots → agent → excel → publish), and — once complete — a severity summary plus short-lived download URLs for the report (JSON) and the Excel workbook. Poll every ~15s while status is pending/running. Only the account that started an audit can read it — or, for a trial (no-account) run, only the caller holding the claimToken start_audit returned. Each poll also spends one trial call, so avoid polling faster than ~15s on a trial run.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe audit id returned by start_audit
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
claimTokenNoTrial (no-account) runs only — the claimToken start_audit returned. Omit if you are logged in.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the authorization model (only the starting account, or the claimToken holder for trial runs), the destructive-ish cost behavior (each poll spends one trial call), and the lifecycle ordering of the pipeline steps. This is richer than the schema provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core purpose before the operational caveats. Slightly long, but every clause carries distinct operational information (poll cadence, auth model, trial cost) without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a polling/status tool with no output schema and no annotations, the description covers all the bases an agent needs to call it safely and effectively: what comes back, when to call, how often, and who is allowed. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning by explaining how id and claimToken relate to the start_audit call and the trial-vs-account distinction, which the schema states only in isolation. It does not add syntax detail for context, llm_model, or conversation_id beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource pair — 'Check an audit started with start_audit' — and enumerates exactly what the caller receives: overall status, per-step progress, severity summary, and download URLs. This distinguishes it from start_audit and from all the other check_* / scan_* siblings unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the alternative workflow entry point (start_audit), instructs polling cadence (~15s while pending/running), and adds a trial-specific warning to avoid polling faster than ~15s because each poll spends a trial call. That is explicit when-to-use and when-to-avoid-in-a-way guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_rulesAInspect

List accessibility rules from both engines — axe-core (104) and the WebAbility detectors (90+) — with optional filters. Every rule carries fixability (mechanical | contextual | visual) and a fix op template, so you can pick the rules worth auto-fixing before scanning. Returns ruleId, engine, description, help, helpUrl, tags/wcag, fixability, fix.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoaxe tag filter (e.g. ["wcag21aa"], ["best-practice"], ["cat.aria"]). WebAbility rules match on their WCAG criterion tag (e.g. "wcag143").
engineNoWhich engine's rules to list (default all)
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
fixabilityNoOnly rules of this fixability tier
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the burden and does well: it enumerates the returned fields (ruleId, engine, description, help, helpUrl, tags/wcag, fixability, fix) and explains the fixability tiers and fix-op template. It omits auth, rate-limit, or pagination details, but the data-model disclosure is genuinely valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose and filters before the enumerated return fields. The closing field list is dense but earns its place given there is no output schema; nothing is wasted, though it is heavier than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, filtering, and the full return shape for a list tool with no output schema, giving the agent enough to call it correctly. The only gaps are the absence of explicit when-not guidance and operational details (pagination), which are minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both filters and required analytics params are already documented in the schema. The description references "optional filters" and the fixability tiers but adds no syntax or meaning beyond what the schema provides, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("List") and resource (accessibility rules), plus the exact scope: axe-core (104) and WebAbility (90+) detectors. It clearly separates itself from the scanning/audit siblings by framing itself as a pre-scan catalog lookup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context — "pick the rules worth auto-fixing before scanning" — which tells the agent when this tool is appropriate. It stops short of naming a specific sibling alternative or stating exclusions, so it does not reach the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_htmlAInspect

Scan a raw HTML snippet or component markup without serving it — IN-PROCESS by default (jsdom + WebAbility detectors + axe-core): milliseconds, no browser, no network, so it fits inside a tight edit loop. Fragments are auto-wrapped into a document. Returns scan_page's three-tier shape (issues / incomplete / summary) with fix.op + fixability on every finding. jsdom has no layout, so visual-tier rules (contrast, target size, focus ring) are NOT evaluated — the dropped count is reported as skippedVisual; pass engine: "browser" to run the axe-core headless-browser path for those (slower, axe rules only, returns axe violations).

ParametersJSON Schema
NameRequiredDescriptionDefault
htmlYesHTML content to test — a full document or a fragment
tagsNoWCAG tags to check (default ["wcag2a","wcag2aa","wcag21aa","wcag22aa"])
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
widthNoViewport width (browser engine only, default 1280)
engineNo"in-process" (default): jsdom, ms, structural rules. "browser": headless Chromium + axe-core, includes contrast.
formatNo"compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json.
heightNoViewport height (browser engine only, default 800)
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so unusually well: it discloses the default execution model, that there is no browser or network, that fragments are auto-wrapped, that jsdom lacks layout so contrast/target-size/focus-ring rules are skipped, and that the skipped count surfaces as skippedVisual. It also describes the alternate browser path's cost and narrower rule set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and default mode are front-loaded in the opening clause, and almost every sentence adds distinct information (execution model, return shape, visual-rule limitation, alternate engine). It is dense and parenthesis-heavy but not padded; a small amount of compression would improve readability without losing content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter tool with no annotations and no output schema, the description covers the essentials: what is scanned, how fragments are handled, the return shape (issues/incomplete/summary with fix.op and fixability), the major limitation of the default engine, and how to opt into visual checking. An agent has enough to call it correctly on the first attempt.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so engine, format, and the other parameters are already documented; the schema also explains the in-process vs browser distinction. The description mainly reinforces engine semantics at a consequence level (visual rules dropped, skippedVisual reported) rather than adding new syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first clause gives a specific verb+resource (scan raw HTML snippet/component markup) plus a distinguishing constraint ("without serving it"), which separates it from the page-serving sibling scan_page. It also names the concrete engines and return shape, so an agent knows exactly what class of operation this is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when the default in-process mode fits (tight edit loop, milliseconds, no browser/network) and when to switch to engine:"browser" for visual-tier rules that jsdom cannot evaluate, including the tradeoff (slower, axe rules only). What is missing is explicit routing versus siblings like scan_page or visual_audit — the agent must infer that those handle served URLs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_pageAInspect

Scan a web page for WCAG accessibility issues. Works on any URL — deployed sites, localhost, staging. Returns the three-tier shape: issues (high-confidence violations safe to fix), incomplete (needs human review — gradient backgrounds, marketing imagery, axe-incomplete results, framer-motion pre-animation states), and a summary. Treat incomplete as questions, never auto-fix them. On React ≤18 / Vue dev builds each issue carries source ({file, line, column, component}) read from the live component tree. Every issue carries a structured fix.op (add-attribute | set-attribute | remove-attribute | add-element | remove-element | add-text-content | suggest) with fix.attribute / fix.value when known, and a fixability tier (mechanical = apply as given; contextual = op known, value needs judgment; visual = needs rendered output, propose only). NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to scan (e.g. https://example.com or http://localhost:3000)
wcagNoOnly these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2").
rulesNoOnly these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules.
formatNo"compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json.
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
viewportNoViewport size (default: desktop)
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
minImpactNoOnly findings at this severity or above (critical > serious > moderate > minor)
sourceRootNoLocal project root (local installs only). Issues without a framework `source` pointer get `sourceCandidates[]` — files whose contents match the selector's id/class/attribute tokens.
rootSelectorNoCSS selector to limit scan scope (optional)
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so unusually well: it documents the three-tier return shape, the semantics of `incomplete` (questions, never auto-fix), the `source` pointer availability by framework/build, the structured `fix.op` contract, and the `fixability` tiers. It also discloses the critical hosting constraint that localhost/private addresses are refused.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and scope, then progressively details return shape, fix metadata, and hosting constraints. It is long and reads as a dense paragraph rather than scannable sections, but nearly every sentence contributes distinct information, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter tool with no output schema and no annotations, the description supplies the missing return-value documentation (issues/incomplete/summary, fix.op, fixability) and the environment caveat. Nothing essential to calling or interpreting the scan is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: it ties `url` and `tunnel_secret` together (secret required for tunnel URLs, otherwise ignored) and explains the localhost/tunnel workflow that makes those parameters usable. It adds little for `rootSelector`, `minImpact`, or `format`, which the schema already covers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scan a web page for WCAG accessibility issues') and immediately scopes it with 'Works on any URL — deployed sites, localhost, staging.' An agent can distinguish it from scan_html, check_aria, check_color_contrast, and visual_audit without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational context: when the hosted server refuses localhost/private addresses, and the two concrete alternatives (run the MCP locally via npx, or open a tunnel and pass tunnel_secret). It also instructs how to treat the `incomplete` bucket. However it never routes among sibling tools (e.g. scan_html vs scan_page), so it stops short of explicit when-not-this-tool guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_auditAInspect

Kick off a FULL accessibility audit deliverable for a URL — a persistent, timestamped artifact, not an inline scan. Runs the server-side pipeline (axe + advanced checks + mobile viewports + annotated screenshots + optional agent spot-check) and produces a downloadable report and a formatted Excel workbook (Cover / Status / Barriers / ADA context sheets) stored durably. Returns immediately with an audit id; poll get_audit for progress and, when complete, download URLs. Use this when someone needs a durable artifact to attach as evidence of testing effort for a compliance officer or legal response — for iterating on code, use scan_page + verify_fix instead. Free without an account for a limited trial (a shared pool of calls across start_audit/get_audit/visual_audit, hosted deploy only) — the response says how many are left and includes a claimToken to pass to get_audit. Past the trial, or on the local/stdio server: authenticate via webability login or set WEBABILITY_API_KEY. Set includeAgent:true to add the (slower, paid) agentic manual-audit pass. To audit a local dev server, open a tunnel (webability-tunnel --port 3000) and pass its URL as url with the printed secret as tunnel_secret; keep the tunnel open until get_audit reports complete (about 5 minutes) — the pipeline loads the page several times.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to audit (a public/staging URL the server can reach, or a webability-tunnel URL — not localhost)
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
includeAgentNoAlso run the agentic manual-audit pass (keyboard/focus/modal exploration). Slower. Default false.
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden and does so richly: it discloses async behavior (returns immediately with an id; poll get_audit), the pipeline's composition, trial/pool limits, auth requirements, the tunnel workflow and its ~5-minute lifetime, and that includeAgent is slower and paid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and tightly packed, with every clause carrying operational value. It is dense and long for one paragraph, but for a complex async tool with trial/auth/tunnel caveats most sentences earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a 6-param async tool, the description covers the full lifecycle: immediate return value (id), polling path (get_audit), download URLs, trial count and claimToken, and auth for non-trial/local usage. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter workflow meaning the schema lacks: tunnel_secret is required when url is a tunnel URL, url must not be localhost, includeAgent is slower/paid, and the tunnel must stay open because 'the pipeline loads the page several times'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Kick off a FULL accessibility audit deliverable for a URL') and immediately scopes it against a sibling behavior ('a persistent, timestamped artifact, not an inline scan'). An agent can distinguish it from scan_page/visual_audit without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('when someone needs a durable artifact to attach as evidence ... for a compliance officer or legal response') and when-not, naming the alternatives ('for iterating on code, use scan_page + verify_fix instead'). Routing conditions are unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_fixAInspect

Re-scan a specific element after applying an accessibility fix and confirm the violation is gone — closes the loop that find-only tools leave open. After you edit the code and serve it (deployed, staging, or http://localhost:3000), call this with the URL and the selector you fixed to get a machine-checked verified: true|false (DOM engines only — visual_audit findings and needs-review items are out of scope). Pass the WCAG criterion (e.g. "1.1.1") or axe rule id (e.g. "color-contrast") to check just that criterion; omit it to require the element be clean of ALL violations. A blocked page (bot-challenge / HTTP error) is reported as unverified, never a pass — verification fails closed. IMPORTANT: if your fix changed the element's class or id, the original selector may no longer match anything, which reads as verified — re-run scan_page or pass the updated selector to be sure. Pair with scan_page → generate_ai_fix → verify_fix for a full find-fix-verify cycle. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL now serving the fix (deployed, staging, or http://localhost:3000)
wcagNoOptional: WCAG criterion (e.g. "1.1.1", "1.4.3") or axe rule id (e.g. "color-contrast") to verify specifically. Omit to require the element be free of ALL violations.
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
selectorYesCSS selector of the element you fixed — use the `selector` from the original scan_page issue
viewportNoViewport size (default: desktop). Use the same viewport the issue was found at.
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses fail-closed behavior (blocked page = unverified, never a pass), engine limits (DOM engines only), the selector-drift gotcha that can produce a false 'verified', and the hosted-server restriction refusing localhost/private addresses. This is well beyond what any structured field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly packed and front-loaded with purpose before caveats; each sentence carries operational value (fail-closed, selector drift, tunnel/localhost). It is dense rather than redundant, though the HOSTED-server and tunnel paragraphs make it heavier than strictly minimal for a quick call.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, no-annotation, no-output-schema tool, the description covers the return contract (verified: true|false), failure semantics, engine scope, and the local-dev access story. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: it clarifies that omitting `wcag` requires the element be clean of ALL violations, explains the criterion-vs-rule-id dual form, and warns that `selector` must be the original or updated selector. These are genuinely useful additions over the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope: 'Re-scan a specific element after applying an accessibility fix and confirm the violation is gone.' It explicitly distinguishes itself from find-only siblings like scan_page and positions itself in the find-fix-verify cycle, so the agent knows exactly what it is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States prerequisites (edit and serve the code), the exact selection workflow (scan_page → generate_ai_fix → verify_fix), and exclusions (visual_audit findings and needs-review items out of scope). It also names alternatives for the local-dev case (run MCP locally vs. use a tunnel), leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visual_auditAInspect

Pixel-level accessibility audit using Claude vision. Catches issues that DOM scanners miss: icon contrast (1.4.11), focus visibility (2.4.7), "looks like a button but isn't" (4.1.2), text rendered as images (1.4.5), visual hierarchy mismatches. Takes a URL, opens it in a headless browser, screenshots, and runs vision-based detection. Complements scan_page — run both for full coverage. Free without an account for a limited trial (shared call pool with start_audit/get_audit, hosted deploy only) — the response says how many are left. Past the trial, or on the local/stdio server, this and start_audit are the paid, server-side tools: authenticate via webability login or set WEBABILITY_API_KEY in your MCP server env before calling. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to audit visually
contextYesExplain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution."
fullPageNoCapture full scrolled page instead of just viewport (default: false)
viewportNoViewport size (default: desktop)
llm_modelYesThe exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess.
brandColorsNoBrand hex colors for context-aware filtering
tunnel_secretNoSecret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise.
conversation_idNoEcho the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: shared trial call pool with start_audit/get_audit, limited free trial, paid server-side otherwise, required authentication, localhost/private-address refusal on the hosted server, and the tunnel_secret dependency. These are exactly the operational traits an agent must know before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the complement-to-scan_page routing are front-loaded, and every subsequent sentence addresses a real invocation risk (auth, trial limits, localhost refusal, tunnel). It is dense and somewhat run-on for a single paragraph, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given eight parameters, no output schema, and no annotations, the description covers the full operational picture: what it does, how it complements siblings, auth needs, rate/trial limits, and networking constraints with workarounds. An agent has everything required to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all eight parameters are already documented, and the description adds little parameter-level meaning beyond echoing the url-plus-tunnel_secret pairing the schema already states. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (pixel-level accessibility audit via vision), the resource (a URL rendered in a headless browser), and the distinct class of issues it catches (icon contrast, focus visibility, 4.1.2 affordance mismatches, text-as-image). It explicitly positions itself against scan_page, so an agent can differentiate it from that sibling without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: 'Complements scan_page — run both for full coverage,' plus a clear when-paid boundary (past trial or on local/stdio) and two concrete alternatives for scanning a local dev server (run MCP locally, or tunnel). It names prerequisites (auth via login/API key) and the hosted-localhost refusal, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updates
    • First observedcheck_aria
    • First observedcheck_color_contrast
    • First observeddetect_framework
    • First observeddiff_scan
    • First observedflow_scan
    • First observedgenerate_ai_fix
    • First observedget_audit
    • First observedget_rules
    • First observedscan_html
    • First observedscan_page
    • First observedstart_audit
    • First observedverify_fix
    • First observedvisual_audit

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.