webability
Server Details
Free WCAG 2.2/ADA/508 accessibility MCP: scan, AI fixes, verify, vision audit, localhost tunnel
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- snayyar00/webability-mcp
- GitHub Stars
- 0
- Server Listing
- @webability/mcp
TDQS
Scored across 13 tools
Most tools have clearly distinct scopes (scan_page vs flow_scan vs start_audit vs visual_audit are each well-differentiated by output type and use case). However, check_aria and scan_html overlap — both validate ARIA/accessibility in raw HTML snippets, with scan_html already running the ARIA rules check_aria targets, so an agent could reasonably misselect between them.
All 13 tools use snake_case with a clear verb_noun pattern (scan_page, start_audit, get_audit, verify_fix, generate_ai_fix, check_aria). Consistent convention throughout with no stylistic mixing.
13 tools is well within the ideal range and each earns its place, covering scanning, fixing, verification, diffing, and audit-reporting distinct functions. No redundancy bloat.
Covers a full find-fix-verify lifecycle (scan → generate_ai_fix → verify_fix), plus diff, multi-page flow, visual audit, and durable audit artifacts — strong coverage. Minor gap: descriptions repeatedly reference a scan_history(id) tool that is not in the surface, forcing agents to work around its absence.
Available Tools
13 toolscheck_ariaAInspect
Validate ARIA attribute + accessible name/role/value usage in an HTML snippet. Runs axe-core cat.aria and cat.name-role-value rules (aria-* attribute correctness, role validity, required parents/children, aria-hidden-focus, accessible names). Returns violations (high-confidence) and incomplete (needs human review, e.g. dangling ARIA references — do NOT auto-fix). Nodes cap at 5 per rule by default — every rule reports nodesTotal + truncated; raise nodeLimit (max 50) or use scan_history(id) for the full set.
| Name | Required | Description | Default |
|---|---|---|---|
| html | Yes | HTML to test for ARIA correctness | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| nodeLimit | No | Max nodes returned per rule (default 5, max 50) | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses the two return buckets, warns that `incomplete` items (dangling ARIA references) must NOT be auto-fixed, and explains the default 5-node cap, per-rule nodesTotal/truncated reporting, and the nodeLimit max of 50. That is unusually rich behavioral disclosure for a validation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what is validated and which rules run, then returns, then the cap escape hatch. No filler and no repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description still tells the agent what comes back (violations vs incomplete), the confidence distinction between them, the truncation behavior, and how to get complete results. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds operational meaning beyond the schema: it explains that nodeLimit controls a per-rule cap, that truncation is reported, and how to escape the cap. The context/llm_model analytics params are left to the schema, which is fine.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (validate) and resource (ARIA attribute + accessible name/role/value usage) scoped to an HTML snippet, and names the exact axe-core rule categories it runs. An agent can distinguish it from scan_html/scan_page because the scope is ARIA/name-role-value correctness rather than a general scan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The snippet scope and the rule categories imply when this is the right tool, and it explicitly routes to scan_history(id) when the capped node set is insufficient. It does not name a sibling to avoid (e.g. scan_html) or state prerequisites for obtaining an id, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_color_contrastAInspect
Check a foreground/background color pair against WCAG contrast thresholds. When it fails, suggests BRAND-aligned replacements — extracts the actual brand palette from a live URL using our scanner (CSS vars + most-used colors), or use a provided brandColors array. No url and no brandColors = ratio + pass/fail only. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Live URL to extract brand palette from (uses our scanner — CSS vars + dominant colors). | |
| isBold | No | Whether text is bold (default false) | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| fontSize | No | Font size in px (default 16) | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| background | Yes | Background color (hex or rgb) | |
| foreground | Yes | Foreground color (hex or rgb) | |
| brandColors | No | Pre-supplied brand palette. Skips URL extraction if provided. | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely succeeds: it discloses the cloud-hosted network restriction, the tunnel_secret requirement for tunnel URLs, and the local-run workaround. It does not cover auth, rate limits, or failure behavior beyond the private-address case, leaving some operational gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Core purpose and the primary usage fork are front-loaded in the first two sentences; the trailing NOTE and tunnel instructions are dense but each clause answers a real question. It is on the long side, yet no sentence is purely filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no output schema and no annotations, the description covers environment constraints, alternatives, and the main invocation modes well. It stops short of describing the return shape (ratio, pass/fail, suggested replacements) beyond passing mention, which is the one remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds combination semantics beyond the schema: the url-vs-brandColors precedence ('skips URL extraction if provided') and the both-empty fallback mode, which is real invocation-relevant meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: checks a foreground/background color pair against WCAG contrast thresholds, and on failure proposes brand-aligned replacements. It also carves out a distinct niche versus siblings like visual_audit, scan_page, and check_aria by focusing solely on color contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit branching rules: supply a `brandColors` array or a `url` to get brand suggestions, while 'No url and no brandColors = ratio + pass/fail only.' It also states when the hosted endpoint is unusable (localhost/private addresses) and names two concrete alternatives for local dev servers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect_frameworkAInspect
Detect which framework/stack a page uses (Tailwind, MUI, Bootstrap, WordPress, Next.js, plain CSS). Use before generate_ai_fix to get framework-appropriate code. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to inspect | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses the hosted-server constraint (localhost/private addresses refused, runs in cloud, cannot reach the user's machine), the two workarounds, and that tunnel_secret is required for tunnel URLs or the relay will refuse. This is exactly the operational context an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the primary usage rule are front-loaded in the first two sentences, with the environment caveat and workarounds following. It is dense but every clause earns its place; the only mild cost is the length of the hosting note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no annotations and no output schema, the description covers purpose, workflow ordering, and environment constraints well. It stops short of describing what the detection result looks like, which is the one remaining gap, but nothing blocking is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema, including the tunnel_secret and conversation_id semantics. The description reinforces the url/tunnel_secret relationship but adds little that the schema does not already state, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (detect) and resource (framework/stack of a page), and enumerates the frameworks it can identify (Tailwind, MUI, Bootstrap, WordPress, Next.js, plain CSS). This clearly distinguishes it from siblings like scan_page or generate_ai_fix.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it 'before generate_ai_fix to get framework-appropriate code', naming the sibling and the ordering condition. It also gives concrete when-to-use routing for the localhost case (run locally vs. tunnel), so an agent knows exactly which path applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_scanAInspect
Compare two scans of the same page and report what changed: fixed[] (in the baseline, gone now), new[] (regressions — not in the baseline, present now), remaining[] (still there). Page-level complement to verify_fix (one element). Baseline is a scan_history id (baselineId, local installs) or a live scan of baselineUrl; current is url (scanned live now) or another history id (currentId). Findings are matched by issue id (rule + element), so a changed class/id on a fixed element reads as fixed AND new — check new[] before calling it a regression. Needs-review findings are diffed separately (incompleteResolved / incompleteNew) and never counted as fixed. Typical loop: scan_page → edit → diff_scan(baselineId=, url=) → confirm new[] is empty. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to scan now as the CURRENT side (deployed, staging, or http://localhost:3000). Omit when passing currentId. | |
| wcag | No | Only these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2"). | |
| rules | No | Only these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules. | |
| format | No | "compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json. | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| viewport | No | Viewport for live scans (default: desktop). Use the same viewport the baseline used. | |
| currentId | No | scan_history id to use as the CURRENT side instead of scanning `url` | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| minImpact | No | Only findings at this severity or above (critical > serious > moderate > minor) | |
| baselineId | No | scan_history id of the BASELINE scan (local installs only) | |
| baselineUrl | No | Scan this URL live as the baseline (e.g. production) — use when there is no stored baseline | |
| rootSelector | No | CSS selector to limit live scans to (optional) | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: findings are matched by issue id (rule + element), so a changed class/id reads as fixed AND new — a non-obvious caveat. It further discloses that needs-review findings are diffed separately and never counted as fixed, and that the hosted server refuses localhost/private addresses, along with the two workarounds. These are behavioral traits well beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Output buckets are front-loaded and the paragraph flows logically from semantics to matching rules to environment constraints to the canonical loop. It is dense and long for a single description, but nearly every clause earns its place; the localhost/tunnel passage is the only section bordering on overload.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter, no-annotation, no-output-schema tool, the description covers the return shape (fixed/new/remaining plus the incomplete buckets), matching rules, environment limits, and the recommended invocation pattern. An agent has enough to call it correctly without opening sibling docs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter is individually documented, so the baseline is 3. The description adds relationship semantics the schema does not: which parameter combinations form the baseline vs current side, that viewport must match the baseline, and that tunnel_secret is mandatory for tunnel URLs. Solid added value, though it stops short of illustrating default resolution when both sides are ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Compare two scans of the same page and report what changed") and immediately enumerates the three output buckets fixed[]/new[]/remaining[], which pins down the semantics. It also explicitly positions itself as the "page-level complement to verify_fix (one element)," so an agent can pick between the two siblings without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit typical loop (scan_page → edit → diff_scan → confirm new[] empty), the conditional selection between baselineId vs baselineUrl and url vs currentId, and the alternative paths when the hosted server cannot reach localhost (run the MCP locally, or use a tunnel). When-to-use and when-to-substitute are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_scanAInspect
Scan a multi-page user journey. Walks startUrl plus the required autoNavigate URLs sequentially (deterministic — one page fully rendered and scanned before the next), then returns ONE consolidated report with issues deduplicated across pages, each carrying the same fix payload / confidence / review flags as scan_page. Every requested URL gets an explicit outcome in pages[] (scanned / nav_failed / scan_failed / redirected_duplicate / duplicate_request / skipped_cap / blocked — bot-challenge, not a clean page) — a page is never silently dropped. Better than per-page scans for journeys (login → checkout etc). For a single page, use scan_page. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| wcag | No | Only these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2"). | |
| rules | No | Only these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules. | |
| format | No | "compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json. | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| maxPages | No | Max pages to scan (default 10) | |
| startUrl | Yes | Starting URL of the journey | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| minImpact | No | Only findings at this severity or above (critical > serious > moderate > minor) | |
| sourceRoot | No | Local project root (local installs only) for `sourceCandidates[]` on issues without a framework `source` pointer. | |
| autoNavigate | Yes | REQUIRED — the URLs to walk after startUrl (the MCP server is headless and cannot discover a journey interactively). For a single page, use scan_page instead. | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it declares deterministic sequential rendering, one consolidated deduplicated report, and a closed enumeration of per-URL outcomes (scanned/nav_failed/scan_failed/redirected_duplicate/duplicate_request/skipped_cap/blocked) with the guarantee that no page is silently dropped. It also discloses the localhost/private-address refusal and the tunnel_secret requirement — significant operational constraints beyond the name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior and the sibling routing rule, with the hosted/local caveat correctly relegated later. It is long and repeats the 'for a single page, use scan_page' guidance twice, but nearly every sentence carries information an agent needs to call it correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description explains the return shape (pages[] with explicit per-URL outcomes, consolidated deduplicated issues each carrying fix payload/confidence/review flags), the traversal semantics, and the environment constraints. An agent has everything needed to invoke and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are already documented; baseline is 3. The description usefully reinforces that autoNavigate is required and walked after startUrl, but it also refers to a `url` parameter for tunnel usage that does not exist in the schema (startUrl does), which slightly muddies parameter mapping rather than clarifying it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (scan a multi-page user journey) and immediately separates itself from the closest sibling: 'For a single page, use scan_page.' The consolidated-report + dedup behavior is described concretely, so an agent can tell exactly what this tool produces versus per-page scans.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules: use flow_scan for journeys (login → checkout), use scan_page for a single page, and two named alternatives for local dev servers (run the MCP locally, or open a tunnel and pass the secret). It also states the hosted-host restriction so the agent knows when the tool simply cannot be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_ai_fixAInspect
Generate framework-aware fix alternatives for a specific accessibility issue. For color contrast issues, returns 3 alternatives (minimal, brand-aligned, high contrast); brand palette is auto-extracted from the live URL using our scanner if brandColors is omitted. For label/ARIA issues, returns 1-2 alternatives. Each alternative includes ready-to-paste code for the detected framework. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Page URL — also used to auto-extract brand palette for contrast issues if `brandColors` is not provided. | |
| html | Yes | The element's outerHTML — send at most ~600 chars | |
| issue | Yes | Issue object from scan_page (with selector, wcag, impact, message, fix.currentValue), or a plain-text issue description | |
| context | No | Parent element outerHTML for context (~400 chars) | |
| framework | Yes | CSS framework — use detect_framework first | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| brandColors | No | Brand palette for brand-aligned suggestions. If omitted on a contrast issue with a `url`, auto-extracted via the scanner. | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses substantial behavior: alternative counts by issue type, ready-to-paste framework code, automatic brand-palette extraction when brandColors is omitted, and the hosted-server restriction that localhost/private addresses are refused. It does not cover rate limits, timeouts, or failure modes, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence, and the issue-type branching follows logically. The lengthy NOTE on localhost/tunnel is verbose but each sentence carries actionable constraint information, so it earns its place even if it could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and nine parameters, the description does a good job explaining return values (3 alternatives named minimal/brand-aligned/high contrast, or 1-2 for label/ARIA issues, each with ready-to-paste code) and the environment constraint. It is nearly complete for invocation, lacking only error/edge-case handling detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already documented, establishing a baseline of 3. The description restates the url/brandColors auto-extraction behavior but adds essentially no syntax or format detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Generate) and resource (framework-aware fix alternatives for an accessibility issue), and it is immediately distinguishable from siblings like verify_fix and scan_page. The description also scopes output by issue type (color contrast vs label/ARIA), so an agent knows exactly what it produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for use (a specific accessibility issue) and practical routing guidance: use detect_framework first for the framework value, and two concrete workarounds when scanning localhost. It stops short of naming verify_fix or other siblings as the follow-up/alternative step, so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_auditAInspect
Check an audit started with start_audit: returns overall status, per-step progress (scan → viewports → screenshots → agent → excel → publish), and — once complete — a severity summary plus short-lived download URLs for the report (JSON) and the Excel workbook. Poll every ~15s while status is pending/running. Only the account that started an audit can read it — or, for a trial (no-account) run, only the caller holding the claimToken start_audit returned. Each poll also spends one trial call, so avoid polling faster than ~15s on a trial run.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The audit id returned by start_audit | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| claimToken | No | Trial (no-account) runs only — the claimToken start_audit returned. Omit if you are logged in. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the authorization model (only the starting account, or the claimToken holder for trial runs), the destructive-ish cost behavior (each poll spends one trial call), and the lifecycle ordering of the pipeline steps. This is richer than the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core purpose before the operational caveats. Slightly long, but every clause carries distinct operational information (poll cadence, auth model, trial cost) without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a polling/status tool with no output schema and no annotations, the description covers all the bases an agent needs to call it safely and effectively: what comes back, when to call, how often, and who is allowed. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning by explaining how id and claimToken relate to the start_audit call and the trial-vs-account distinction, which the schema states only in isolation. It does not add syntax detail for context, llm_model, or conversation_id beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource pair — 'Check an audit started with start_audit' — and enumerates exactly what the caller receives: overall status, per-step progress, severity summary, and download URLs. This distinguishes it from start_audit and from all the other check_* / scan_* siblings unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative workflow entry point (start_audit), instructs polling cadence (~15s while pending/running), and adds a trial-specific warning to avoid polling faster than ~15s because each poll spends a trial call. That is explicit when-to-use and when-to-avoid-in-a-way guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_rulesAInspect
List accessibility rules from both engines — axe-core (104) and the WebAbility detectors (90+) — with optional filters. Every rule carries fixability (mechanical | contextual | visual) and a fix op template, so you can pick the rules worth auto-fixing before scanning. Returns ruleId, engine, description, help, helpUrl, tags/wcag, fixability, fix.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | axe tag filter (e.g. ["wcag21aa"], ["best-practice"], ["cat.aria"]). WebAbility rules match on their WCAG criterion tag (e.g. "wcag143"). | |
| engine | No | Which engine's rules to list (default all) | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| fixability | No | Only rules of this fixability tier | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the burden and does well: it enumerates the returned fields (ruleId, engine, description, help, helpUrl, tags/wcag, fixability, fix) and explains the fixability tiers and fix-op template. It omits auth, rate-limit, or pagination details, but the data-model disclosure is genuinely valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and filters before the enumerated return fields. The closing field list is dense but earns its place given there is no output schema; nothing is wasted, though it is heavier than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, filtering, and the full return shape for a list tool with no output schema, giving the agent enough to call it correctly. The only gaps are the absence of explicit when-not guidance and operational details (pagination), which are minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both filters and required analytics params are already documented in the schema. The description references "optional filters" and the fixability tiers but adds no syntax or meaning beyond what the schema provides, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("List") and resource (accessibility rules), plus the exact scope: axe-core (104) and WebAbility (90+) detectors. It clearly separates itself from the scanning/audit siblings by framing itself as a pre-scan catalog lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context — "pick the rules worth auto-fixing before scanning" — which tells the agent when this tool is appropriate. It stops short of naming a specific sibling alternative or stating exclusions, so it does not reach the top tier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_htmlAInspect
Scan a raw HTML snippet or component markup without serving it — IN-PROCESS by default (jsdom + WebAbility detectors + axe-core): milliseconds, no browser, no network, so it fits inside a tight edit loop. Fragments are auto-wrapped into a document. Returns scan_page's three-tier shape (issues / incomplete / summary) with fix.op + fixability on every finding. jsdom has no layout, so visual-tier rules (contrast, target size, focus ring) are NOT evaluated — the dropped count is reported as skippedVisual; pass engine: "browser" to run the axe-core headless-browser path for those (slower, axe rules only, returns axe violations).
| Name | Required | Description | Default |
|---|---|---|---|
| html | Yes | HTML content to test — a full document or a fragment | |
| tags | No | WCAG tags to check (default ["wcag2a","wcag2aa","wcag21aa","wcag22aa"]) | |
| wcag | No | Only these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2"). | |
| rules | No | Only these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules. | |
| width | No | Viewport width (browser engine only, default 1280) | |
| engine | No | "in-process" (default): jsdom, ms, structural rules. "browser": headless Chromium + axe-core, includes contrast. | |
| format | No | "compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json. | |
| height | No | Viewport height (browser engine only, default 800) | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| minImpact | No | Only findings at this severity or above (critical > serious > moderate > minor) | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so unusually well: it discloses the default execution model, that there is no browser or network, that fragments are auto-wrapped, that jsdom lacks layout so contrast/target-size/focus-ring rules are skipped, and that the skipped count surfaces as skippedVisual. It also describes the alternate browser path's cost and narrower rule set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and default mode are front-loaded in the opening clause, and almost every sentence adds distinct information (execution model, return shape, visual-rule limitation, alternate engine). It is dense and parenthesis-heavy but not padded; a small amount of compression would improve readability without losing content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with no annotations and no output schema, the description covers the essentials: what is scanned, how fragments are handled, the return shape (issues/incomplete/summary with fix.op and fixability), the major limitation of the default engine, and how to opt into visual checking. An agent has enough to call it correctly on the first attempt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so engine, format, and the other parameters are already documented; the schema also explains the in-process vs browser distinction. The description mainly reinforces engine semantics at a consequence level (visual rules dropped, skippedVisual reported) rather than adding new syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first clause gives a specific verb+resource (scan raw HTML snippet/component markup) plus a distinguishing constraint ("without serving it"), which separates it from the page-serving sibling scan_page. It also names the concrete engines and return shape, so an agent knows exactly what class of operation this is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when the default in-process mode fits (tight edit loop, milliseconds, no browser/network) and when to switch to engine:"browser" for visual-tier rules that jsdom cannot evaluate, including the tradeoff (slower, axe rules only). What is missing is explicit routing versus siblings like scan_page or visual_audit — the agent must infer that those handle served URLs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_pageAInspect
Scan a web page for WCAG accessibility issues. Works on any URL — deployed sites, localhost, staging. Returns the three-tier shape: issues (high-confidence violations safe to fix), incomplete (needs human review — gradient backgrounds, marketing imagery, axe-incomplete results, framer-motion pre-animation states), and a summary. Treat incomplete as questions, never auto-fix them. On React ≤18 / Vue dev builds each issue carries source ({file, line, column, component}) read from the live component tree. Every issue carries a structured fix.op (add-attribute | set-attribute | remove-attribute | add-element | remove-element | add-text-content | suggest) with fix.attribute / fix.value when known, and a fixability tier (mechanical = apply as given; contextual = op known, value needs judgment; visual = needs rendered output, propose only). NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to scan (e.g. https://example.com or http://localhost:3000) | |
| wcag | No | Only these WCAG criteria. A prefix selects the whole guideline ("1.4") or principle ("2"). | |
| rules | No | Only these rule ids (WebAbility type such as "missing_alt" or axe rule id such as "image-alt"). See get_rules. | |
| format | No | "compact" prints one line per element with rule metadata once — far fewer tokens than the default JSON. Default json. | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| viewport | No | Viewport size (default: desktop) | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| minImpact | No | Only findings at this severity or above (critical > serious > moderate > minor) | |
| sourceRoot | No | Local project root (local installs only). Issues without a framework `source` pointer get `sourceCandidates[]` — files whose contents match the selector's id/class/attribute tokens. | |
| rootSelector | No | CSS selector to limit scan scope (optional) | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so unusually well: it documents the three-tier return shape, the semantics of `incomplete` (questions, never auto-fix), the `source` pointer availability by framework/build, the structured `fix.op` contract, and the `fixability` tiers. It also discloses the critical hosting constraint that localhost/private addresses are refused.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and scope, then progressively details return shape, fix metadata, and hosting constraints. It is long and reads as a dense paragraph rather than scannable sections, but nearly every sentence contributes distinct information, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with no output schema and no annotations, the description supplies the missing return-value documentation (issues/incomplete/summary, fix.op, fixability) and the environment caveat. Nothing essential to calling or interpreting the scan is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: it ties `url` and `tunnel_secret` together (secret required for tunnel URLs, otherwise ignored) and explains the localhost/tunnel workflow that makes those parameters usable. It adds little for `rootSelector`, `minImpact`, or `format`, which the schema already covers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Scan a web page for WCAG accessibility issues') and immediately scopes it with 'Works on any URL — deployed sites, localhost, staging.' An agent can distinguish it from scan_html, check_aria, check_color_contrast, and visual_audit without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: when the hosted server refuses localhost/private addresses, and the two concrete alternatives (run the MCP locally via npx, or open a tunnel and pass tunnel_secret). It also instructs how to treat the `incomplete` bucket. However it never routes among sibling tools (e.g. scan_html vs scan_page), so it stops short of explicit when-not-this-tool guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_auditAInspect
Kick off a FULL accessibility audit deliverable for a URL — a persistent, timestamped artifact, not an inline scan. Runs the server-side pipeline (axe + advanced checks + mobile viewports + annotated screenshots + optional agent spot-check) and produces a downloadable report and a formatted Excel workbook (Cover / Status / Barriers / ADA context sheets) stored durably. Returns immediately with an audit id; poll get_audit for progress and, when complete, download URLs. Use this when someone needs a durable artifact to attach as evidence of testing effort for a compliance officer or legal response — for iterating on code, use scan_page + verify_fix instead. Free without an account for a limited trial (a shared pool of calls across start_audit/get_audit/visual_audit, hosted deploy only) — the response says how many are left and includes a claimToken to pass to get_audit. Past the trial, or on the local/stdio server: authenticate via webability login or set WEBABILITY_API_KEY. Set includeAgent:true to add the (slower, paid) agentic manual-audit pass. To audit a local dev server, open a tunnel (webability-tunnel --port 3000) and pass its URL as url with the printed secret as tunnel_secret; keep the tunnel open until get_audit reports complete (about 5 minutes) — the pipeline loads the page several times.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to audit (a public/staging URL the server can reach, or a webability-tunnel URL — not localhost) | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| includeAgent | No | Also run the agentic manual-audit pass (keyboard/focus/modal exploration). Slower. Default false. | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden and does so richly: it discloses async behavior (returns immediately with an id; poll get_audit), the pipeline's composition, trial/pool limits, auth requirements, the tunnel workflow and its ~5-minute lifetime, and that includeAgent is slower and paid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and tightly packed, with every clause carrying operational value. It is dense and long for one paragraph, but for a complex async tool with trial/auth/tunnel caveats most sentences earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a 6-param async tool, the description covers the full lifecycle: immediate return value (id), polling path (get_audit), download URLs, trial count and claimToken, and auth for non-trial/local usage. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter workflow meaning the schema lacks: tunnel_secret is required when url is a tunnel URL, url must not be localhost, includeAgent is slower/paid, and the tunnel must stay open because 'the pipeline loads the page several times'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Kick off a FULL accessibility audit deliverable for a URL') and immediately scopes it against a sibling behavior ('a persistent, timestamped artifact, not an inline scan'). An agent can distinguish it from scan_page/visual_audit without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('when someone needs a durable artifact to attach as evidence ... for a compliance officer or legal response') and when-not, naming the alternatives ('for iterating on code, use scan_page + verify_fix instead'). Routing conditions are unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_fixAInspect
Re-scan a specific element after applying an accessibility fix and confirm the violation is gone — closes the loop that find-only tools leave open. After you edit the code and serve it (deployed, staging, or http://localhost:3000), call this with the URL and the selector you fixed to get a machine-checked verified: true|false (DOM engines only — visual_audit findings and needs-review items are out of scope). Pass the WCAG criterion (e.g. "1.1.1") or axe rule id (e.g. "color-contrast") to check just that criterion; omit it to require the element be clean of ALL violations. A blocked page (bot-challenge / HTTP error) is reported as unverified, never a pass — verification fails closed. IMPORTANT: if your fix changed the element's class or id, the original selector may no longer match anything, which reads as verified — re-run scan_page or pass the updated selector to be sure. Pair with scan_page → generate_ai_fix → verify_fix for a full find-fix-verify cycle. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL now serving the fix (deployed, staging, or http://localhost:3000) | |
| wcag | No | Optional: WCAG criterion (e.g. "1.1.1", "1.4.3") or axe rule id (e.g. "color-contrast") to verify specifically. Omit to require the element be free of ALL violations. | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| selector | Yes | CSS selector of the element you fixed — use the `selector` from the original scan_page issue | |
| viewport | No | Viewport size (default: desktop). Use the same viewport the issue was found at. | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses fail-closed behavior (blocked page = unverified, never a pass), engine limits (DOM engines only), the selector-drift gotcha that can produce a false 'verified', and the hosted-server restriction refusing localhost/private addresses. This is well beyond what any structured field provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but tightly packed and front-loaded with purpose before caveats; each sentence carries operational value (fail-closed, selector drift, tunnel/localhost). It is dense rather than redundant, though the HOSTED-server and tunnel paragraphs make it heavier than strictly minimal for a quick call.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, no-annotation, no-output-schema tool, the description covers the return contract (verified: true|false), failure semantics, engine scope, and the local-dev access story. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: it clarifies that omitting `wcag` requires the element be clean of ALL violations, explains the criterion-vs-rule-id dual form, and warns that `selector` must be the original or updated selector. These are genuinely useful additions over the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'Re-scan a specific element after applying an accessibility fix and confirm the violation is gone.' It explicitly distinguishes itself from find-only siblings like scan_page and positions itself in the find-fix-verify cycle, so the agent knows exactly what it is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States prerequisites (edit and serve the code), the exact selection workflow (scan_page → generate_ai_fix → verify_fix), and exclusions (visual_audit findings and needs-review items out of scope). It also names alternatives for the local-dev case (run MCP locally vs. use a tunnel), leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visual_auditAInspect
Pixel-level accessibility audit using Claude vision. Catches issues that DOM scanners miss: icon contrast (1.4.11), focus visibility (2.4.7), "looks like a button but isn't" (4.1.2), text rendered as images (1.4.5), visual hierarchy mismatches. Takes a URL, opens it in a headless browser, screenshots, and runs vision-based detection. Complements scan_page — run both for full coverage. Free without an account for a limited trial (shared call pool with start_audit/get_audit, hosted deploy only) — the response says how many are left. Past the trial, or on the local/stdio server, this and start_audit are the paid, server-side tools: authenticate via webability login or set WEBABILITY_API_KEY in your MCP server env before calling. NOTE: on this HOSTED server, localhost and private addresses are refused — it runs in our cloud and cannot reach your machine. Two ways to scan a local dev server: run the MCP locally (npx -y @webability/mcp, simplest — nothing leaves the machine), or open a tunnel (webability-tunnel --port 3000) and pass its URL as url together with the printed secret as tunnel_secret.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to audit visually | |
| context | Yes | Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as "a user", "the customer", or "an account". Example: "Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution." | |
| fullPage | No | Capture full scrolled page instead of just viewport (default: false) | |
| viewport | No | Viewport size (default: desktop) | |
| llm_model | Yes | The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. "claude-opus-4-8", "gpt-5.2"). Used for analytics only. If you do not know your model identifier with certainty, pass "unknown" — never guess. | |
| brandColors | No | Brand hex colors for context-aware filtering | |
| tunnel_secret | No | Secret printed by `webability-tunnel`. Required when `url` is a tunnel URL; the URL alone will be refused by the relay. Ignored otherwise. | |
| conversation_id | No | Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: shared trial call pool with start_audit/get_audit, limited free trial, paid server-side otherwise, required authentication, localhost/private-address refusal on the hosted server, and the tunnel_secret dependency. These are exactly the operational traits an agent must know before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the complement-to-scan_page routing are front-loaded, and every subsequent sentence addresses a real invocation risk (auth, trial limits, localhost refusal, tunnel). It is dense and somewhat run-on for a single paragraph, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given eight parameters, no output schema, and no annotations, the description covers the full operational picture: what it does, how it complements siblings, auth needs, rate/trial limits, and networking constraints with workarounds. An agent has everything required to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all eight parameters are already documented, and the description adds little parameter-level meaning beyond echoing the url-plus-tunnel_secret pairing the schema already states. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (pixel-level accessibility audit via vision), the resource (a URL rendered in a headless browser), and the distinct class of issues it catches (icon contrast, focus visibility, 4.1.2 affordance mismatches, text-as-image). It explicitly positions itself against scan_page, so an agent can differentiate it from that sibling without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: 'Complements scan_page — run both for full coverage,' plus a clear when-paid boundary (past trial or on local/stdio) and two concrete alternatives for scanning a local dev server (run MCP locally, or tunnel). It names prerequisites (auth via login/API key) and the hosted-localhost refusal, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
- First observed
check_aria - First observed
check_color_contrast - First observed
detect_framework - First observed
diff_scan - First observed
flow_scan - First observed
generate_ai_fix - First observed
get_audit - First observed
get_rules - First observed
scan_html - First observed
scan_page - First observed
start_audit - First observed
verify_fix - First observed
visual_audit
Related MCP Connectors
Scan URLs or HTML for WCAG 2.2 violations. 75-rule manifest, weighted score, shareable reports.
Security, SEO and AI-visibility scanner for web apps · free scans and focused checks via MCP.
Scan URLs for WCAG 2.1 violations, generate AI fixes, and produce VPAT 2.5 compliance reports.
Scan a web page for accessibility, security, privacy, quality and SEO issues, with fixes.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAutonomous WCAG 2.1 accessibility auditor that scans, fixes, re-verifies, and generates VPAT 2.5 EN 301 549 reports using AI vision analysis + DOM scanning.9,148 npm1MIT
- AlicenseNot gradedqualityCmaintenanceVisual frontend accessibility inspector MCP server. WCAG contrast checking, touch target validation, heading hierarchy audits, responsive screenshots, and S+ grading across mobile and desktop viewports.25 PyPIMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI coding assistants to test web accessibility by scanning URLs, detecting violations, and running focused audits on keyboard navigation, screen reader compatibility, and WCAG criteria — all within the assistant's loop.MIT
- AlicenseAqualityCmaintenanceProvides conversational, actionable accessibility testing for AI agents, including auditing, prioritization, and code-level fixes.2218 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.