Skip to main content
Glama

AccessLens MCP

AccessLens is a local MCP server for accessibility evidence, fix verification, and consistency checks across an agent session. It scans a rendered page with axe-core, stores DOM fingerprints rather than brittle selectors alone, and compares the evidence after a fix.

Install

npx accesslens-mcp --install-browser

After it is published to npm, a client can launch it directly with npx -y accesslens-mcp. The setup command downloads Playwright Chromium once; the browser is required to scan rendered pages.

Until the npm package is published, use the GitHub package spec instead:

npx -y github:AbhishekX-dev/AccessLens-mcp --install-browser

Related MCP server: Accessibility MCP Server

Add to an MCP client

{
  "mcpServers": {
    "accesslens": {
      "command": "npx",
      "args": ["-y", "accesslens-mcp"]
    }
  }
}

Evidence Fusion Protocol

Every finding carries its raw engine evidence, DOM fingerprint, methodology, and fusion explanation. AccessLens only labels a result cross-method corroborated when separate methodologies report the same criterion on the same target. It does not convert a vote count into a false probability of truth. Verification outcomes feed a per-engine calibration record, making the fusion policy empirically auditable over time.

Human-review gate

AccessLens follows report → human review → approved remediation → verify. After every scan, an agent must call run_accessibility_review (or get_review_report after a legacy scan), show it to the human, and wait. Only explicit human decisions may be recorded with record_human_review. verify_fix refuses to run for a finding without an approved review record. This makes the intended MCP workflow human-in-the-loop; the coding agent's own system prompt should separately prohibit file edits until that approval is received.

The installed engines are axe-core, QualWeb ACT Rules (URL scans), and the AccessLens semantic analyzer. When a scan cannot preserve equivalent rendering context, that engine is explicitly marked unavailable rather than silently excluded or treated as agreement.

Tools

  • run_accessibility_review: recommended one-call workflow. Renders the target once with Playwright and returns engine status, traceable grouped issues, session checks, advisory checks, and a pending-human-review report.

  • crawl_accessibility_review: crawls same-origin anchor links breadth-first (default: 10 pages, depth 2), runs the full review on every discovered page, and returns one combined pending-human-review report. It never follows external links or invents unlinked SPA routes.

  • analyze_accessibility_evidence: scan a URL or HTML, collect evidence, and store a snapshot in a session.

  • verify_fix: re-scan an explicit updated URL/HTML and compare the original finding plus its affected subtree.

  • check_session_consistency: detect duplicate IDs, heading-level jumps, invalid landmark counts, and inconsistent accessible names.

  • get_session_report: returns outcome metrics and open findings.

  • analyze_reading_order, analyze_alt_text, analyze_control_names, and analyze_keyboard_paths: advisory modules for the roadmap checks.

Findings are returned in a versioned normalized schema designed for multiple engines and a provenance-preserving Evidence Fusion Protocol.

Security model

This is intended for local development targets. By default it accepts only http(s) URLs and blocks localhost/private-network targets unless ACCESSLENS_ALLOW_PRIVATE_TARGETS=true is set. HTML input never makes a network request.

Available Tools

12 tools
analyze_accessibility_evidenceA

Render a URL or HTML document, run axe-core plus AccessLens semantic checks (and QualWeb ACT Rules for URL scans), preserve raw per-engine provenance, and persist a pending-human-review report. Do not modify code: present get_review_report to the human first.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
url_or_htmlYesA public http(s) URL, or an HTML document.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the side effect of persisting a pending-human-review report, the non-code-modifying guardrail, and the preservation of raw per-engine provenance. It does not mention auth or rate limits, but the core behavioral traits are clearly covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; the first sentence front-loads the diagnostic pipeline and the second adds a necessary workflow guard. The first sentence is somewhat dense but every clause earns its place, so just short of top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex multi-engine tool with no output schema, it covers the workflow and next step (get_review_report), but fails to explain the role of session_id or what the function returns after persisting the report, leaving notable invocation ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, with session_id lacking any schema description. The tool description likewise omits session_id, only reinforcing url_or_html (already described in schema) and adding the URL-only QualWeb nuance. This leaves one of two parameters semantically opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific pipeline: render a URL/HTML document, run axe-core plus AccessLens semantic checks, QualWeb ACT Rules for URL scans, preserve per-engine provenance, and persist a pending-human-review report. This clearly distinguishes it from sibling review tools by its evidence-preservation and human-review workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies a workflow by instructing the agent to present get_review_report to the human first and not modify code, but never explicitly states when to choose this tool over run_accessibility_review or crawl_accessibility_review, nor any exclusions or alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_alt_textB

Advisory check for missing or generic alt text. It cannot assess semantic truthfulness without a vision model.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
url_or_htmlYesA public http(s) URL, or an HTML document.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals one important limitation (cannot assess semantic truthfulness without a vision model), which is valuable. However, it does not disclose whether the check is read-only, what it returns, or any side effects. For an advisory check, this is partially sufficient but lacks depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. The primary purpose is front-loaded, followed by a concise limitation. It is exactly as long as needed and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, and the description does not hint at what the advisory check returns (e.g., findings, scores, recommendations). It also omits how session_id is used and when this tool should be invoked among the many accessibility checks. For a tool that likely returns structured results, this is a significant gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (only url_or_html has a description). The tool description does not mention parameters at all, so it adds no meaning beyond the schema. The required parameter's format is already in the schema, but session_id is left undocumented in both schema and description. The description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: an advisory check for missing or generic alt text. It also specifies a limitation, which further clarifies its scope and distinguishes it from broader semantic analysis. This is a specific verb-resource pair with no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus its siblings (e.g., analyze_control_names, analyze_reading_order). It implies it's for alt text issues, but offers no guidance on context or exclusions. An agent would have to infer usage from the name and purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_control_namesC

Advisory inventory of repeated control names and roles for human review.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
url_or_htmlYesA public http(s) URL, or an HTML document.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It hints at non-destructiveness ('Advisory') and a report-like output, but does not disclose side effects, input processing behavior, or output format. The 'advisory' wording suggests read-only, yet this is not confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words. It front-loads the core purpose and immediately conveys the tool's advisory nature. Strikingly concise and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool in an accessibility suite with sibling analysis tools, this description is thin. It does not explain the expected output structure, the relationship to sessions, or how it fits into the review workflow. Agents cannot predict the return value or know when to rely on it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 50% coverage (url_or_html is described, session_id is not). The tool description adds nothing about parameters—no mention of how url_or_html is used or the role of session_id. It fails to compensate for the half of parameters left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('analyze' via 'inventory'), a resource ('control names and roles'), and a clear intent ('for human review'). It distinguishes from sibling analysis tools by focusing on 'repeated' instances, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus any sibling, no context on prerequisites (e.g., session), and no mention of when it is appropriate to invoke. The phrase 'for human review' implies it feeds into a review workflow, but that link is not made explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_keyboard_pathsB

Advisory static keyboard reachability check. It does not simulate full interaction flows.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
url_or_htmlYesA public http(s) URL, or an HTML document.

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the check is advisory and static, and that it avoids simulating full flows. It does not address side effects, network behavior, or output format, so it is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the key 'static/advisory' scope is front-loaded. The exclusion of interaction simulation is useful and earned. It is concise, though slightly under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, this is incomplete: it does not describe what the check returns, what 'advisory' means in practice, or how session_id affects the analysis. An agent has enough to infer basic behavior but not enough to confidently use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% and the description adds no parameter-level meaning. url_or_html is documented in the schema, but session_id has no description and the tool description does not clarify it or how the two parameters interact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the tool's specific function as a keyboard reachability check, and 'static'/'advisory' conveys its scope. It is distinct from sibling analysis tools like analyze_alt_text or analyze_control_names, though it does not name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The statement 'It does not simulate full interaction flows' provides an implied exclusion for dynamic interaction testing. However, it does not explicitly say when to prefer this tool over siblings such as run_accessibility_review or provide concrete use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_reading_orderB

Advisory check: compare focus/tab order with rendered bounding-box order. It is not a substitute for a human reading-order assessment.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
url_or_htmlYesA public http(s) URL, or an HTML document.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does disclose that the check is advisory and not a substitute for human assessment, which is useful. However, it does not explain what the operation actually returns, whether it needs a rendered page, or what limitations apply beyond the human-substitute caveat, so coverage is only moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no redundant content. The core purpose is front-loaded, and the important limitation is stated immediately after. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool, the description covers the main purpose and one key limitation. However, it lacks any mention of return/output behavior and leaves session_id semantically unexplained. It is adequate for basic invocation but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: url_or_html is documented in the schema, but session_id has no schema description. The tool description itself provides no parameter-level explanation, so it does not compensate for the gap. An agent cannot infer what session_id is for or how it affects the analysis.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'compare focus/tab order with rendered bounding-box order.' It clearly identifies this as an advisory check about reading order, which distinguishes it from most sibling tools. However, it does not explicitly differentiate it from closely related tools like analyze_keyboard_paths, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage as an advisory, non-authoritative check and explicitly says it is 'not a substitute for a human reading-order assessment.' This gives some when-not-to-use guidance, but it does not name alternative tools or state when this tool should be preferred over siblings. The usage context is implied rather than fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_session_consistencyC

Review stored page snapshots for duplicate IDs, heading-level gaps, multiple main landmarks, and controls whose stable DOM path acquired a different accessible name across scans.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does not state whether the operation is read-only, whether it requires existing session data, what happens if the session_id is invalid, or what the output format is. 'Review' implies a non-mutating action, but this is weak and unsupported by explicit statements. Significant behavioral aspects are left undocumented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the primary action ('Review stored page snapshots') and then lists specific checks in a compact list. It avoids redundancy and is reasonably scannable, though the enumeration of four checks could be broken out for clarity. Overall, it is efficient without being overly terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description is incomplete. It does not explain the expected return value, how results are presented, or how session_id relates to the stored snapshots. The agent cannot fully anticipate what happens after invoking the tool, nor does it know whether the tool is safe to call without side effects. The complexity of the checks (multiple landmarks, DOM path analysis) warrants more context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0% and the description does not mention the session_id parameter at all. An agent has no information about the parameter's meaning, format, constraints, or whether it is required (the schema marks it as not required, which is misleading for a tool that likely needs a valid session). The description fails to compensate for the schema's lack of detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Review') on a defined resource ('stored page snapshots') and enumerates concrete checks (duplicate IDs, heading-level gaps, multiple main landmarks, control name changes). It clearly distinguishes itself from siblings like analyze_reading_order or analyze_alt_text by focusing on cross-scan consistency issues, leaving no ambiguity about its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the sibling analysis tools. There is no mention of scenarios that would select this over analyze_control_names or get_session_report, nor any exclusions. The phrase 'across scans' implies a comparative context, but the description does not explicitly state when a user or agent should invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crawl_accessibility_reviewA

Crawl ordinary same-origin anchor links breadth-first from a start URL, within explicit page/depth limits. Runs the same Playwright, axe-core, QualWeb, structural, and advisory review on every discovered page, then returns one combined pending-human-review report. External links and unlinked SPA routes are not guessed.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_depthNo
max_pagesNo
start_urlYesThe http(s) URL where same-origin crawling begins.
session_idNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It discloses the crawl strategy (breadth-first, same-origin anchors), limits (page/depth), the review engines used, the output status (pending-human-review), and what it deliberately avoids. It lacks details on auth, rate limits, or side effects, but the core behavior is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three focused sentences with no filler. The first sentence front-loads the action and scope, the second explains the review process and output, and the third clarifies exclusions. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives a solid overview for a crawl tool, but with no output schema and only 25% parameter coverage, it leaves gaps: max_depth/max_pages are not defined, session_id is unexplained, and the report structure is only named, not described. It is adequate but not complete enough for an agent to use all parameters confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only start_url is described). The description generically mentions 'page/depth limits' but does not explain max_depth or max_pages semantics, and session_id is completely undocumented in both schema and description. The description adds some context for start_url but does not compensate for the other three parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Crawl') and a precise resource ('ordinary same-origin anchor links'), then states the outcome ('returns one combined pending-human-review report'). It clearly distinguishes this from siblings like run_accessibility_review by emphasizing the breadth-first multi-page crawl.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use case is implied: crawl same-origin anchor links from a start URL. The description gives negative guidance ('External links and unlinked SPA routes are not guessed') but never explicitly says when to prefer this tool over run_accessibility_review or another sibling, leaving the alternative selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_review_reportC

Create the human-facing review packet: each open finding, evidence, proposed safe remediation category, and its approval status. An agent must show this report and ask for explicit human approval before editing code.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It mentions the report is human-facing and requires explicit approval, which is a key behavior. However, it doesn't specify what happens if approval is not granted, whether the report is persisted, or the exact structure of the approval status. The description is minimal and leaves much to inference.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, fitting in two sentences. The first sentence lists the report contents, and the second states the critical usage requirement. It is front-loaded with the purpose and includes an action-oriented note. No wasted words, but it could be slightly more structured if it included param semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one parameter, no annotations, and no output schema, the description needs to cover parameter meaning and return value expectations. It lacks both. The description explains the report contents partially but doesn't clarify what the output looks like or how it connects to other tools like record_human_review. Incomplete for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and there is one parameter (session_id) with no description in the schema. The description doesn't mention session_id at all, so the agent has to guess its purpose. The description should explain that session_id identifies the session for which the report is generated, but it doesn't, leaving the agent without essential semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates a human-facing review packet with specific contents (open findings, evidence, proposed safe remediation category, approval status). It distinguishes itself by emphasizing the human-facing aspect and the approval workflow, which sets it apart from internal analysis tools like analyze_accessibility_evidence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: before editing code and when explicit human approval is needed. However, it doesn't explicitly state when not to use it or name alternatives such as get_session_report, which might also provide report-like output. The context is clear but lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_session_reportB

Return all recorded verification outcomes, the verified-fix-rate, current open findings, and scan totals.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. 'Return' signals a read-only retrieval and the output contents are listed, which is helpful. However, it does not mention prerequisites like an existing session, failure behavior, or whether any state changes occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with the action front-loaded and concrete data elements listed. Every word earns its place; there is no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the return contents but omits the role of session_id, the fact that the schema marks no parameters as required, and how this report relates to sibling report tools. With no annotations or output schema, an agent still has unanswered setup and selection questions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the session_id parameter, and the description never mentions session_id or explains how it selects the report. The property name is mildly self-explanatory, but the description adds no semantic value and fails to compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and enumerates concrete report contents: verification outcomes, verified-fix-rate, open findings, and scan totals. It clearly identifies what the tool produces, though it does not explicitly contrast itself with sibling tools like get_review_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use get_session_report versus get_review_report or the other sibling tools. The only hint is the tool's name, so an agent must infer usage context without any explicit conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_human_reviewA

Record an explicit human decision after the human has reviewed the report. Only call this tool in response to a clear human approval or rejection; it is not permission for the agent to decide on its own.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
decisionYes
session_idNo
finding_idsYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It usefully clarifies that the tool records a human decision and must not be used for autonomous agent decisions, but it does not disclose side effects, persistence, validation behavior, or return values. This is adequate but leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no filler. The core purpose and usage constraint are front-loaded, making the tool's intent immediately clear without requiring the reader to parse unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 4 parameters, no output schema, and no annotations, yet the description only covers the decision aspect and the triggering condition. An agent still lacks guidance on what finding_ids refer to, what session_id means, how the note field is used, and what the tool returns or changes. This is incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the bare parameter names. It only hints at 'approval or rejection' corresponding to the decision enum, but gives no meaning for finding_ids, session_id, or note. The description does not sufficiently explain how to populate the required parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Record') and resource ('explicit human decision') and clearly distinguishes the tool from the sibling analysis/review tools by stating it logs human approval or rejection rather than performing analysis. This gives an agent a precise understanding of the tool's role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to call the tool ('Only call this tool in response to a clear human approval or rejection') and what it is not for ('not permission for the agent to decide on its own'). This is direct, actionable guidance with clear exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_accessibility_reviewA

Preferred one-call workflow. Render a URL or HTML once with Playwright; run axe-core, AccessLens structural/session checks, QualWeb ACT Rules when an URL is supplied, and static reading-order, alt-text, control-name, and keyboard advisories. Returns one human-review report. Do not edit code; present the report and wait for explicit approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idNo
url_or_htmlYesA public http(s) URL, or an HTML document.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the tool renders with Playwright, runs multiple backend checks (axe-core, AccessLens, QualWeb, static advisories), returns 'one human-review report', and explicitly states 'Do not edit code; present the report and wait for explicit approval.' This conveys that it is a read-only analysis action with no side effects on code, and that it will not proceed without user consent. Minor omissions like session_id behavior or potential rate limits, but the core behavioral traits are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one dense but efficient sentence, front-loading the 'Preferred one-call workflow' tagline and then listing the checks and outcomes. It avoids redundancy and stays within a reasonable length for a compound tool. It could be slightly clearer with line breaks, but it is appropriately sized without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the lack of an output schema, the description covers the essential workflow: what it renders, which checks it runs, what it returns (a human-review report), and the required behavior (no code edits, wait for approval). It does not describe the report's structure or potential errors, but for an agent to invoke and handle the result, the main steps are covered. The session_id parameter is unaddressed, which is a small gap in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: url_or_html is described, but session_id is not. The description adds value by explaining that url_or_html can be a public HTTP URL or an HTML document, and that QualWeb ACT Rules are only run 'when an URL is supplied.' However, it does not explain the purpose or usage of session_id (e.g., for continuing an existing review session), leaving an agent uncertain about when and why to pass it. This partial compensation is not fully sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific, compound action: render a URL/HTML once, run multiple accessibility checks (axe-core, AccessLens, QualWeb ACT Rules, and static advisories), and return a single human-review report. It explicitly differentiates itself as the 'preferred one-call workflow' and lists the checks, distinguishing it from the sibling analysis tools that focus on single aspects (e.g., analyze_reading_order, analyze_alt_text).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description positions itself as the 'preferred one-call workflow' for a full accessibility review, and clarifies that code should not be edited; the agent should present the report and wait for explicit approval. It implies this is the go-to for comprehensive reviews, but does not explicitly name alternatives or state when to use the more granular sibling tools (e.g., 'use analyze_reading_order for a focused check'). This is clear context but lacks explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_fixA

Re-scan a fixed page and determine whether a previously observed finding cleared. This tool rejects verification unless that finding has a recorded human approval. Also reports new rules introduced since the original scan.

ParametersJSON Schema
NameRequiredDescriptionDefault
finding_idYes
session_idNo
url_or_htmlNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals important behavior: verification is rejected unless a human approval is recorded, and new rules introduced since the original scan are reported. However, it does not disclose whether the tool has side effects, writes any state, or what happens when verification succeeds or fails beyond the rejection condition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It leads with the primary action, then states the key rejection condition, and finishes with the additional reporting behavior. Every sentence adds meaningful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and three undocumented parameters, the description is not complete enough. It omits crucial details such as what constitutes a valid finding_id, how session_id should be obtained, what format url_or_html should take, what the returned result looks like, and what side effects (if any) the verification process causes. An agent would likely need to guess or consult external sources to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no descriptions for the parameters (0% coverage), and the description does not explicitly define them. It only offers indirect context: 'previously observed finding' hints at finding_id, 'fixed page' suggests url_or_html, and 'original scan' implies session_id. This is insufficient for an agent to confidently map parameters to their intended values without additional inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states a specific action ('re-scan a fixed page') and the core determination (whether a previously observed finding cleared), and it also mentions the additional behavior of reporting new rules. This clearly distinguishes it from sibling tools like run_accessibility_review, which would perform a fresh review rather than verifying a prior finding.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the intended use case—verifying a previously observed finding after a fix—and states a key precondition (human approval must be recorded). However, it does not explicitly mention when not to use it or name alternative tools for related scenarios, such as running a fresh accessibility review instead of verifying a specific finding.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 12 tool updatesv0.1.0
    • First observedanalyze_accessibility_evidence
    • First observedanalyze_alt_text
    • First observedanalyze_control_names
    • First observedanalyze_keyboard_paths
    • First observedanalyze_reading_order
    • First observedcheck_session_consistency
    • First observedcrawl_accessibility_review
    • First observedget_review_report
    • First observedget_session_report
    • First observedrecord_human_review
    • First observedrun_accessibility_review
    • First observedverify_fix

TDQS

B3.4/5.0

Scored across 12 tools

Disambiguation3/5

run_accessibility_review, analyze_accessibility_evidence, and crawl_accessibility_review all trigger overlapping accessibility scans, and crawl is mainly differentiated by multi-page breadth. The descriptions help clarify intent, especially with 'Preferred one-call workflow' on run_accessibility_review, but several advisory tools also duplicate checks already included in the main review.

Naming Consistency5/5

Every tool uses a consistent lowercase snake_case verb_noun pattern: crawl_, run_, analyze_, get_, record_, verify_, and check_. Report, session, review, and fix nouns are used predictably, making the set highly navigable.

Tool Count4/5

Twelve tools is within a reasonable range for an accessibility audit server, but the count feels slightly padded by overlapping review entry points and standalone advisory checks that duplicate capabilities already present in run_accessibility_review. Consolidation would make the set tighter without losing coverage.

Completeness4/5

The scan-to-report-to-human-approval-to-verification lifecycle is well covered, with crawl, session consistency, and aggregate reporting adding useful breadth. Minor gaps like explicit session management or scan-scope configuration exist, but the core workflow has no dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers