Skip to main content
Glama
mya11yreport

mya11yreport-mcp

by mya11yreport

MyA11yReport MCP server

What continuous accessibility monitoring? Check out MyA11yReport MyA11yReport is an automated accessibility scanner that uses AI to filter out false positives and explain genuine WCAG issues in plain English, featuring a centralized dashboard to track active issue counts, severities, and site progress over time.

Try it for free today

About this MCP

An open-source Model Context Protocol server that gives AI agents real accessibility-auditing abilities. It runs axe-core audits and drives a real browser through Playwright (navigate, click, type, check/uncheck, scroll, screenshots, aria snapshots, page JavaScript), plus page reviews for alt text, structure and tab order, and a session-free WCAG 2.2 color-contrast checker.

Built for agents: sessions are explicit, every action is logged, and audit history is keyed by URL. It speaks MCP over stdio and needs no account, no API key and no network service of its own.

Related MCP server: mcp-a11y-tools

What you get

  • Automated audits — run_a11y_audit runs axe-core on the current page and accumulates results per URL.

  • Browser control — navigate, click, type, press_key, check, uncheck, select_option, scroll_by, screenshot, get_page_snapshot, evaluate, get_viewport_size, get_scroll_position.

  • Page reviews — list_images (alt text), get_structure (landmarks, headings, lists, frames), get_tab_order (focus order).

  • Contrast maths — check_color_contrast computes the WCAG 2.2 ratio for any two colors, with no session or browser.

  • Audit guide — get_audit_guide returns the full step-by-step audit workflow as markdown, including which checks must be done by a human.

Requirements

  • Node.js 20 or newer (node --version)

  • Chromium, installed once with one command (see below). It is ~150 MB.

  • For headless: false sessions, a machine with a real display.

Install

npm i -g mya11yreport-mcp
mya11yreport-mcp install chromium

The second step downloads the Chromium build that matches the server's Playwright version. It is a separate step rather than a post-install script because some machines block npm lifecycle scripts. If you skip it, the server still starts and simply tells you to run the command the first time a browser tool is used.

Prefer not to install globally? Use npx:

npx -y mya11yreport-mcp install chromium

To confirm the server starts on your machine:

echo '{}' | mya11yreport-mcp

It should exit cleanly; EOF on stdin closes the server.

From source

Clone the repository and build it locally:

git clone <repository-url> mya11yreport-mcp
cd mya11yreport-mcp
npm install
npm run build
node dist/index.js install chromium

Then run the server from the checkout:

node /path/to/mya11yreport-mcp/dist/index.js

Optionally put the mya11yreport-mcp command on your PATH with npm link, so you can use it anywhere:

npm link
mya11yreport-mcp install chromium

Connect it to your MCP client

Point your MCP client at the server over stdio. In opencode, add it to your project opencode.json or your global ~/.config/opencode/opencode.json.

After a global install or npm link:

{
  "mcp": {
    "mya11y-audit": {
      "type": "local",
      "command": [
        "mya11yreport-mcp"
      ],
      "enabled": true
    }
  }
}

Without a global install, use npx:

"command": [
  "npx",
  "-y",
  "mya11yreport-mcp"
]

When running from a local checkout, point at the built entry file directly:

"command": [
  "node",
  "/path/to/mya11yreport-mcp/dist/index.js"
]

Other MCP clients use the same idea — run the server executable with no arguments over stdio. Restart the client after changing its configuration.

Using the server

You do not call tools yourself: once the server is connected, your agent does. The typical flow is:

  1. get_audit_guide — needs no session: returns the full step-by-step audit workflow as markdown (sitemap enumeration, the automated pass, the structure/tab-order/images/contrast reviews, report format, and the human-only checks). Call it before a full audit. The same guide is published at https://mya11y.report/mcp/auditor-skill.

  2. start_session — returns a sessionId (pass an optional free-form alias, max 225 chars, sanitized to [A-Za-z0-9_-]; omit for a hex id). Also picks headless: true|false.

  3. Pass sessionId to every call: navigate, click, type, press_key, check, uncheck, scroll_by, screenshot, get_page_snapshot, evaluate, get_viewport_size, get_scroll_position, run_a11y_audit, list_images, get_structure, get_tab_order.

  4. Targets are {"selector": "#id"} or {"fn": {"name": "getByRole", "args": ["button", {"name": "Sign in"}]}} (getByRole | getByText | getByLabel | getByPlaceholder | getByAltText | getByTitle).

  5. run_a11y_audit audits the current session page (or navigates to url first) and returns the session's audit history: {pages: string[], audits: [{pageId, audit}]} — one hex pageId per URL, audits accumulate per page.

  6. evaluate runs a JavaScript expression or function in the page via page.evaluate and returns the serialized value (non-serializable or undefined results come back as null with serializable: false). Use it to read computed styles, geometry or custom element state.

  7. check_color_contrast needs no session and no browser: pass two colors (hex or rgb()/rgba()) and optionally fontSizePx/fontWeight, and it returns the WCAG 2.2 ratio, the AA/AAA grades, and whether the pair counts as large text.

  8. list_images returns every <img> / inline <svg> with its accessible name and source, src/currentSrc, sanitized SVG markup, a decorative flag and derived missingAlt/longAlt flags plus counts. Filter with filter: all | decorative | missing-alt | has-alt.

  9. get_structure returns landmark regions, headings, nested lists and iframes, plus a heading outline grouped by region with skipped-level / repeated-H1 flags. Regions and headings also carry their raw ariaLabel / ariaLabelledby attributes. Narrow with include: [...].

  10. get_tab_order returns the Tab order (positive tabindex first, then document order) with each stop's tag, accessible name (text), effective tabindex, shadow-piercing selector, raw ariaLabel / ariaLabelledby, nameSource, hasLabel and outOfOrder flags.

  11. close_session when done.

The three review tools (list_images, get_structure, get_tab_order) inspect the top frame and open shadow roots only. get_tab_order uses the same tabbable engine, vendored into the build.

Sessions and logs

  • Sessions auto-close after 5 minutes of inactivity.

  • Every session writes an action history to .mya11yreport-mcp/logs/<sessionId>/<sessionId>.json. Screenshots land in the same directory as <actionId>.png.

  • Nothing is sent off your machine; the logs are plain local files you can inspect or delete at any time.

Configuration

Variable

Default

Purpose

MYA11Y_MCP_IDLE_CLOSE_MS

300000 (5 min)

Idle auto-close delay (min 1000)

MYA11Y_MCP_LOG_DIR

<cwd>/.mya11yreport-mcp/logs

Session log directory

PLAYWRIGHT_HEADLESS

true

Default headless mode (param overrides)

headless: false needs a real display on the host machine.

Troubleshooting

  • "Run mya11yreport-mcp install chromium" — you have not installed the browser yet. Run that command once.

  • Timeouts or "element not found" — the target was not actionable in time. Ask the agent to call get_page_snapshot to discover targets, or retry with a larger timeoutMs.

  • Server exits immediately — that is expected when stdin closes (for example in the smoke test above). It is not an error.

Available Tools

21 tools
checkCheck checkbox/radioA

Checks a checkbox or radio input in the session page. No-op if already checked. Targets support selectors or getBy* locators (e.g. getByRole with name).

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesHow to locate the element: {"selector": "#id"} or {"fn": {"name": "getByRole", "args": ["button", {"name": "Sign in"}]}}.
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
timeoutMsNoAction timeout in milliseconds (default 10000).

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does disclose a meaningful behavior: 'No-op if already checked', and clarifies acceptable target forms. However, it omits other relevant behavioral traits such as what happens on failure, whether it waits for the element, return value, or side effects like unchecking a radio's siblings. It is minimally adequate but has clear gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler. It front-loads the core action, then adds the no-op behavior and targeting options. Every sentence contributes essential information, making it appropriately sized for the tool's simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the detailed input schema, and the absence of an output schema, the description is largely complete for correct invocation. It covers the action, the no-op case, and target specification. It could be more explicit about return behavior or error conditions, but those are minor for this type of action tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a small amount of value by summarizing that targets support selectors or getBy* locators and giving an example (getByRole with name), but it does not materially extend parameter semantics beyond what the schema already documents for sessionId, target, or timeoutMs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Checks a checkbox or radio input in the session page.' This distinguishes it from the sibling tool 'uncheck' and other interaction tools like 'click' by naming the exact input types it acts on. It also adds a useful behavioral fact ('No-op if already checked') that reinforces its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by naming the target input types and notes that it is a no-op when already checked, which tells the agent when the call is unnecessary. However, it does not explicitly contrast with alternatives like 'click' or 'uncheck', nor does it state when to prefer this tool over them. The targeting guidance (selectors or getBy* locators) is practical but not usage-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_color_contrastCheck color contrastA

Computes the WCAG 2.2 contrast ratio between two colors and which of the four AA/AAA thresholds they clear. Accepts hex or rgb()/rgba() colors. Needs no session and no browser. Pass fontSizePx and fontWeight to also classify the pair as WCAG large text.

ParametersJSON Schema
NameRequiredDescriptionDefault
backgroundYesBackground color. Accepts hex (#RGB/#RRGGBB) or rgb()/rgba().
fontSizePxNoOptional font size in CSS pixels; used with fontWeight to classify WCAG large text.
fontWeightNoOptional font weight (e.g. 700 or "bold"); used with fontSizePx to classify WCAG large text.
foregroundYesForeground (text) color. Accepts hex (#RGB/#RRGGBB) or rgb()/rgba().

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for behavioral disclosure. It explicitly states that it needs no session/browser, indicating a stateless, side-effect-free computation. It also specifies accepted input formats (hex or rgb()/rgba()), which is behavioral detail beyond the schema. It does not mention error handling or return shape, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The primary function is front-loaded, followed by input formats, then the optional parameters. Every sentence earns its place, and the description is appropriately compact for a straightforward utility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateless computation tool with full schema coverage and no output schema, the description is complete enough. It explains the inputs, the computation, and the optional classification. It does not explicitly describe the return format (e.g., whether it is a ratio or a boolean for each threshold), but the phrase 'computes... and which thresholds they clear' strongly implies the output; given the tool's simplicity, this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all parameters. The description adds meaningful semantics by explaining that fontSizePx and fontWeight are used together to classify WCAG large text, which is not evident from the schema alone. This goes beyond repeating parameter names and aids correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('computes'), a precise resource (WCAG 2.2 contrast ratio), and the output (which of the four AA/AAA thresholds are cleared). It also clarifies accepted color formats and the optional large-text classification. This clearly differentiates it from sibling tools that operate on the browser/session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: 'Needs no session and no browser,' implying this is a standalone utility to be used without a browser context. It does not explicitly name alternatives or exclusions, but among the siblings this is the only contrast-calculation tool, and the guidance is sufficient to know when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickClick elementA

Clicks an element in the session page. Auto-waits for the element to be visible, stable, and actionable. Targets support CSS selectors or accessibility-first locators (getByRole, getByText, getByLabel, ...).

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesHow to locate the element: {"selector": "#id"} or {"fn": {"name": "getByRole", "args": ["button", {"name": "Sign in"}]}}.
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
timeoutMsNoAction timeout in milliseconds (default 10000).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It adds useful transparency by stating that the tool auto-waits for the element to be visible, stable, and actionable. However, it does not disclose possible side effects such as navigation, page changes, or failure behavior on timeout, so coverage is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences place the action first, then the key behavioral guarantee, then the targeting options. There is no filler, repetition of schema details, or unnecessary context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The combination of a rich parameter schema and this focused description covers the core knowledge an agent needs to call the tool correctly. The main gaps are the lack of alternative routing to state-change siblings and no statement of return/error behavior, but neither is critical for a straightforward click action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents sessionId, target, and timeoutMs in detail. The description adds only a light framing around CSS selectors and accessibility-first locators, which largely mirrors the schema's anyOf structure. It meets the baseline but does not add significant meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Clicks an element in the session page.' It clearly identifies both the action and the target context. However, it does not explicitly differentiate from sibling state-change tools such as check, uncheck, or select_option, so differentiation is mostly implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The auto-wait behavior and supported locator types suggest when the tool is appropriate, but the description gives no explicit when-to-use or when-not-to-use guidance. It does not mention alternatives like check/uncheck for checkbox states or select_option for dropdowns, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_sessionClose sessionA

Stops a session and terminates its browser context (page state and cookies are gone; the session's history file is finalized with a closedAt timestamp). Returns closed:false plus the active sessions when the id is unknown. A closed session can be started again later; its previous history is archived, not overwritten.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It transparently describes destructive effects: page state and cookies are gone, the history file is finalized, and the session's history is archived rather than overwritten. It also explains the behavior for an unknown id, which is valuable edge-case information.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: primary effect, error return behavior, and lifecycle caveat. No filler or repetition of the schema. The most important action is stated first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is thorough for a single-parameter tool with no output schema, covering effects, error behavior, and post-close lifecycle. The only minor gap is that it does not explicitly state the success return value or shape, though it is largely inferable from the unknown-id behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful parameter-related behavior by explaining what happens when the sessionId is unknown (returns closed:false and active sessions). This goes beyond the schema's simple type and description, aiding correct invocation and error handling.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Stops', 'terminates') and identifies the exact resource (a session and its browser context). It clearly distinguishes itself from sibling tools like start_session by defining the lifecycle relationship. This is unambiguous and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about when to use the tool: when a session should be ended and its browser context cleaned up. It also notes that a closed session can be restarted later, which helps the agent understand the lifecycle. It does not explicitly name alternatives or exclusions, but no close alternative exists among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluateEvaluate JavaScriptA

Evaluates a JavaScript expression or function in the session page via Playwright page.evaluate and returns its value. Use it to inspect the live DOM (computed styles, geometry, custom element state) when get_page_snapshot is not enough. The result must be serializable across the Playwright boundary; non-serializable or undefined values come back as null with serializable:false. This runs arbitrary page JavaScript — only evaluate code you intend to run.

ParametersJSON Schema
NameRequiredDescriptionDefault
argNoOptional JSON-serializable value passed as the first argument when `expression` is a function. Ignored for plain expressions.
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
expressionYesJavaScript evaluated in the page context. Either a plain expression (e.g. "document.title") or a function/arrow-function body (e.g. "() => document.querySelectorAll(\"a\").length", or "(n) => n + 1"). Functions are invoked with `arg`; plain expressions ignore it. Multi-statement logic needs a function body.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that results must be serializable, that non-serializable or undefined values return as null with a serializable:false flag, and that the tool runs arbitrary page JavaScript. These are essential behavioral traits not otherwise communicated, making the description highly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with zero wasted words. The primary purpose is first, followed by a concrete use case and a critical warning. Every sentence adds distinct value, and the structure front-loads the most important information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that executes arbitrary JavaScript, the description covers the key expectations: return format (serializability and null behavior), the use case, and the safety warning. Since there is no output schema, the description fully explains what the agent will receive. Combined with the complete parameter schema, nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters in detail, including the role of `arg` and functions vs. plain expressions. The description adds only the serializability requirement context, which is useful but not necessary for parameter understanding. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (evaluates), the resource (JavaScript expression or function in the session page), and the mechanism (Playwright page.evaluate). It also distinguishes from the sibling get_page_snapshot by explicitly noting when to use one over the other, so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance: 'Use it to inspect the live DOM (computed styles, geometry, custom element state) when get_page_snapshot is not enough.' This names the alternative and the condition that selects it. It also warns about arbitrary JavaScript execution, which is a critical usage consideration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audit_guideGet audit guideA

Returns the step-by-step accessibility audit workflow for agents as markdown: how to enumerate URLs from a sitemap, run the automated axe-core pass, the structure/tab-order/images/contrast reviews, the output-report format, and exactly which checks (keyboard navigation, screen-reader compatibility, image meaning) must be done by a human. Call this before starting an audit. Needs no session or browser. Also published at https://mya11y.report/mcp/auditor-skill.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explains the return format (markdown), the content of the returned guide, and the lack of session/browser prerequisites. This gives the agent a solid mental model of a side-effect-free information retrieval operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core behavior, uses a compact colon-list to convey rich content detail, and includes only two short supporting sentences about when to call and prerequisites. Every sentence adds distinct value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only guide-retrieval tool, this description is complete: it states what is returned, what topics are covered, when to call it, and what is not needed. The sibling list further confirms the tool's role relative to audit execution and browser-control tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is empty, so there is nothing for the description to clarify beyond confirming no inputs are needed. The description adds context by noting that no session or browser is required, which is more valuable than restating parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it 'returns the step-by-step accessibility audit workflow for agents as markdown.' It clearly distinguishes itself from execution tools like run_a11y_audit by framing itself as the preparatory guide rather than the audit itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear when-to-use guidance: 'Call this before starting an audit.' It also specifies that 'Needs no session or browser,' which tells the agent no setup is required. It does not explicitly name alternatives or when not to use it, but the context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_page_snapshotGet page snapshotA

Returns a YAML accessibility-tree snapshot of the session page (roles, names, states). Use it to discover roles/names to target with getByRole/getByText/getByLabel in click/type/check.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It transparently states that the tool returns a YAML snapshot with roles, names, and states, implying a read-only operation. It does not discuss edge cases like stale snapshots or errors, but the core behavioral contract is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence is front-loaded with what the tool returns and its format; the second explains the practical use case. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with no output schema, the description is complete: it specifies return format, content, and intended usage. The agent has enough information to select and invoke the tool correctly without needing additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the single sessionId parameter fully documented in the input schema. The description adds no additional parameter semantics, but none are needed because the schema already explains the parameter clearly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Returns'), a specific resource ('YAML accessibility-tree snapshot of the session page'), and the content included ('roles, names, states'). This clearly distinguishes it from sibling tools like get_structure by focusing on the accessibility tree rather than generic page structure.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: use it to discover roles/names to target with getByRole/getByText/getByLabel in click/type/check operations. It does not explicitly name alternatives or exclusions, but the intended workflow is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_scroll_positionGet scroll positionA

Returns the current scroll position of the session page plus document scroll dimensions and the viewport size — useful for planning further scroll_by deltas.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses that the tool returns scroll position, document scroll dimensions, and viewport size, which is helpful. However, it doesn't mention whether this is a read-only operation, any side effects, or what the exact return shape is. Since it's a getter, the read-only nature is implied but not explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the main purpose, and the use-case hint is appended. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with one parameter and no output schema, the description is mostly complete. It could mention that the return values are in pixels or that it's read-only, but the core information an agent needs to call it correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the sessionId parameter. The description doesn't add any parameter-specific meaning beyond what the schema provides, which is fine given the high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Returns') and resource ('current scroll position of the session page') and adds the extra dimensions and viewport size. It clearly distinguishes itself from get_viewport_size and scroll_by, which are siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says it is 'useful for planning further scroll_by deltas', which gives clear context for when to use it. It doesn't explicitly mention when not to use it or name alternatives like get_viewport_size, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_structureGet page structureA

Reads the page structure: landmark regions, headings, lists (nested) and iframes, each with a shadow-piercing selector. A heading outline grouped by region with skipped-level and repeated-H1 flags is included when both regions and headings are returned. Top frame only, open shadow roots only.

ParametersJSON Schema
NameRequiredDescriptionDefault
includeNoOptional subset of categories to return: 'regions' | 'headings' | 'lists' | 'frames'. Defaults to all four. Counts always cover every category.
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden, and it does well: it discloses shadow-piercing selector behavior, the conditional nature of the heading outline, and the top-frame/open-shadow-roots limitations. It does not describe the exact response shape or error behavior, but the key behavioral constraints are stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly scoped sentences: the first defines the resource and output elements, the second explains a conditional output detail, and the third states limitations. Every sentence earns its place with no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficient for a simple read-only structure tool, especially given the 100% schema coverage for parameters. It explains the main content, selector behavior, conditional outline, and scope limits. The main gap is the lack of explicit return shape, but the enumerated output categories largely compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies. The description aligns with the include categories (regions, headings, lists, frames) but adds no parameter-level detail beyond what the schema already provides. It does not need to duplicate the schema's defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Reads the page structure') and enumerates exactly what is included: landmark regions, headings, nested lists, iframes, and shadow-piercing selectors. This is clear and non-generic, but it does not explicitly contrast the tool with siblings like get_page_snapshot, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives meaningful context: 'Top frame only, open shadow roots only' tells an agent when the tool's scope applies. However, it gives no explicit when-to-use or when-not-to-use guidance against sibling tools such as get_page_snapshot or run_a11y_audit, so alternatives are left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tab_orderGet tab orderA

Lists the page's tab stops in the order the Tab key reaches them, using the tabbable algorithm: positive tabindex values first (ascending), then everything else in document order. Each stop reports its position, tag, label, effective tabindex, shadow-piercing selector and an outOfOrder flag when tabindex is positive. Top frame only; hidden/disabled/inert elements are excluded.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it details ordering rules, output fields, the outOfOrder flag, top-frame-only scope, and exclusions for hidden/disabled/inert elements. This goes well beyond a minimal 'gets tab order' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, each earning its place: the first gives the core purpose, the second details algorithm and output fields, the third covers scope and exclusions. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema or annotations, the description is complete enough for correct invocation: it explains what is returned, the ordering semantics, frame scope, and excluded elements. An agent has what it needs to call and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, sessionId, is already fully described in the schema, including why it is required. The description adds no parameter-specific information, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Lists the page's tab stops in the order the Tab key reaches them.' It also clarifies the algorithm and scope, making it clearly distinct from generic structure or audit tools among its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the description—inspecting keyboard tab order for accessibility—but it never explicitly says when to use this tool versus an alternative. No sibling routing or exclusion conditions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_viewport_sizeGet viewport sizeA

Returns the session page's viewport size in CSS pixels.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It states the output is in CSS pixels, which is useful, but it does not disclose whether the viewport size is for the current page or the session's default viewport, nor does it mention any side effects or limitations. The description is accurate but minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that is appropriately sized and front-loaded with the key information. Every word earns its place, and there is no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one parameter and no output schema, the description is nearly complete. It states the return value and unit, which is sufficient for an agent to call it correctly. It could optionally mention that the viewport size is for the current page, but this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the sessionId parameter fully. The description adds no additional parameter semantics beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the session page's viewport size in CSS pixels, with a specific verb ('Returns') and resource ('session page's viewport size'). It is distinct from siblings like get_scroll_position and get_page_snapshot, so an agent can understand its purpose without confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when an agent needs viewport dimensions, but it does not explicitly state when to use this tool versus alternatives like get_scroll_position or get_page_snapshot. There is no mention of when not to use it or any prerequisites beyond the sessionId parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_imagesList images and alt textA

Lists every and inline on the session page with its accessible name, the source that name came from (aria-label, aria-labelledby, alt, title, desc or none), its src/currentSrc and sanitized SVG markup, a decorative flag, and a shadow-piercing selector. Derived flags mark genuinely missing alt text and overly long alt text (>250 chars). Top frame only, open shadow roots only.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoWhich images to return: 'all' (default), 'decorative', 'missing-alt' (not decorative and no accessible name), or 'has-alt' (not decorative and named). Counts always cover every image.all
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure and does so thoroughly: it describes output fields, sanitization of SVG markup, derived flags for missing/overly long alt text, and scoping to top frame and open shadow roots. The agent knows what to expect without needing additional annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence with no filler. It front-loads the core purpose and then packs only value-adding details: output fields, flags, and scope limitations. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description fully covers what the agent will receive, including specific fields and flags. Combined with the schema's complete parameter documentation, this is sufficient for correct invocation and interpretation of results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds a bit of behavioral context (e.g., 'Counts always cover every image'), but it does not meaningfully expand on parameter meaning beyond the enum descriptions already present.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Lists') and a clear resource ('every <img> and inline <svg> on the session page'), and expands on exactly what is returned, including accessible name, source, and derived flags. This clearly separates it from general page-structure tools like get_structure or get_page_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes its scope explicit ('Top frame only, open shadow roots only') and the filter parameter clarifies result subsets. It does not explicitly name alternative tools or state when not to use it, but the purpose and limitations are clear enough for an agent to select it appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_keyPress keyA

Presses a key or chord (Enter, Tab, Escape, arrows, Space, Shift+Tab, Control+A, ...) in the session page. With a target, focuses that element and sends the key to it; without one, sends the key to the currently focused element. Use it for keyboard navigation, submitting a form with Enter, closing a dialog with Escape, or exercising keyboard-only interactions.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesKey or chord to press, using Playwright key names: "Enter", "Tab", "Escape", "ArrowDown", "Space", "Shift+Tab", "Control+A", "Meta+Enter", "F1", "Backspace", "PageDown". To type a literal character use "type" instead.
targetNoOptional element to focus first and send the key to (Playwright locator.press). Omit to send the key to the page's currently focused element (page.keyboard.press).
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
timeoutMsNoAction timeout in milliseconds when a target is given (default 10000).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It clearly explains the two operating modes (with target vs. without) and that it sends the key to the focused element. It doesn't mention edge cases like waiting or navigation side effects, but for a simple key press this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two well-structured sentences, front-loaded with the primary action and followed by mode behavior and use cases. No wasted words, and the key information is immediately visible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main behavior, both modes, and typical use cases. It does not discuss return values (no output schema exists, and it's an action tool), which is acceptable. The required sessionId is in the schema, so completeness is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all parameters in detail. The description's explanation of target behavior largely mirrors the schema's target description, adding no new semantic value. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it presses a key or chord in the session page, listing specific examples and explicitly distinguishing from the 'type' tool for literal characters. This makes the purpose unambiguous and separates it from the sibling 'type' tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit use cases (keyboard navigation, submitting forms, closing dialogs) and directly says to use 'type' for literal characters, giving a clear alternative. This is strong guidance on when and when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_a11y_auditRun accessibility auditA

Runs an axe-core accessibility audit on the session page (or on url if provided, navigating first). Returns the session audit history: pages (hex ids, one per URL) and audits ({pageId, audit} entries) including the fresh result. Each audit contains violation counts by impact, per-violation details (rule id, help, selectors, failure summaries), and pass/incomplete rule counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoOptional URL to navigate to before auditing. Omit to audit the current session page.
tagsNoOptional axe-core rule tag filters, e.g. ["wcag2a", "wcag2aa"]. Defaults to all rules.
detailNo'summary' returns violation details plus rule counts (default). 'full' additionally embeds the complete raw axe-core JSON report inside each audit entry.summary
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
timeoutMsNoPage-load timeout in milliseconds when url is given (default 30000).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It transparently states that providing a url causes navigation first, and that the tool returns the full session audit history including the fresh result. It also details the audit payload structure, covering violation counts, per-violation details, and pass/incomplete counts. This gives an agent a clear model of what happens and what to expect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long and every sentence contributes: the first states the action and target, the second explains the return shape, and the third specifies the audit result contents. It is front-loaded with the core behavior and contains no filler or redundant restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description sufficiently explains the return structure: pages as hex ids and audits as pageId/audit entries including the fresh result. It also covers the optional URL navigation behavior and the level of detail in each audit. For a tool with five parameters and no annotations, this description provides enough context for an agent to invoke it correctly and interpret its output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter, including url, tags, detail, sessionId, and timeoutMs. The description adds context about the audit result and the URL navigation behavior, but it does not meaningfully enhance parameter-level understanding beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific action and resource: running an axe-core accessibility audit on the session page or a provided URL. This clearly differentiates it from sibling tools like check_color_contrast, which targets a single contrast check, and get_audit_guide, which is guidance-oriented. The verb 'runs' plus the resource 'axe-core accessibility audit' makes the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: audit the current session page, or provide a url to navigate first. However, it does not explicitly discuss when to prefer this tool over alternatives such as check_color_contrast or get_audit_guide, nor does it state exclusions. The usage context is useful but the comparison to sibling tools is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotTake screenshotA

Captures the current viewport of the session page (what is on screen after scrolling — not the full page). Saved as .png inside the session log directory; returns the file path plus the scroll position and viewport at capture time.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses key behaviors: viewport-only capture (not full page), saving as <actionId>.png in the session log directory, and returning file path plus scroll position and viewport. This is substantial, though it omits details like error handling or whether the operation is synchronous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundancy. The core action is front-loaded, and the details on saving and return are concise and relevant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description adequately explains the return value (file path, scroll position, viewport) and the file location. It could mention potential failure modes or how to access the file, but for a simple tool with one parameter, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers the only parameter (sessionId) with a clear description, so schema coverage is 100%. The tool description adds no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('captures') and resource ('current viewport of the session page'), and explicitly distinguishes from full-page capture. It also details the output (saved .png, file path, scroll position, viewport), which differentiates it from sibling tools like get_page_snapshot or get_viewport_size.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for visual capture but does not explicitly say when to use this tool versus alternatives like get_page_snapshot or get_viewport_size. There are no conditions, exclusions, or references to sibling tools, leaving usage guidance to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_byScroll pageA

Scrolls the session page window by the given deltas (window.scrollBy) and returns the new scroll position. For lazy-loaded content, scroll then re-run get_page_snapshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoHorizontal scroll delta in pixels (negative scrolls left). Default 0.
yNoVertical scroll delta in pixels (negative scrolls up). Default 0.
behaviorNoScroll behavior (default 'auto').auto
sessionIdYesSession id returned by start_session. Required on every stateful tool call.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does a decent job: it reveals the operation is a relative window.scrollBy, that it mutates scroll position, and that it returns the new position. It also flags the lazy-loading caveat, though it does not describe return shape or async behavior of smooth scrolling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The core action and result are front-loaded, and the lazy-loading tip is the only additional sentence, earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple scroll tool without an output schema, the description covers the action, result, and a practical follow-up (get_page_snapshot). It falls slightly short only because the return value's exact shape is unspecified and there is no mention of edge cases like coordinates outside viewport boundaries or waiting for smooth scroll to settle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents x, y, behavior, and sessionId. The description adds value by framing x/y as deltas and tying the operation to window.scrollBy, making the relative semantics explicit beyond the individual parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation ('Scrolls the session page window'), the input mechanism ('by the given deltas'), and the result ('returns the new scroll position'). This makes it clearly distinct from read-only siblings like get_scroll_position or get_viewport_size.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides one concrete usage pattern: for lazy-loaded content, scroll then re-run get_page_snapshot. However, it does not explicitly state when to choose this over alternatives (e.g., get_scroll_position or navigate), nor does it state exclusions or prerequisites beyond the schema-covered sessionId.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_optionSelect optionA

Selects an inside a element on the session page, matched by the option's value attribute. Waits for the element to be actionable. Targets support selectors or getBy* locators, same as type.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYesThe option to select, matched against the <option>'s value attribute.
targetYesHow to locate the element: {"selector": "#id"} or {"fn": {"name": "getByRole", "args": ["button", {"name": "Sign in"}]}}.
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
timeoutMsNoAction timeout in milliseconds (default 10000).

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It adds useful context by stating that the tool 'waits for the element to be actionable' and that matching is by the value attribute. However, it does not disclose behavior on missing options, whether change/input events fire, or what the tool returns, so it is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loads the primary action, and contains no filler. Every sentence adds meaningful information: what is selected, how matching works, waiting behavior, and target locator compatibility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a relatively simple browser action with a rich schema, the description covers the essential operational details: target element type, matching mechanism, waiting behavior, and locator options. It does not explain return values or error cases, but with no output schema and a straightforward action, the provided context is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by clarifying that 'value' is matched against the option's value attribute and that 'target' supports the same selector/getBy* locator forms as 'type'. This extra semantic guidance justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Selects an <option> inside a <select> element'. It also clarifies the matching mechanism ('matched by the option's value attribute'), which clearly distinguishes this tool from siblings like click, type, check, and uncheck.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use this tool: whenever an option inside a select element needs to be chosen. It also provides locator guidance ('same as type'), which helps the agent reuse known patterns, though it does not explicitly state when not to use alternatives like click.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_sessionStart sessionA

Starts a browser session and returns its sessionId — pass that sessionId to every subsequent tool call (navigate, click, type, run_a11y_audit, ...). Each session is an isolated page with its own cookies, audit history, and action log. Sessions auto-close after 5 minutes of inactivity (MYA11Y_MCP_IDLE_CLOSE_MS overrides).

ParametersJSON Schema
NameRequiredDescriptionDefault
aliasNoFree-form session alias (max 225 chars). Becomes the sessionId after sanitizing to [A-Za-z0-9_-]; omit for a generated hex id. Starting an id that already exists returns it.
headlessNoRun the browser headless (default true) or headed (false, visible window).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it delivers: it discloses that each session is isolated, has its own cookies/audit history/action log, auto-closes after 5 minutes, and honors an env override. This goes well beyond what the schema alone conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tight sentences, front-loading the core purpose and return value, then adding isolation and lifecycle behavior. Every sentence contributes useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a session-starting tool with only two optional parameters and no output schema, the description is complete: it explains what is returned, how to use it, session isolation semantics, and auto-close behavior. An agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The main description adds no new parameter-level detail, but the schema already thoroughly explains alias sanitization, existing-id behavior, and the headless default. The description's mention of passing sessionId concerns the return value rather than the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Starts a browser session') and immediately states the key output (sessionId) and its role for subsequent calls. It clearly differentiates this tool from the many sibling action tools by establishing it as the session initiator.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs the agent to pass the returned sessionId to every subsequent tool call and lists representative siblings (navigate, click, type, run_a11y_audit). It also explains session isolation and idle auto-close, giving clear context for when starting a session is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

typeType textA

Types text into an element (input, textarea, contenteditable). By default replaces the current value; set clear:false to append. Targets support selectors or getBy* locators.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to type into the element.
clearNotrue (default) replaces the current value (fill). false appends keystroke-by-keystroke (pressSequentially).
targetYesHow to locate the element: {"selector": "#id"} or {"fn": {"name": "getByRole", "args": ["button", {"name": "Sign in"}]}}.
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
timeoutMsNoAction timeout in milliseconds (default 10000).
pressEnterNoPress Enter after typing (e.g. to submit).

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It covers the default fill vs. append behavior and target locator support, but it does not mention waiting, error conditions, or side effects like form submission when pressEnter is used, leaving a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences carry exactly the required information: what the tool does, its default vs. append behavior, and how targets can be specified. There is no repetition or filler, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the schema covers all parameters and the tool is conceptually simple, the description provides enough context for the core invocation. It could mention return behavior or edge cases, but for a text-entry action the current level is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter. The description mainly restates the clear behavior and locator options that are already present in the schema, adding little value beyond the structured definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Types text into an element (input, textarea, contenteditable)'. It immediately distinguishes itself from sibling tools like click, check, press_key, and navigate, so an agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the two key invocation modes: default replaces the current value, while clear:false appends keystroke-by-keystroke. It doesn't explicitly name alternatives or exclusions, but the element types and text-entry behavior make the intended use case reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

uncheckUncheck checkboxB

Unchecks a checkbox in the session page. No-op if already unchecked.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesHow to locate the element: {"selector": "#id"} or {"fn": {"name": "getByRole", "args": ["button", {"name": "Sign in"}]}}.
sessionIdYesSession id returned by start_session. Required on every stateful tool call.
timeoutMsNoAction timeout in milliseconds (default 10000).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool is a no-op if already unchecked, which is useful behavior. However, it doesn't state any side effects like waiting for the action to complete, what happens if the element is not found, or if it triggers any events like 'change'. For a mutation tool, this is some but not full disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that states the action and the no-op behavior. It is front-loaded with the core purpose and adds the no-op detail efficiently. No extra words, though it could incorporate more usage guidance, but as it stands it's concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the 'target' parameter with its anyOf structures, the schema handles that well. The tool's behavior is simple, and the description covers the main purpose and no-op case. However, for a mutation tool with no annotations, it might be expected to mention what happens on failure or whether it waits for navigation, but overall it's adequate for an agent to call it, especially considering the schema richness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage, so all parameters are documented in the schema. The description adds nothing about parameters. The 'target' parameter description in the schema is comprehensive, including examples, so the baseline is 3. The description doesn't need to add more, but it also doesn't exceed the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: it unchecks a checkbox in the session page, and also specifies it's a no-op if already unchecked. This differentiates it from its sibling 'check', which does the opposite. However, it doesn't explicitly mention that it is a mutation, but the verb 'unchecks' implies that clearly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context: it acts on the session page and is a no-op if already unchecked, which hints at when to use it (when a checkbox should be unchecked). However, it doesn't explicitly mention its opposite 'check' as the alternative for checking, nor does it provide explicit conditions for when to prefer one over the other. The sibling name is visible, so the agent can infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 21 tool updatesv1.0.3
    • First observedcheck
    • First observedcheck_color_contrast
    • First observedclick
    • First observedclose_session
    • First observedevaluate
    • First observedget_audit_guide
    • First observedget_page_snapshot
    • First observedget_scroll_position
    • First observedget_structure
    • First observedget_tab_order
    • First observedget_viewport_size
    • First observedlist_images
    • First observednavigate
    • First observedpress_key
    • First observedrun_a11y_audit
    • First observedscreenshot
    • First observedscroll_by
    • First observedselect_option
    • First observedstart_session
    • First observedtype
    • First observeduncheck

TDQS

A3.8/5.0

Scored across 21 tools

Disambiguation4/5

The tools split into clear lifecycle, inspection, and interaction groups, and each inspection tool targets a distinct artifact (a11y tree, landmarks, tab stops, images). A few minor overlaps exist, such as get_viewport_size and get_scroll_position both returning viewport dimensions, and get_page_snapshot versus get_structure both describing page structure, but descriptions clarify the intended use.

Naming Consistency4/5

Most tools follow a verb_noun snake_case pattern (start_session, get_tab_order, run_a11y_audit). The interaction tools deviate as bare verbs (navigate, click, type, check, uncheck, evaluate, screenshot) and scroll_by uses a preposition, so the pattern is not uniform, but the categories are still readable and predictable.

Tool Count3/5

21 tools is on the heavy side for a single server, landing in the 16-25 range that feels bloated at first glance. That said, the count is justified by the combination of browser automation, page inspection, session management, and audit functionality; it is broad but not redundant.

Completeness4/5

The surface covers the full accessibility audit loop: session lifecycle, navigation, automated axe-core audits, contrast checks, structural inspection, and keyboard/form interactions for manual checks. Minor gaps remain, such as no explicit viewport resizing, hover action, or full-page screenshot option, but agents can work around them with evaluate or scrolling.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers