Skip to main content
Glama
helmif

semantic-dom-mcp

by helmif

semantic-dom-mcp

Local MCP server (stdio, Node.js + TypeScript) that drives a real Chromium browser via Playwright and gives a coding agent what it needs to write a Playwright test it did not have to guess: verified locators (counted by Playwright's own engine, scoped the way a QA engineer scopes them), structured assertion data (tables as rows and cells, dialogs as label/value pairs), the behavior the page showed (navigations, requests, console errors), and diffs between steps of a flow. Same page → same output → same conventions → same test style across the team.

How it keeps tokens low. An agent starts with an outline of the page (regions, tables, dialogs, alerts; a few hundred to a few thousand characters, smaller than the native accessibility tree), then extracts one region, and acts with the snapshot in the same call. Measured on a real seller dashboard, the outline was 82–95% smaller than Playwright's aria snapshot of the same page, and a whole authenticated add-to-cart flow ran in about 10k tokens. A full-page extraction is still 2–3x the size of the aria snapshot, because it carries what that tree does not: executable locators, uniqueness verdicts and state. Ask for the whole page only when you need it. Numbers: benchmark/RESULTS.md, FLOW-VALIDATION.md. Docs: How it works (deep dive) · Team guide (setup + connecting your agent) · Benchmark methodology · v0.5 flow validation · Roadmap · Changelog

Quickstart

No clone, no build — the package is on npm. One-time browser setup (installs the Chromium build matching the package's bundled Playwright):

npx -y -p semantic-dom-mcp playwright install chromium

Then add the server to your MCP client:

{
  "mcpServers": {
    "semantic-dom": {
      "command": "npx",
      "args": ["-y", "semantic-dom-mcp"],
      "env": {
        "QA_MCP_ALLOWED_HOSTS": "staging.yourapp.internal,staging.admin.internal",
        "QA_MCP_STORAGE_STATE": "./.auth/staging.json"
      }
    }
  }
}

That's the whole setup. Verify by asking your agent to list its MCP tools — you should see extract_semantic_dom. See docs/GUIDE.md for per-client config locations (Claude Code, Claude Desktop, Cursor, Windsurf), authenticated staging, and troubleshooting. To run from a clone instead (contributors), see Development below.

Related MCP server: Browser MCP Server

Workflow

Single page. Ask your agent: "extract the checkout page and write a success-path test."

  1. The agent calls extract_semantic_dom({ url }). The server navigates a real Chromium page, runs the extractor inside it, and returns Semantic JSON: every interactive node with a ready-to-paste Playwright locator, uniqueness verified by Playwright's own engine.

  2. The agent uses the write_playwright_test prompt (scenario + the JSON), which injects the team conventions.

  3. The result is a Playwright test in team style, grounded in real locators, never guessed ones.

Multi-step flow (v0.8 way). Ask: "write a test for login → add to cart → checkout."

  1. session_open({ url }) opens a persistent page, then session_extract({ mode: "outline" }) maps it: regions with a selector, tables with row identity and cells, dialogs, alerts.

  2. session_extract({ scope, roles, visible_only }) pulls only the region the step needs, with locators verified unique inside it.

  3. session_act({ actions, then_extract: { mode: "diff", scope } }) runs declared actions (paste any returned playwright expression as the locator) and returns, in the same call, what the page did (observed: redirects, requests, console errors) and what changed (added nodes, removed nodes, state transitions). That is the wait list and the assertion list for the step.

  4. Repeat per step, then session_verify_locators with the expressions in the written spec, then session_close.

Every fact a test needs (locator, redirect, API path, state transition, cell value) comes from the page, not from the agent's memory.

MCP surface

Kind

Name

Purpose

Tool

extract_outline

The page as a map: regions (each with a selector for scope), structured tables, open dialogs, alerts. Start here.

Tool

extract_semantic_dom

Extract a URL into Semantic JSON (url, wait_for default auto, wait_selector, scope, roles, visible_only, max_output_chars, include_tables, include_hidden, max_nodes, viewport, include_click_targets). Read-only, never touches the page.

Tool

verify_locators

Count every Playwright expression from a written spec against the live page; matches, uniqueness, first match, summary.

Tool

get_conventions

The team conventions as text, for clients that hide MCP prompts.

Tool

extract_semantic_dom_after

Same, but first runs a short declared action list (fill/click/press/select/goto/wait, max 20) in the main frame and snapshots the resulting state, plus an observed block of what the page did meanwhile. Refuses to extract if the actions navigated off the allowlist.

Tool

session_open

Open a persistent page for a multi-step flow (url, wait_for, wait_selector, viewport). Returns a session_id. Sessions are capped and expire when idle.

Tool

session_act

Run declared actions in an open session; then_extract returns the diff, a scoped extraction or an outline in the same call. Action locators take a returned playwright expression verbatim. Returns the resulting URL/title and observed: main-frame navigations, xhr/fetch requests (method, path, status), console errors, dialogs (dismissed), popups (closed).

Tool

session_extract

Snapshot the session's current state (snapshot_id included), or mode: "outline". Takes scope/roles/visible_only/max_output_chars/include_tables. With diff_against (a snapshot id or "previous") returns a diff: added, removed, changed nodes and the behavior observed in between.

Tool

session_verify_locators

verify_locators against the session's current page.

Tool

session_close

Release the session's browser context.

Tool

session_list

Diagnostic: open sessions with URL, expiry and counts.

Tool

check_auth

Diagnostic: navigates with the configured storageState and reports whether the session bounced to a login-looking page (expired auth shows up as an answer, not a mystery).

Tool

list_frames

Diagnostic frame tree with same-origin/reachability classification.

Prompt

write_playwright_test

Team-standard test-writing prompt (scenario, extract_json, team_name?, framework_note?).

Resource

conventions://playwright

The same team conventions as read-only text.

Errors (navigation failure, denied host, missing selector, expired session) come back as structured JSON in the tool result, so the agent can react instead of crashing. Every tool carries MCP annotations (readOnlyHint, openWorldHint: false) so clients can auto-approve the read-only ones.

Configuration (environment variables)

Variable

Meaning

QA_MCP_ALLOWED_HOSTS

Required. Comma-separated hostnames the server may navigate to. Navigation is denied by default. Supports host, host:port, and *.domain entries.

QA_MCP_STORAGE_STATE

Optional path to a Playwright storageState JSON for pre-authenticated staging sessions. This file holds a live session — it is gitignored; never commit it.

QA_MCP_TEAM_NAME

Optional team name used in the write_playwright_test prompt (default QA).

QA_MCP_SESSION_TTL_MS

Idle time before a session is closed automatically (default 600000, 10 minutes).

QA_MCP_MAX_SESSIONS

Max concurrently open sessions (default 3).

Security posture

  • Tool inputs are untrusted (they arrive via an LLM): strict schemas (additionalProperties: false), http/https only, host allowlist enforced before any navigation.

  • extract_semantic_dom only reads the DOM. It never clicks, submits, or mutates the page. The sanctioned exceptions are extract_semantic_dom_after and session_act, which execute only an explicit, bounded, schema-validated action list, never log fill values, and refuse to extract if the page leaves the allowlisted hosts (a session in that state is closed).

  • Sessions add their own limits: an idle TTL, a cap on open sessions, one in-flight call per session, and a bounded snapshot history kept in memory only.

  • Behavior capture is observation only. The server never issues requests of its own. Request and response bodies are never read, query strings are stripped from recorded URLs (they may carry tokens), console text is capped, dialogs are dismissed and popups closed at once.

  • Sessions check the allowlist before any action, after every action, after a failed action and before every snapshot. A session found off the allowlist is closed on the spot.

  • No network egress beyond navigating the browser to allowlisted URLs. No telemetry. Page contents are never logged (stderr carries only high-level events) and are not stored beyond the current call or session.

Semantics worth knowing

  • Snapshot honesty: the JSON is a single moment. A disabled submit button is reported is_disabled: true with a note. The conventions instruct the model to write the interactions that change state, not to assume it stays disabled. In a session, the diff shows the transition itself (is_disabled: false → true), so the test asserts a fact rather than an assumption.

  • Scoped locators (schema 1.4): a nameless or repeated control is located inside its container the way a QA engineer writes it by hand: getByRole('row', { name: 'Charizard' }).getByRole('radio'), getByRole('listitem').filter({ hasText: 'Puthera' }).getByRole('button', { name: 'Pilih Pembeli Ini' }), getByTestId('customer-checkbox-flex').getByRole('checkbox'). The node carries within so an action can target the same element; the expression is verified unique by Playwright like every other locator.

  • Compact wire format (schema 1.3): results are compact JSON and a node field that carries no information is omitted: null fields, frame_path: [], in_shadow: false, kind: "element", empty fallback_locators, and text_content equal to accessible_name. An absent property is null (not applicable, never false); absent structure means the default (main document, light DOM, nothing worth listing). is_visible is always present. Fallbacks appear only when the primary is ambiguous or brittle. Same facts, about a third of the tokens.

  • Assertable state (since schema 1.2): every node reports value (never for password fields), aria_expanded, aria_selected, aria_invalid, described_by (the text of the elements aria-describedby points at, where validation messages live), validation_message (browser constraint validation) and, for <select>, options. Absent state is null, never false.

  • Observed behavior: observed.requests lists xhr/fetch/document requests as method + path + status; static assets are counted in dropped, not listed. observed.navigations lists main-frame URL changes in order. Neither is ever guessed; if a redirect or API path is not in observed, the conventions tell the agent not to wait on it.

  • Diff identity: nodes pair across snapshots by the most stable fact available: test-id, else id, else placeholder, else tag + role + accessible name, plus frame path and a document-order index for non-unique nodes. A hidden menu item that becomes visible is therefore a changed entry, with its primary locator switch (getByText → getByRole) listed as one of the changes. A renamed node with no stable attribute shows as removed + added; the diff says so in its notes.

  • Visibility is Playwright's: is_visible predicts toBeVisible(), so it uses Playwright's rule (not display:none, visibility not hidden, width and height both > 0). Opacity and aria-hidden do not hide an element for Playwright and do not here either.

  • Secrets never leave the server: a value typed into a password field, or a fill marked secret: true, is scrubbed from every string in every result (values, text, console, dialogs, errors) when it is 8+ characters. The set is process-wide and lasts until the server exits, so mark only real secrets: a redacted string blanks that text everywhere, and a locator whose text was redacted is flagged not unique. Password and credential-autocomplete fields never report a value at all.

  • Hidden nodes are included and flagged is_visible: false (tests often assert hidden-ness); pass include_hidden: false to drop them (the count dropped is noted, never silent).

  • Open shadow DOM is traversed and flagged in_shadow — locators pierce it natively, so no >>>/::shadow CSS is ever emitted. Closed shadow roots appear as shadow_boundary marker nodes (detected via pre-navigation attachShadow instrumentation; closed roots created by declarative shadow DOM parse before scripts run and cannot be detected).

  • Same-origin iframes are extracted per-frame with frame_path set (chain frameLocator() in that order). Cross-origin iframes are recorded as opaque cross_origin_frame nodes with URL/name only — their DOM is never touched.

  • Notification & dialog surfaces (role="alert", role="status", dialogs) are extracted like interactive nodes. When a toast library keeps the live region empty and renders the message in a sibling (a common pattern across UI libraries), the message text is pulled from the enclosing container and flagged. For UI that renders late after an interaction, wait_selector_after on extract_semantic_dom_after waits deterministically instead of guessing settle_ms. Since those ARIA roles take names from the author (not contents), their role locator is getByRole('alert') — or with the aria-label name when one exists. For UI that only appears after an interaction (login-success toast), use extract_semantic_dom_after.

  • JS-click cards (product tiles with no anchor/role/test-id) are invisible to the factual rules by design — pass include_click_targets: true to include cursor-pointer boundary elements with content, flagged as heuristic and located by their heading text.

  • Links carry href (schema 1.1) so agents can discover which page to extract next without scraping. Framework-generated ids (rc_select_*, React useId, Radix, MUI...) are detected and demoted to last-resort with a note — they change between builds and must never be primary.

  • viewport: "mobile" (375×812, touch) snapshots responsive states; visibility flags reflect the active media queries.

  • Truncation is loud: max_nodes / depth caps set truncated: true plus a note. Non-unique locators carry is_unique: false and disambiguation guidance.

Development

git clone https://github.com/helmif/semantic-dom-mcp.git && cd semantic-dom-mcp
npm install
npx playwright install chromium
npm run dev        # run the server over stdio via tsx
npm run typecheck  # tsc --noEmit (strict)
npm test           # Vitest suites against real fixture pages in headless Chromium
npm run build      # compile to dist/ (clients can then use "command": "node", "args": ["<path>/dist/index.js"])

Repo layout: src/index.ts (bootstrap) · src/server.ts (MCP surface) · src/browser.ts (Playwright layer + single-shot orchestration) · src/session.ts (persistent sessions) · src/observe.ts (behavior capture) · src/diff.ts (snapshot diff) · src/extractor/ (in-page engine + locator resolution) · src/types.ts (frozen contract, schema 1.3) · src/compact.ts (wire rules) · src/conventions.ts (single source of team conventions).

Available Tools

13 tools
check_authA
Read-onlyIdempotent

Diagnostic: navigates with the configured QA_MCP_STORAGE_STATE session and reports whether the page bounced to a login-looking path (session likely expired). Use when extractions unexpectedly return login forms instead of the requested page.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page whose frame tree to report. Must be http/https and allowlisted.
wait_forNoNavigation wait; see extract_semantic_dom.auto

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive, so the bar is lower. The description adds genuine context beyond annotations: it discloses the authentication mechanism (QA_MCP_STORAGE_STATE session) and the failure semantics (bounce-to-login indicates expired session). It does not describe what the report format looks like, but that is partially an output-schema concern.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the diagnostic purpose then the trigger condition. Every clause earns its place — no filler, no restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param read-only diagnostic with full schema coverage and annotations carrying the safety profile, the description supplies everything an agent needs: what it does, how it authenticates, and when to reach for it. No output schema exists, but the description adequately frames the result (whether the page bounced to login).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% — both parameters are fully described in the schema (url with allowlist/http-https constraint, wait_for enum with default). The description adds no parameter detail beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific diagnostic verb ('navigates with the configured QA_MCP_STORAGE_STATE session and reports whether the page bounced to a login-looking path') and clearly names what it distinguishes from siblings — it's not an extraction tool but a session-health probe. An agent can tell it apart from extract_semantic_dom or session_extract without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use: 'Use when extractions unexpectedly return login forms instead of the requested page.' This gives a concrete trigger condition that routes the agent here rather than to a re-extraction. No alternative tool is named, but the trigger is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_outlineA
Read-onlyIdempotent

The page as a MAP, a few thousand characters: landmark regions (header/nav/main/forms/tables/lists/dialogs) each with a selector you can pass as scope to extract_semantic_dom / session_extract, structured tables (headers, row identity for getByRole('row', { name }), cells) and open dialogs (label/value fields), plus current alert text. Start here on any page you have not seen; then extract one region.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to extract. Must be http/https and on an allowlisted host.
scopeNoCSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped.
viewportNoViewport preset — 'mobile' is 375x812 with touch, for responsive states.desktop
wait_forNoNavigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages.auto
wait_selectorNoOptional selector to await before extracting (for SPA content).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: the output is bounded ('a few thousand characters'), which matters for context budgeting, and the region selectors are reusable downstream. It does not cover failure modes (e.g. non-allowlisted hosts, timeouts), which the schema touches on.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the output payload and ending on the usage directive. The first sentence is densely packed with parentheticals, but every clause carries distinct information (region types, selector handoff, table/dialog/alert contents), so nothing is wasted. Slightly heavy to parse in one pass.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the return-value burden and describes the map contents well. It also explains the intended next step. It omits any mention of output size limits, truncation behavior, or error conditions for a tool that fetches an arbitrary URL, which is a minor gap for a read tool with full annotation coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by linking output to input: it explains that the `selector` in each outline region is what you pass as `scope`, which is a relationship the schema alone does not convey. It adds no new detail on viewport/wait_for, hence not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool returns — a page 'MAP' of landmark regions, structured tables, open dialogs, and alert text — and enumerates the region types (header/nav/main/forms/tables/lists/dialogs). It also implicitly distinguishes itself from siblings by positioning itself as the first-pass survey before extract_semantic_dom/session_extract. An agent knows the resource and output shape without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit workflow rule: 'Start here on any page you have not seen; then extract one region.' It also names the follow-up tools (extract_semantic_dom / session_extract) and explains the handoff via the region `selector` passed as `scope`. When-to-use and the alternative path are both stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_semantic_domA
Read-onlyIdempotent

Navigate to a staging URL and return factual Semantic JSON of all interactive/test-relevant elements with Playwright-native locators and live state. Use this before writing any Playwright test so selectors are real, not guessed.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to extract. Must be http/https and on an allowlisted host.
rolesNoKeep only nodes with these roles or tags (e.g. ['button','textbox','row']).
scopeNoCSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped.
viewportNoViewport preset — 'mobile' is 375x812 with touch, for responsive states.desktop
wait_forNoNavigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages.auto
max_nodesNoCap on extracted nodes; truncation is flagged, never silent.
visible_onlyNoSkip hidden nodes entirely.
wait_selectorNoOptional selector to await before extracting (for SPA content).
include_hiddenNoKeep hidden nodes flagged rather than dropping them.
include_tablesNoAttach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions.
max_output_charsNoBudget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent).
include_click_targetsNoOpt-in heuristic: also include cursor:pointer elements with content that match no other rule (JS-click product cards without anchors/roles/test-ids). Heuristic nodes carry a context_note.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that it navigates to a staging URL and captures 'live state', but it discloses nothing about auth/permission requirements, rate limits, timeout behavior, or execution cost. Useful context, but thin beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly constructed sentences: capability first, usage guidance second, no filler. Both sentences earn their place and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does sketch the return ('factual Semantic JSON... with Playwright-native locators and live state'), which is the key thing an agent needs to interpret results. It omits the output shape details (nodes, omitted counts, context_note) that the params hint at, so it is strong but not fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across 12 params, and the schema itself carries rich per-parameter guidance (wait_for semantics, scope selector sourcing, truncation flags). The description contributes no parameter detail at all, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Navigate... and return') plus a well-defined resource: Semantic JSON of interactive/test-relevant elements with Playwright-native locators. It is clearly distinguishable from extract_outline by emphasizing test-relevant elements and live state, but it never explicitly names or contrasts a sibling (e.g. extract_semantic_dom_after), so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this before writing any Playwright test so selectors are real, not guessed' gives a clear triggering context and rationale. However it offers no when-not guidance and does not point to the obvious alternative, extract_semantic_dom_after, for post-interaction states.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_semantic_dom_afterA

Like extract_semantic_dom, but first performs a short DECLARED list of actions (fill/click/press/select/goto/wait) in the main frame, then returns Semantic JSON of the RESULTING state. Use it for post-interaction UI a plain snapshot cannot see: success/error toasts, validation messages, opened dialogs. Derive action locators from a prior extract_semantic_dom call. The page must remain on allowlisted hosts after the actions, or nothing is extracted. Uniqueness reflects capture time — accumulating UI (chat threads, lists) can multiply matches later. The result's observed block lists navigations, xhr/fetch requests (method, path, status), console errors, dialogs and popups seen while the actions ran — use them for waitForURL/waitForResponse.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to extract. Must be http/https and on an allowlisted host.
rolesNoKeep only nodes with these roles or tags (e.g. ['button','textbox','row']).
scopeNoCSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped.
actionsYesDeclared actions (fill/click/press/select/goto/wait) executed in order in the MAIN frame after navigation.
viewportNoViewport preset — 'mobile' is 375x812 with touch, for responsive states.desktop
wait_forNoNavigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages.auto
max_nodesNoCap on extracted nodes; truncation is flagged, never silent.
settle_msNoWait after the last action before snapshotting (for toasts/animations).
visible_onlyNoSkip hidden nodes entirely.
wait_selectorNoOptional selector to await before extracting (for SPA content).
include_hiddenNoKeep hidden nodes flagged rather than dropping them.
include_tablesNoAttach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions.
max_output_charsNoBudget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent).
wait_selector_afterNoSelector to await (visible) AFTER the actions, before snapshotting — deterministic wait for late-rendering toasts/modals instead of guessing settle_ms.
include_click_targetsNoOpt-in heuristic: also include cursor:pointer elements with content that match no other rule (JS-click product cards without anchors/roles/test-ids). Heuristic nodes carry a context_note.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so safety is partly covered. The description adds substantial behavior beyond them: allowlist enforcement causing silent non-extraction, the timing-dependent uniqueness caveat for accumulating UI, and the contents of the `observed` block (navigations, xhr/fetch method/path/status, console errors, dialogs, popups) with guidance to use them for waitForURL/waitForResponse.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the 'Like extract_semantic_dom, but...' framing, then usage, then caveats, then the observed-block payoff. Every sentence carries information, though the prose is dense and a couple of clauses (uniqueness note, observed block) are packed into long sentences that could be split for scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter mutation-adjacent extraction tool with no output schema, the description covers the action model, locator sourcing, allowlist constraints, timing caveats, and the shape of the returned observation metadata. Nothing an agent needs to invoke it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by listing the declared action verbs (fill/click/press/select/goto/wait), naming the `observed` block fields for follow-up waits, and instructing where locators should come from (a prior extraction). That is meaningful semantic guidance layered on top of fully documented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names the base operation it extends (extract_semantic_dom), states the specific added behavior (executes a declared action list in the main frame first), and states what is returned (Semantic JSON of the resulting state). An agent can distinguish it from extract_semantic_dom and the session_* tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('post-interaction UI a plain snapshot cannot see: success/error toasts, validation messages, opened dialogs') and how to prepare ('Derive action locators from a prior extract_semantic_dom call'). It also states the failure condition (page must remain on allowlisted hosts or nothing is extracted), which is a clear 'when this won't work' qualifier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_conventionsA
Read-onlyIdempotent

The team's Playwright test-writing conventions as text (same content as the write_playwright_test prompt), for clients that do not surface MCP prompts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety and idempotency are covered. The description adds that the output is text identical to a prompt's content, which is useful, but it does not disclose return format details, size, or any rate limits. With annotations covering the behavioral profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-structured sentence that front-loads the resource and the distinguishing context. Every clause earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only text-getter with annotations already covering safety and idempotency, the description is nearly complete. It explains what is returned and why the tool exists. It could be slightly more complete by noting that the content always matches the prompt, implying updates to the prompt affect this output, but that is a minor nuance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so parameter semantics are trivially satisfied. The baseline for zero parameters is 4, and the description adds no misleading parameter information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (team's Playwright test-writing conventions) delivered as text, with the verb implicit in 'get'. It also relates the tool to the write_playwright_test prompt, but does not differentiate it from the sibling tools listed, which center on semantic DOM extraction and session management.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states the condition for use: clients that do not surface MCP prompts. This is a specific, actionable when-to-use cue. However, it does not mention when not to use the tool (e.g., if prompts are available, use write_playwright_test instead), so it stops short of explicitly naming alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_framesA
Read-onlyIdempotent

Diagnostic: navigate to a URL and return its frame tree (frame_path, url, name, same_origin, reachable). Useful for debugging cross-origin iframe boundaries before extraction.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page whose frame tree to report. Must be http/https and allowlisted.
wait_forNoNavigation wait; see extract_semantic_dom.auto

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world, so the safety profile is covered. The description adds that this is a diagnostic navigation, but says nothing about failure/timeout behavior or what 'reachable: false' means in practice, leaving room for more behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the 'Diagnostic:' qualifier front-loaded, then the purpose and the debugging use case. Zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned frame fields, and the 2 parameters are fully documented in the schema. Only minor gaps remain — no mention of how navigation failures are surfaced or the cost/latency of loading a page.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the url allowlist/constraint and the wait_for enum documented in the schema itself. The description restates the URL concept but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource ('navigate to a URL and return its frame tree') and enumerates the returned fields (frame_path, url, name, same_origin, reachable), which distinguishes it from the extract_* siblings that consume the tree rather than report it. It stops short of explicitly naming a sibling, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear use context — 'debugging cross-origin iframe boundaries before extraction' — which implicitly positions it ahead of the extraction tools. There is no explicit when-not guidance or named alternative, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_actA

Perform a short DECLARED action list (fill/click/press/select/goto/wait, max 20) in an open session's main frame. Returns the page's resulting URL/title and observed: main-frame navigations, xhr/fetch requests (method, path, status — bodies and query strings never captured), console errors, dialogs (auto-dismissed) and popups (recorded, closed). These are the facts for waitForURL/waitForResponse. Derive action locators from a prior extraction (paste its playwright expression). Pass then_extract to get the diff (or a scoped extraction / outline) in the same call. If the actions leave the allowlisted hosts the session is closed and nothing further is extracted.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsYesDeclared actions (fill/click/press/select/goto/wait) executed in order in the MAIN frame after navigation.
settle_msNoWait after the last action before snapshotting (for toasts/animations).
session_idYesFrom session_open.
then_extractNoAlso snapshot in this call — one round trip per step. Default mode 'diff' against the previous snapshot; supports scope/roles/visible_only/max_output_chars/include_tables.
wait_selector_afterNoSelector to await (visible) AFTER the actions, before snapshotting — deterministic wait for late-rendering toasts/modals instead of guessing settle_ms.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover readOnly/destructive/idempotent/openWorld flags. The description goes well beyond them: it enumerates what `observed` returns, states that request bodies and query strings are never captured, that dialogs are auto-dismissed and popups recorded then closed, and that leaving allowlisted hosts closes the session and halts extraction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then the return contract, then the locator-derivation and side-effect warnings. Dense but every sentence carries operational information; the single long paragraph is slightly harder to scan than a bulleted form would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter, nested-schema mutation tool with no output schema, the description supplies the return shape (URL/title plus `observed` categories), the locator provenance rule, the allowlist failure mode, and the extraction-folding option. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3; the description still adds usable semantics by explaining that action locators should be sourced from a prior extraction's `playwright` expression and that `then_extract` defaults to a diff in the same call. It adds routing/derivation guidance rather than restating field docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (perform a declared action list of fill/click/press/select/goto/wait) on a specific resource (an open session's main frame) with scope limits (max 20). An agent can distinguish this from extraction/verification siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: derive locators from a prior extraction and paste its `playwright` expression, pass `then_extract` to fold the snapshot into the same call, and the allowlist consequence. It does not explicitly name alternative siblings (e.g. when to prefer session_extract or extract_semantic_dom_after), so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_closeA
Idempotent

Close an open session and release its browser context. Always call this when the flow is done. Idempotent: closing an unknown or already-closed id succeeds with was_open: false.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and destructiveHint=false, so the description's idempotency claim is partly redundant. However it adds real detail beyond the annotations: that closing an unknown/already-closed id succeeds and returns was_open: false — a behavioral guarantee that matters when no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and effect, then the usage rule, then the edge-case guarantee. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the action, the resource released, the usage trigger, and idempotent edge-case behavior including the was_open field, which compensates for the absent output schema. It stops short of explaining where session_id comes from or what errors (other than unknown ids) can occur.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter session_id has no description in either the schema or the description. The description implies session_id identifies a live session but gives no format, source, or example. With one parameter and no schema help, this is the minimum viable level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Close an open session') plus the concrete effect ('release its browser context'), which clearly separates it from session_open, session_act, and session_list. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Always call this when the flow is done' gives an explicit trigger condition for use, which is stronger than most sibling definitions. It does not name an alternative or a when-not case, but for a lifecycle terminator there is no plausible alternative to route to.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_extractA
Read-only

Snapshot the CURRENT state of an open session as Semantic JSON (same shape as extract_semantic_dom, plus snapshot_id). Pass diff_against: 'previous' (or a snapshot_id) to receive only what changed: added nodes (new toasts/dialogs/fields), removed nodes, changed properties (value, is_disabled, aria_invalid, described_by…) and the behavior observed in between — far smaller than a full re-extraction and exactly the assertion list for the step. Diff identity: frame + test-id, else id, else placeholder, else tag+role+accessible name (+ document-order index); a renamed node with no stable attribute shows as removed + added; primary_locator.playwright changes say which locator is valid in which state.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'outline': the page as a map (regions with `selector`, structured tables, dialogs, alerts) in a few thousand characters; does not consume a snapshot id.nodes
rolesNoKeep only nodes with these roles or tags (e.g. ['button','textbox','row']).
scopeNoCSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped.
max_nodesNoCap on extracted nodes; truncation is flagged, never silent.
session_idYesFrom session_open.
diff_againstNoReturn only what changed since that snapshot_id (or 'previous' = the last snapshot in this session) instead of the full extraction — added/removed/changed nodes plus the behavior observed in between.
visible_onlyNoSkip hidden nodes entirely.
include_hiddenNoKeep hidden nodes flagged rather than dropping them.
include_tablesNoAttach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions.
max_output_charsNoBudget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent).
include_click_targetsNoOpt-in heuristic: also include cursor:pointer elements with content that match no other rule (JS-click product cards without anchors/roles/test-ids). Heuristic nodes carry a context_note.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, non-idempotent), and the description adds real behavioral context beyond them: truncation/omission is 'flagged, never silent', diffs are 'far smaller than a full re-extraction', the observed-behavior window between snapshots, and the diff-identity rules (frame+test-id → id → placeholder → tag+role+name) including that a renamed node appears as removed+added. The idempotentHint=false is implicitly explained by snapshot consumption. It stops short of auth/permission or rate-limit notes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and output shape, then the diff mode and its payoff, then the identity rules needed to interpret a diff. Dense but every sentence carries non-redundant information, and the long identity clause is the one piece that genuinely cannot be inferred from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 11 parameters and no output schema, the description carries the return-value burden and largely discharges it: it names the output shape, the diff payload, the snapshot_id field, and primary_locator.playwright state validity. It leaves the full node/table/dialog output structure to the reader's knowledge of extract_semantic_dom, which is a minor gap for a tool this rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema for the most consequential parameter: what diff_against actually returns (added/removed/changed nodes plus intervening behavior) and the identity rules that determine what counts as changed. It also clarifies that outline mode 'does not consume a snapshot id'. The remaining 9 parameters are left to their schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (snapshot/extract) on a specific resource (the CURRENT state of an open session), names the exact output shape (Semantic JSON) and explicitly differentiates it from the sibling extract_semantic_dom by the added `snapshot_id` and the diff capability. An agent can tell it apart from extract_semantic_dom/extract_semantic_dom_after without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context: pass diff_against: 'previous' to get 'exactly the assertion list for the step', implying the verification-step workflow, and the schema documents outline mode as a low-cost alternative. It does not, however, state when to prefer this over extract_semantic_dom or extract_semantic_dom_after, so the alternative-selection guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_listA
Read-onlyIdempotent

Diagnostic: list open sessions (id, URL, expiry, counts) — recover a session id after losing context.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds value beyond them by enumerating what the response contains (id, URL, expiry, counts), which matters since there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the purpose and then the payoff, with no redundant or filler clauses. Every element (diagnostic framing, resource, return fields, use case) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only listing tool with no output schema, the description covers purpose, usage trigger, and the shape of returned data. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The parenthetical describes returned data rather than inputs, which is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('list open sessions') plus the fields returned (id, URL, expiry, counts), and the 'Diagnostic' label frames intent. The verb/resource pair is distinct from session_open/session_close, though it never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger for use: 'recover a session id after losing context.' That is a clear usage context, but there is no statement of when not to use it or an explicit pointer to session_open/session_close as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_openA

Open a persistent browser session at a staging URL for a MULTI-STEP flow (login → cart → checkout). The page stays open across calls: use session_act to perform declared actions and session_extract to snapshot or diff, then session_close. Fresh context per session (storageState applied if configured). Sessions expire after an idle TTL and are capped in number; the allowlist is re-checked after every step.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to extract. Must be http/https and on an allowlisted host.
viewportNoViewport preset — 'mobile' is 375x812 with touch, for responsive states.desktop
wait_forNoNavigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages.auto
wait_selectorNoOptional selector to await before extracting (for SPA content).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare openWorldHint=false, readOnlyHint=false, non-idempotent, non-destructive. The description adds substantial context beyond that: persistent-across-calls page, fresh context per session with storageState applied if configured, idle-TTL expiry, session count cap, and allowlist re-check after every step. These are non-obvious operational constraints an agent must know, though the description doesn't cover reset/cleanup semantics explicitly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then the lifecycle (open → act → extract → close), then operational constraints. Every sentence carries a distinct fact; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the lifecycle, state model (fresh context, storageState), lifecycle limits (TTL, cap), and security re-check — enough to call correctly without an output schema. Minor gaps remain (does the description note what session_open returns, e.g., a session id, and what triggers immediate expiration?), but for a tool with full schema coverage this is strong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema already documents url, viewport, wait_for, and wait_selector with rich enum semantics. The description adds little parameter-level meaning (URL must be staging, allowlisted is echoed by schema), so baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (open a persistent browser session) with an explicit scope constraint (staging URL, MULTI-STEP flow). Clearly distinguishes itself from sibling session_act/session_extract/session_close by naming them as the follow-on workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use ('MULTI-STEP flow') and names the exact alternatives to chain (session_act to act, session_extract to snapshot/diff, then session_close). It also implies the boundary: this is the entry point of a multi-step session rather than a one-shot extract.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_verify_locatorsB
Read-onlyIdempotent

verify_locators against an open session's current page (after acting, without a fresh navigation).

ParametersJSON Schema
NameRequiredDescriptionDefault
locatorsYesPlaywright expressions as written in the test (getByRole(...), getByTestId(...), scoped forms, .nth(i)).
session_idYesFrom session_open.

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that it operates on the current page without fresh navigation, which is useful behavioral context, but it does not explain what verification produces, what happens on failure, or how results are reported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence, front-loaded with the action and followed by key context. No wasted sentences. The opening term repeats the tool name, but the overall structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description carries the burden of explaining return values and what 'verify' yields (e.g., pass/fail per locator, resolved elements, errors). It omits this entirely, leaving an agent unsure how to interpret results, which is a significant gap for a verification tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the locators array documents Playwright expression syntax and session_id documents its source. The description adds no parameter-level meaning beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (verify_locators) and adds meaningful context (against an open session's current page, after acting, without a fresh navigation). However, it largely restates the tool name and never clarifies what 'verify' actually checks (existence, visibility, count, resolution), nor does it differentiate from the sibling tool named verify_locators.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(after acting, without a fresh navigation)' gives an implied timing condition for when this tool is appropriate, but it never names the alternative tool (e.g., verify_locators) or states explicit exclusions. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_locatorsA
Read-onlyIdempotent

Count every given Playwright expression against the live page (same engine that verified the extraction). Use after writing a test: paste the spec's getBy*/locator expressions and get matches, uniqueness, and the first matched element per expression, plus a summary. Also the drift check to run in CI against a page.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to extract. Must be http/https and on an allowlisted host.
locatorsYesPlaywright expressions as written in the test (getByRole(...), getByTestId(...), scoped forms, .nth(i)).
viewportNoViewport preset — 'mobile' is 375x812 with touch, for responsive states.desktop
wait_forNoNavigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages.auto
wait_selectorNoOptional selector to await before extracting (for SPA content).

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a safe, read-only, idempotent, non-open-world operation, so the bar is lower. The description adds real value beyond them by disclosing the return shape (matches, uniqueness, first matched element per expression, plus a summary) and the CI drift-check scenario.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then usage and a secondary CI scenario. Dense but every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing what the tool returns. Combined with fully documented parameters and safety annotations, an agent has enough to call it correctly, though the relationship to session_verify_locators remains unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters including url, locators, viewport, wait_for, and wait_selector are already documented with defaults, enums, and constraints. The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: counting Playwright expressions against a live page, and notes it uses the same engine that verified the extraction. The purpose is unambiguous, but it does not explicitly distinguish itself from the sibling session_verify_locators, leaving the agent to infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context: 'Use after writing a test' with instructions to paste the spec's getBy*/locator expressions, and a second use case as a CI drift check. No exclusions or explicit sibling comparison are provided, but the triggering conditions are concrete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.8.0
    • Changedcheck_auth3 fields changed
      • changedInput schema / properties / wait_for / default
        Previous value: -"networkidle"New value: +"auto"
      • changedInput schema / properties / wait_for / description
        Previous value: -"Navigation wait condition."New value: +"Navigation wait; see extract_semantic_dom."
      • changedInput schema / properties / wait_for / enum
        Previous value: -[
        -  "load",
        -  "domcontentloaded",
        -  "networkidle"
        -]New value: +[
        +  "auto",
        +  "load",
        +  "domcontentloaded",
        +  "networkidle"
        +]
    • Addedextract_outline
    • Changedextract_semantic_dom8 fields changed
      • addedInput schema / properties / include_tables
        Added value: +{
        +  "default": false,
        +  "description": "Attach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / max_output_chars
        Added value: +{
        +  "description": "Budget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent).",
        +  "exclusiveMinimum": 0,
        +  "type": "integer"
        +}
      • addedInput schema / properties / roles
        Added value: +{
        +  "description": "Keep only nodes with these roles or tags (e.g. ['button','textbox','row']).",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / scope
        Added value: +{
        +  "description": "CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role=\"dialog\"]'). Everything outside is skipped.",
        +  "type": "string"
        +}
      • addedInput schema / properties / visible_only
        Added value: +{
        +  "default": false,
        +  "description": "Skip hidden nodes entirely.",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / wait_for / default
        Previous value: -"networkidle"New value: +"auto"
      • changedInput schema / properties / wait_for / description
        Previous value: -"Navigation wait condition."New value: +"Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages."
      • changedInput schema / properties / wait_for / enum
        Previous value: -[
        -  "load",
        -  "domcontentloaded",
        -  "networkidle"
        -]New value: +[
        +  "auto",
        +  "load",
        +  "domcontentloaded",
        +  "networkidle"
        +]
    • Changedextract_semantic_dom_after10 fields changed
      • changedInput schema / properties / actions / description
        Previous value: -"Declared actions executed in order in the MAIN frame after navigation."New value: +"Declared actions (fill/click/press/select/goto/wait) executed in order in the MAIN frame after navigation."
      • changedInput schema / properties / actions / items / anyOf
        Previous value: -[
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "locator": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "nth": {
        -            "description": "Optional .nth(i) index from the extraction's disambiguation guidance.",
        -            "minimum": 0,
        -            "type": "integer"
        -          },
        -          "role": {
        -            "description": "ARIA role — required when strategy is 'role'.",
        -            "type": "string"
        -          },
        -          "strategy": {
        -            "description": "Locator strategy, matching the strategies in extraction output.",
        -            "enum": [
        -              "test-id",
        -              "role",
        -              "label",
        -              "placeholder",
        -              "text",
        -              "id",
        -              "css"
        -            ],
        -            "type": "string"
        -          },
        -          "value": {
        -            "description": "The locator value (test id, accessible name, label, selector...).",
        -            "minLength": 1,
        -            "type": "string"
        -          }
        -        },
        -        "required": [
        -          "strategy",
        -          "value"
        -        ],
        -        "type": "object"
        -      },
        -      "type": {
        -        "const": "fill",
        -        "type": "string"
        -      },
        -      "value": {
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "type",
        -      "locator",
        -      "value"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "locator": {
        -        "$ref": "#/properties/actions/items/anyOf/0/properties/locator"
        -      },
        -      "type": {
        -        "const": "click",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "type",
        -      "locator"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "key": {
        -        "maxLength": 30,
        -        "type": "string"
        -      },
        -      "locator": {
        -        "$ref": "#/properties/actions/items/anyOf/0/properties/locator"
        -      },
        -      "type": {
        -        "const": "press",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "type",
        -      "locator",
        -      "key"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "ms": {
        -        "exclusiveMinimum": 0,
        -        "maximum": 10000,
        -        "type": "integer"
        -      },
        -      "type": {
        -        "const": "wait",
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "type",
        -      "ms"
        -    ],
        -    "type": "object"
        -  }
        -]New value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "locator": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "nth": {
        +            "description": "Optional .nth(i) index from the extraction's disambiguation guidance.",
        +            "minimum": 0,
        +            "type": "integer"
        +          },
        +          "playwright": {
        +            "description": "A `playwright` expression exactly as an extraction returned it (scoped forms and .nth included). When given, strategy/value are not needed.",
        +            "minLength": 1,
        +            "type": "string"
        +          },
        +          "role": {
        +            "description": "ARIA role — required when strategy is 'role'.",
        +            "type": "string"
        +          },
        +          "strategy": {
        +            "default": "css",
        +            "description": "Locator strategy, matching the strategies in extraction output (ignored when `playwright` is given).",
        +            "enum": [
        +              "test-id",
        +              "role",
        +              "label",
        +              "placeholder",
        +              "text",
        +              "id",
        +              "css"
        +            ],
        +            "type": "string"
        +          },
        +          "value": {
        +            "default": "",
        +            "description": "The locator value (test id, accessible name, label, selector...). Empty for a bare role inside `within`.",
        +            "type": "string"
        +          },
        +          "within": {
        +            "additionalProperties": false,
        +            "description": "Scope to a container first; copy the extraction's `within` verbatim.",
        +            "properties": {
        +              "kind": {
        +                "enum": [
        +                  "row",
        +                  "listitem",
        +                  "test-id",
        +                  "css"
        +                ],
        +                "type": "string"
        +              },
        +              "value": {
        +                "minLength": 1,
        +                "type": "string"
        +              }
        +            },
        +            "required": [
        +              "kind",
        +              "value"
        +            ],
        +            "type": "object"
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "secret": {
        +        "description": "Mark the value as a secret: it is scrubbed from every string the server returns. Password fields are detected automatically.",
        +        "type": "boolean"
        +      },
        +      "type": {
        +        "const": "fill",
        +        "type": "string"
        +      },
        +      "value": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "locator",
        +      "value"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "locator": {
        +        "$ref": "#/properties/actions/items/anyOf/0/properties/locator"
        +      },
        +      "type": {
        +        "const": "click",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "locator"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "key": {
        +        "maxLength": 30,
        +        "type": "string"
        +      },
        +      "locator": {
        +        "$ref": "#/properties/actions/items/anyOf/0/properties/locator"
        +      },
        +      "type": {
        +        "const": "press",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "locator",
        +      "key"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "description": "Choose a <select> option by value or label.",
        +    "properties": {
        +      "locator": {
        +        "$ref": "#/properties/actions/items/anyOf/0/properties/locator"
        +      },
        +      "type": {
        +        "const": "select",
        +        "type": "string"
        +      },
        +      "value": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "locator",
        +      "value"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "description": "Navigate within the flow (must be http/https and allowlisted).",
        +    "properties": {
        +      "type": {
        +        "const": "goto",
        +        "type": "string"
        +      },
        +      "url": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "url"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "ms": {
        +        "exclusiveMinimum": 0,
        +        "maximum": 10000,
        +        "type": "integer"
        +      },
        +      "type": {
        +        "const": "wait",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "ms"
        +    ],
        +    "type": "object"
        +  }
        +]
      • addedInput schema / properties / include_tables
        Added value: +{
        +  "default": false,
        +  "description": "Attach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / max_output_chars
        Added value: +{
        +  "description": "Budget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent).",
        +  "exclusiveMinimum": 0,
        +  "type": "integer"
        +}
      • addedInput schema / properties / roles
        Added value: +{
        +  "description": "Keep only nodes with these roles or tags (e.g. ['button','textbox','row']).",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / scope
        Added value: +{
        +  "description": "CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role=\"dialog\"]'). Everything outside is skipped.",
        +  "type": "string"
        +}
      • addedInput schema / properties / visible_only
        Added value: +{
        +  "default": false,
        +  "description": "Skip hidden nodes entirely.",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / wait_for / default
        Previous value: -"networkidle"New value: +"auto"
      • changedInput schema / properties / wait_for / description
        Previous value: -"Navigation wait condition."New value: +"Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages."
      • changedInput schema / properties / wait_for / enum
        Previous value: -[
        -  "load",
        -  "domcontentloaded",
        -  "networkidle"
        -]New value: +[
        +  "auto",
        +  "load",
        +  "domcontentloaded",
        +  "networkidle"
        +]
    • Addedget_conventions
    • Changedlist_frames3 fields changed
      • changedInput schema / properties / wait_for / default
        Previous value: -"networkidle"New value: +"auto"
      • changedInput schema / properties / wait_for / description
        Previous value: -"Navigation wait condition."New value: +"Navigation wait; see extract_semantic_dom."
      • changedInput schema / properties / wait_for / enum
        Previous value: -[
        -  "load",
        -  "domcontentloaded",
        -  "networkidle"
        -]New value: +[
        +  "auto",
        +  "load",
        +  "domcontentloaded",
        +  "networkidle"
        +]
    • Addedsession_act
    • Addedsession_close
    • Addedsession_extract
    • Addedsession_list
    • Addedsession_open
    • Addedsession_verify_locators
    • Addedverify_locators
  2. 4 tool updatesv0.4.0
    • First observedcheck_auth
    • First observedextract_semantic_dom
    • First observedextract_semantic_dom_after
    • First observedlist_frames

TDQS

A3.9/5.0

Scored across 13 tools

Disambiguation4/5

The set has a clear two-tier structure (one-shot extraction/verification vs. persistent session tools), and descriptions carefully separate them. However, extract_semantic_dom_after overlaps conceptually with the session_open/session_act/session_extract flow, and extract_semantic_dom vs session_extract require reading descriptions carefully to pick correctly.

Naming Consistency4/5

Almost everything is consistent snake_case verb_noun, and the session_* prefix gives a clean, predictable grouping. Minor deviation: extract_semantic_dom_after uses a temporal suffix rather than a distinct verb, which slightly breaks the pattern.

Tool Count5/5

13 tools is well within the sweet spot and each one earns its place, covering extraction, diagnostics, verification, conventions, and full session lifecycle without filler.

Completeness5/5

The surface covers the full Playwright test-authoring workflow: page mapping, snapshot extraction, post-interaction observation, locator verification, auth/frame diagnostics, conventions, and end-to-end session management with open/act/extract/verify/close/list. No obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to understand web page structure and content through structured data extraction and element discovery using Playwright, eliminating the need for screenshots.
    4
    7 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Gives AI agents a compact, semantic interface to the browser, returning structured page snapshots with stable element IDs instead of raw DOM. Enables agents to navigate, interact, and extract information from web pages efficiently.
    6
    26
    353 npm
    17
    MIT