semantic-dom-mcp
Summary: This server drives a real Chromium browser via Playwright to give coding agents verified Playwright locators, structured page/assertion data, and observed page behavior so they can write real (not guessed) Playwright tests.
Extract a page's semantic DOM (
extract_semantic_dom): navigate an allowlisted staging URL (desktop or mobile viewport) and get factual Semantic JSON of all interactive/test-relevant elements with ready-to-use, uniqueness-verified Playwright locators and live state.Extract after interactions (
extract_semantic_dom_after): run a short declared action list (fill/click/press/wait, max 20) in the main frame, then snapshot the resulting state — good for success/error toasts, validation messages, and opened dialogs — with deterministic waits viawait_selector_after.Diagnose frames (
list_frames): inspect a page's frame tree (frame_path, url, name, same-origin/reachability) to debug cross-origin iframe boundaries before extraction.Check authentication (
check_auth): navigate using the configured storageState and detect if the session bounced to a login page, revealing expired auth instead of mysterious results.Tune extractions with options for
wait_for,wait_selector,max_nodes,include_hidden, andinclude_click_targets(heuristic JS-click cards).Use these alongside the README's session tools (
session_open/session_act/session_extract/etc.), prompts (write_playwright_test), and resources (conventions://playwright) to build multi-step flows grounded in real locators, redirects, API paths, and state transitions.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@semantic-dom-mcpextract the checkout page and write a success-path test"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
semantic-dom-mcp
Local MCP server (stdio, Node.js + TypeScript) that drives a real Chromium browser via Playwright and gives a coding agent what it needs to write a Playwright test it did not have to guess: verified locators (counted by Playwright's own engine, scoped the way a QA engineer scopes them), structured assertion data (tables as rows and cells, dialogs as label/value pairs), the behavior the page showed (navigations, requests, console errors), and diffs between steps of a flow. Same page → same output → same conventions → same test style across the team.
How it keeps tokens low. An agent starts with an outline of the page (regions, tables, dialogs, alerts; a few hundred to a few thousand characters, smaller than the native accessibility tree), then extracts one region, and acts with the snapshot in the same call. Measured on a real seller dashboard, the outline was 82–95% smaller than Playwright's aria snapshot of the same page, and a whole authenticated add-to-cart flow ran in about 10k tokens. A full-page extraction is still 2–3x the size of the aria snapshot, because it carries what that tree does not: executable locators, uniqueness verdicts and state. Ask for the whole page only when you need it. Numbers: benchmark/RESULTS.md, FLOW-VALIDATION.md. Docs: How it works (deep dive) · Team guide (setup + connecting your agent) · Benchmark methodology · v0.5 flow validation · Roadmap · Changelog
Quickstart
No clone, no build — the package is on npm. One-time browser setup (installs the Chromium build matching the package's bundled Playwright):
npx -y -p semantic-dom-mcp playwright install chromiumThen add the server to your MCP client:
{
"mcpServers": {
"semantic-dom": {
"command": "npx",
"args": ["-y", "semantic-dom-mcp"],
"env": {
"QA_MCP_ALLOWED_HOSTS": "staging.yourapp.internal,staging.admin.internal",
"QA_MCP_STORAGE_STATE": "./.auth/staging.json"
}
}
}
}That's the whole setup. Verify by asking your agent to list its MCP tools — you should
see extract_semantic_dom. See docs/GUIDE.md for per-client config
locations (Claude Code, Claude Desktop, Cursor, Windsurf), authenticated staging, and
troubleshooting. To run from a clone instead (contributors), see Development below.
Related MCP server: Browser MCP Server
Workflow
Single page. Ask your agent: "extract the checkout page and write a success-path test."
The agent calls
extract_semantic_dom({ url }). The server navigates a real Chromium page, runs the extractor inside it, and returns Semantic JSON: every interactive node with a ready-to-paste Playwright locator, uniqueness verified by Playwright's own engine.The agent uses the
write_playwright_testprompt (scenario + the JSON), which injects the team conventions.The result is a Playwright test in team style, grounded in real locators, never guessed ones.
Multi-step flow (v0.8 way). Ask: "write a test for login → add to cart → checkout."
session_open({ url })opens a persistent page, thensession_extract({ mode: "outline" })maps it: regions with aselector, tables with row identity and cells, dialogs, alerts.session_extract({ scope, roles, visible_only })pulls only the region the step needs, with locators verified unique inside it.session_act({ actions, then_extract: { mode: "diff", scope } })runs declared actions (paste any returnedplaywrightexpression as the locator) and returns, in the same call, what the page did (observed: redirects, requests, console errors) and what changed (added nodes, removed nodes, state transitions). That is the wait list and the assertion list for the step.Repeat per step, then
session_verify_locatorswith the expressions in the written spec, thensession_close.
Every fact a test needs (locator, redirect, API path, state transition, cell value) comes from the page, not from the agent's memory.
MCP surface
Kind | Name | Purpose |
Tool |
| The page as a map: regions (each with a |
Tool |
| Extract a URL into Semantic JSON ( |
Tool |
| Count every Playwright expression from a written spec against the live page; matches, uniqueness, first match, summary. |
Tool |
| The team conventions as text, for clients that hide MCP prompts. |
Tool |
| Same, but first runs a short declared action list (fill/click/press/select/goto/wait, max 20) in the main frame and snapshots the resulting state, plus an |
Tool |
| Open a persistent page for a multi-step flow ( |
Tool |
| Run declared actions in an open session; |
Tool |
| Snapshot the session's current state ( |
Tool |
|
|
Tool |
| Release the session's browser context. |
Tool |
| Diagnostic: open sessions with URL, expiry and counts. |
Tool |
| Diagnostic: navigates with the configured storageState and reports whether the session bounced to a login-looking page (expired auth shows up as an answer, not a mystery). |
Tool |
| Diagnostic frame tree with same-origin/reachability classification. |
Prompt |
| Team-standard test-writing prompt ( |
Resource |
| The same team conventions as read-only text. |
Errors (navigation failure, denied host, missing selector, expired session) come back as
structured JSON in the tool result, so the agent can react instead of crashing. Every tool carries
MCP annotations (readOnlyHint, openWorldHint: false) so clients can auto-approve the read-only
ones.
Configuration (environment variables)
Variable | Meaning |
| Required. Comma-separated hostnames the server may navigate to. Navigation is denied by default. Supports |
| Optional path to a Playwright |
| Optional team name used in the |
| Idle time before a session is closed automatically (default |
| Max concurrently open sessions (default |
Security posture
Tool inputs are untrusted (they arrive via an LLM): strict schemas (
additionalProperties: false), http/https only, host allowlist enforced before any navigation.extract_semantic_domonly reads the DOM. It never clicks, submits, or mutates the page. The sanctioned exceptions areextract_semantic_dom_afterandsession_act, which execute only an explicit, bounded, schema-validated action list, never log fill values, and refuse to extract if the page leaves the allowlisted hosts (a session in that state is closed).Sessions add their own limits: an idle TTL, a cap on open sessions, one in-flight call per session, and a bounded snapshot history kept in memory only.
Behavior capture is observation only. The server never issues requests of its own. Request and response bodies are never read, query strings are stripped from recorded URLs (they may carry tokens), console text is capped, dialogs are dismissed and popups closed at once.
Sessions check the allowlist before any action, after every action, after a failed action and before every snapshot. A session found off the allowlist is closed on the spot.
No network egress beyond navigating the browser to allowlisted URLs. No telemetry. Page contents are never logged (stderr carries only high-level events) and are not stored beyond the current call or session.
Semantics worth knowing
Snapshot honesty: the JSON is a single moment. A disabled submit button is reported
is_disabled: truewith a note. The conventions instruct the model to write the interactions that change state, not to assume it stays disabled. In a session, the diff shows the transition itself (is_disabled: false → true), so the test asserts a fact rather than an assumption.Scoped locators (schema 1.4): a nameless or repeated control is located inside its container the way a QA engineer writes it by hand:
getByRole('row', { name: 'Charizard' }).getByRole('radio'),getByRole('listitem').filter({ hasText: 'Puthera' }).getByRole('button', { name: 'Pilih Pembeli Ini' }),getByTestId('customer-checkbox-flex').getByRole('checkbox'). The node carrieswithinso an action can target the same element; the expression is verified unique by Playwright like every other locator.Compact wire format (schema 1.3): results are compact JSON and a node field that carries no information is omitted:
nullfields,frame_path: [],in_shadow: false,kind: "element", emptyfallback_locators, andtext_contentequal toaccessible_name. An absent property is null (not applicable, never false); absent structure means the default (main document, light DOM, nothing worth listing).is_visibleis always present. Fallbacks appear only when the primary is ambiguous or brittle. Same facts, about a third of the tokens.Assertable state (since schema 1.2): every node reports
value(never for password fields),aria_expanded,aria_selected,aria_invalid,described_by(the text of the elementsaria-describedbypoints at, where validation messages live),validation_message(browser constraint validation) and, for<select>,options. Absent state isnull, neverfalse.Observed behavior:
observed.requestslists xhr/fetch/document requests as method + path + status; static assets are counted indropped, not listed.observed.navigationslists main-frame URL changes in order. Neither is ever guessed; if a redirect or API path is not inobserved, the conventions tell the agent not to wait on it.Diff identity: nodes pair across snapshots by the most stable fact available: test-id, else id, else placeholder, else tag + role + accessible name, plus frame path and a document-order index for non-unique nodes. A hidden menu item that becomes visible is therefore a
changedentry, with its primary locator switch (getByText→getByRole) listed as one of the changes. A renamed node with no stable attribute shows as removed + added; the diff says so in its notes.Visibility is Playwright's:
is_visiblepredictstoBeVisible(), so it uses Playwright's rule (notdisplay:none,visibilitynot hidden, width and height both > 0). Opacity andaria-hiddendo not hide an element for Playwright and do not here either.Secrets never leave the server: a value typed into a password field, or a
fillmarkedsecret: true, is scrubbed from every string in every result (values, text, console, dialogs, errors) when it is 8+ characters. The set is process-wide and lasts until the server exits, so mark only real secrets: a redacted string blanks that text everywhere, and a locator whose text was redacted is flagged not unique. Password and credential-autocomplete fields never report avalueat all.Hidden nodes are included and flagged
is_visible: false(tests often assert hidden-ness); passinclude_hidden: falseto drop them (the count dropped is noted, never silent).Open shadow DOM is traversed and flagged
in_shadow— locators pierce it natively, so no>>>/::shadowCSS is ever emitted. Closed shadow roots appear asshadow_boundarymarker nodes (detected via pre-navigationattachShadowinstrumentation; closed roots created by declarative shadow DOM parse before scripts run and cannot be detected).Same-origin iframes are extracted per-frame with
frame_pathset (chainframeLocator()in that order). Cross-origin iframes are recorded as opaquecross_origin_framenodes with URL/name only — their DOM is never touched.Notification & dialog surfaces (
role="alert",role="status", dialogs) are extracted like interactive nodes. When a toast library keeps the live region empty and renders the message in a sibling (a common pattern across UI libraries), the message text is pulled from the enclosing container and flagged. For UI that renders late after an interaction,wait_selector_afteronextract_semantic_dom_afterwaits deterministically instead of guessingsettle_ms. Since those ARIA roles take names from the author (not contents), their role locator isgetByRole('alert')— or with thearia-labelname when one exists. For UI that only appears after an interaction (login-success toast), useextract_semantic_dom_after.JS-click cards (product tiles with no anchor/role/test-id) are invisible to the factual rules by design — pass
include_click_targets: trueto include cursor-pointer boundary elements with content, flagged as heuristic and located by their heading text.Links carry
href(schema 1.1) so agents can discover which page to extract next without scraping. Framework-generated ids (rc_select_*, ReactuseId, Radix, MUI...) are detected and demoted to last-resort with a note — they change between builds and must never be primary.viewport: "mobile"(375×812, touch) snapshots responsive states; visibility flags reflect the active media queries.Truncation is loud:
max_nodes/ depth caps settruncated: trueplus a note. Non-unique locators carryis_unique: falseanddisambiguationguidance.
Development
git clone https://github.com/helmif/semantic-dom-mcp.git && cd semantic-dom-mcp
npm install
npx playwright install chromium
npm run dev # run the server over stdio via tsx
npm run typecheck # tsc --noEmit (strict)
npm test # Vitest suites against real fixture pages in headless Chromium
npm run build # compile to dist/ (clients can then use "command": "node", "args": ["<path>/dist/index.js"])Repo layout: src/index.ts (bootstrap) · src/server.ts (MCP surface) · src/browser.ts
(Playwright layer + single-shot orchestration) · src/session.ts (persistent sessions) ·
src/observe.ts (behavior capture) · src/diff.ts (snapshot diff) · src/extractor/ (in-page
engine + locator resolution) · src/types.ts (frozen contract, schema 1.3) · src/compact.ts (wire rules) · src/conventions.ts
(single source of team conventions).
Available Tools
13 toolscheck_authARead-onlyIdempotent
Diagnostic: navigates with the configured QA_MCP_STORAGE_STATE session and reports whether the page bounced to a login-looking path (session likely expired). Use when extractions unexpectedly return login forms instead of the requested page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page whose frame tree to report. Must be http/https and allowlisted. | |
| wait_for | No | Navigation wait; see extract_semantic_dom. | auto |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, so the bar is lower. The description adds genuine context beyond annotations: it discloses the authentication mechanism (QA_MCP_STORAGE_STATE session) and the failure semantics (bounce-to-login indicates expired session). It does not describe what the report format looks like, but that is partially an output-schema concern.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the diagnostic purpose then the trigger condition. Every clause earns its place — no filler, no restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param read-only diagnostic with full schema coverage and annotations carrying the safety profile, the description supplies everything an agent needs: what it does, how it authenticates, and when to reach for it. No output schema exists, but the description adequately frames the result (whether the page bounced to login).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — both parameters are fully described in the schema (url with allowlist/http-https constraint, wait_for enum with default). The description adds no parameter detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific diagnostic verb ('navigates with the configured QA_MCP_STORAGE_STATE session and reports whether the page bounced to a login-looking path') and clearly names what it distinguishes from siblings — it's not an extraction tool but a session-health probe. An agent can tell it apart from extract_semantic_dom or session_extract without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use: 'Use when extractions unexpectedly return login forms instead of the requested page.' This gives a concrete trigger condition that routes the agent here rather than to a re-extraction. No alternative tool is named, but the trigger is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_outlineARead-onlyIdempotent
The page as a MAP, a few thousand characters: landmark regions (header/nav/main/forms/tables/lists/dialogs) each with a selector you can pass as scope to extract_semantic_dom / session_extract, structured tables (headers, row identity for getByRole('row', { name }), cells) and open dialogs (label/value fields), plus current alert text. Start here on any page you have not seen; then extract one region.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page to extract. Must be http/https and on an allowlisted host. | |
| scope | No | CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped. | |
| viewport | No | Viewport preset — 'mobile' is 375x812 with touch, for responsive states. | desktop |
| wait_for | No | Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages. | auto |
| wait_selector | No | Optional selector to await before extracting (for SPA content). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: the output is bounded ('a few thousand characters'), which matters for context budgeting, and the region selectors are reusable downstream. It does not cover failure modes (e.g. non-allowlisted hosts, timeouts), which the schema touches on.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the output payload and ending on the usage directive. The first sentence is densely packed with parentheticals, but every clause carries distinct information (region types, selector handoff, table/dialog/alert contents), so nothing is wasted. Slightly heavy to parse in one pass.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return-value burden and describes the map contents well. It also explains the intended next step. It omits any mention of output size limits, truncation behavior, or error conditions for a tool that fetches an arbitrary URL, which is a minor gap for a read tool with full annotation coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by linking output to input: it explains that the `selector` in each outline region is what you pass as `scope`, which is a relationship the schema alone does not convey. It adds no new detail on viewport/wait_for, hence not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool returns — a page 'MAP' of landmark regions, structured tables, open dialogs, and alert text — and enumerates the region types (header/nav/main/forms/tables/lists/dialogs). It also implicitly distinguishes itself from siblings by positioning itself as the first-pass survey before extract_semantic_dom/session_extract. An agent knows the resource and output shape without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit workflow rule: 'Start here on any page you have not seen; then extract one region.' It also names the follow-up tools (extract_semantic_dom / session_extract) and explains the handoff via the region `selector` passed as `scope`. When-to-use and the alternative path are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_semantic_domARead-onlyIdempotent
Navigate to a staging URL and return factual Semantic JSON of all interactive/test-relevant elements with Playwright-native locators and live state. Use this before writing any Playwright test so selectors are real, not guessed.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page to extract. Must be http/https and on an allowlisted host. | |
| roles | No | Keep only nodes with these roles or tags (e.g. ['button','textbox','row']). | |
| scope | No | CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped. | |
| viewport | No | Viewport preset — 'mobile' is 375x812 with touch, for responsive states. | desktop |
| wait_for | No | Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages. | auto |
| max_nodes | No | Cap on extracted nodes; truncation is flagged, never silent. | |
| visible_only | No | Skip hidden nodes entirely. | |
| wait_selector | No | Optional selector to await before extracting (for SPA content). | |
| include_hidden | No | Keep hidden nodes flagged rather than dropping them. | |
| include_tables | No | Attach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions. | |
| max_output_chars | No | Budget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent). | |
| include_click_targets | No | Opt-in heuristic: also include cursor:pointer elements with content that match no other rule (JS-click product cards without anchors/roles/test-ids). Heuristic nodes carry a context_note. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds that it navigates to a staging URL and captures 'live state', but it discloses nothing about auth/permission requirements, rate limits, timeout behavior, or execution cost. Useful context, but thin beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly constructed sentences: capability first, usage guidance second, no filler. Both sentences earn their place and nothing is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does sketch the return ('factual Semantic JSON... with Playwright-native locators and live state'), which is the key thing an agent needs to interpret results. It omits the output shape details (nodes, omitted counts, context_note) that the params hint at, so it is strong but not fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across 12 params, and the schema itself carries rich per-parameter guidance (wait_for semantics, scope selector sourcing, truncation flags). The description contributes no parameter detail at all, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Navigate... and return') plus a well-defined resource: Semantic JSON of interactive/test-relevant elements with Playwright-native locators. It is clearly distinguishable from extract_outline by emphasizing test-relevant elements and live state, but it never explicitly names or contrasts a sibling (e.g. extract_semantic_dom_after), so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this before writing any Playwright test so selectors are real, not guessed' gives a clear triggering context and rationale. However it offers no when-not guidance and does not point to the obvious alternative, extract_semantic_dom_after, for post-interaction states.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_semantic_dom_afterA
Like extract_semantic_dom, but first performs a short DECLARED list of actions (fill/click/press/select/goto/wait) in the main frame, then returns Semantic JSON of the RESULTING state. Use it for post-interaction UI a plain snapshot cannot see: success/error toasts, validation messages, opened dialogs. Derive action locators from a prior extract_semantic_dom call. The page must remain on allowlisted hosts after the actions, or nothing is extracted. Uniqueness reflects capture time — accumulating UI (chat threads, lists) can multiply matches later. The result's observed block lists navigations, xhr/fetch requests (method, path, status), console errors, dialogs and popups seen while the actions ran — use them for waitForURL/waitForResponse.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page to extract. Must be http/https and on an allowlisted host. | |
| roles | No | Keep only nodes with these roles or tags (e.g. ['button','textbox','row']). | |
| scope | No | CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped. | |
| actions | Yes | Declared actions (fill/click/press/select/goto/wait) executed in order in the MAIN frame after navigation. | |
| viewport | No | Viewport preset — 'mobile' is 375x812 with touch, for responsive states. | desktop |
| wait_for | No | Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages. | auto |
| max_nodes | No | Cap on extracted nodes; truncation is flagged, never silent. | |
| settle_ms | No | Wait after the last action before snapshotting (for toasts/animations). | |
| visible_only | No | Skip hidden nodes entirely. | |
| wait_selector | No | Optional selector to await before extracting (for SPA content). | |
| include_hidden | No | Keep hidden nodes flagged rather than dropping them. | |
| include_tables | No | Attach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions. | |
| max_output_chars | No | Budget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent). | |
| wait_selector_after | No | Selector to await (visible) AFTER the actions, before snapshotting — deterministic wait for late-rendering toasts/modals instead of guessing settle_ms. | |
| include_click_targets | No | Opt-in heuristic: also include cursor:pointer elements with content that match no other rule (JS-click product cards without anchors/roles/test-ids). Heuristic nodes carry a context_note. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so safety is partly covered. The description adds substantial behavior beyond them: allowlist enforcement causing silent non-extraction, the timing-dependent uniqueness caveat for accumulating UI, and the contents of the `observed` block (navigations, xhr/fetch method/path/status, console errors, dialogs, popups) with guidance to use them for waitForURL/waitForResponse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the 'Like extract_semantic_dom, but...' framing, then usage, then caveats, then the observed-block payoff. Every sentence carries information, though the prose is dense and a couple of clauses (uniqueness note, observed block) are packed into long sentences that could be split for scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter mutation-adjacent extraction tool with no output schema, the description covers the action model, locator sourcing, allowlist constraints, timing caveats, and the shape of the returned observation metadata. Nothing an agent needs to invoke it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by listing the declared action verbs (fill/click/press/select/goto/wait), naming the `observed` block fields for follow-up waits, and instructing where locators should come from (a prior extraction). That is meaningful semantic guidance layered on top of fully documented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names the base operation it extends (extract_semantic_dom), states the specific added behavior (executes a declared action list in the main frame first), and states what is returned (Semantic JSON of the resulting state). An agent can distinguish it from extract_semantic_dom and the session_* tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('post-interaction UI a plain snapshot cannot see: success/error toasts, validation messages, opened dialogs') and how to prepare ('Derive action locators from a prior extract_semantic_dom call'). It also states the failure condition (page must remain on allowlisted hosts or nothing is extracted), which is a clear 'when this won't work' qualifier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_conventionsARead-onlyIdempotent
The team's Playwright test-writing conventions as text (same content as the write_playwright_test prompt), for clients that do not surface MCP prompts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety and idempotency are covered. The description adds that the output is text identical to a prompt's content, which is useful, but it does not disclose return format details, size, or any rate limits. With annotations covering the behavioral profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-structured sentence that front-loads the resource and the distinguishing context. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only text-getter with annotations already covering safety and idempotency, the description is nearly complete. It explains what is returned and why the tool exists. It could be slightly more complete by noting that the content always matches the prompt, implying updates to the prompt affect this output, but that is a minor nuance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so parameter semantics are trivially satisfied. The baseline for zero parameters is 4, and the description adds no misleading parameter information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (team's Playwright test-writing conventions) delivered as text, with the verb implicit in 'get'. It also relates the tool to the write_playwright_test prompt, but does not differentiate it from the sibling tools listed, which center on semantic DOM extraction and session management.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the condition for use: clients that do not surface MCP prompts. This is a specific, actionable when-to-use cue. However, it does not mention when not to use the tool (e.g., if prompts are available, use write_playwright_test instead), so it stops short of explicitly naming alternatives and exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_framesARead-onlyIdempotent
Diagnostic: navigate to a URL and return its frame tree (frame_path, url, name, same_origin, reachable). Useful for debugging cross-origin iframe boundaries before extraction.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page whose frame tree to report. Must be http/https and allowlisted. | |
| wait_for | No | Navigation wait; see extract_semantic_dom. | auto |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and closed-world, so the safety profile is covered. The description adds that this is a diagnostic navigation, but says nothing about failure/timeout behavior or what 'reachable: false' means in practice, leaving room for more behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the 'Diagnostic:' qualifier front-loaded, then the purpose and the debugging use case. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned frame fields, and the 2 parameters are fully documented in the schema. Only minor gaps remain — no mention of how navigation failures are surfaced or the cost/latency of loading a page.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the url allowlist/constraint and the wait_for enum documented in the schema itself. The description restates the URL concept but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('navigate to a URL and return its frame tree') and enumerates the returned fields (frame_path, url, name, same_origin, reachable), which distinguishes it from the extract_* siblings that consume the tree rather than report it. It stops short of explicitly naming a sibling, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use context — 'debugging cross-origin iframe boundaries before extraction' — which implicitly positions it ahead of the extraction tools. There is no explicit when-not guidance or named alternative, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_actA
Perform a short DECLARED action list (fill/click/press/select/goto/wait, max 20) in an open session's main frame. Returns the page's resulting URL/title and observed: main-frame navigations, xhr/fetch requests (method, path, status — bodies and query strings never captured), console errors, dialogs (auto-dismissed) and popups (recorded, closed). These are the facts for waitForURL/waitForResponse. Derive action locators from a prior extraction (paste its playwright expression). Pass then_extract to get the diff (or a scoped extraction / outline) in the same call. If the actions leave the allowlisted hosts the session is closed and nothing further is extracted.
| Name | Required | Description | Default |
|---|---|---|---|
| actions | Yes | Declared actions (fill/click/press/select/goto/wait) executed in order in the MAIN frame after navigation. | |
| settle_ms | No | Wait after the last action before snapshotting (for toasts/animations). | |
| session_id | Yes | From session_open. | |
| then_extract | No | Also snapshot in this call — one round trip per step. Default mode 'diff' against the previous snapshot; supports scope/roles/visible_only/max_output_chars/include_tables. | |
| wait_selector_after | No | Selector to await (visible) AFTER the actions, before snapshotting — deterministic wait for late-rendering toasts/modals instead of guessing settle_ms. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover readOnly/destructive/idempotent/openWorld flags. The description goes well beyond them: it enumerates what `observed` returns, states that request bodies and query strings are never captured, that dialogs are auto-dismissed and popups recorded then closed, and that leaving allowlisted hosts closes the session and halts extraction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, then the return contract, then the locator-derivation and side-effect warnings. Dense but every sentence carries operational information; the single long paragraph is slightly harder to scan than a bulleted form would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter, nested-schema mutation tool with no output schema, the description supplies the return shape (URL/title plus `observed` categories), the locator provenance rule, the allowlist failure mode, and the extraction-folding option. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description still adds usable semantics by explaining that action locators should be sourced from a prior extraction's `playwright` expression and that `then_extract` defaults to a diff in the same call. It adds routing/derivation guidance rather than restating field docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (perform a declared action list of fill/click/press/select/goto/wait) on a specific resource (an open session's main frame) with scope limits (max 20). An agent can distinguish this from extraction/verification siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: derive locators from a prior extraction and paste its `playwright` expression, pass `then_extract` to fold the snapshot into the same call, and the allowlist consequence. It does not explicitly name alternative siblings (e.g. when to prefer session_extract or extract_semantic_dom_after), so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_closeAIdempotent
Close an open session and release its browser context. Always call this when the flow is done. Idempotent: closing an unknown or already-closed id succeeds with was_open: false.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so the description's idempotency claim is partly redundant. However it adds real detail beyond the annotations: that closing an unknown/already-closed id succeeds and returns was_open: false — a behavioral guarantee that matters when no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and effect, then the usage rule, then the edge-case guarantee. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the action, the resource released, the usage trigger, and idempotent edge-case behavior including the was_open field, which compensates for the absent output schema. It stops short of explaining where session_id comes from or what errors (other than unknown ids) can occur.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter session_id has no description in either the schema or the description. The description implies session_id identifies a live session but gives no format, source, or example. With one parameter and no schema help, this is the minimum viable level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Close an open session') plus the concrete effect ('release its browser context'), which clearly separates it from session_open, session_act, and session_list. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Always call this when the flow is done' gives an explicit trigger condition for use, which is stronger than most sibling definitions. It does not name an alternative or a when-not case, but for a lifecycle terminator there is no plausible alternative to route to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_extractARead-only
Snapshot the CURRENT state of an open session as Semantic JSON (same shape as extract_semantic_dom, plus snapshot_id). Pass diff_against: 'previous' (or a snapshot_id) to receive only what changed: added nodes (new toasts/dialogs/fields), removed nodes, changed properties (value, is_disabled, aria_invalid, described_by…) and the behavior observed in between — far smaller than a full re-extraction and exactly the assertion list for the step. Diff identity: frame + test-id, else id, else placeholder, else tag+role+accessible name (+ document-order index); a renamed node with no stable attribute shows as removed + added; primary_locator.playwright changes say which locator is valid in which state.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'outline': the page as a map (regions with `selector`, structured tables, dialogs, alerts) in a few thousand characters; does not consume a snapshot id. | nodes |
| roles | No | Keep only nodes with these roles or tags (e.g. ['button','textbox','row']). | |
| scope | No | CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role="dialog"]'). Everything outside is skipped. | |
| max_nodes | No | Cap on extracted nodes; truncation is flagged, never silent. | |
| session_id | Yes | From session_open. | |
| diff_against | No | Return only what changed since that snapshot_id (or 'previous' = the last snapshot in this session) instead of the full extraction — added/removed/changed nodes plus the behavior observed in between. | |
| visible_only | No | Skip hidden nodes entirely. | |
| include_hidden | No | Keep hidden nodes flagged rather than dropping them. | |
| include_tables | No | Attach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions. | |
| max_output_chars | No | Budget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent). | |
| include_click_targets | No | Opt-in heuristic: also include cursor:pointer elements with content that match no other rule (JS-click product cards without anchors/roles/test-ids). Heuristic nodes carry a context_note. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, non-idempotent), and the description adds real behavioral context beyond them: truncation/omission is 'flagged, never silent', diffs are 'far smaller than a full re-extraction', the observed-behavior window between snapshots, and the diff-identity rules (frame+test-id → id → placeholder → tag+role+name) including that a renamed node appears as removed+added. The idempotentHint=false is implicitly explained by snapshot consumption. It stops short of auth/permission or rate-limit notes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and output shape, then the diff mode and its payoff, then the identity rules needed to interpret a diff. Dense but every sentence carries non-redundant information, and the long identity clause is the one piece that genuinely cannot be inferred from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 11 parameters and no output schema, the description carries the return-value burden and largely discharges it: it names the output shape, the diff payload, the snapshot_id field, and primary_locator.playwright state validity. It leaves the full node/table/dialog output structure to the reader's knowledge of extract_semantic_dom, which is a minor gap for a tool this rich.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema for the most consequential parameter: what diff_against actually returns (added/removed/changed nodes plus intervening behavior) and the identity rules that determine what counts as changed. It also clarifies that outline mode 'does not consume a snapshot id'. The remaining 9 parameters are left to their schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (snapshot/extract) on a specific resource (the CURRENT state of an open session), names the exact output shape (Semantic JSON) and explicitly differentiates it from the sibling extract_semantic_dom by the added `snapshot_id` and the diff capability. An agent can tell it apart from extract_semantic_dom/extract_semantic_dom_after without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage context: pass diff_against: 'previous' to get 'exactly the assertion list for the step', implying the verification-step workflow, and the schema documents outline mode as a low-cost alternative. It does not, however, state when to prefer this over extract_semantic_dom or extract_semantic_dom_after, so the alternative-selection guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_listARead-onlyIdempotent
Diagnostic: list open sessions (id, URL, expiry, counts) — recover a session id after losing context.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds value beyond them by enumerating what the response contains (id, URL, expiry, counts), which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the purpose and then the payoff, with no redundant or filler clauses. Every element (diagnostic framing, resource, return fields, use case) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only listing tool with no output schema, the description covers purpose, usage trigger, and the shape of returned data. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The parenthetical describes returned data rather than inputs, which is appropriate here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('list open sessions') plus the fields returned (id, URL, expiry, counts), and the 'Diagnostic' label frames intent. The verb/resource pair is distinct from session_open/session_close, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete trigger for use: 'recover a session id after losing context.' That is a clear usage context, but there is no statement of when not to use it or an explicit pointer to session_open/session_close as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_openA
Open a persistent browser session at a staging URL for a MULTI-STEP flow (login → cart → checkout). The page stays open across calls: use session_act to perform declared actions and session_extract to snapshot or diff, then session_close. Fresh context per session (storageState applied if configured). Sessions expire after an idle TTL and are capped in number; the allowlist is re-checked after every step.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page to extract. Must be http/https and on an allowlisted host. | |
| viewport | No | Viewport preset — 'mobile' is 375x812 with touch, for responsive states. | desktop |
| wait_for | No | Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages. | auto |
| wait_selector | No | Optional selector to await before extracting (for SPA content). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare openWorldHint=false, readOnlyHint=false, non-idempotent, non-destructive. The description adds substantial context beyond that: persistent-across-calls page, fresh context per session with storageState applied if configured, idle-TTL expiry, session count cap, and allowlist re-check after every step. These are non-obvious operational constraints an agent must know, though the description doesn't cover reset/cleanup semantics explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then the lifecycle (open → act → extract → close), then operational constraints. Every sentence carries a distinct fact; no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the lifecycle, state model (fresh context, storageState), lifecycle limits (TTL, cap), and security re-check — enough to call correctly without an output schema. Minor gaps remain (does the description note what session_open returns, e.g., a session id, and what triggers immediate expiration?), but for a tool with full schema coverage this is strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema already documents url, viewport, wait_for, and wait_selector with rich enum semantics. The description adds little parameter-level meaning (URL must be staging, allowlisted is echoed by schema), so baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (open a persistent browser session) with an explicit scope constraint (staging URL, MULTI-STEP flow). Clearly distinguishes itself from sibling session_act/session_extract/session_close by naming them as the follow-on workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use ('MULTI-STEP flow') and names the exact alternatives to chain (session_act to act, session_extract to snapshot/diff, then session_close). It also implies the boundary: this is the entry point of a multi-step session rather than a one-shot extract.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_verify_locatorsBRead-onlyIdempotent
verify_locators against an open session's current page (after acting, without a fresh navigation).
| Name | Required | Description | Default |
|---|---|---|---|
| locators | Yes | Playwright expressions as written in the test (getByRole(...), getByTestId(...), scoped forms, .nth(i)). | |
| session_id | Yes | From session_open. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that it operates on the current page without fresh navigation, which is useful behavioral context, but it does not explain what verification produces, what happens on failure, or how results are reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single concise sentence, front-loaded with the action and followed by key context. No wasted sentences. The opening term repeats the tool name, but the overall structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description carries the burden of explaining return values and what 'verify' yields (e.g., pass/fail per locator, resolved elements, errors). It omits this entirely, leaving an agent unsure how to interpret results, which is a significant gap for a verification tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the locators array documents Playwright expression syntax and session_id documents its source. The description adds no parameter-level meaning beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (verify_locators) and adds meaningful context (against an open session's current page, after acting, without a fresh navigation). However, it largely restates the tool name and never clarifies what 'verify' actually checks (existence, visibility, count, resolution), nor does it differentiate from the sibling tool named verify_locators.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(after acting, without a fresh navigation)' gives an implied timing condition for when this tool is appropriate, but it never names the alternative tool (e.g., verify_locators) or states explicit exclusions. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_locatorsARead-onlyIdempotent
Count every given Playwright expression against the live page (same engine that verified the extraction). Use after writing a test: paste the spec's getBy*/locator expressions and get matches, uniqueness, and the first matched element per expression, plus a summary. Also the drift check to run in CI against a page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page to extract. Must be http/https and on an allowlisted host. | |
| locators | Yes | Playwright expressions as written in the test (getByRole(...), getByTestId(...), scoped forms, .nth(i)). | |
| viewport | No | Viewport preset — 'mobile' is 375x812 with touch, for responsive states. | desktop |
| wait_for | No | Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages. | auto |
| wait_selector | No | Optional selector to await before extracting (for SPA content). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a safe, read-only, idempotent, non-open-world operation, so the bar is lower. The description adds real value beyond them by disclosing the return shape (matches, uniqueness, first matched element per expression, plus a summary) and the CI drift-check scenario.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then usage and a secondary CI scenario. Dense but every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing what the tool returns. Combined with fully documented parameters and safety annotations, an agent has enough to call it correctly, though the relationship to session_verify_locators remains unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters including url, locators, viewport, wait_for, and wait_selector are already documented with defaults, enums, and constraints. The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: counting Playwright expressions against a live page, and notes it uses the same engine that verified the extraction. The purpose is unambiguous, but it does not explicitly distinguish itself from the sibling session_verify_locators, leaving the agent to infer the difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use context: 'Use after writing a test' with instructions to paste the spec's getBy*/locator expressions, and a second use case as a CI drift check. No exclusions or explicit sibling comparison are provided, but the triggering conditions are concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.8.0- Changed
check_auth3 fields changed- changed
Input schema / properties / wait_for / defaultPrevious value: -"networkidle"New value: +"auto" - changed
Input schema / properties / wait_for / descriptionPrevious value: -"Navigation wait condition."New value: +"Navigation wait; see extract_semantic_dom." - changed
Input schema / properties / wait_for / enumPrevious value: -[ - "load", - "domcontentloaded", - "networkidle" -]New value: +[ + "auto", + "load", + "domcontentloaded", + "networkidle" +]
- Added
extract_outline - Changed
extract_semantic_dom8 fields changed- added
Input schema / properties / include_tablesAdded value: +{ + "default": false, + "description": "Attach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions.", + "type": "boolean" +} - added
Input schema / properties / max_output_charsAdded value: +{ + "description": "Budget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent).", + "exclusiveMinimum": 0, + "type": "integer" +} - added
Input schema / properties / rolesAdded value: +{ + "description": "Keep only nodes with these roles or tags (e.g. ['button','textbox','row']).", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / scopeAdded value: +{ + "description": "CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role=\"dialog\"]'). Everything outside is skipped.", + "type": "string" +} - added
Input schema / properties / visible_onlyAdded value: +{ + "default": false, + "description": "Skip hidden nodes entirely.", + "type": "boolean" +} - changed
Input schema / properties / wait_for / defaultPrevious value: -"networkidle"New value: +"auto" - changed
Input schema / properties / wait_for / descriptionPrevious value: -"Navigation wait condition."New value: +"Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages." - changed
Input schema / properties / wait_for / enumPrevious value: -[ - "load", - "domcontentloaded", - "networkidle" -]New value: +[ + "auto", + "load", + "domcontentloaded", + "networkidle" +]
- Changed
extract_semantic_dom_after10 fields changed- changed
Input schema / properties / actions / descriptionPrevious value: -"Declared actions executed in order in the MAIN frame after navigation."New value: +"Declared actions (fill/click/press/select/goto/wait) executed in order in the MAIN frame after navigation." - changed
Input schema / properties / actions / items / anyOfPrevious value: -[ - { - "additionalProperties": false, - "properties": { - "locator": { - "additionalProperties": false, - "properties": { - "nth": { - "description": "Optional .nth(i) index from the extraction's disambiguation guidance.", - "minimum": 0, - "type": "integer" - }, - "role": { - "description": "ARIA role — required when strategy is 'role'.", - "type": "string" - }, - "strategy": { - "description": "Locator strategy, matching the strategies in extraction output.", - "enum": [ - "test-id", - "role", - "label", - "placeholder", - "text", - "id", - "css" - ], - "type": "string" - }, - "value": { - "description": "The locator value (test id, accessible name, label, selector...).", - "minLength": 1, - "type": "string" - } - }, - "required": [ - "strategy", - "value" - ], - "type": "object" - }, - "type": { - "const": "fill", - "type": "string" - }, - "value": { - "type": "string" - } - }, - "required": [ - "type", - "locator", - "value" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "locator": { - "$ref": "#/properties/actions/items/anyOf/0/properties/locator" - }, - "type": { - "const": "click", - "type": "string" - } - }, - "required": [ - "type", - "locator" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "key": { - "maxLength": 30, - "type": "string" - }, - "locator": { - "$ref": "#/properties/actions/items/anyOf/0/properties/locator" - }, - "type": { - "const": "press", - "type": "string" - } - }, - "required": [ - "type", - "locator", - "key" - ], - "type": "object" - }, - { - "additionalProperties": false, - "properties": { - "ms": { - "exclusiveMinimum": 0, - "maximum": 10000, - "type": "integer" - }, - "type": { - "const": "wait", - "type": "string" - } - }, - "required": [ - "type", - "ms" - ], - "type": "object" - } -]New value: +[ + { + "additionalProperties": false, + "properties": { + "locator": { + "additionalProperties": false, + "properties": { + "nth": { + "description": "Optional .nth(i) index from the extraction's disambiguation guidance.", + "minimum": 0, + "type": "integer" + }, + "playwright": { + "description": "A `playwright` expression exactly as an extraction returned it (scoped forms and .nth included). When given, strategy/value are not needed.", + "minLength": 1, + "type": "string" + }, + "role": { + "description": "ARIA role — required when strategy is 'role'.", + "type": "string" + }, + "strategy": { + "default": "css", + "description": "Locator strategy, matching the strategies in extraction output (ignored when `playwright` is given).", + "enum": [ + "test-id", + "role", + "label", + "placeholder", + "text", + "id", + "css" + ], + "type": "string" + }, + "value": { + "default": "", + "description": "The locator value (test id, accessible name, label, selector...). Empty for a bare role inside `within`.", + "type": "string" + }, + "within": { + "additionalProperties": false, + "description": "Scope to a container first; copy the extraction's `within` verbatim.", + "properties": { + "kind": { + "enum": [ + "row", + "listitem", + "test-id", + "css" + ], + "type": "string" + }, + "value": { + "minLength": 1, + "type": "string" + } + }, + "required": [ + "kind", + "value" + ], + "type": "object" + } + }, + "type": "object" + }, + "secret": { + "description": "Mark the value as a secret: it is scrubbed from every string the server returns. Password fields are detected automatically.", + "type": "boolean" + }, + "type": { + "const": "fill", + "type": "string" + }, + "value": { + "type": "string" + } + }, + "required": [ + "type", + "locator", + "value" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "locator": { + "$ref": "#/properties/actions/items/anyOf/0/properties/locator" + }, + "type": { + "const": "click", + "type": "string" + } + }, + "required": [ + "type", + "locator" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "key": { + "maxLength": 30, + "type": "string" + }, + "locator": { + "$ref": "#/properties/actions/items/anyOf/0/properties/locator" + }, + "type": { + "const": "press", + "type": "string" + } + }, + "required": [ + "type", + "locator", + "key" + ], + "type": "object" + }, + { + "additionalProperties": false, + "description": "Choose a <select> option by value or label.", + "properties": { + "locator": { + "$ref": "#/properties/actions/items/anyOf/0/properties/locator" + }, + "type": { + "const": "select", + "type": "string" + }, + "value": { + "type": "string" + } + }, + "required": [ + "type", + "locator", + "value" + ], + "type": "object" + }, + { + "additionalProperties": false, + "description": "Navigate within the flow (must be http/https and allowlisted).", + "properties": { + "type": { + "const": "goto", + "type": "string" + }, + "url": { + "type": "string" + } + }, + "required": [ + "type", + "url" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "ms": { + "exclusiveMinimum": 0, + "maximum": 10000, + "type": "integer" + }, + "type": { + "const": "wait", + "type": "string" + } + }, + "required": [ + "type", + "ms" + ], + "type": "object" + } +] - added
Input schema / properties / include_tablesAdded value: +{ + "default": false, + "description": "Attach structured `tables` (headers, row identity, cells) and `dialogs` (label/value fields) inside the scope, for value assertions.", + "type": "boolean" +} - added
Input schema / properties / max_output_charsAdded value: +{ + "description": "Budget for the node list; nodes past it are dropped in document order and counted in `omitted` (never silent).", + "exclusiveMinimum": 0, + "type": "integer" +} - added
Input schema / properties / rolesAdded value: +{ + "description": "Keep only nodes with these roles or tags (e.g. ['button','textbox','row']).", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / scopeAdded value: +{ + "description": "CSS selector to extract within — take it from an outline region's `selector` (e.g. 'main table', '[role=\"dialog\"]'). Everything outside is skipped.", + "type": "string" +} - added
Input schema / properties / visible_onlyAdded value: +{ + "default": false, + "description": "Skip hidden nodes entirely.", + "type": "boolean" +} - changed
Input schema / properties / wait_for / defaultPrevious value: -"networkidle"New value: +"auto" - changed
Input schema / properties / wait_for / descriptionPrevious value: -"Navigation wait condition."New value: +"Navigation wait. 'auto' (default) waits for load, then until the DOM has been quiet for 500ms (max 6s) — works on SPAs that render after load and on pages whose analytics never let the network go idle. 'networkidle' times out on such pages." - changed
Input schema / properties / wait_for / enumPrevious value: -[ - "load", - "domcontentloaded", - "networkidle" -]New value: +[ + "auto", + "load", + "domcontentloaded", + "networkidle" +]
- Added
get_conventions - Changed
list_frames3 fields changed- changed
Input schema / properties / wait_for / defaultPrevious value: -"networkidle"New value: +"auto" - changed
Input schema / properties / wait_for / descriptionPrevious value: -"Navigation wait condition."New value: +"Navigation wait; see extract_semantic_dom." - changed
Input schema / properties / wait_for / enumPrevious value: -[ - "load", - "domcontentloaded", - "networkidle" -]New value: +[ + "auto", + "load", + "domcontentloaded", + "networkidle" +]
- Added
session_act - Added
session_close - Added
session_extract - Added
session_list - Added
session_open - Added
session_verify_locators - Added
verify_locators
4 tool updates
v0.4.0- First observed
check_auth - First observed
extract_semantic_dom - First observed
extract_semantic_dom_after - First observed
list_frames
TDQS
Scored across 13 tools
The set has a clear two-tier structure (one-shot extraction/verification vs. persistent session tools), and descriptions carefully separate them. However, extract_semantic_dom_after overlaps conceptually with the session_open/session_act/session_extract flow, and extract_semantic_dom vs session_extract require reading descriptions carefully to pick correctly.
Almost everything is consistent snake_case verb_noun, and the session_* prefix gives a clean, predictable grouping. Minor deviation: extract_semantic_dom_after uses a temporal suffix rather than a distinct verb, which slightly breaks the pattern.
13 tools is well within the sweet spot and each one earns its place, covering extraction, diagnostics, verification, conventions, and full session lifecycle without filler.
The surface covers the full Playwright test-authoring workflow: page mapping, snapshot extraction, post-interaction observation, locator verification, auth/frame diagnostics, conventions, and end-to-end session management with open/act/extract/verify/close/list. No obvious dead ends.
Maintenance
Related MCP Connectors
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
Headless browser primitives for AI agents when sites need real JS rendering.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables AI assistants to inspect, debug, and test web pages using Playwright. Provides comprehensive DOM inspection, visibility debugging, layout validation, and element finding capabilities in real browser environments.34434 npm5MIT
- AlicenseAqualityDmaintenanceEnables AI agents to understand web page structure and content through structured data extraction and element discovery using Playwright, eliminating the need for screenshots.47 npmMIT
- AlicenseAqualityAmaintenanceGives AI agents a compact, semantic interface to the browser, returning structured page snapshots with stable element IDs instead of raw DOM. Enables agents to navigate, interact, and extract information from web pages efficiently.626353 npm17MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to drive Playwright-based browser automation for UI testing, returning JSON/HTML reports with screenshots without server-side LLM or test scripts.1MIT