playwright-e2e-mcp
Runs Playwright end-to-end tests against Firefox: the browser argument selects Firefox (matched against the Playwright config project names), and Firefox can be launched headed or headless for test runs, failure diagnoses, live DOM inspection (inspect-page/validate-selector) and visual comparisons.
Reads the project's Git working tree (git status, falling back to git diff HEAD~1 and recent mtimes) to discover the agent's recent changes and extract the locators those files actually declare, so generated Playwright tests are built from the project's real selectors.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@playwright-e2e-mcprun my checkout e2e tests and tell me why they failed"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
playwright-e2e-mcp
An MCP server that lets AI agents run, debug, and inspect Playwright end-to-end tests — with structured results, actionable failure diagnostics, and live DOM inspection.
run-test ──▶ get-failure ──▶ inspect-page ──▶ validate-selector ──▶ fix ──▶ re-run
▲ │
└──────────────────────── list-tests ◀────────────────────────────────┘Instead of handing an agent raw Playwright output, this server turns every run into
machinable results: pass/fail stats, per-failure messages with file:line, a failure
kind (assertion, timeout, browser crash, syntax error, dead dev server, full disk…),
and a concrete "how to fix" hint. When a test fails because a selector no longer
matches, the agent can open the live page in a headless browser, see the real DOM
with unique CSS selectors, and validate the replacement selector before re-running.
Tools
Tool | Purpose |
| Run Playwright tests and return stats, failures, diagnostics and hints |
| Deep analysis of one failure: stack, expected/actual, DOM snapshot at failure (from the Playwright trace), next steps |
| Open a URL headlessly and return the rendered DOM: selectors, visibility, boxes, text, console output, HTML |
| List available tests ( |
| Check a CSS selector against a live page: validity, match count, sample matches |
| Scaffold a Playwright test from a description using the project's real selectors, discovered from recent file changes |
| Visual regression: screenshot before/after a change and report what moved and how colors shifted |
| Run a failing test 2–10 times with retries disabled and return an evidence verdict: |
run-test
Argument | Type | Description |
| string | Project directory (default: server working directory) |
| string[] | Files/directories relative to the root; |
| string | Only run tests whose title matches this regex |
|
| Playwright project to run (matched against config project names) |
| boolean | Visible browser window |
| number | Hard wall-clock limit for the run (default |
| number | Per-test timeout passed to Playwright |
| number | Passed through to Playwright |
| string |
|
| boolean | Auto-retry failures once before reporting them (default |
| boolean | Only re-run tests that failed in the previous run (Playwright |
| string[] | Extra CLI flags (shell metacharacters are rejected) |
Flakiness handling: by default the server injects --retries=1 (unless the config
already sets retries), so a test that passes on the retry is reported as flaky,
not failed. Traces are captured automatically (--trace=retain-on-failure) so
get-failure can show the DOM at the moment of failure.
Example result:
## Playwright run — ❌ FAILED
**Command:** `playwright test --config playwright.config.ts tests/checkout.spec.ts --reporter=json`
**duration 4.2s · exit 1 · config `playwright.config.ts`**
| passed | failed | flaky | skipped | duration |
| ---: | ---: | ---: | ---: | ---: |
| 0 | 1 | 0 | 0 | 1.1s |
### ❌ 1 failing test(s)
### 1 of 1. checkout.spec.ts › pays with card
**File:** `checkout.spec.ts:5` | **failed · server-unreachable**
### ⚠️ SERVER_NOT_RUNNING
Your app (dev server) does not appear to be reachable. Start it in another terminal
(e.g. npm run dev / npm start), keep it running, then retry — or configure `webServer`
in playwright.config.* so Playwright starts it automatically.get-failure
Argument | Type | Description |
| number | 1-based failure index from the last run (default |
| string | Only used when re-reading the stored report |
Returns the message/code frame, expected vs actual, stack, failure kind with a
diagnosis, the test's console output, the DOM snapshot from the Playwright trace
(plus the failed action, its selector, and the action log leading up to it),
the network requests that failed (4xx/5xx, dead endpoints, no-response — with
method, URL, status and resource type), the console errors/warnings the page
logged before the failure, and
numbered next steps (re-run this single test by file:line, headed/debug mode,
validate-selector when the message mentions a locator, …).
inspect-page
Argument | Type | Description |
| string | Full http(s) URL to open (required) |
| string | Project whose Playwright launches the browser |
| string | Inspect matches of this CSS selector instead of the whole DOM |
| string | Wait for a selector (CSS or |
|
| Navigation wait condition |
| boolean | Include the rendered HTML (capped) |
| number | HTML cap, default |
| number | Overall limit, default |
Returns each element's unique CSS selector, tag, visibility, bounding box, text and attributes, plus captured console messages (errors first).
list-tests
Argument | Type | Description |
| string | Project directory |
| string | Config path or 1-based index |
| string | Restrict scanning to a directory (must stay inside the project) |
| string | Case-insensitive substring filter on |
| number | Max tests returned, default |
Uses playwright test --list when Playwright works, and falls back to a source scan
(keeping the reason) when the install or a spec file is broken.
validate-selector
Argument | Type | Description |
| string | Live page to test against (required) |
| string | CSS selector to validate (required) |
| string | Project whose Playwright launches the browser |
| number | Overall limit, default |
Verdicts: ✅ VALID — N matches (with a sample of matches), ✅ VALID — 0 matches
(with debugging advice), ❌ INVALID (parse error + fix), or a warning when the input
uses a Playwright-only engine (text=, xpath=, >>, :has-text()), which is not
plain CSS.
generate-e2e-test
Argument | Type | Description |
| string | What the test should cover (required) |
| string | Page the test starts on (default: |
| string | Where to write the spec (default: detected |
| boolean | Write the file to disk (default |
| boolean | Replace an existing file at the target path |
| boolean | Cross-check selectors against the live page (default on when a URL is known) |
| string | As with the other tools |
Reads the agent's recent changes (git status, falling back to git diff HEAD~1,
then recent mtimes), extracts the locators those files actually declare
(data-testid, getByRole, aria-label, placeholder, id, name, element text),
ranks verified-live selectors first, writes a spec built from them, and reports each
selector with its source file:line.
compare-visual-state
Argument | Type | Description |
| string | Page to capture (required) |
| string | Baseline id, e.g. |
|
|
|
| string | Capture just this element |
| boolean | Capture the full scrollable page |
| number | Percent of pixels that may differ (default |
| number | Per-pixel channel delta considered different (default |
| — | As with |
The first call saves a baseline under .pw-mcp/visual/ (add that to .gitignore, or
commit it for CI comparisons). Later calls report changed-pixel counts, merged
regions ((x, y) 120×40 — 1,200 px), the average color shift ("blue → red"),
and write a red-highlighted diff image for review.
diagnose-flaky
Argument | Type | Description |
| string[] | Tests to diagnose ( |
| number | Times to run them, |
| — | As with |
| number | Hard wall-clock limit per run (default |
| string | Project directory |
Every run executes with --retries=0 and auto-retry disabled, so each result is
honest evidence. The response contains a per-run table (status, duration, first
failure), the count of distinct normalized error signatures, and one of:
❌ CONSISTENTLY FAILING — failed every run (same error → reproducible bug, different errors → still broken, just noisy). Fix it; it is not flaky.
⚠️ FLAKY — some runs passed. Includes
N of Mcounts and whether the failures share one signature (real intermittent bug) or vary (timing/environment instability).✅ NOT REPRODUCING — passed every re-run; the original failure was one-off.
The last run is stored, so get-failure can analyze it immediately afterwards.
Related MCP server: Limetest MCP Server
Installation
Requirements:
Node.js ≥ 20 (the server is built on MCP SDK v2 — the
2026-07-28spec line)A project with
@playwright/testinstalled and browsers available (npx playwright install chromium)
No npm account needed — install straight from GitHub (the prepare script
builds dist/ automatically on install):
npx -y github:trajectiq-ai/E2E
npm install -D github:trajectiq-ai/E2E @playwright/test # or as a project dependency
npx playwright install chromiumOr grab the packaged tarball from the repo's GitHub Releases page and install it locally:
npm install -D https://github.com/trajectiq-ai/E2E/releases/download/v0.1.0/playwright-e2e-mcp-0.1.0.tgzListed in the official MCP Registry as
io.github.trajectiq-ai/E2E — registry-aware clients discover it there, and every
v* release tag republishes the entry from CI via server.json.
MCP client configuration
Claude Code / generic (project-scoped):
{
"mcpServers": {
"playwright-e2e": {
"command": "npx",
"args": ["-y", "github:trajectiq-ai/E2E"],
"env": { "PW_MCP_PROJECT_ROOT": "/absolute/path/to/your/project" }
}
}
}Claude Desktop / Cursor / Windsurf: add the same block to their MCP config file.
The server uses its working directory as the project root; set PW_MCP_PROJECT_ROOT
when the client launches it somewhere else (e.g. your home directory).
Codex / VS Code / Copilot CLIs:
codex mcp add playwright-e2e -- npx -y github:trajectiq-ai/E2E
code --add-mcp '{"name":"playwright-e2e","command":"npx","args":["-y","github:trajectiq-ai/E2E"]}'Claude Desktop (one-click): download and double-click the .mcpb Desktop
Extension attached to the latest release —
the bundle ships its own dependencies, so no Node setup is required.
Claude Code:
claude mcp add playwright-e2e -- npx -y github:trajectiq-ai/E2EGemini CLI / Qwen Code: paste the mcpServers block above into
.gemini/settings.json (Qwen Code: .qwen/settings.json) — both speak the same
MCP settings format.
All tools ship MCP tool annotations (readOnlyHint, destructiveHint,
idempotentHint, openWorldHint), so clients can show accurate safety prompts
before running anything.
From a local checkout:
{
"mcpServers": {
"playwright-e2e": {
"command": "node",
"args": ["/path/to/playwright-e2e-mcp/dist/index.js"],
"env": { "PW_MCP_PROJECT_ROOT": "/path/to/your/project" }
}
}
}Configuration
Environment variable | Default | Purpose |
| server cwd | Default project root for every tool |
|
|
|
|
|
|
Logs always go to stderr — stdout is reserved for the MCP protocol.
Typical workflow
generate-e2e-test{ "description": "checkout with a saved card" }— scaffolds a spec from your real selectors (skipped if you write the test yourself).list-tests— see what exists (tests/checkout.spec.ts:5 checkout › pays with card).run-test{ "testFiles": ["tests/checkout.spec.ts"] }— run it; get stats + failures (flaky tests are auto-retried once before being called failures).get-failure{ "index": 1 }— read the code frame, expected/actual, the DOM snapshot at failure from the trace, the failed network requests, the page's console errors, and next steps.If it looks selector-related:
inspect-page{ "url": "http://localhost:3000/checkout" }to see the real DOM, thenvalidate-selectorto prove the replacement selector works.After changing CSS/components:
compare-visual-state{ "url": "…", "name": "checkout" }to catch unintended visual regressions.If a failure looks intermittent:
diagnose-flaky{ "runs": 3 }— get the evidence verdict (flaky vs consistently broken) before deciding what to fix.Fix the spec or the app, then re-run only what failed:
run-test{ "lastFailed": true }, and repeat until green.
Edge cases handled
Situation | Behaviour |
No Playwright installed |
|
Dev server not running | Failure classified |
Test exceeds | Process group is killed (SIGINT→SIGKILL on POSIX, |
Browser crashes | Classified |
Windows backslashes | All paths normalized lexically ( |
Flaky tests | Failing tests are automatically retried once ( |
Trace/DOM context |
|
Slow re-runs after a fix |
|
Several | Returns a numbered menu ( |
Syntax error in a spec |
|
MCP client disconnects | Per-request |
Disk full |
|
Malicious paths |
|
Security notes
No shell: Playwright is spawned as
node <playwright/cli.js> …with an argument array — no command interpolation.Path sandbox: user paths are resolved lexically and must stay inside the project root.
Cleanup: temp report/script files are written to the OS temp dir and removed; child processes are tracked and killed on shutdown.
Development
src/
├── index.ts # bin entry point (--version/--help, main-module guard)
├── server.ts # McpServer setup, tool registration, shutdown handling
├── tools/ # the eight tools + shared plumbing
├── utils/ # playwright-runner, report-parser, project-detector, path-utils,
│ # logger, trace-reader (trace.zip → DOM/network/console),
│ # image-diff (PNG codec + pixel diff), change-analyzer
└── types/ # shared interfaces and the ErrorKind taxonomyBuilt on @modelcontextprotocol/server v2 (the 2026-07-28 MCP spec line) with
Zod v4 standard schemas; every tool declares spec tool annotations.
npm install
npm run build # tsc → dist/ (zero errors)
npm test # build + test/run-tests.mjs (70 unit tests, any Node ≥20)
npm run e2e # build + e2e/run.mjs: live MCP ↔ Playwright integration suiteTests cover the report parser (sample Playwright JSON, trace attachments), path utils
(Windows and macOS paths, sandboxing), the project detector (temp-dir fixtures: config
discovery, multiple configs, missing install, test-file scanning), the shared tool
helpers, the trace reader (synthetic trace.zip: error, failed action, DOM snapshot,
*.network failed-request parsing, console error/warning events), the image diff
(PNG round-trip, regions, color shift, dimension changes), the change analyzer
(selector extraction, git + mtime paths) and the flaky verdict logic
(failure signatures, CONSISTENTLY FAILING / FLAKY / NOT REPRODUCING / NO TESTS RAN).
Integration suite (npm run e2e)
Unit tests prove the logic; the integration suite proves the loop. It boots the real
server over stdio against a live fixture app and a Playwright project under
e2e/fixture/, then drives it exactly like an MCP client and asserts ~40 behaviours
that only appear end-to-end:
initialize handshake, 8 tools, spec tool annotations and object input schemas,
live DOM inspection, CSS selector validation (matches, zero matches, engine syntax, parse errors), dead-server detection,
visual regression: baseline → unchanged compare →
blue → reddiff detection,run-testpass/fail/lastFailedstats and meta lines,get-failuretrace diagnostics: DOM at failure, the 404 network request, theconsole.errormessage, diagnosis and next steps,auto-retry turning a first-run failure into
PASSED (1 flaky),diagnose-flakyverdict FLAKY (2 of 3 runs) with retries disabled,generate-e2e-testwriting its scaffold, plus error paths (missing test path, unknown tool, unreachable server).
First run needs the browser once: npx playwright install chromium.
CI runs the suite on Ubuntu and Windows (see .github/workflows/ci.yml).
Try the example
With Playwright installed in your project:
npx playwright test examples/sample-test.spec.tsor ask your agent to call run-test with
"testFiles": ["examples/sample-test.spec.ts"] — it hits the public
example.com page, so it verifies browsers, network and the MCP
pipeline in one shot.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Approved test intent, reviewed Playwright automation and run evidence, inside your editor.
Direct access to Cypress tests results and accessibility reports in your AI workflow.
Agentic testing: HyperExecute jobs, test failure triage, SmartUI visual diffs, a11y audits
Discover Playwright workflows, start runs, and inspect results in Playrunner Cloud.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables AI assistants to inspect, debug, and test web pages using Playwright. Provides comprehensive DOM inspection, visibility debugging, layout validation, and element finding capabilities in real browser environments.34109 npm5MIT
- AlicenseNot gradedqualityCmaintenanceEnables automated end-to-end testing powered by Playwright where test cases are defined in natural language and executed by AI. Uses lightweight snapshot analysis with vision mode fallback for sophisticated testing scenarios.3Apache 2.0
- AlicenseAqualityDmaintenanceRun real Playwright E2E tests from your AI coding agent.413 npmMIT
- FlicenseAqualityDmaintenanceEnables AI agents to run reusable Playwright test fixtures against live deployments, providing structured test results, screenshots, and assertions.3-