playwright-e2e-mcp
MCP serverWeb & browsingRun, debug and inspect Playwright E2E tests from any AI agent: diagnostics, live DOM, selectors.
Use playwright-e2e-mcp in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add playwright-e2e-mcp and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use playwright-e2e-mcp
No other account needed.
Details
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Install playwright-e2e-mcp
The server’s own address, for the clients that take one directly. Or connect ahel once and every client you use reads it from one address, with the account kept on ahel rather than in each client’s config.
Claude Code
claude mcp add --transport http --scope user playwright-e2e-mcp 'https://playwright-e2e-mcp.vercel.app/api/mcp'Run it once in your project, then open /mcp to approve any sign-in the server asks for.
Claude Desktop
https://playwright-e2e-mcp.vercel.app/api/mcpAdd a custom connector in Settings, paste this address, and approve the sign-in.
Cursor
cursor://anysphere.cursor-deeplink/mcp/install?name=playwright-e2e-mcp&config=eyJ1cmwiOiJodHRwczovL3BsYXl3cmlnaHQtZTJlLW1jcC52ZXJjZWwuYXBwL2FwaS9tY3AifQ==Open the link and Cursor adds the server at that address.
ChatGPT
https://playwright-e2e-mcp.vercel.app/api/mcpIn Settings, enable Developer mode, create an MCP app, and paste this address. Your plan and workspace must allow custom apps.
Codex
codex mcp add playwright-e2e-mcp --url 'https://playwright-e2e-mcp.vercel.app/api/mcp'Run it once, then sign in with codex mcp login playwright-e2e-mcp if the server asks for an account.
From the project's README
As published by trajectiq-ai/E2E in README.md.
An MCP server that lets AI agents run, debug, and inspect Playwright end-to-end tests — with structured results, actionable failure diagnostics, and live DOM inspection.
run-test ──▶ get-failure ──▶ inspect-page ──▶ validate-selector ──▶ fix ──▶ re-run
▲ │
└──────────────────────── list-tests ◀────────────────────────────────┘
Instead of handing an agent raw Playwright output, this server turns every run into
machinable results: pass/fail stats, per-failure messages with file:line, a failure
kind (assertion, timeout, browser crash, syntax error, dead dev server, full disk…),
and a concrete "how to fix" hint. When a test fails because a selector no longer
matches, the agent can open the live page in a headless browser, see the real DOM
with unique CSS selectors, and validate the replacement selector before re-running.
Demo
Hosted endpoint — what an initialize + tools/list round-trip against
https://playwright-e2e-mcp.vercel.app/api/mcp returns for a client that sends the
bearer token (without one, only list-tests and get-failure are listed):
A real test run — run-test served over stdio by npx -y playwright-e2e-mcp
against the bundled examples/sample-test.spec.ts (actual output, unedited):
Images are rendered with node scripts/gen-demo-images.mjs: the run-test card is real
captured output; the endpoint card is an illustration of the authenticated listing.
Install
Works with Claude Desktop, Claude Code, Cursor, Windsurf, Codex, Gemini CLI, Freebuff and every other MCP client — pick whichever route fits:
| Route | How |
|---|---|
| npm (canonical, fastest) | npx -y playwright-e2e-mcp |
| MCP Registry (registry-aware clients discover it automatically) | io.github.trajectiq-ai/E2E — listing |
| Any client, no npm account needed | npx -y github:trajectiq-ai/E2E#v0.1.2 (pin a release tag) |
| Claude Desktop, zero Node setup | double-click the .mcpb extension |
| Remote-only clients (ChatGPT connectors) | https://playwright-e2e-mcp.vercel.app/api/mcp |
Details and per-client config: Installation · MCP client configuration.
Tools
| Tool | Purpose |
|---|---|
run-test | Run Playwright tests and return stats, failures, diagnostics and hints |
get-failure | Deep analysis of one failure: stack, expected/actual, DOM snapshot at failure (from the Playwright trace), next steps |
inspect-page | Open a URL headlessly and return the rendered DOM: selectors, visibility, boxes, text, console output, HTML |
list-tests | List available tests (file, line, full title, projects) with filtering |
validate-selector | Check a CSS selector against a live page: validity, match count, sample matches |
generate-e2e-test | Scaffold a Playwright test from a description using the project's real selectors, discovered from recent file changes |
compare-visual-state | Visual regression: screenshot before/after a change and report what moved and how colors shifted |
diagnose-flaky | Run a failing test 2–10 times with retries disabled and return an evidence verdict: CONSISTENTLY FAILING, FLAKY or NOT REPRODUCING |
run-test
| Argument | Type | Description |
|---|---|---|
projectRoot | string | Project directory inside the configured root (default: server working directory) |
testFiles | string[] | Files/directories relative to the root; file:line supported. Omit to run everything |
grep | string | Only run tests whose title matches this regex |
browser | chromium | firefox | webkit | Playwright project to run (matched against config project names) |
headed | boolean | Visible browser window |
timeoutMs | number | Hard wall-clock limit for the run (default 120000); the whole process tree is killed past it and partial results are returned |
testTimeoutMs | number | Per-test timeout passed to Playwright |
workers / retries | number | Passed through to Playwright |
config | string | playwright.config path or 1-based index when the project has several |
retryOnFailure | boolean | Auto-retry failures once before reporting them (default true; ignored when retries is set) |
lastFailed | boolean | Only re-run tests that failed in the previous run (Playwright --last-failed) — the fast fix → re-run loop |
args | string[] | Extra Playwright flags from an allowlist (--repeat-each=N, --max-failures=N, --update-snapshots, --shard=1/3, --trace=on, …); values go after = and are checked, and flags that take a path, such as --config or --output, are rejected |
Flakiness handling: by default the server injects --retries=1 (unless the config
already sets retries), so a test that passes on the retry is reported as flaky,
not failed. Traces are captured automatically (--trace=retain-on-failure) so
get-failure can show the DOM at the moment of failure.
Example result:
## Playwright run — ❌ FAILED
**Command:** `playwright test --config playwright.config.ts tests/checkout.spec.ts --reporter=json`
**duration 4.2s · exit 1 · config `playwright.config.ts`**
| passed | failed | flaky | skipped | duration |
| ---: | ---: | ---: | ---: | ---: |
| 0 | 1 | 0 | 0 | 1.1s |
### ❌ 1 failing test(s)
### 1 of 1. checkout.spec.ts › pays with card
**File:** `checkout.spec.ts:5` | **failed · server-unreachable**
### ⚠️ SERVER_NOT_RUNNING
Your app (dev server) does not appear to be reachable. Start it in another terminal
(e.g. npm run dev / npm start), keep it running, then retry — or configure `webServer`
in playwright.config.* so Playwright starts it automatically.
get-failure
| Argument | Type | Description |
|---|---|---|
index | number | 1-based failure index from the last run (default 1) |
projectRoot | string | Only used when re-reading the stored report |
Returns the message/code frame, expected vs actual, stack, failure kind with a
diagnosis, the test's console output, the DOM snapshot from the Playwright trace
(plus the failed action, its selector, and the action log leading up to it),
the network requests that failed (4xx/5xx, dead endpoints, no-response — with
method, URL, status and resource type), the console errors/warnings the page
logged before the failure, and
numbered next steps (re-run this single test by file:line, headed/debug mode,
validate-selector when the message mentions a locator, …).
inspect-page
| Argument | Type | Description |
|---|---|---|
url | string | Full http(s) URL to open (required) |
projectRoot | string | Project whose Playwright launches the browser |
selector | string | Inspect matches of this CSS selector instead of the whole DOM |
waitFor | string | Wait for a selector (CSS or text=…) before inspecting |
waitUntil | load | domcontentloaded | networkidle | Navigation wait condition |
includeHtml | boolean | Include the rendered HTML (capped) |
maxHtmlChars | number | HTML cap, default 20000 |
timeoutMs | number | Overall limit, default 45000 |
Returns each element's unique CSS selector, tag, visibility, bounding box, text and attributes, plus captured console messages (errors first).
list-tests
| Argument | Type | Description |
|---|---|---|
projectRoot | string | Project directory |
config | string | Config path or 1-based index |
testDir | string | Restrict scanning to a directory (must stay inside the project) |
filter | string | Case-insensitive substring filter on file › title |
limit | number | Max tests returned, default 500 |
Uses playwright test --list when Playwright works, and falls back to a source scan
(keeping the reason) when the install or a spec file is broken.
validate-selector
| Argument | Type | Description |
|---|---|---|
url | string | Live page to test against (required) |
selector | string | CSS selector to validate (required) |
projectRoot | string | Project whose Playwright launches the browser |
timeoutMs | number | Overall limit, default 45000 |
Verdicts: ✅ VALID — N matches (with a sample of matches), ✅ VALID — 0 matches
(with debugging advice), ❌ INVALID (parse error + fix), or a warning when the input
uses a Playwright-only engine (text=, xpath=, >>, :has-text()), which is not
plain CSS.
generate-e2e-test
| Argument | Type | Description |
|---|---|---|
description | string | What the test should cover (required) |
pageUrl | string | Page the test starts on (default: baseURL / webServer.url from config) |
testDir / file | string | Where to write the spec (default: detected testDir + generated/<slug>.spec.ts); file must end in .spec.* or .test.* |
write | boolean | Write the file to disk (default true) |
overwrite | boolean | Replace an existing spec at the target path; only specs this tool generated can be replaced |
liveInspect | boolean | Cross-check selectors against the live page (default on when a URL is known) |
projectRoot / config | string | As with the other tools |
Reads the agent's recent changes (git status, falling back to git diff HEAD~1,
then recent mtimes), extracts the locators those files actually declare
(data-testid, getByRole, aria-label, placeholder, id, name, element text),
ranks verified-live selectors first, writes a spec built from them, and reports each
selector with its source file:line.
compare-visual-state
| Argument | Type | Description |
|---|---|---|
url | string | Page to capture (required) |
name | string | Baseline id, e.g. checkout-page (letters, digits, . _ -) |
action | compare | baseline | compare (default) diffs; baseline re-captures the reference |
selector | string | Capture just this element |
fullPage | boolean | Capture the full scrollable page |
tolerance | number | Percent of pixels that may differ (default 0.1) |
pixelThreshold | number | Per-pixel channel delta considered different (default 60) |
waitUntil / waitFor / timeoutMs | — | As with inspect-page |
The first call saves a baseline under .pw-mcp/visual/ (add that to .gitignore, or
commit it for CI comparisons). Later calls report changed-pixel counts, merged
regions ((x, y) 120×40 — 1,200 px), the average color shift ("blue → red"),
and write a red-highlighted diff image for review.
diagnose-flaky
| Argument | Type | Description |
|---|---|---|
testFiles | string[] | Tests to diagnose (file:line supported). Defaults to the tests that failed in the most recent run |
runs | number | Times to run them, 2–10 (default 3) |
browser / headed / workers / config | — | As with run-test |
timeoutMs | number | Hard wall-clock limit per run (default 120000) |
projectRoot | string | Project directory |
Every run executes with --retries=0 and auto-retry disabled, so each result is
honest evidence. The response contains a per-run table (status, duration, first
failure), the count of distinct normalized error signatures, and one of:
- ❌ CONSISTENTLY FAILING — failed every run (same error → reproducible bug, different errors → still broken, just noisy). Fix it; it is not flaky.
- ⚠️ FLAKY — some runs passed. Includes
N of Mcounts and whether the failures share one signature (real intermittent bug) or vary (timing/environment instability). - ✅ NOT REPRODUCING — passed every re-run; the original failure was one-off.
The last run is stored, so get-failure can analyze it immediately afterwards.
Installation
Requirements:
- Node.js ≥ 20 (the server is built on MCP SDK v2 — the
2026-07-28spec line) - A project with
@playwright/testinstalled and browsers available (npx playwright install chromium)
No npm account needed — install straight from GitHub (the prepare script
builds dist/ automatically on install). Pin a release tag: an unpinned
github:trajectiq-ai/E2E runs whatever is on the default branch at that moment.
npx -y github:trajectiq-ai/E2E#v0.1.2
npm install -D github:trajectiq-ai/E2E#v0.1.2 @playwright/test # or as a project dependency
npx playwright install chromium
Or grab the packaged tarball from the repo's GitHub Releases page and install it locally:
npm install -D https://github.com/trajectiq-ai/E2E/releases/download/v0.1.2/playwright-e2e-mcp-0.1.2.tgz
Listed in the official MCP Registry as
io.github.trajectiq-ai/E2E — registry-aware clients discover it there, and every
v* release tag republishes the entry from CI via server.json.
MCP client configuration
Claude Code / generic (project-scoped):
{
"mcpServers": {
"playwright-e2e": {
"command": "npx",
"args": ["-y", "github:trajectiq-ai/E2E#v0.1.2"],
"env": { "PW_MCP_PROJECT_ROOT": "/absolute/path/to/your/project" }
}
}
}
Claude Desktop / Cursor / Windsurf: add the same block to their MCP config file.
The server uses its working directory as the project root; set PW_MCP_PROJECT_ROOT
when the client launches it somewhere else (e.g. your home directory).
Codex / VS Code / Copilot CLIs:
codex mcp add playwright-e2e -- npx -y github:trajectiq-ai/E2E#v0.1.2
code --add-mcp '{"name":"playwright-e2e","command":"npx","args":["-y","github:trajectiq-ai/E2E#v0.1.2"]}'
Codex's defaults fight this server: the first launch clones the repo and runs
tsc (measured 30 s on a cold npx cache, against a 10 s
startup_timeout_sec default), and a Playwright run with retries beats the 60 s
tool_timeout_sec default. Raise both in ~/.codex/config.toml:
[mcp_servers.playwright-e2e]
command = "npx"
args = ["-y", "github:trajectiq-ai/E2E#v0.1.2"]
startup_timeout_sec = 60
tool_timeout_sec = 600
Claude Desktop (one-click): download and double-click the .mcpb Desktop
Extension attached to the latest release —
the bundle ships its own dependencies, so no Node setup is required. On install it
asks you to pick your project root (required, no default: choose the project
folder, not your home directory) and wires it
into PW_MCP_PROJECT_ROOT, so the tools point at a real project from the first call.
Claude Code:
claude mcp add playwright-e2e -- npx -y github:trajectiq-ai/E2E#v0.1.2
Gemini CLI / Qwen Code: paste the mcpServers block above into
.gemini/settings.json (Qwen Code: .qwen/settings.json) — both speak the same
MCP settings format.
Freebuff / Codebuff (project-scoped): this repo ships a committed
.agents/mcp.json, so opening the checkout in Freebuff
attaches the server workspace-wide — no global config needed. Your own projects
can do the same: drop an mcp.json with the block above into their .agents/
directory. Freebuff asks you to trust a repository's .agents/ on first run.
All tools ship MCP tool annotations (readOnlyHint, destructiveHint,
idempotentHint, openWorldHint), so clients can show accurate safety prompts
before running anything.
From a local checkout:
{
"mcpServers": {
"playwright-e2e": {
"command": "node",
"args": ["/path/to/playwright-e2e-mcp/dist/index.js"],
"env": { "PW_MCP_PROJECT_ROOT": "/path/to/your/project" }
}
}
}
Hosted endpoint (ChatGPT & remote clients)
Some clients — ChatGPT custom connectors especially — only accept remote
HTTPS MCP servers and refuse to spawn a local npx process. This repo ships a
Streamable HTTP bridge for exactly that case:
| Endpoint | https://playwright-e2e-mcp.vercel.app/api/mcp |
| Transport | MCP Streamable HTTP (POST JSON in, JSON or SSE out) |
| Auth | optional bearer token (PW_MCP_HTTP_TOKEN); without one only read-only tools are served |
| Source | api/mcp.ts → src/http.ts |
The bridge runs the same createServer() as the stdio transport; the SDK
serves every request with a fresh server instance, which is what a serverless
function wants. test/http-bridge.test.mjs drives the real Node adapter over
node:http so a broken bridge fails in CI, not in ChatGPT.
Open vs. token-protected. When the deployment has no PW_MCP_HTTP_TOKEN,
anyone can reach the URL, so the bridge serves only list-tests and
get-failure, in restricted mode: list-tests scans sources instead of running
playwright test --list (which would execute the project's config), callers
cannot pick another projectRoot, and nothing spawns a process, drives a browser
or writes a file. Set PW_MCP_HTTP_TOKEN (at least 16 characters; use a random
value) to serve all eight tools to clients that send Authorization: Bearer <token>;
other requests get 401. Child processes started over HTTP get only an allowlisted
environment. PW_MCP_ALLOWED_HOSTS (comma separated, * for any) limits the
accepted Host header; without a token and without that variable, only
localhost names and the deployment's own Vercel hostnames are accepted
(DNS-rebinding protection).
Add it to ChatGPT: Settings → Connectors → turn on Advanced → Developer mode → Create custom connector → paste the endpoint above → authentication None (read-only tools).
Codex can also take the remote transport instead of spawning npx, if you'd
rather not ship Playwright to every machine:
codex mcp add playwright-e2e-remote --url https://playwright-e2e-mcp.vercel.app/api/mcp
What to expect: list-tests works and reports the specs bundled with the
deployment. Even with a token, tools that spawn a browser (run-test,
inspect-page, validate-selector, diagnose-flaky, …) cannot download
Chromium in a serverless function, so they return their normal NO_PLAYWRIGHT
hint. Use the stdio install for real runs; the hosted endpoint is for discovery
and for clients that cannot run local processes.
# verify the handshake without any client
curl -X POST https://playwright-e2e-mcp.vercel.app/api/mcp \
-H 'content-type: application/json' \
-H 'accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"1.0"}}}'
Redeploy after a change: merge to main. Vercel's Git integration deploys
every push, so there is no token to manage and no CLI step — watch the
Vercel commit status for the deployment result.
Configuration
| Environment variable | Default | Purpose |
|---|---|---|
PW_MCP_PROJECT_ROOT | server cwd | Default project root for every tool |
PW_MCP_ALLOWED_ROOTS | — | Extra directories a caller may pass as projectRoot (:-separated, ; on Windows). Anything outside these and the default root is rejected |
PW_MCP_HTTP_TOKEN | — | HTTP bridge only: bearer token (16+ characters) that unlocks all tools (see above) |
PW_MCP_ALLOWED_HOSTS | localhost + Vercel hostnames when there is no token | HTTP bridge only: comma-separated Host allowlist; * accepts any |
PW_MCP_PASSTHROUGH_ENV | — | HTTP bridge only: comma-separated extra variables passed to test runs (e.g. BASE_URL) |
PW_MCP_MAX_CHILDREN | 4 | HTTP bridge only: how many test runs and browser probes may run at once |
PW_MCP_BLOCK_PRIVATE_URLS | off (on for the HTTP bridge) | 1 makes the URL tools refuse loopback, private-network and cloud-metadata addresses, including redirects and subresources |
LOG_LEVEL | info | debug | info | warn | error | silent |
LOG_FORMAT | text | text or json (structured) |
Logs always go to stderr — stdout is reserved for the MCP protocol.
Typical workflow
generate-e2e-test{ "description": "checkout with a saved card" }— scaffolds a spec from your real selectors (skipped if you write the test yourself).list-tests— see what exists (tests/checkout.spec.ts:5 checkout › pays with card).run-test{ "testFiles": ["tests/checkout.spec.ts"] }— run it; get stats + failures (flaky tests are auto-retried once before being called failures).get-failure{ "index": 1 }— read the code frame, expected/actual, the DOM snapshot at failure from the trace, the failed network requests, the page's console errors, and next steps.- If it looks selector-related:
inspect-page{ "url": "http://localhost:3000/checkout" }to see the real DOM, thenvalidate-selectorto prove the replacement selector works. - After changing CSS/components:
compare-visual-state{ "url": "…", "name": "checkout" }to catch unintended visual regressions. - If a failure looks intermittent:
diagnose-flaky{ "runs": 3 }— get the evidence verdict (flaky vs consistently broken) before deciding what to fix. - Fix the spec or the app, then re-run only what failed:
run-test{ "lastFailed": true }, and repeat until green.
Edge cases handled
Shortened here. Read the whole README on GitHub.
Advanced
- Delivery
- E2E MCP server → your ahel connector (mcp.ahel.ai) → your AI.
- Item type
- mcp-server
- Key
io-github-trajectiq-ai-e2e- Source
- github.com/trajectiq-ai/E2E
- Hosted endpoint
https://playwright-e2e-mcp.vercel.app/api/mcp
More in Web & browsing
MCP server
More in Web & browsingfirecrawl-mcp-server
MCP server · firecrawl
More in Web & browsingapify-mcp-server
MCP server · apify
More in Web & browsingBrowserbase
MCP server · browserbase
More in Web & browsinginspo
MCP server · nutlope
More in Web & browsingplaywright-mcp
MCP server · microsoft
More in Web & browsing