playwright-e2e-mcp

MCP serverWeb & browsing

Run, debug and inspect Playwright E2E tests from any AI agent: diagnostics, live DOM, selectors.

Use playwright-e2e-mcp in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add playwright-e2e-mcp and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use playwright-e2e-mcp

Details

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

playwright-e2e-mcpStart free

Install playwright-e2e-mcp

The server’s own address, for the clients that take one directly. Or connect ahel once and every client you use reads it from one address, with the account kept on ahel rather than in each client’s config.

  • Claude Code

    claude mcp add --transport http --scope user playwright-e2e-mcp 'https://playwright-e2e-mcp.vercel.app/api/mcp'

    Run it once in your project, then open /mcp to approve any sign-in the server asks for.

  • Claude Desktop

    https://playwright-e2e-mcp.vercel.app/api/mcp

    Add a custom connector in Settings, paste this address, and approve the sign-in.

  • Cursor

    cursor://anysphere.cursor-deeplink/mcp/install?name=playwright-e2e-mcp&config=eyJ1cmwiOiJodHRwczovL3BsYXl3cmlnaHQtZTJlLW1jcC52ZXJjZWwuYXBwL2FwaS9tY3AifQ==

    Open the link and Cursor adds the server at that address.

  • ChatGPT

    https://playwright-e2e-mcp.vercel.app/api/mcp

    In Settings, enable Developer mode, create an MCP app, and paste this address. Your plan and workspace must allow custom apps.

  • Codex

    codex mcp add playwright-e2e-mcp --url 'https://playwright-e2e-mcp.vercel.app/api/mcp'

    Run it once, then sign in with codex mcp login playwright-e2e-mcp if the server asks for an account.

From the project's README

As published by trajectiq-ai/E2E in README.md.

An MCP server that lets AI agents run, debug, and inspect Playwright end-to-end tests — with structured results, actionable failure diagnostics, and live DOM inspection.

run-test ──▶ get-failure ──▶ inspect-page ──▶ validate-selector ──▶ fix ──▶ re-run
   ▲                                                                    │
   └──────────────────────── list-tests ◀────────────────────────────────┘

Instead of handing an agent raw Playwright output, this server turns every run into machinable results: pass/fail stats, per-failure messages with file:line, a failure kind (assertion, timeout, browser crash, syntax error, dead dev server, full disk…), and a concrete "how to fix" hint. When a test fails because a selector no longer matches, the agent can open the live page in a headless browser, see the real DOM with unique CSS selectors, and validate the replacement selector before re-running.

Demo

Hosted endpoint — what an initialize + tools/list round-trip against https://playwright-e2e-mcp.vercel.app/api/mcp returns for a client that sends the bearer token (without one, only list-tests and get-failure are listed):

A real test run — run-test served over stdio by npx -y playwright-e2e-mcp against the bundled examples/sample-test.spec.ts (actual output, unedited):

Images are rendered with node scripts/gen-demo-images.mjs: the run-test card is real captured output; the endpoint card is an illustration of the authenticated listing.

Install

Works with Claude Desktop, Claude Code, Cursor, Windsurf, Codex, Gemini CLI, Freebuff and every other MCP client — pick whichever route fits:

RouteHow
npm (canonical, fastest)npx -y playwright-e2e-mcp
MCP Registry (registry-aware clients discover it automatically)io.github.trajectiq-ai/E2E — listing
Any client, no npm account needednpx -y github:trajectiq-ai/E2E#v0.1.2 (pin a release tag)
Claude Desktop, zero Node setupdouble-click the .mcpb extension
Remote-only clients (ChatGPT connectors)https://playwright-e2e-mcp.vercel.app/api/mcp

Details and per-client config: Installation · MCP client configuration.


Tools

ToolPurpose
run-testRun Playwright tests and return stats, failures, diagnostics and hints
get-failureDeep analysis of one failure: stack, expected/actual, DOM snapshot at failure (from the Playwright trace), next steps
inspect-pageOpen a URL headlessly and return the rendered DOM: selectors, visibility, boxes, text, console output, HTML
list-testsList available tests (file, line, full title, projects) with filtering
validate-selectorCheck a CSS selector against a live page: validity, match count, sample matches
generate-e2e-testScaffold a Playwright test from a description using the project's real selectors, discovered from recent file changes
compare-visual-stateVisual regression: screenshot before/after a change and report what moved and how colors shifted
diagnose-flakyRun a failing test 2–10 times with retries disabled and return an evidence verdict: CONSISTENTLY FAILING, FLAKY or NOT REPRODUCING

run-test

ArgumentTypeDescription
projectRootstringProject directory inside the configured root (default: server working directory)
testFilesstring[]Files/directories relative to the root; file:line supported. Omit to run everything
grepstringOnly run tests whose title matches this regex
browserchromium | firefox | webkitPlaywright project to run (matched against config project names)
headedbooleanVisible browser window
timeoutMsnumberHard wall-clock limit for the run (default 120000); the whole process tree is killed past it and partial results are returned
testTimeoutMsnumberPer-test timeout passed to Playwright
workers / retriesnumberPassed through to Playwright
configstringplaywright.config path or 1-based index when the project has several
retryOnFailurebooleanAuto-retry failures once before reporting them (default true; ignored when retries is set)
lastFailedbooleanOnly re-run tests that failed in the previous run (Playwright --last-failed) — the fast fix → re-run loop
argsstring[]Extra Playwright flags from an allowlist (--repeat-each=N, --max-failures=N, --update-snapshots, --shard=1/3, --trace=on, …); values go after = and are checked, and flags that take a path, such as --config or --output, are rejected

Flakiness handling: by default the server injects --retries=1 (unless the config already sets retries), so a test that passes on the retry is reported as flaky, not failed. Traces are captured automatically (--trace=retain-on-failure) so get-failure can show the DOM at the moment of failure.

Example result:

## Playwright run — ❌ FAILED

**Command:** `playwright test --config playwright.config.ts tests/checkout.spec.ts --reporter=json`
**duration 4.2s · exit 1 · config `playwright.config.ts`**

| passed | failed | flaky | skipped | duration |
| ---: | ---: | ---: | ---: | ---: |
| 0 | 1 | 0 | 0 | 1.1s |

### ❌ 1 failing test(s)

### 1 of 1. checkout.spec.ts › pays with card
**File:** `checkout.spec.ts:5`  |  **failed · server-unreachable**

### ⚠️ SERVER_NOT_RUNNING
Your app (dev server) does not appear to be reachable. Start it in another terminal
(e.g. npm run dev / npm start), keep it running, then retry — or configure `webServer`
in playwright.config.* so Playwright starts it automatically.

get-failure

ArgumentTypeDescription
indexnumber1-based failure index from the last run (default 1)
projectRootstringOnly used when re-reading the stored report

Returns the message/code frame, expected vs actual, stack, failure kind with a diagnosis, the test's console output, the DOM snapshot from the Playwright trace (plus the failed action, its selector, and the action log leading up to it), the network requests that failed (4xx/5xx, dead endpoints, no-response — with method, URL, status and resource type), the console errors/warnings the page logged before the failure, and numbered next steps (re-run this single test by file:line, headed/debug mode, validate-selector when the message mentions a locator, …).

inspect-page

ArgumentTypeDescription
urlstringFull http(s) URL to open (required)
projectRootstringProject whose Playwright launches the browser
selectorstringInspect matches of this CSS selector instead of the whole DOM
waitForstringWait for a selector (CSS or text=…) before inspecting
waitUntilload | domcontentloaded | networkidleNavigation wait condition
includeHtmlbooleanInclude the rendered HTML (capped)
maxHtmlCharsnumberHTML cap, default 20000
timeoutMsnumberOverall limit, default 45000

Returns each element's unique CSS selector, tag, visibility, bounding box, text and attributes, plus captured console messages (errors first).

list-tests

ArgumentTypeDescription
projectRootstringProject directory
configstringConfig path or 1-based index
testDirstringRestrict scanning to a directory (must stay inside the project)
filterstringCase-insensitive substring filter on file › title
limitnumberMax tests returned, default 500

Uses playwright test --list when Playwright works, and falls back to a source scan (keeping the reason) when the install or a spec file is broken.

validate-selector

ArgumentTypeDescription
urlstringLive page to test against (required)
selectorstringCSS selector to validate (required)
projectRootstringProject whose Playwright launches the browser
timeoutMsnumberOverall limit, default 45000

Verdicts: ✅ VALID — N matches (with a sample of matches), ✅ VALID — 0 matches (with debugging advice), ❌ INVALID (parse error + fix), or a warning when the input uses a Playwright-only engine (text=, xpath=, >>, :has-text()), which is not plain CSS.

generate-e2e-test

ArgumentTypeDescription
descriptionstringWhat the test should cover (required)
pageUrlstringPage the test starts on (default: baseURL / webServer.url from config)
testDir / filestringWhere to write the spec (default: detected testDir + generated/<slug>.spec.ts); file must end in .spec.* or .test.*
writebooleanWrite the file to disk (default true)
overwritebooleanReplace an existing spec at the target path; only specs this tool generated can be replaced
liveInspectbooleanCross-check selectors against the live page (default on when a URL is known)
projectRoot / configstringAs with the other tools

Reads the agent's recent changes (git status, falling back to git diff HEAD~1, then recent mtimes), extracts the locators those files actually declare (data-testid, getByRole, aria-label, placeholder, id, name, element text), ranks verified-live selectors first, writes a spec built from them, and reports each selector with its source file:line.

compare-visual-state

ArgumentTypeDescription
urlstringPage to capture (required)
namestringBaseline id, e.g. checkout-page (letters, digits, . _ -)
actioncompare | baselinecompare (default) diffs; baseline re-captures the reference
selectorstringCapture just this element
fullPagebooleanCapture the full scrollable page
tolerancenumberPercent of pixels that may differ (default 0.1)
pixelThresholdnumberPer-pixel channel delta considered different (default 60)
waitUntil / waitFor / timeoutMs—As with inspect-page

The first call saves a baseline under .pw-mcp/visual/ (add that to .gitignore, or commit it for CI comparisons). Later calls report changed-pixel counts, merged regions ((x, y) 120×40 — 1,200 px), the average color shift ("blue → red"), and write a red-highlighted diff image for review.

diagnose-flaky

ArgumentTypeDescription
testFilesstring[]Tests to diagnose (file:line supported). Defaults to the tests that failed in the most recent run
runsnumberTimes to run them, 2–10 (default 3)
browser / headed / workers / config—As with run-test
timeoutMsnumberHard wall-clock limit per run (default 120000)
projectRootstringProject directory

Every run executes with --retries=0 and auto-retry disabled, so each result is honest evidence. The response contains a per-run table (status, duration, first failure), the count of distinct normalized error signatures, and one of:

  • ❌ CONSISTENTLY FAILING — failed every run (same error → reproducible bug, different errors → still broken, just noisy). Fix it; it is not flaky.
  • ⚠️ FLAKY — some runs passed. Includes N of M counts and whether the failures share one signature (real intermittent bug) or vary (timing/environment instability).
  • ✅ NOT REPRODUCING — passed every re-run; the original failure was one-off.

The last run is stored, so get-failure can analyze it immediately afterwards.


Installation

Requirements:

  • Node.js ≥ 20 (the server is built on MCP SDK v2 — the 2026-07-28 spec line)
  • A project with @playwright/test installed and browsers available (npx playwright install chromium)

No npm account needed — install straight from GitHub (the prepare script builds dist/ automatically on install). Pin a release tag: an unpinned github:trajectiq-ai/E2E runs whatever is on the default branch at that moment.

npx -y github:trajectiq-ai/E2E#v0.1.2
npm install -D github:trajectiq-ai/E2E#v0.1.2 @playwright/test   # or as a project dependency
npx playwright install chromium

Or grab the packaged tarball from the repo's GitHub Releases page and install it locally:

npm install -D https://github.com/trajectiq-ai/E2E/releases/download/v0.1.2/playwright-e2e-mcp-0.1.2.tgz

Listed in the official MCP Registry as io.github.trajectiq-ai/E2E — registry-aware clients discover it there, and every v* release tag republishes the entry from CI via server.json.

MCP client configuration

Claude Code / generic (project-scoped):

{
  "mcpServers": {
    "playwright-e2e": {
      "command": "npx",
      "args": ["-y", "github:trajectiq-ai/E2E#v0.1.2"],
      "env": { "PW_MCP_PROJECT_ROOT": "/absolute/path/to/your/project" }
    }
  }
}

Claude Desktop / Cursor / Windsurf: add the same block to their MCP config file. The server uses its working directory as the project root; set PW_MCP_PROJECT_ROOT when the client launches it somewhere else (e.g. your home directory).

Codex / VS Code / Copilot CLIs:

codex mcp add playwright-e2e -- npx -y github:trajectiq-ai/E2E#v0.1.2
code --add-mcp '{"name":"playwright-e2e","command":"npx","args":["-y","github:trajectiq-ai/E2E#v0.1.2"]}'

Codex's defaults fight this server: the first launch clones the repo and runs tsc (measured 30 s on a cold npx cache, against a 10 s startup_timeout_sec default), and a Playwright run with retries beats the 60 s tool_timeout_sec default. Raise both in ~/.codex/config.toml:

[mcp_servers.playwright-e2e]
command = "npx"
args = ["-y", "github:trajectiq-ai/E2E#v0.1.2"]
startup_timeout_sec = 60
tool_timeout_sec = 600

Claude Desktop (one-click): download and double-click the .mcpb Desktop Extension attached to the latest release — the bundle ships its own dependencies, so no Node setup is required. On install it asks you to pick your project root (required, no default: choose the project folder, not your home directory) and wires it into PW_MCP_PROJECT_ROOT, so the tools point at a real project from the first call.

Claude Code:

claude mcp add playwright-e2e -- npx -y github:trajectiq-ai/E2E#v0.1.2

Gemini CLI / Qwen Code: paste the mcpServers block above into .gemini/settings.json (Qwen Code: .qwen/settings.json) — both speak the same MCP settings format.

Freebuff / Codebuff (project-scoped): this repo ships a committed .agents/mcp.json, so opening the checkout in Freebuff attaches the server workspace-wide — no global config needed. Your own projects can do the same: drop an mcp.json with the block above into their .agents/ directory. Freebuff asks you to trust a repository's .agents/ on first run.

All tools ship MCP tool annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint), so clients can show accurate safety prompts before running anything.

From a local checkout:

{
  "mcpServers": {
    "playwright-e2e": {
      "command": "node",
      "args": ["/path/to/playwright-e2e-mcp/dist/index.js"],
      "env": { "PW_MCP_PROJECT_ROOT": "/path/to/your/project" }
    }
  }
}

Hosted endpoint (ChatGPT & remote clients)

Some clients — ChatGPT custom connectors especially — only accept remote HTTPS MCP servers and refuse to spawn a local npx process. This repo ships a Streamable HTTP bridge for exactly that case:

Endpointhttps://playwright-e2e-mcp.vercel.app/api/mcp
TransportMCP Streamable HTTP (POST JSON in, JSON or SSE out)
Authoptional bearer token (PW_MCP_HTTP_TOKEN); without one only read-only tools are served
Sourceapi/mcp.ts → src/http.ts

The bridge runs the same createServer() as the stdio transport; the SDK serves every request with a fresh server instance, which is what a serverless function wants. test/http-bridge.test.mjs drives the real Node adapter over node:http so a broken bridge fails in CI, not in ChatGPT.

Open vs. token-protected. When the deployment has no PW_MCP_HTTP_TOKEN, anyone can reach the URL, so the bridge serves only list-tests and get-failure, in restricted mode: list-tests scans sources instead of running playwright test --list (which would execute the project's config), callers cannot pick another projectRoot, and nothing spawns a process, drives a browser or writes a file. Set PW_MCP_HTTP_TOKEN (at least 16 characters; use a random value) to serve all eight tools to clients that send Authorization: Bearer <token>; other requests get 401. Child processes started over HTTP get only an allowlisted environment. PW_MCP_ALLOWED_HOSTS (comma separated, * for any) limits the accepted Host header; without a token and without that variable, only localhost names and the deployment's own Vercel hostnames are accepted (DNS-rebinding protection).

Add it to ChatGPT: Settings → Connectors → turn on Advanced → Developer mode → Create custom connector → paste the endpoint above → authentication None (read-only tools).

Codex can also take the remote transport instead of spawning npx, if you'd rather not ship Playwright to every machine:

codex mcp add playwright-e2e-remote --url https://playwright-e2e-mcp.vercel.app/api/mcp

What to expect: list-tests works and reports the specs bundled with the deployment. Even with a token, tools that spawn a browser (run-test, inspect-page, validate-selector, diagnose-flaky, …) cannot download Chromium in a serverless function, so they return their normal NO_PLAYWRIGHT hint. Use the stdio install for real runs; the hosted endpoint is for discovery and for clients that cannot run local processes.

# verify the handshake without any client
curl -X POST https://playwright-e2e-mcp.vercel.app/api/mcp \
  -H 'content-type: application/json' \
  -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"1.0"}}}'

Redeploy after a change: merge to main. Vercel's Git integration deploys every push, so there is no token to manage and no CLI step — watch the Vercel commit status for the deployment result.

Configuration

Environment variableDefaultPurpose
PW_MCP_PROJECT_ROOTserver cwdDefault project root for every tool
PW_MCP_ALLOWED_ROOTS—Extra directories a caller may pass as projectRoot (:-separated, ; on Windows). Anything outside these and the default root is rejected
PW_MCP_HTTP_TOKEN—HTTP bridge only: bearer token (16+ characters) that unlocks all tools (see above)
PW_MCP_ALLOWED_HOSTSlocalhost + Vercel hostnames when there is no tokenHTTP bridge only: comma-separated Host allowlist; * accepts any
PW_MCP_PASSTHROUGH_ENV—HTTP bridge only: comma-separated extra variables passed to test runs (e.g. BASE_URL)
PW_MCP_MAX_CHILDREN4HTTP bridge only: how many test runs and browser probes may run at once
PW_MCP_BLOCK_PRIVATE_URLSoff (on for the HTTP bridge)1 makes the URL tools refuse loopback, private-network and cloud-metadata addresses, including redirects and subresources
LOG_LEVELinfodebug | info | warn | error | silent
LOG_FORMATtexttext or json (structured)

Logs always go to stderr — stdout is reserved for the MCP protocol.

Typical workflow

  1. generate-e2e-test { "description": "checkout with a saved card" } — scaffolds a spec from your real selectors (skipped if you write the test yourself).
  2. list-tests — see what exists (tests/checkout.spec.ts:5 checkout › pays with card).
  3. run-test { "testFiles": ["tests/checkout.spec.ts"] } — run it; get stats + failures (flaky tests are auto-retried once before being called failures).
  4. get-failure { "index": 1 } — read the code frame, expected/actual, the DOM snapshot at failure from the trace, the failed network requests, the page's console errors, and next steps.
  5. If it looks selector-related: inspect-page { "url": "http://localhost:3000/checkout" } to see the real DOM, then validate-selector to prove the replacement selector works.
  6. After changing CSS/components: compare-visual-state { "url": "…", "name": "checkout" } to catch unintended visual regressions.
  7. If a failure looks intermittent: diagnose-flaky { "runs": 3 } — get the evidence verdict (flaky vs consistently broken) before deciding what to fix.
  8. Fix the spec or the app, then re-run only what failed: run-test { "lastFailed": true }, and repeat until green.

Edge cases handled

Shortened here. Read the whole README on GitHub.

Advanced
Delivery
E2E MCP server → your ahel connector (mcp.ahel.ai) → your AI.
Item type
mcp-server
Key
io-github-trajectiq-ai-e2e
Source
github.com/trajectiq-ai/E2E
Hosted endpoint
https://playwright-e2e-mcp.vercel.app/api/mcp