super-qa — BFS Route-Crawler that Builds the Spec Suite
SkillWeb & browsingSuper QA canonical workflow for functional bug-bash, evidence capture, and fix-ready issue filing. Crawls routes from `docs/super-qa/queue.md` with Playwright, records screenshots/logs/HARs, files product-readable issues with clear type/priority/area labels, and continues until traversal is complete or only human-gated blockers remain. Crawling is always browser-driven, but the regression test it leaves behind is written at whichever layer the defect actually lives, unit or integration in Vitest, e2e in Playwright only when the browser is the thing that broke. Use when the user says "run super-qa", "QA review", "bug bash", "verify and fix", or invokes `/super-qa`.
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the super-qa skill
What this skill tells your AI
The instructions your AI receives, as published by erictechpro/super-board in skills/super-qa/SKILL.md and read by ahel’s review.
Auth bootstrap — auto-discover, never ask
Before iter 1, resolve a test login. Never ask the user; discover or create. Order:
- Cached creds at
.claude/settings.local.json.e2e.user/.e2e.pass— use them as-is. - Env-file fallback — read in order:
.env.production.local,.env.local,.env. PullSUPABASE_URL+SUPABASE_SERVICE_ROLE_KEY. If both present, auto-createe2e-test@<owner-domain>(owner fromgit config user.emaildomain, fallbackerictech.ca) via:
Generate a 24-char random password (curl -sS -X POST "$SUPABASE_URL/auth/v1/admin/users" \ -H "apikey: $SUPABASE_SERVICE_ROLE_KEY" \ -H "Authorization: Bearer $SUPABASE_SERVICE_ROLE_KEY" \ -H "Content-Type: application/json" \ -d "{\"email\":\"$E2E_USER\",\"password\":\"$E2E_PASS\",\"email_confirm\":true,\"user_metadata\":{\"role\":\"e2e-test\",\"created_by\":\"super-qa-bootstrap\"}}"openssl rand -base64 24 | tr -d '/+=' | cut -c1-24). Save to.claude/settings.local.json.e2eso the next run skips bootstrap. Verify by logging in (POST /auth/v1/token?grant_type=passwordreturns access_token). - Unauth-only fallback — if neither cached creds nor env-file admin key exist, log "no admin key — restricting to public routes" and proceed with
/login,/signup,/forgot-passwordonly. Do NOT halt.
The bootstrap is non-interactive. If the curl returns an existing-user error (User already registered), reset the password via the same admin endpoint and continue. Save creds either way.
QA loop state board — GitHub Project as the state store
State for the QA↔Build loop lives in a GitHub Project named Super Ultimate QA (auto-discovered by title at the user/org level). It has six Status columns:
| Column | Meaning |
|---|---|
| Queue | Work the next iteration will pick up |
| Testing | Feature currently being explored or verified |
| Done | Spec passing, no action needed |
| Bug | Failing spec; Super Build picks these up next |
| Flaky | Only passes on retry; quarantine + investigate |
| Skip | Out of scope; documented and parked |
This board is the durable, machine-readable state for the loop. docs/super-qa/queue.md remains the BFS route seed and audit log, but all actionable findings land on the project board so the build lane can pick them up without parsing markdown.
Project resolution
Resolution order (used by scripts/super-qa-file-bug.sh and the orchestrator):
- If
SUPER_QA_PROJECT_TITLEis set, querygh project list --owner $SUPER_QA_PROJECT_OWNER --format jsonand pick the project whose title matches (case-insensitive). - Otherwise default to title
Super Ultimate QA. - Owner defaults to
$(gh repo view --json owner -q .owner.login)orSUPER_QA_PROJECT_OWNERif set.
If no matching project exists, the loop halts with a one-line error and instructs the operator to create one. Do not silently fall back to the repo's primary project — column semantics differ.
Filing semantics
When the worker finds a red spec, it calls scripts/super-qa-file-bug.sh. The script:
- Creates the issue with
source:qalabel and full evidence body. - Adds the issue to the resolved Super Ultimate QA project.
- Sets the Status field to
Bug(override withSUPER_QA_TARGET_OPTION_NAME).
A Flaky retry-only finding goes to Flaky; out-of-scope routes go to Skip. The same issue may also live on a repo-level project for normal triage — Status on the QA project is the loop state; status elsewhere is independent.
Algorithm overview
You are the orchestrator. Your job: drive a BFS crawl of the target app
that builds e2e/paths/ into a comprehensive Playwright suite while
fixing any bugs it stumbles on. Each iteration is a fresh worker — a sub-agent you launch in this
session with the Agent tool — that does ONE iter and exits. There is no
worktree: the worker commits directly to the active branch (default main).
This pack ships no headless dispatcher for this loop.
Autonomous mode — NO AskUserQuestion mid-loop
/super-qa is invoked by an operator who walks away (often overnight). The loop must run end-to-end without the
orchestrator pausing on AskUserQuestion.
Decide-and-proceed, don't ask, on these classes of issue:
- Harness script bugs (parser errors, broken xargs, missing
chmod +x, stale lock files). Fix locally, commit as🐛 [fix] super-qa: <what>(writing-standard.md § 1), then continue. - Missing env auto-loads, leftover MCP zombies, log-dir creation.
- Lint / formatting nits introduced by the worker that block its own commit.
- Choosing between equivalent dispatch modes (sequential vs sequential — parallel is unsafe, never picked).
Still halt on these (real forks):
- Critical-path HUMAN GATE (worker status 4) — production money flow red >2 iters. Mandatory halt per skill contract.
- Worker status 2 / 3 with an error that isn't a script bug (e.g. Anthropic API down, a production 503, auth credential rejected).
- Working tree dirty with changes the orchestrator didn't make (could be user's in-flight work).
- Anything that would delete or rewrite history (
git reset --hard, force-push, branch deletion).
Status updates instead of questions: when you make an autonomous fix, report it in the next status message — what was broken, what you committed (with SHA), and a one-line revert path, in the session output.
After this loop has run for a while, e2e/paths/ is the deliverable. CI runs
npm run test:e2e on every PR. No AI involvement in steady-state
regression. AI is only needed to extend the suite when new pages ship.
AI loop (build phase) CI (steady state)
───────────────────── ──────────────────
/super-qa → e2e/paths/ → npm run test:e2e on PR
(drains queue.md, grows toward catches regressions
fixes bugs) comprehensive without any AI tokens
This skill is INDEPENDENT of /super-build. They share no state.
super-build executes new sessions from the GitHub Project board (forward
progress); super-qa hardens what already ships by exhaustively
walking the route tree.
The queue is the curation surface
docs/super-qa/queue.md is a markdown checklist the loop reads from
and writes to. Four states per line:
[ ]queued — will be popped next[x]green spec exists (links to the spec file + iter discovered)[!]skipped permanently (with reason)[b]bug found here, children NOT pushed; will be re-tried after fix[?]flaky — passed on retry, not a bug yet (see flaky policy below)
The queue is append-only at the back during normal exploration. The user can hand-edit it to reorder (turn BFS into DFS for one branch, push something to the front, etc.). The skill respects whatever the user wrote.
The seed is client/src/routes.ts — every static route becomes a [ ]
entry under "Level 0".
Findings land in clear GitHub Issues, NOT in markdown
Every [b] cell triggers the worker to call scripts/super-qa-file-bug.sh.
The resulting Project card must be readable from the board without opening it.
Do not use legacy [VFL], verify-fix, or severity:P* wording in new issue titles.
Title format:
<emoji> <Type> <route?> — <short action/result>
Examples:
🐛 Bug /imports — CSV upload fails after submit
🎨 UX /settings/users — Add user button is a no-op
🧪 Tests /orders — missing coverage for failed payment state
📝 Docs /delivery — SPEC section missing for dispatch flow
Priority labels use plain words:
priority:high— urgent, release-blocking, data-loss, security/auth, money, or critical feature broken.priority:medium— important feature degraded, important edge case, a11y/i18n issue, or user confusion that should be fixed soon.priority:low— polish, cosmetic, copy, documentation, testability, or cleanup.
Required labels:
- Type:
bug,feature,ux,tests,docs, ortech-debt. - Source:
source:qa. - Priority:
priority:high,priority:medium, orpriority:low. - Area:
area:<product-area>when known (area:settings,area:imports,area:admin-shell, etc.). - QA category when relevant:
qa:functional,qa:visual,qa:network,qa:console,qa:i18n,qa:a11y,qa:data,qa:testability. - Suggested skill owner when helpful:
skill:super-build,skill:super-qa,skill:ui-refine-loop, orskill:super-review.skill:ui-refine-loopmarks a UI ticket for a human to run/ui-refine-loopon; the board never triggers it.
The script adds the issue to the resolved Super Ultimate QA project and moves it into the Bug column (override with SUPER_QA_TARGET_OPTION_NAME) so a human (or /super-build with BUILD_LOOP_SOURCE_COLUMN=Bug) can pick it up immediately. The repo's standalone feature project is not touched by this flow.
The iteration-N.md Section 3 is the per-iter audit (with gh_issue: <N> back-references); the GH issue is the durable tracker. The fix-commit message includes (closes #<N>) so the issue auto-closes on merge.
Triage all loop-filed findings:
gh issue list -l source:qa --state open
This is a carved exception to the project rule "ask before gh issue create": the loop is autonomous and unattended, so it is authorized to auto-file — but ONLY with the clear source:qa label and the evidence template below.
Required issue body template
Every auto-filed finding must give a future headless Claude/Super Build session enough context to fix it without re-discovering the bug from scratch:
## Summary
<one sentence: what is wrong and where>
## Repro steps
1. <exact login/route/action>
2. <next action>
3. <observed failure>
## Expected behavior
<what should happen, citing SPEC.md, DESIGN.md, or product intent when possible>
## Actual behavior
<what happened instead>
## Evidence
- Screenshot: <embedded image or repo/raw path>
- Console: <0 errors or paste relevant lines / path to console.log>
- Network: <failed requests or path to network.json/HAR>
- Page errors: <0 errors or path to pageerrors.log>
- Spec: `<e2e/paths/...spec.ts>`
## Suggested fix path
- Suggested owner: `super-build` | `super-qa` | `super-review` | `ui-refine-loop` (UI ticket — a human can run `/ui-refine-loop`; the board never triggers it)
- Suggested skills: `mattpocock-skills:diagnosing-bugs`, `mattpocock-skills:tdd`, `verification-before-completion`
- Notes for implementer: <first suspected file/function, if known>
## Acceptance criteria
- [ ] <user-visible behavior fixed>
- [ ] <Playwright/regression coverage added or updated>
- [ ] <Super QA rerun passes this route>
If a screenshot/network/console artifact does not exist, explicitly write not captured and explain why. Do not leave evidence ambiguous.
Filing guardrails
super-qa-file-bug.sh validates issue bodies before filing. It must reject weak tickets that are missing required sections or still contain placeholders like TBD, <...>, or TODO:. Do not bypass this with SUPER_QA_ALLOW_WEAK_BODY=1 during normal autonomous runs.
Every filed issue must include:
## Board summarynear the top for quick project-card context.- A hidden
super-qa-metablock with route, spec, iteration, area, category, priority, type, andfingerprint. - A deterministic fingerprint. Use a meaningful key when known, such as
settings-users|add-user|no-op; otherwise the script derives one from kind/route/category/title.
Dedupe policy: before creating a new issue, the script searches open source:qa issues for the fingerprint. If a match exists, it comments with the new evidence and returns the existing issue number instead of creating a duplicate Ready card.
What counts as "green" — non-blank guards + forensics
A 200 response with a blank body is a bug, not a green test. Every spec
under e2e/paths/ enforces:
- No console errors (
level === 'error'only) - No uncaught page errors (
page.on('pageerror', ...)) - No 5xx network responses during the run
- No 401/403 on auth-required pages
- Type-specific non-blank guard (e.g., list-page: ≥1 row OR
empty-state element; dashboard: ≥1 widget with non-
—data) — seedocs/super-qa/page-types.md"Non-blank guard" column
If any of those fire → cell is [b] → bug filed.
Per-spec forensics captured by e2e/lib/report-fixture.ts:
| Artifact | Path |
|---|---|
Screenshot per report.step() | docs/super-qa/report/<slug>/tc-N/<locale>/*.jpg |
| Console errors | docs/super-qa/report/<slug>/tc-N/<locale>/console.log |
| Page errors | docs/super-qa/report/<slug>/tc-N/<locale>/pageerrors.log |
| Network HAR | docs/super-qa/report/<slug>/tc-N/<locale>/network.har |
| Network summary | docs/super-qa/report/<slug>/tc-N/<locale>/network.json |
| Sentry probe | docs/super-qa/report/<slug>/tc-N/sentry-events.json |
The fixture (e2e/lib/report-fixture.ts) already captures console
errors, page errors, failed requests, network summary, and Sentry events,
plus HAR via recordHar. Disk writes are gated behind
SUPER_QA_FORENSICS=1; set it in every worker's environment. Fixture
exposes everything via report.forensics.* (e.g.
report.forensics.consoleErrors).
Algorithm
1. Determine iteration count
- If user provides arg (e.g.
/super-qa 5), use that count. - If
--resume, start atmax(existing iteration-*.md) + 1. - Default: 10 iterations, sequential. (Concurrency is unsafe — workers
share
queue.md. To run parallel work, branch and run a second orchestrator.)
Report once: 🐛 Super QA starting — N iterations.
2. For each iteration N (sequential)
2a. Pre-flight
next_n = max(existing iteration-*.md numbers in docs/super-qa/iter/) + 1. If none exist, start at 1.- Confirm working tree is clean (
git statushas no staged/unstaged tracked changes — untracked files are OK). If dirty → halt and notify. - Verify
docs/super-qa/queue.mdexists. If not, refuse to dispatch and report — the queue is hand-seeded once via this skill's setup (or already by the commit that introduced the skill).
2b. Launch the iteration worker
Launch one sub-agent (Agent tool) in the repo root, no worktree, with
prompt = references/iteration-preamble.md + a per-iteration footer (iter
num, base SHA, mandatory final-commit format) and SUPER_QA_FORENSICS=1.
When it returns, classify the iteration:
0iter complete — a🧪 [test] super-qa: iter Ncommit is on the current branch.2worker failed — it errored and left no done or WIP marker.3no done-commit — it finished but no🧪 [test] super-qa: iter Ncommit exists.4HUMAN GATE — its output containsHUMAN GATE TRIPPED:.5WIP-CHECKPOINT — it hit its budget mid-fix and left a🚧 [wip] super-qa:commit; the next iter picks it up.
Report: 🔍 Iter N launched.
2c. Wait, then advance
On status 0:
- Read the close-out commit subject to extract
(X bugs, Y items). - Report:
✅ Iter N done — X bugs, Y items processed. - Loop to next iteration.
On status 2 / 3 / 4 / 5:
- Report the worker's last output (its final 50 lines).
- Halt the loop (unless
--continue-on-errorwas passed). - Exit 4 (HUMAN GATE) is mandatory halt regardless — see "critical paths" below.
- Exit 5 (WIP) is not halt; the next iter's regression phase finds the red spec and finishes the fix. Continue.
3. Termination check (after every iter)
Stop when any of:
- Queue has no
[ ]items left — natural completion. - User-supplied iteration count
Nis reached (theNin/super-qa N). - A halt gate fires (HUMAN GATE, dirty tree, worker status 2/3/4).
- User interrupts.
Notification trigger (NOT a stop): after 3 consecutive iters with zero
bugs found, report a summary like "diminishing returns: 3 iters, 0
bugs, N items still in queue — continuing". The loop continues until the
queue actually drains (or N is reached). The user can interrupt manually
if they accept the diminishing returns.
Rationale: zero bugs in recent iters does not prove the unexplored remainder of the queue is clean. The only way to know is to keep going. The previous "3 zero-bug iters" auto-halt was budget pragmatism dressed up as completeness.
Coverage % is reported every iter as a progress indicator, never a gate.
4. Final report
After termination (or on halt):
- Aggregate: total iters, items moved from
[ ]→[x]/[b]/[!], total bugs found, total fixed, queue size now. - Report a summary linking to all
docs/super-qa/iter/iteration-*.mdand the currentdocs/super-qa/report/QA-REPORT.md.
One iteration, in plain English
The worker (per iter N) runs three phases:
Phase 1 — Regression. npx playwright test e2e/paths/ against
everything. Any red spec? Apply the retry policy (re-run once in isolation).
Real reds get filed as bugs in iteration-N.md, fixed via TDD, re-run until
green. This is the "verify what we already have" step.
Phase 2 — Explore. Until the budget is hit (default: 5 cells popped OR 30 min wall-clock):
- Pop the top
[ ]item fromqueue.md. - Classify via the URL-pattern heuristic in
docs/super-qa/page-types.md. - Write a spec at
e2e/paths/<slug>.spec.tsusing the test recipe for that type. - Run it.
- Green → mark
[x]. Walk the rendered page; push newly-discovered children (links / buttons / dialogs / tabs — see preamble for the verbatim "what counts as a child" rule) to the back of the queue. Cap children pushed per page at 10. - Red → file the bug to
iteration-N.md. Mark[b]. Do NOT push children. Try to fix via TDD if budget allows; otherwise leave for next iter. After fix, re-mark[ ]so a future iter expands the subtree.
- Green → mark
Phase 3 — Report. Write docs/super-qa/iter/iteration-N.md (bugs
found, items processed, queue size before/after, coverage snapshot). Run
npm run qa:report:render. Commit:
🧪 [test] super-qa: iter N (X bugs, Y items, Z PRs opened) (writing-standard.md § 1).
The bug-handling rule (non-blocking)
When a spec goes red, the worker logs the bug, marks the item
[b], stops expanding into that subtree, and continues with the rest of the iter's batch. Other siblings in the queue still get tested.
A bug on /customers/:id shouldn't block exploration of /orders or
/dashboard. Maximum coverage per iter; one bug never halts the loop.
Fix can happen:
- Same iter — if budget allows, the worker runs TDD on the bug after the explore phase.
- Next iter — if budget exhausted, the bug sits in
iteration-N.mdand Phase 1 of iter N+1 finds the red spec and finishes the fix. - Manually by human — fix it whenever; the loop picks it up on the next regression pass.
After fix, the [b] item gets re-marked [ ] and a future iter processes
it (and pushes its children).
Page-type taxonomy
Classification gives the worker a test recipe per kind of place. See
docs/super-qa/page-types.md for the canonical list:
list-page form-page settings-page
detail-page dashboard-page modal/drawer
public-page import-flow wizard
settings-tab
Classification order: URL-pattern heuristic first (cheap default) → AI
override if the heuristic looks wrong → hand-curated override in
page-types.md if the user wants to lock something in.
Critical paths (HUMAN GATE)
docs/super-qa/critical-paths.md is a hand-curated list of
money-flow specs that MUST always be green (login, create order, take
payment, run cutoff snapshot, send driver email). The user maintains it;
the skill never auto-edits it.
If any critical-path spec has been red for >2 consecutive iters, the the iteration ends with status 4 (HUMAN GATE). The loop halts. A human investigates. This protects production-critical flows from being silently broken by in-flight queue items.
Multi-step user journeys — docs/super-qa/flows.md
critical-paths.md is the money-flow guard (login, checkout, payment) —
small, hand-curated, mandatory human gate if red >2 iters. It is not a
coverage layer.
docs/super-qa/flows.md is a separate, broader hand-curated list of
multi-step user journeys the BFS crawler doesn't reliably exercise on its
own. Examples:
- create user → set role → user logs in → sees their assigned region
- import CSV → review preview → confirm import → verify rows in
/orders - create order → assign driver → driver email mock fires → order moves to dispatched
Each entry in flows.md is one line per flow with a slug. The worker
generates or updates e2e/flows/<slug>.spec.ts for each entry. Specs in
e2e/flows/ run alongside e2e/paths/ in Phase 1 regression — same fixture,
same forensics, same [b] rules apply. They are not pushed via BFS
expansion; they exist as hand-curated chains.
Flows take precedence over individual route specs when both exist for the
same endpoint — a green e2e/paths/orders.spec.ts doesn't override a red
e2e/flows/create-order-end-to-end.spec.ts.
The user maintains flows.md; the skill never auto-edits it. To add a new
flow, the user hand-edits the file; the next iter's Phase 1 picks up the
new slug and generates the spec from a flow recipe.
flows.md template (the worker creates this on first encounter if missing)
# super-qa multi-step user journeys
One line per flow. Format: `- slug — short description (route1 → route2 → ...)`
The worker generates / updates `e2e/flows/<slug>.spec.ts` per entry.
These run in Phase 1 regression alongside `e2e/paths/`.
## Flows
- onboard-user — create user, set role, user logs in, sees their region
(/admin/users/new → /login → /dashboard)
- import-csv-end-to-end — import CSV, review preview, confirm, verify rows
(/imports → /imports/preview → /orders)
Test-gap check (after build)
Issue-scoped mode only (super-board Tester, or any run handed a finished diff). The
Builder's tests answer "does it do what I built?"; this answers "what breaks silently?".
Scope is the diff: git diff --name-only $(git merge-base HEAD origin/<base>)...HEAD.
On auth, money, data-loss or security changes, hand steps 1-3 to a read-only sub-agent
given only the ACs, the range and the changed-file list, so it is not anchored on the
Builder's reasoning.
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 30
- Forks
- 12
- Last commit
- Oct 2026
ahel review
K4binfo
destructive-scoped (in references/iteration-preamble.md)
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
super-qa- Source
- github.com/erictechpro/super-board
github.com/erictechpro/super-board
Related picks
Skill · thedaviddias
The pick for JavaScriptmodern-javascript-patterns
Skill · wshobson
The pick for JavaScriptsetup-ts-deep-modules
Skill · mattpocock
The pick for TypeScripttypescript-pro
Skill · jeffallan
The pick for TypeScriptlogs-nextjs
Skill · posthog
The pick for Next.jsclerk-nextjs-patterns
Skill · clerk
The pick for Next.js