cxc-qa — Manual Surface QA Gate

SkillDev tools

MUST USE after building or changing any user-facing surface (web UI, TUI, CLI, HTTP API) before claiming done — manual, surface-driving QA: real invocations on real surfaces, captured artifacts, adversarial classes, and teardown receipts feeding the PABCD C gate. Automated suites are dev-testing's job; this skill proves the surface actually works when driven. Triggers: manual QA, QA this, does it actually work, drive the UI, smoke test, visual QA, screenshot check, TUI alignment, CJK clipping, 수동 QA, 실제로 되는지 확인, 동작 확인, 직접 돌려봐.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the cxc-qa — Manual Surface QA Gate skill

What this skill tells your AI

The instructions your AI receives, as published by lidge-jun/codexclaw in plugins/codexclaw/skills/qa/SKILL.md and read by ahel’s review.

Prove a built surface works by DRIVING it, not by inferring from green tests. No scenario closes on a status string; it closes on a captured artifact from a real surface, an adversarial pass, and a teardown receipt. Everything in this skill is E7 discipline (agent-followed, not hook-enforced); the one shipped E2 touchpoint is noted in §7. Lineage: lazycodex visual-qa / review-work / lazycodex-qa-executor (vendored at devlog/.lazycodex/), translated to codexclaw's no-server, Codex-native-tool model.

Modular References

FileWhen to ReadWhat It Covers
references/visual-qa.mdANY visual surface verdict (web UI, TUI)Visual artifact rubric: companion-skill grounding (QA-VISUAL-COMPANION-01), objective-metrics-first (QA-VISUAL-METRIC-01), oracle judge limits, extended adversarial classes, TUI harness options
references/http-api-qa.mdHTTP API surface scenariosWire-driving procedure (QA-HTTP-01): curl capture discipline, auth-state/contract/idempotency/boundary/content-negotiation/CORS axes, status-code truth table
references/cli-tui-qa.mdCLI or TUI surface scenariosSession mechanics (QA-CLI-01): stdout/stderr separation, TTY-vs-pipe, env/config precedence, signal cleanup, tmux lifecycle + wait-for-marker driving

Browser selection (QA-TOOL-LADDER-01) is owned by portable browser routing. Suitable available Aside, agbrowse and native browser capabilities may drive built UI; none is required.

0. Scope split (single ownership)

  • cxc-dev-testing owns AUTOMATED verification: unit, contract, E2E, Playwright suites and CI gates. Browser selection lives in the shared dev policy; dev-testing §4.7 connects that policy to exploratory tests.
  • cxc-qa (this skill) owns the manual QA PROCEDURE: scenario matrix, faithful channels, evidence contract, adversarial classes, oracle passes, teardown receipts.
  • Both feed PABCD C. Neither replaces the other: a green suite without a driven surface can still ship a broken border or a dead route; a driven surface without a suite has no regression guard. Promote a QA flow to a deterministic test (dev-testing §4.1) when it must stay guarded.

1. Trust nothing

Executor claims, previous logs, transcript memory, and evidence summaries are untrusted until you inspect or reproduce them. A remembered "it worked" is not evidence (FAMILY-PROOF-01). Every PASS points at a non-empty artifact you (or your dispatched QA worker) captured THIS round.

2. Surface detection + faithful channels

State the exact surface and invocation BEFORE running each scenario. Use the channel faithful to the surface — CLI output parsed from the wrong layer is not evidence of the layer you changed:

SurfaceFaithful channelArtifact
HTTP APIcurl -i (headers + body) — scenarios in references/http-api-qa.mdresponse capture
CLIreal invocation, stdout/stderr split + exit code — discipline in references/cli-tui-qa.mdterminal capture
TUItmux session at FIXED dims — mechanics in references/cli-tui-qa.md, visual rubric in references/visual-qa.mdplain + ANSI captures
Web UIbrowser screenshot at a STATED viewport, read via view_image — workflow in references/visual-qa.mdscreenshot(s)
Desktop GUIcomputer-use + screenshots (per-app approval; never drive terminals/Codex itself)screenshots + action log

Tool choice for the browser/CU rows follows the shared portable browser policy; the inspect -> act -> re-inspect protocol applies. Data-shaped behavior may use parsed CLI/data output as its channel.

3. Evidence contract

Artifacts live under .codexclaw/evidence/<sessionId>/qa/<scenario-id>/:

  • invocation.txt — the exact command(s)/steps, copy-pasteable.
  • the artifact(s) — capture, screenshot, response, transcript.
  • verdict.json{ "scenario": "<id>", "criterion": "<what this proves>", "surface": "http|cli|tui|web|gui", "verdict": "PASS|FAIL|NA", "artifactRefs": ["<relative paths>"], "note": "<one line>", "capturedAt": "<RFC3339>", "sourceSnapshotAt": { "kind": "resolved|unavailable", "commitSha": "<sha>", "dirty": <bool>, "capturedAt": "<RFC3339>", "treeHash": "<sha256, dirty only>" } }. On web and gui only, also "captureChecks": { "signature": <bool>, "nonEmpty": <bool>, "dimensionsMatch": <bool>, "composited": <bool> } — all four keys (QA-CAPTURE-INTEGRITY-01 in references/visual-qa.md).

Rules:

  • Every PASS names at least one non-empty artifact in artifactRefs.
  • capturedAt and sourceSnapshotAt are required on every surface: knowing when evidence was captured, and against which tree, does not depend on what kind of surface produced it (QA-EVIDENCE-FRESHNESS-01).
  • inferred and partial verdicts DO NOT EXIST — a scenario either ran against the real surface or it did not.
  • A scenario that cannot run is a FAIL carrying the blocker and the missing prerequisite, not a skip.
  • NA is legal only when the class structurally cannot apply to the surface (e.g. viewport class on a headless API), always with a recorded reason.
  • This directory is shared with the SubagentStop receipt gate's root (.codexclaw/evidence/); main-session QA artifacts do not interact with worker receipts (the gate validates only the worker's own EVIDENCE_RECORDED: marker path).
  • A worker that CANNOT write there was dispatched wrong, not gated wrong: read-only lanes belong on agent_type:"explorer", which the gate never touches. The retry budget is terminal, so a mis-routed worker is released rather than trapped — but its verdict stays unresolved and blocks goal completion until a receipt settles it.

After every scenario is done, emit the aggregate receipt:

node plugins/codexclaw/skills/qa/scripts/validate-evidence.mjs \
  .codexclaw/evidence/<sessionId>/qa/ --emit-receipt

It validates every verdict.json, confirms they all describe the same tree, and writes .codexclaw/evidence/<sessionId>/qa-receipt.json. Any failure leaves no receipt behind, including deleting one an earlier run produced — a receipt that outlives the QA it attests to is worse than none.

Current limitation: the final gate does not pick this file up on its own. Until the final-gate lifecycle CLI exists, point finalGate.qaReceiptPath at that path by editing the goalplan directly. Remove this paragraph when that CLI lands.

4. Adversarial classes

For each scenario, probe every APPLICABLE class and record the observable result per class (own row in the matrix):

  1. Empty/absent input (no args, empty body, blank field).
  2. Malformed input (bad flag, invalid JSON, wrong type, unknown enum).
  3. Boundary size (longest plausible value, zero-length list, max count).
  4. Repeat/concurrent invocation (double-submit, rerun idempotence).
  5. (web/TUI only) Narrow viewport / narrow terminal width + CJK text (clipping, baseline drop, wide-char column drift, border misalignment).

Class 5 findings feed the visual pass in §5. N/A classes follow the §3 rule. Surface-specialized class details: http-api-qa.md (wire), cli-tui-qa.md (terminal), visual-qa.md (extended visual classes).

5. Oracle passes (depth scales by work class)

  • C2: run the matrix yourself; one self-review pass over the artifacts before writing verdicts.
  • C3+ with a visual surface: dispatch TWO parallel read-only reviewer passes (spawn_agent, explorer role, DISPATCH-TASK-01 packet; paste the captures/screenshot paths + script/tool observations into each prompt — do not make the oracle re-derive context): Both passes are rubric-bound: attach cxc-dev-frontend (anti-slop + visual-verification checklist ARE the rubric) and, for design-direction judgments only, cxc-dev-uiux-design; every PASS cites the rule ids it checked (QA-VISUAL-COMPANION-01, references/visual-qa.md). Capture the objective evidence layer FIRST (viewport matrix, DOM text extraction for text/CJK claims, console errors — QA-VISUAL-METRIC-01).
    • Pass A — design-system + functional integrity: is this a real, coherent implementation driven by reused primitives (not a mock-only screen or a pasted raster), and do the intended features actually work?
    • Pass B — visual fidelity + CJK precision: open the screenshots/captures directly; hunt clipping, baseline drop, glyph breakage, KO/JA/ZH precision, border/box-drawing drift. Synthesize both into one PASS / REVISE / FAIL in the MAIN session. On FAIL, REVIEW-SYNTHESIS-01 applies before any re-dispatch; revision rounds reuse the same oracle (DISPATCH-ACTOR-01), and the final C gate gets a fresh one.
  • Long passes: require WORKING: <task> - <phase> progress messages and BLOCKED: <reason> from dispatched QA workers; a wait timeout is not a failure signal (a running child is alive). A lane that crashed or returned no deliverable is INCONCLUSIVE — never counted as PASS.

6. Teardown receipts

Every resource the QA pass spawned — dev-server PIDs, ports, tmux sessions, browser tabs, containers, temp dirs — gets its own teardown line WITH proof: lsof -i :<port> empty, tmux ls clean, ps check, rm + absent check. A stateless pass records an explicit "no resources spawned" line with what was checked. No QA asset is left running after the verdict.

7. Binding to PABCD C

A user-facing surface change closes C with BOTH: the automated gate (dev-testing) AND this skill's QA matrix — any FAIL verdict blocks the C>D claim until repaired (LOOP-REPAIR-01 counts apply) or the criterion is re-scoped through a P-phase amendment, never silently. This is E7 discipline: no hook reads verdict.json. The E2 touchpoint: QA delegated to a worker subagent rides the existing SubagentStop receipt gate — the worker cannot finish without a non-empty receipt under .codexclaw/evidence/.

v2 candidates (deliberately not shipped)

  • Node-only ports of lazycodex's image-diff / tui-check scoring scripts (objective similarity/overflow JSON as oracle reference input). v1 uses direct view_image inspection + inline width checks; vendor scripts only if field friction shows the inline checks miss real defects.
  • A review-work-style multi-lane orchestrator: not planned — the A gate, C adversarial review, REVIEW-SYNTHESIS-01, and this skill already cover those lanes without a 5-spawn token bill.

Signals

GitHub stars
37
Forks
7
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
cxc-qa
Source
github.com/lidge-jun/codexclaw