Sigil: UX Evidence Validator

SkillWeb & browsing

Use when: translating UX research, accessibility standards, market practice, and Playwright browser evidence into validator-safe checks, fixture plans, or evidence reports for finished frontend interfaces.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Sigil: UX Evidence Validator skill

What this skill tells your AI

The instructions your AI receives, as published by cyberalchemyai/arcanum in arcana/ux-evidence-validator/development/candidates/20260820-impeccable-method-adoption/SKILL.md and read by ahel’s review.

  • a finished frontend needs browser evidence rather than prose-only review,
  • UX references need to become validator-safe claims,
  • accessibility, cognitive science, perception research, market heuristics, and Playwright evidence must be kept in separate authority lanes,
  • hard gates need fixture calibration before promotion,
  • screenshot review and human-study residues must be explicit,
  • a future implementation should produce screenshots, traces, ARIA snapshots, accessibility output, DOM measurements, findings JSON, and a residue ledger.
  • the user only wants a quick subjective design critique,
  • no frontend route, HTML artifact, screenshot, or scenario exists to inspect,
  • the task is pure WCAG auditing with no UX evidence synthesis,
  • a human usability study is required but browser evidence is irrelevant,
  • the request would collapse subjective quality into a deterministic score.
  • target URL, local route, HTML artifact, or interface screenshot set,
  • scenario file or user-task description,
  • optional surface_mode describing whether success is persuasion, operation, reading, or experience; this is planning context, never gate authority,
  • optional design_authority distinguishing observed incumbent values from owner-confirmed normative contracts,
  • optional content_profiles covering empty, short, typical, long, localized, bidirectional, numeric, and collection-size ranges,
  • optional context_matrix covering viewport, orientation, pointer/hover, keyboard, touch, zoom, motion preference, contrast/theme, and connection,
  • optional state_contract naming required loading, empty, error, success, recovery, persistence, permission, and interruption outcomes,
  • viewport and browser matrix,
  • product domain tags such as dashboard, ecommerce, authoring, service, marketing, game, or data tool,
  • source-card set or research artifact path,
  • fixture corpus path when calibrating,
  • desired output root, defaulting to output/playwright/ux-validator/<run-id>/.
output/playwright/ux-validator/<run-id>/

Required outputs:

  • run-metadata.json,
  • summary.md,
  • findings.json,
  • console-network.json,
  • accessibility/*.json,
  • aria/*.yml,
  • screenshots/*.png,
  • measurements/*.json,
  • traces/*.zip when interaction evidence is required,
  • residue-ledger.yml.
  • keep deterministic browser failures separate from UX risk proxies,
  • cite source cards for every non-trivial UX claim,
  • identify whether each expected outcome comes from a standard, an explicit product contract, a calibrated fixture contract, or a non-canonical method,
  • name what Playwright observed and what it cannot prove,
  • use hard gates only for deterministic or standards-backed failures,
  • require fixture calibration before promoting reusable hard gates,
  • capture screenshots, traces, accessibility output, ARIA snapshots, DOM measurements, console/network summaries, and residues when running browser validation,
  • keep domain-specific market rules behind explicit domain tags,
  • report human-study claims instead of pretending automation can measure them,
  • preserve development status when fixture or implementation evidence is missing.
  • keep observed incumbent values separate from owner-confirmed normative contracts, and never infer authority from repetition alone.
  • producing a single universal UX score,
  • claiming automated accessibility scans prove complete accessibility,
  • converting neuroscience or cognitive science references into deterministic browser truth,
  • treating market heuristics as universal rules outside their domain,
  • blocking dense expert interfaces without false-positive calibration,
  • turning external design taste, aggregate scores, detector output, or synthetic persona walkthroughs into evidence authority,
  • importing an external method's runtime, hooks, root files, visual system, or numeric heuristics when source cards and scenario declarations are sufficient,
  • accepting screenshot diffs without stable fonts, data, motion, and viewport controls,
  • treating a good spec as implementation or promotion evidence.
  • mode,
  • target URL or artifact type,
  • scenario count,
  • generated output count,
  • validator layer coverage L0-L6,
  • hard gate count,
  • soft flag count,
  • screenshot review count,
  • human-study residue count,
  • source-card coverage,
  • external-method card count and hard-gate-ceiling violations,
  • content, context, state, and fault-profile coverage,
  • fixture calibration status,
  • quality bar status,
  • anti-pattern hits,
  • reflection trigger recommendation.
  • implemented fixture corpus,
  • at least one known-good fixture,
  • accessibility, keyboard, layout, interaction, cognitive-risk, domain, and false-positive fixtures,
  • Playwright evidence reports for fixture runs,
  • hard gates catching expected deterministic failures,
  • L4/L5 claims remaining explainable and non-blocking unless independently justified,
  • external expert methods remaining planning or review inputs unless an independent source and calibrated fixtures justify a deterministic rule,
  • Experiment Harness report,
  • Sigil Development review.
## UX Evidence Validator Result

- Status: pass | flag | block | seed-only
- Mode: research | spec | fixture-plan | calibrate | validate-interface | report
- Target: <url, artifact, scenario, or research path>
- Evidence cards: <path or status>
- Claim classes: <hard gates, soft flags, screenshot review, human study, not automatable>
- Browser evidence: <output root or not run>
- Fixture calibration: pass | flag | block | not run
- Findings: <summary or path>
- Residue: <human-review or user-study claims>
- Validation: <checks performed>
- Next lifecycle step: <step>

Signals

GitHub stars
25
Forks
3
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
ux-evidence-validator
Source
github.com/cyberalchemyai/arcanum