Lofn QA — Codex-backed adversarial gate
SkillMediaAudit, validate, and classify completed Lofn pipeline outputs against the strict quality gates, backed by Codex. Use after a lofn-music / lofn-image / lofn-video / lofn-story run, or on a suspicious partial run, to get a SHIP / REPAIR / FAIL verdict and a repair brief. Do NOT use for creative generation — this is the adversarial auditor, not the artist.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Lofn QA — Codex-backed adversarial gate skill
What this skill tells your AI
The instructions your AI receives, as published by localsymmetry/lofn in .agents/skills/lofn-qa/SKILL.md and read by ahel’s review.
The competitive auditor. It proves two things at once: a listener/viewer can grasp the surface, and a second pass reveals the cathedral. It fails both extremes — impressive obscurity AND competent blandness. QA stays strict: it does not loosen structural gates to protect "creative freedom."
Replaces the OpenClaw Python validators with Codex judgment + the same checklists. Adopt the auditor stance from skills/orchestration/references/adversarial_qa_stance.md first.
Procedure
- Run in a fresh, clean-context judgment subagent. The QA / Somatic / Step-11-reject judgment MUST execute in a clean Agent-tool subagent — never appended to the thread that generated the artifact. A generator grading its own homework is the conflict that ships corpses past the Andon Cord. Music uses two sealed stages: Stage A receives only the shuffled plain-lyric plate +
music_absolute_force_gate.mdquestions and saves its receipt; only after that file is frozen may Stage B receive the artifact + verbatim ICB + full gate spec +GATE_REPORT.json. Do not put prompts, scores,[Theme:],[SONG FORM:], EMO/Disc headers, rationales, or Golden identities into Stage A. Other modalities receive the artifact + ICB + gate evidence directly. Tier follows the role (judge, not generate), never the step number; this is a dedicated clean-context spawn, not a per-spawn flag that can silently fall back. - Identify the run directory + modality.
- Load the gates just-in-time:
- All modalities:
skills/qa/references/qa_full_legacy.md(the full tuned procedure — authoritative). - Music: read
skills/qa/references/music_absolute_force_gate.mdfirst, thenskills/qa/references/suno_15_point_qa.md+skills/qa/references/eligibility_7_properties.md. Freeze the plain-lyric absolute-force receipt before package/ICB evidence; then classify ACCESSIBLE vs AMBITIOUS and run the 16-point gate. Eligibility classification cannot clear Absolute Force. - Image:
vault/VISION_QA_DEPTH_AUDIT.md(Visual Somatic Gate + 7-element density checklist). - Video/Animation:
vault/DIRECTOR_QA_DEPTH_AUDIT.md(Cinematic Somatic Gate + 5-element shot checklist). - Thresholds: numeric bands are read from
vault/gates.yamlwhere present (the single source for the 850–1000 /<5000/ 70–120 / ≥80-word numbers);EXECUTION.md§4 is authoritative if the two disagree.
- All modalities:
- Music: freeze Stage A before deterministic evidence. Other modalities: read deterministic evidence first. Once the music plain-lyric receipt exists, or immediately for another modality, load
GATE_REPORT.jsonifscripts/validate_step.pywrote one and paste its measured{pair, step, check, expected, actual, pass}rows verbatim as proof-of-fix evidence in the Evidence column. These are computed values, not QA guesses, and they never retroactively alter the frozen listener read. The helper is fail-open: if absent or warning, fall back to in-prose measurement and note it; never block a valid run on a missing/broken helper. - Verify ICB integrity (cheap proof, not a soul substitute). Confirm the canonical ICB / CREATIVE_CONTEXT prefix appears as an unbroken verbatim substring at the head of each step's prompt (additions below pass; edits above the block fail), and that the count of
(afterspeaker tags == 18. A voice count ≠ 18 or a broken prefix is caught asREPAIR — THREAD LOSS. The count proves presence, not fidelity — a paraphrase can match length and voice-count. So this is only the cheap tripwire for a gross drop; the human personality-fidelity read (below) stays the real guarantee, and the count never substitutes for the soul read. Never summarize, trim, or "optimize" the ICB to make a count easier. 4.5. Music Absolute Force bracket — plain lyrics, actual house bar. The coordinator (not the judge) builds one shuffled Stage-A plate per finalist with: the candidate's plain sung words, all three core controls (Triple Arch Over Me,Spare Hair Tie,Burning Bush With a Barcode), and one known-mediocre decoy. Usescripts/build_blind_set.py --mode lyrics; keep the key outside the blind directory. AddASK FOR IT BY NAMEandThe Wedding Takes It Louderas challenge controls when cleverness, formal machinery, moral debt, or dense constraint is central. The judge freezes the six-question receipt and rank before identity reveal. Readings:- decoy above any core control → judge broken; halt and re-run;
- candidate above decoy → zero positive SHIP evidence;
- candidate below all three core controls →
REPAIR — BELOW ABSOLUTE BAR; - mixed result → at most
HOLD, with exact win/loss evidence; - above all three → may continue, never automatic SHIP.
Full-package golden+decoy comparison may still audit package judgment, but cannot contribute positive evidence to Absolute Force or Overall SHIP. The complete contract is
music_absolute_force_gate.md.
- Run the gate without weakening any check. For music, run the pressure/mechanism proxy-capture sentences and evidence-boundary check before Somatic voting. Apply the Codex-native self-check gates in
.agents/skills/lofn/EXECUTION.md§4 for structural/pipeline integrity. - Write
QA_REPORT.mdin the run directory, including the structured evidence block (below) BESIDE the verdict. Music reports also include the frozen Plain-Lyric First Read, Absolute High-Bar Bracket, Evidence Boundary, Proxy-Capture Check, and anAbsolute Force: CLEAR / HOLD / REPAIRverdict; without them Overall cannot be SHIP. If failures require rerun, write the repair brief in the rerun format fromqa_full_legacy.mdand route to the failing step (09/07 for music thread loss, 08 for prompt-format, etc.). REDIRECT — mandatory when the gate is stuck (EXECUTION.md§7.3): if the specific failed gate's value has not moved across attempts (the no-progress predicate), the repair brief MUST carry a sideways PROPOSAL beside the return target — one of: promote a step-05 cut-ledger reserve concept, re-derive that pair's variation angles, or re-run the panel's skeptic transformation for that pair's slice. The Skeptics' COUNTER-MOVES (Somatic Gate, below) are the raw material for this proposal. The brief proposes; executing sideways is a coordinator decision surfaced to the human — the frozen ICB is never edited mid-run, a sideways route spawns a NEW pair artifact chain. - On SHIP (every shipped/selected piece): append ONE curated failure-ledger entry to
vault/COMPETITION_LEARNINGS.md— see "Failure-ledger write-back" below.
What every PASS must clear
- Pipeline integrity / granularity: coordinator steps exist as separate files (
step00…step05), and steps 06–10 exist as separate per-pair files (pair_{NN}_step06…step10). A collapsedpair_{NN}_steps_06_10.mdrollup or a single batch run is a blocking failure even if filenames look present. 6 pairs × 4 = 24 outputs unless the Scientist downsized. - ⛔ NO-SKIP / NON-CANONICAL (
EXECUTION.md§4): steps 07, 09, and 10 artifacts exist for every non-quarantined pair. A run missing its editorial spine is NON-CANONICAL — the Overall verdict can never be SHIP and the run cannot be published under Lofn's name, no matter how clean every other gate reads. (2026-06-28: a run that skipped 07/09/10 shipped 6/6 because the gates measured structure and couldn't see nobody wrote the arrangements. Never again.) Mark the report headerNON-CANONICAL RUNwhen this fires. - Continuity / ICB: every step cites the continuity payload (Special Flairs + all 18 panel voices + Golden Seed + active personality + previous artifact). Missing →
REPAIR — THREAD LOSS, even if formatting passes. - Personality fidelity: the piece proves which personality made it (sonic-world/voice sentence + signature device + seed-derived weirdness). If any competent prompt could have produced it →
REPAIR — SOUL LOSS. - Music Absolute Force: the frozen plain-lyric receipt names a listener-ownable pressure, first necessity, first-cycle surface, apparatus-removal result, section continuity, and one borrowable line; the candidate completes the actual three-control bracket. Package compliance and successful semantic rekey cannot substitute. Missing receipt → Overall SHIP impossible.
- Somatic Gate: the 3 Hyper-Skeptics freeze votes independently before discussion. For music they vote separately on Recognition / Necessity / Evidence Boundary; for other modalities they retain "distinctive enough to be Lofn, or generic?" 2 of 3 NO on any axis = BLOCKED. One supported music NO on Recognition or Evidence Boundary makes Overall at most HOLD until direct lyric-surface counterevidence is recorded — scores, majority enthusiasm, compliance, and promised production cannot erase it. This is the primary gate. No structured evidence block, measured count, or helper report below ever overrides, substitutes for, or pre-empts the somatic read. Quantify the corpses; never pretend to quantify the soul. Every Skeptic NO vote carries a one-line COUNTER-MOVE — "the one change that would make this unmistakably Lofn" — a new angle, a register rotation, or a step-05 cut-ledger reserve concept. Hard critique proposes, it never just vetoes; the counter-moves feed the repair brief's REDIRECT/sideways proposal (procedure step 6). A NO with no counter-move is an incomplete vote.
Named-corpse Andon checklist (prompts for the 3-Hyper-Skeptic bloc)
These are enumerated reject-conditions worded as prompts the Hyper-Skeptics carry into the Somatic Gate — not scores, not auto-reject floors, not a Python detector. Pulling the cord is never failure; shipping past it is. A Skeptic who flags one must cite the named condition with evidence ("emotion goes nowhere — no second movement") rather than a bare "feels generic." Under-flagging is acceptable; the conditions exist to give the human a sharper language, not to auto-fail. Walk each:
- One-note emotional arc — does the piece reach a second movement, or does it hold a single register start to finish with nowhere to go?
- Single dominant repeated section — is one section (chorus/refrain/stanza) doing all the load-bearing work while the rest is filler around it?
- A motif that never transforms — does the central image/phrase/hook change across the piece (recontextualized, inverted, paid off), or just recur unchanged?
- Repeated-line collapse — has repetition stopped being a deliberate device and become the piece running out of things to say?
- Late person / procedural first cycle — do thesis, policy, device setup, or procedure occupy the first hook cycle while the affected person or recognizable pressure arrives only later?
- Proxy capture — is the formal mechanism easier to explain, and more consequential, than why the song matters?
- Production-dependent text — does the decisive event exist only in an unrendered pan move, missing beat, bass reversal, formant event, or other prompt intention?
- Concrete proof set — are body/place/object details accumulating as evidence for a proposition rather than deepening one stable human pressure?
A deliberate refrain, villanelle, incantatory INDIGNATION dirge, or sustained-register elegy is NOT a corpse. These prompts ask the Skeptic to tell intentional return from creative exhaustion — a judgment only a human read makes. Never convert any of these into a numeric threshold that gates SHIP/FAIL.
Modality hard gates (non-waivable — a custom/looser parent checklist cannot waive these)
Phrase each check as a natural-language assertion that names the concrete checkable value — e.g. "the EXCLUDE field is present and its content is under cap", "the lyric carries a 3+3 emotional duality", "the MUSIC PROMPT is a dense paragraph of 850–1000 chars, NOT bracket tag-soup". Evaluate the assertion by meaning, not by an exact-string grep that false-fails on minor format drift; the concrete band / EMO-header shape stays inside the assertion so "robust" never decays into vibes.
- Music:
music_absolute_force_gate.mdis non-waivable and runs before the package read. Then check a standalone## 1. MUSIC PROMPT(a dense paragraph of 850–1000 chars — NOT bracketed key:value tag-soup — naming no real artist) + a separate EXCLUDE field present and under cap; lyrics open[Theme:…]then[SONG FORM:…]; full EMO headers[Section - EMO:<emotion> - <Role> - <cue>]use taxonomy emotions, not bare AWE/INDIGNATION; ≥1 SFX cue; 70–120 sung lines; the Suno lyrics field stays<5000chars (target ≤4800); verse-structure diversity across the 6; Lineage & Credit block on living-scene genres. ⭐ THE DEDICATION GATE (vault/DEDICATION_STANDARD.md) — a missing player-credit is a REPAIR, not a note: every shipped piece has a dedication measured ≤500 chars; and ⛔ every source-sound figure named in a style field is credited to the PERSON WHO PLAYED IT — sweep the shipped prompts foramen/funky drummer/think/apacheand named historical figures of sound (amen → Gregory C. Coleman / The Winstons). A block naming three living artists and omitting the drummer whose four bars are the floor passes the old lineage check and fails this one — that is the 2026-08-13 finding, caught pre-render on our own work. Also verify the honesty clause: a generated track is not a sample, and must never claim to be. Run the full Absolute-Force + 7 Singer-Surface + 5 Cathedral-Engine + 3 Suno-Package + Lineage gates. (Note: the music SKILL's 2026-06-09 mandate is dense-paragraph prompts; wheresuno_15_point_qa.mdstill says "categorized key:value", the paragraph mandate is the newer authority — flag the conflict, do NOT fail a correct paragraph prompt for the stale bracket rule.) - Image: noun-first present-tense scene description, no imperative openers, medium named early, emotion shown not named, no Storybook-Assassin ban-words ("ethereal/dreamlike/whimsical/gentle light/soft glow/magical/delicate"), no living-artist names, subject legible at first glance, ≥80 words (Flux) / five-slot directive (GPT_I2).
- Video/Animation:
[CAMERA]+[SUBJECT]+[ACTION]+[SETTING]+[STYLE & AUDIO]; audio directed explicitly; loop logic stated (animation); distinct camera grammar across pairs. - Story: standalone distinct voice per pair, body-before-thesis, world coherence, no generic AI cadence.
Human-Subject identifiability backstop (HOLD-FOR-HUMAN)
QA enforces the identifiability backstop from vault/HUMAN_SUBJECT_STANDARD.md. The line forbidden is identifiability, not subject matter: a real, named, or recognizably-depicted private individual rendered in a way that ties the piece to that person. On any reasonable suspicion of an identifiable real human subject, the verdict is HOLD-FOR-HUMAN — not an auto-FAIL, not a silent SHIP. The piece does not ship and is surfaced to the human for an explicit decision. This backstop sits above every numeric and somatic gate and is non-waivable by a looser parent checklist.
Structured QA evidence block (BESIDE the verdict, never replacing it)
Surface the measured values QA already reasons about as a structured block alongside SHIP / REPAIR / FAIL so verdicts are trace-auditable and trendable across runs — a floor under the judgment, never a substitute for it. Quantify the corpses; do not pretend to quantify the soul. Report, for the modality:
- char counts vs caps — MUSIC PROMPT chars vs 850–1000; Suno lyrics field vs
<5000(target ≤4800); Flux ≥80 words. - sung-line count vs 70–120 (music).
- EMO-tag balance — taxonomy-emotion distribution; flag a single dominant emotion or bare AWE/INDIGNATION tags.
- repeated-line / n-gram collapse ratio — reported as a FLAG only, with chorus/refrain exempt; a deliberate refrain must never auto-fail on this number. It routes to the human/Somatic read, it does not decide.
- boundary-hugging + house-lexicon + one-fact FLAGs (
gates.yaml2026-07-01 rows) — prompt chars pinned ≥985 / sung lines pinned at the floor; anyhouse_lexiconhit (calcified golden-output phrasing); >1 sung numeric-fact line. All FLAG-only, all routed to the Somatic read — a house-lexicon hit is primeSOUL LOSSevidence (self-copying is a soul failure even when every count passes). WhereGATE_REPORT.jsonexists, populate these from its measured rows. A piece that passes every count can still be REPAIR — SOUL LOSS by the somatic read. The structured block sits beside the 3-Hyper-Skeptic verdict; it never displaces it.
Failure-ledger write-back (advisory, capped, INDIGNATION-exempt)
On every shipped/selected piece, append exactly ONE curated entry to vault/COMPETITION_LEARNINGS.md capturing: theme-tags (theme-type / venue / modality), what the gate caught, one transferable rule, and a confidence %. Discipline:
- Advisory, never a constraint. These are failure-ledger / process notes a future run reads as advice, confidence-stamped, LOW until human- or second-run-corroborated. An auto-written entry may only feed the corpse/process checklist — it must NEVER become an aesthetic constraint (no "warm palette mandatory", no "INDIGNATION underperforms"). Only a human promotes a lesson to a hard creative constraint.
- Venue/modality-scoped. A lesson is tagged to its venue+modality and must not leak into music or non-competition runs.
- Triggered-INDIGNATION is explicitly EXEMPT from suppression. Never write an entry, and never let one stand, that would suppress or tune away triggered-INDIGNATION work toward what a voting venue rewards.
- Hard-capped (~25 live curated entries). Growth is by disciplined append-and-prune, not unbounded accumulation; prune the lowest-confidence stale entry when over cap.
- Apply the mandatory "would this lesson have hurt our best past entry?" check before recording — if yes, do not record it as a rule.
Operational-ledger write-back (infra/process only — FIREWALLED from the aesthetic ledger above)
If the run hit an operational failure — a dropped/quarantined pair, a stale path, a quota/API error, a timeout, a gate that false-failed a valid artifact, a resume/manifest disagreement — append ONE row to vault/RUN_LEDGER.md with the fixed schema {date · run_id · what_broke · root_cause · infra_fix · status}. This is the operational ledger, not the aesthetic one: it records ONLY plumbing facts and never a taste claim (taste lessons stay in COMPETITION_LEARNINGS.md, INDIGNATION-exempt). Nothing is written if the run had no operational failure. It is read back at Phase −1 as a heads-up, never as a creative constraint.
Release the run lock (last action, after the INDEX)
python3 scripts/run_lock.py release <run-dir> — see EXECUTION.md §5.1. The file is kept as the run's record, and the directory stays claimed: a finished run's artifacts are at least as destroyable as a live run's, so a different run is still refused and must use its own directory. If release reports the lock is not yours, that is an operational failure — write the row, and say plainly in the report that this run may have been writing into another run's directory.
Verdicts (every report)
- Pipeline Integrity Verdict: PASS / REPAIR REQUIRED / FAIL / NON-CANONICAL (no-skip rule fired — SHIP impossible)
- Package Verdict (modality contract): PASS / REPAIR REQUIRED / FAIL
- Human-Subject Verdict: CLEAR / HOLD-FOR-HUMAN
- Absolute Force Verdict (music): CLEAR / HOLD / REPAIR
- Overall: SHIP / HOLD / REPAIR / FAIL (the Somatic Gate is decisive; music also requires Absolute Force CLEAR; the structured evidence block is a floor beside it, never the verdict)
The publish bar (standards never float)
- Borderline defaults to HOLD. SHIP means unambiguously good enough to publish under Lofn's name. A piece the Skeptics or the Scientist would call "borderline" gets HOLD, not SHIP — held pieces wait for repair, a stronger run, or a deliberate human override recorded in the report. An empty publish day is an acceptable outcome; a lowered bar is not. The 2026-07-01 lesson: the standard slipped exactly once — when nothing in the run was good and something "had to" ship. Nothing has to ship.
- Actual controls, not a soft proxy. For music, candidate > decoy proves only non-filler. The core absolute bracket is
Triple Arch Over Me+Spare Hair Tie+Burning Bush With a Barcode; failure to clear them routes to retry/HOLD undermusic_absolute_force_gate.md. Cleverness-heavy or constraint-heavy runs also useASK FOR IT BY NAMEandThe Wedding Takes It Louderas challenge controls. Model agreement never overrides The Scientist's ear. - The zero-rejection tripwire (
EXECUTION.md§7.5). A healthy 6-pair run produces ≥1 REPAIR or substantive escalated FLAG. A full run reporting 0 repairs + 0 holds + 0 quarantines across 24 artifacts triggers an audit of the judge (re-run the plain-lyric Absolute Force bracket on a sample), not a celebration. QA that never says no is decorative.
Report format (QA_REPORT.md)
# Lofn QA Report — <run> (<modality>)
## Verdicts
- Pipeline Integrity: …
- Package: …
- Human-Subject: CLEAR | HOLD-FOR-HUMAN
- Absolute Force (music): CLEAR | HOLD | REPAIR
- Overall: …
## Plain-Lyric First Read (music; frozen before package/ICB reveal)
## Absolute High-Bar Bracket (music; 3 core controls + decoy)
## Evidence Boundary (music; text vs production intent vs rendered fact)
## Proxy-Capture Check (music; pressure sentence vs mechanism sentence)
## ICB Integrity
- prefix verbatim substring: yes/no · `(after ` voice count: N (==18?) · injected bytes: N
## Structured Evidence Block (measured values, BESIDE the verdict — from GATE_REPORT.json where present)
| Metric | Measured | Band/Cap | Pass/Flag |
(char counts vs caps · sung-line count vs 70–120 · EMO-tag balance · repeated-line ratio = FLAG only, chorus-exempt)
## Score Table
| # | Gate | Verdict | Evidence | Repair |
## Blocking Fails
## Somatic Gate (3 Hyper-Skeptics) — PRIMARY; cite named-corpse conditions with evidence
## Required Repairs (routed to step N)
## Failure-ledger entry appended (on SHIP) — theme-tags · what the gate caught · transferable rule · confidence
## Final Recommendation
A strong artifact with weak structure still fails; a soulful artifact missing required pieces still fails structurally. A piece that passes every measured count still fails if the Somatic Gate reads SOUL LOSS — the numbers are a floor, the somatic read is the guarantee. Fix, then re-QA before delivery.
Signals
- GitHub stars
- 22
- Forks
- 1
- Last commit
- Aug 2026
Advanced
- Catalog kind
- skill
- Gateway key
lofn-qa- Source
- github.com/localsymmetry/lofn