comic-blind-comparison-review

SkillMedia

Phase-1 comic-author step (post-final eval) — a DOUBLE-BLIND A/B of two FINAL whole comics: our cross-model-audited progressive render (the comic-author + comic-director output) vs a naive single-shot baseline. A single sealed coin-flip hides which is which; two cross-model reviewers (Codex + Gemini) score both on a fixed rubric reading only a SHARED blind spec (intent + ART_BIBLE); only AFTER both reviews land do we unseal, re-label, and write a Chinese comparison.md + the A/B verdict nodes. editability/traceability is the structural wedge that can win even when the baseline looks prettier. Use when the user says "和 baseline 比", "blind comparison", "A/B 评测", "对比 baseline", "盲评", "whole-comic vs one-shot", or a finished comic needs a baseline-relative honest verdict. NOT an authoring skill and NOT the per-panel panel_gate / assembly_gate — those run DURING production; this runs AFTER, on two complete works.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the comic-blind-comparison-review skill

What this skill tells your AI

The instructions your AI receives, as published by wanshuiyin/aris-movie-director in skills/comic-blind-comparison-review/SKILL.md and read by ahel’s review.

The "vs baseline" capability the comic side otherwise lacks. The per-panel comic-director panel_gate and the cross-panel assembly_gate operate during production on single units; this skill operates after production, on two complete works — our cross-model-audited progressive comic (authored by the comic-author suite, baked by comic-director) vs a naive single-prompt one-shot baseline — and proves the spiral pipeline beats the one-shot, under a blind so neither reviewer can favor the home team. It is a direct, faithful port of aris_movie's blind-comparison-review from video to a whole-comic A/B (outputs/aris_movie_progressive.mp4 vs outputs/baseline_sd2_naive.mp4 → two rendered comic page sets), and it is an evaluation skill, not an authoring one: it never edits a comic, it judges two finished ones.

Cardinal lesson baked in as a gate, not prose: the order of operations IS the integrity guarantee. A single coin-flip seals .ab_mapping.json and chmod 600s it before any reviewer is invoked; the seal's mtime MUST precede both reviews; from seal-time until UNSEAL the orchestrator (Claude) may not read or quote the mapping; any home-team token in a reviewer prompt is a HARD-ABORT, not a warning. If the orchestrator reads the mapping before both reviews land, it may subconsciously phrase the synthesis to favor the home team — so the timestamp ordering is the audit trail, and "blinding integrity" is REPORTED, never assumed.

  progressive comic  ┐                                            ┌─▶ comparison.md (Chinese deliverable)
  (our audited)      ├─▶ ⓪ SEAL .ab_mapping.json (1 coin-flip,    │
  baseline comic     ┘       chmod 600, mtime BEFORE any reviewer) │
  (naive one-shot)   │            ▼                                │
                     │      ① STAGE pages → A/ and B/ (sealed ids) │
                     │            ▼                                │
                     │      ② CODEX blind review (xhigh, RO)  ─────┤  (no home-team token; reads SHARED
                     │            ▼                                │   blind spec = intent + ART_BIBLE only)
                     │      ③ GEMINI blind review (per-page + ─────┤
                     │            synthesis, auto-gemini-3)        │
                     │            ▼                                │
                     └─▶  ④ UNSEAL (verify seal_ts < BOTH reviews) ┘
                                ▼
                          ⑤ WRITE comparison.md  +  ⑥ WIKI A/B verdict nodes + edges

Constants

  • ARTIFACTS = two FINAL whole comics. progressive = our spiral output (comic-director's baked frames + the single-file viewer, or comic.json → render); baseline = a naive single-prompt one-shot comic of the SAME story+style (no spiral, no per-panel gate). For the worked example the contract boundary is examples/comic_m3_audit/comic.json (the structured-IR artifact under test) + examples/comic_m3_audit/ART_BIBLE.md (the SHARED blind spec both reviewers read).
  • MAPPING_FILE = outputs/.ab_mapping.json — the seal, chmod 600. Written by ⓪ before any reviewer.
  • REVIEWERS (two families, the load-bearing diversity) = Codex via a fresh mcp__codex__codex call, model_reasoning_effort: xhigh, sandbox: "read-only" (no model pin — the call follows the local codex CLI config; only the bake pins a model, via run_comic.get_bake_plan()) ‖ Gemini auto-gemini-3 (per-page mcp__gemini__analyzeFile + a text-only mcp__gemini-cli__ask-gemini synthesis). Never downgrade the effort tier (reviewer-routing). Optionally fold in the project's CC-narrative vote to make it the same tri-reviewer panel as panel_gate (CC ‖ Gemini ‖ Codex) — see Adaptation.
  • BLIND TOKENS = the only identifiers a reviewer ever sees are comic_A / comic_B (and per-page A_p01.png / B_p01.png). Banned from every reviewer prompt: progressive, baseline, ARIS, ARIS-Movie, SD2, naive, one-shot, spiral, our system, system under test, and any rhetorical "show that A is better" framing.
  • A/B RUBRIC DIMENSIONS (comparative, each scored 0–5 for comic_A AND comic_B) = story_comprehension, visual_consistency, overlay_readability, editability_traceability, reproducibility, polish. editability_traceability is the structural wedge (see the gate below).
  • NO PASS/FAIL THRESHOLD — this is comparative A-vs-B, not a gate. Each reviewer emits overall_winner ∈ {comic_A, comic_B, tie} + confidence 0.0–1.0; the consensus is ∈ {progressive, baseline, tie, disagree}.
  • OUTPUTS = outputs/.ab_mapping.json (sealed, 600); outputs/blind_review_codex_raw.json; outputs/blind_review_gemini_raw.json (+ outputs/gemini_perpage.json); comparison.md (the Chinese human deliverable, per output-language); the wiki A/B verdict nodes + edges (§ below). Stdout JSON {comparison_decision_id, consensus_winner, codex_winner, gemini_winner, blinding_integrity, comparison_md_path}.

Input contract — two FINISHED comics + ONE shared blind spec

This skill is pure evaluation; it never authors, edits, or re-bakes a comic.

  • Two final artifacts, both complete. Refuse (HALT) if either path is missing, if either comic is not a whole finished work (a half-baked spiral run is not a fair A/B subject), or if both paths are the same artifact. The progressive comic must be the shippable output (nothing escalated/needs_human/flagged per comic-director's run-report); a non-shippable progressive comic is not eligible for the head-to-head.
  • One SHARED blind spec read by BOTH reviewers — the story intent (the comic's logline/script, e.g. the intent_spec's logline + narrative_beats) + the style bible (ART_BIBLE.md). The spec is the identity-stripped ground truth ("what this comic is supposed to be"); it must NOT name which artifact is the home team. This is the comic mapping of the video skill's MOVIE_BRIEF.md + style_bible.md.
  • Reviewer-independence ≠ reviewer-blinding — BOTH apply (the lesson that names this skill). Independence = the reviewer sees no Claude prose (the panel_gate rule). Blinding = the reviewer sees no method IDENTITY of which artifact is the home team. The per-panel gate already does independence; this adds the second axis (identity-blind) for the whole-artifact A/B.

Procedure (followable, fail-closed, in strict order)

⓪ SEAL the A/B mapping (MUST complete BEFORE any reviewer)

Inputs (skill args): PROG_PATH = the progressive comic dir/artifact, BASE_PATH = the baseline comic dir/artifact, PROJECT_DIR = the project root the wiki lives under (where comic.json + wiki/ sit, e.g. examples/comic_m3_audit). Bind and FAIL FAST first (the "fail fast if either missing or identical" the prose promises must be ACTUAL CODE — the seal is the load-bearing integrity step and must not crash). Then a single coin-flip and write the seal — its mtime MUST precede the first reviewer call:

# --- bind the three skill args explicitly (these are THIS skill's inputs) ---
PROG_PATH="$1"; BASE_PATH="$2"; PROJECT_DIR="${3:-.}"
# --- fail fast: non-empty, both exist, and the two artifacts differ (prose promise → real code) ---
[ -n "$PROG_PATH" ] && [ -n "$BASE_PATH" ] || { echo "HALT: PROG_PATH and BASE_PATH are required args"; exit 2; }
[ -e "$PROG_PATH" ] || { echo "HALT: progressive artifact not found: $PROG_PATH"; exit 2; }
[ -e "$BASE_PATH" ] || { echo "HALT: baseline artifact not found: $BASE_PATH"; exit 2; }
[ "$(cd "$(dirname "$PROG_PATH")" && pwd)/$(basename "$PROG_PATH")" != \
  "$(cd "$(dirname "$BASE_PATH")" && pwd)/$(basename "$BASE_PATH")" ] \
  || { echo "HALT: the two A/B artifacts are the SAME path — not a fair head-to-head"; exit 2; }
mkdir -p outputs
if [ $(( $(od -An -N2 -tu2 /dev/urandom) % 2 )) -eq 0 ]; then PROG=A; BASE=B; else PROG=B; BASE=A; fi
# pass the two real paths as argv (NOT os.environ — they were never exported) so the heredoc can never KeyError
python3 - "$PROG" "$BASE" "$PROG_PATH" "$BASE_PATH" <<'PY'
import json, sys, datetime, os
prog, base, prog_path, base_path = sys.argv[1], sys.argv[2], sys.argv[3], sys.argv[4]
seal = {
  "sealed_at": datetime.datetime.now(datetime.timezone.utc).isoformat(),
  "comic_A": os.path.abspath(prog_path if prog=="A" else base_path),
  "comic_B": os.path.abspath(base_path if prog=="A" else prog_path),
  "progressive_label": prog, "baseline_label": base,
}
json.dump(seal, open("outputs/.ab_mapping.json","w"), ensure_ascii=False, indent=2)
PY
chmod 600 outputs/.ab_mapping.json

(The outputs/.ab_mapping.json seal stores os.path.abspath(...) and that is FINE — cli/validate_wiki.py only scans wiki/nodes/*.json, never the seal file. The abs-path discipline applies to the WIKI NODE payloads written in ⑥, not to this seal.) From here until ④ UNSEAL the orchestrator MUST NOT read or quote outputs/.ab_mapping.json. Reviewers see only comic_A / comic_B. (Procedural, not a technical guarantee — but it produces a paper trail sufficient for a research artefact, and that trail IS the integrity claim.)

① STAGE pages into A/ and B/ (the video frame-extraction collapses)

The video skill extracted ~10 frames/video into comparison-frames/{A,B}/; for comics this collapses — the comic pages ARE the discrete units, deterministic renders, so no extraction or regeneration is needed. Stage each comic's already-rendered page PNGs into the sealed letter-dirs with identity-free labels (A_pNN.png / B_pNN.png) so reviewers cite by page, never by home-team identity. The premise is deterministic pages, don't regenerate — so this skill stages existing PNGs and never bakes. Use the deterministic recipe for whichever input form each side ships as (apply it independently to comic_A's source and comic_B's source — neither side knows which is which):

Form (a) — a comic.json + WHOLE-PAGE renders (the structured-IR case). The canonical page order is comic.json pages[] (a LIST, each page = a discrete unit). Stage one page-level render per page (page.page_image / page.rendered_path — the whole page as the reader sees it: panels + HTML bubbles + narration), never a raw per-panel panel_attempt.image_path (that drops the bubbles/narration and, on a multi-panel page, the other panels — see the code's HALT). Know that the SHIPPED reference IR does NOT qualify: examples/comic_m3_audit/comic.json carries NO page_image/rendered_path on any of its 18 pages (verified), so stage() HALTs on it by design — Form (a) applies only after a separate whole-page render step has populated those fields; on the shipped IR (and any comic.json like it) route to Form (b) (rasterize the viewer). Copy into the sealed dir as one PNG per page, renumbered A_pNN.png/B_pNN.png in pages[] order (never re-bake):

mkdir -p outputs/comparison-pages/A outputs/comparison-pages/B
stage() {  # $1=comic.json FILE or a dir containing one   $2=letter (A|B)
  python3 - "$1" "$2" <<'PY'
import json, sys, shutil, os
arg, letter = sys.argv[1], sys.argv[2]
# NORMALIZE the artifact path: accept a comic.json file OR a dir holding one (never assume it's a dir).
cpath = arg if (os.path.isfile(arg) and arg.endswith(".json")) else os.path.join(arg, "comic.json")
if not os.path.isfile(cpath): sys.exit(f"HALT: no comic.json at {arg}")
cdir = os.path.dirname(os.path.abspath(cpath))
cj = json.load(open(cpath))
panels = cj["panels"]                 # dict {"S01": {...}, ...}
pages  = cj.get("pages") or []        # list, the canonical page ORDER; each has panel_ids
if not pages: sys.exit(f"HALT: {cpath} has no pages[] — cannot determine page order")
out = os.path.join("outputs/comparison-pages", letter)
def resolve(rel): return rel if os.path.isabs(rel) else os.path.join(cdir, rel)
n = 0
for pg in pages:                      # pages[] order IS the comic order (don't sort the dict keys)
    pids = pg.get("panel_ids") or []
    if not pids: sys.exit(f"HALT: page {pg.get('id')} has no panel_ids")
    # A fair A/B compares WHOLE PAGES (each panel + its HTML bubbles/narration as the reader sees it). EVERY
    # page needs a page-level render — a raw per-panel PNG is NOT a page (it drops bubbles/narration, and on a
    # multi-panel page the other panels). No single-panel shortcut: even a 1-panel page is rendered as a page.
    src = pg.get("page_image") or pg.get("rendered_path")
    if not src:
        sys.exit(f"HALT: page {pg.get('id')} has no page-level render (page_image/rendered_path) — rasterize the "
                 f"viewer to whole-page PNGs (Form b) before A/B; a per-panel PNG is not a page (no regen here)")
    src = resolve(src)
    if not os.path.isfile(src): sys.exit(f"HALT: missing rendered page image {src} for page {pg.get('id')}")
    n += 1; shutil.copyfile(src, os.path.join(out, f"{letter}_p{n:02d}.png"))
print(f"{letter}: staged {n} pages")
PY
}
# DO NOT call stage() blindly — it is Form-(a) ONLY (a comic.json with page-level renders). Dispatch each
# artifact by TYPE via stage_any() (defined after Form (b) below): comic.json → stage(); .html viewer →
# rasterize_viewer(). Both route by the SEALED coin-flip letter ($PROG/$BASE from ⓪), never by reading the seal.

(Sealing caveat — keep the blind intact.) The two stage calls above route each artifact to its sealed letter via the coin-flip variables $PROG/$BASE from ⓪ — not by reading comic_A/comic_B from the seal file, which is FORBIDDEN until ④. The orchestrator knows the letters (it flipped the coin) without ever opening the mapping, so the blind holds; the reviewer still only ever sees A_*/B_* filenames.

Form (b) — a single-file HTML viewer with NO pre-rendered PNGs (the common comic-director ship form). This skill does not own a bake; it is fail-closed by design. Prefer pre-rendered PNGs; if absent, the caller must rasterize first via the repo's own deterministic page renderer, then re-invoke this skill pointing at the PNG dir. If you must rasterize inline, the ONLY sanctioned path is a headless-Chromium screenshot at a FIXED width/DPI (deterministic), one PNG per page, e.g.:

# headless rasterize the single-file viewer → one PNG per page (deterministic fixed window so both sides are
# pixel-comparable). The viewer's per-page anchor is `?p=N` (the comic-director viewer convention).
rasterize_viewer() {  # $1=viewer .html   $2=letter   $3=NPAGES (from comic.json pages[])
  CHROME="$(command -v chromium || command -v google-chrome || command -v chromium-browser || true)"
  [ -n "$CHROME" ] || { echo "HALT: no headless chromium to rasterize the viewer — supply pre-rendered page PNGs instead"; exit 2; }
  local v abs; abs="$(cd "$(dirname "$1")" && pwd)/$(basename "$1")"
  mkdir -p "outputs/comparison-pages/$2"
  for i in $(seq 1 "$3"); do
    pp=$(printf "%02d" "$i")
    "$CHROME" --headless --disable-gpu --window-size=1200,1600 \
      --screenshot="outputs/comparison-pages/$2/$2_p${pp}.png" "file://${abs}?p=${i}" >/dev/null 2>&1 \
      || { echo "HALT: screenshot of page $i failed"; exit 2; }
  done
  echo "$2: rasterized $3 pages from the viewer"
}
# e.g.  rasterize_viewer "$PROG_PATH" "$PROG" "$NPAGES"   (NPAGES = len(comic.json pages[]))

If neither pre-rendered PNGs nor a headless renderer is available → HALT with that message; never feed the reviewers a partial or wrong-DPI rasterize (a corrupted page set makes the A/B score garbage).

Stage dispatch — the actual control flow (picks the form per artifact; never calls stage() blindly).

stage_any() {  # $1=artifact (a DIR of pre-rendered *_pNN.png / a comic.json / a single-file .html viewer)   $2=sealed letter
  case "$1" in
    *.html)  # the SINGLE-FILE viewer (the common ship form, NO sibling comic.json): rasterize whole-page PNGs.
             # derive NPAGES from the viewer's EMBEDDED JSON island (the build inlines comic.json into a <script>).
      NP=$(python3 - "$1" <<'PY'
import json, re, sys
html = open(sys.argv[1], encoding="utf-8").read()
n = ""
for blob in re.findall(r'<script[^>]*>\s*(\{.*?\})\s*</script>', html, re.S):   # the inlined comic-IR island
    try:
        d = json.loads(blob)
        if isinstance(d.get("pages"), list): n = len(d["pages"]); break
    except Exception:
        pass
print(n)
PY
)
      [ -n "$NP" ] || { echo "HALT: cannot read page count from the viewer's embedded JSON island in $1"; exit 2; } ;
      rasterize_viewer "$1" "$2" "$NP" ;;
    *)
      if [ -d "$1" ] && ls "$1"/*_p[0-9]*.png >/dev/null 2>&1; then   # a DIR of pre-rendered whole-page PNGs
        mkdir -p "outputs/comparison-pages/$2"; i=0
        for f in $(ls "$1"/*_p[0-9]*.png | sort); do i=$((i+1)); cp "$f" "outputs/comparison-pages/$2/$2_p$(printf '%02d' "$i").png"; done
        echo "$2: staged $i pre-rendered page PNGs"
      else
        stage "$1" "$2"   # a comic.json file or a project dir — stage() normalizes + requires page-level renders
      fi ;;
  esac
}
stage_any "$PROG_PATH" "$PROG"   # route by the SEALED letter; no read of the seal file → blind holds until ④
stage_any "$BASE_PATH" "$BASE"

Page-count parity gate (fail-closed). After staging both sides, assert A and B staged the same number of pages — an unequal A/B is not a fair head-to-head:

nA=$(ls outputs/comparison-pages/A/*.png 2>/dev/null | wc -l | tr -d ' ')
nB=$(ls outputs/comparison-pages/B/*.png 2>/dev/null | wc -l | tr -d ' ')
[ "$nA" -gt 0 ] && [ "$nA" = "$nB" ] || { echo "HALT: page-count parity failed (A=$nA B=$nB) — not a fair A/B"; exit 2; }

Page filenames embed only the sealed letter + a page index (A_p03.png) so reviewers cite by page, never by home-team identity.

② CODEX blind review (read-only, xhigh)

A fresh mcp__codex__codex call (NOT codex-reply), config {"model_reasoning_effort":"xhigh"}, sandbox: "read-only". It sees the two page dirs (comparison-pages/A, comparison-pages/B) + the SHARED blind spec (the identity-stripped intent + ART_BIBLE.md). Scrub the prompt with the banned-token scan (Constants) BEFORE submission. The prompt asks Codex to read the pages itself and score each comic 0–5 on every A/B rubric dimension, cite page filenames as evidence, and output JSON only. Save verbatim to outputs/blind_review_codex_raw.json. Required JSON shape:

{"comic_A": {"<dim>": {"score": 0-5, "evidence": "...cites A_p0X.png..."}, ...},
 "comic_B": {"<dim>": {"score": 0-5, "evidence": "...cites B_p0X.png..."}, ...},
 "overall_winner": "comic_A|comic_B|tie", "confidence": 0.0-1.0, "rationale": "..."}

If Codex returns non-JSON → retry once, stricter; still non-JSON → write a {"parse_failed": true} stub and continue Gemini-only, noting the degradation in comparison.md.

③ GEMINI blind review (independent, same blind rules)

3.1 per-pagemcp__gemini__analyzeFile, model auto-gemini-3, one call per image (every page in A/ and B/), each page NOT seeing the other comic. Per-page JSON {label, page_filename, character_appearance ≤150, scene_signature ≤80, style_features ≤80, artifact_severity_0_to_5 (int, 0=clean 5=catastrophic), notable_anomalies[]}; persist to outputs/gemini_perpage.json. 3.2 synthesismcp__gemini-cli__ask-gemini (text-only), fed both per-page observation lists + the blind spec summaries (≤1000 chars each), scoring the same 0–5 A/B rubric, JSON only. Save verbatim to outputs/blind_review_gemini_raw.json. Gemini is always auto-gemini-3. Fail fast if fewer than the staged pages succeed per comic (the video floor was <6 frames/video → abort; for comics use "every staged page must return, else re-run that page once then abort").

④ UNSEAL — only NOW read the mapping

Read outputs/.ab_mapping.json for the first time. Verify the seal predates BOTH reviews (its mtime is older than both raw review files):

# portable mtime probe — BSD stat is `-f %m`, GNU stat is `-c %Y`; python3 works identically on both:
mt() { python3 -c 'import os,sys; print(int(os.path.getmtime(sys.argv[1])))' "$1"; }
SEAL=$(mt outputs/.ab_mapping.json)
CODEX=$(mt outputs/blind_review_codex_raw.json)
GEM=$(mt outputs/blind_review_gemini_raw.json)
[ "$SEAL" -lt "$CODEX" ] && [ "$SEAL" -lt "$GEM" ] && echo "intact" || echo "compromised"

If out of order → blinding_integrity = "compromised": still produce comparison.md, but flag it prominently, never silently elide. Re-label comic_A/comic_B back to progressive/baseline in the PARSED reviews; keep _A/_B in the RAW files for audit (never edit a raw file after it is written).

⑤ WRITE comparison.md (Chinese — per output-language)

Six fixed sections, in Chinese:

  1. 执行裁定 / Executive verdict — each reviewer's winner + confidence; the consensus.
  2. 逐维度评分表 / Per-dimension table维度 | Progressive (Codex / Gemini) | Baseline (Codex / Gemini) | Δ across all six A/B dimensions.
  3. 证据 / Evidence — verbatim per-dimension evidence + page citations (A_p0X.png re-labeled to which artifact it actually was, after unseal).
  4. 可编辑性 / 架构差异表 / Editability table — the wedge made concrete (single-unit regen cost, failure localizability via wiki edges, minimum human-patch unit, audit record). See the worked-example table below.
  5. 盲评审计 / Blinding audit — mapping path, seal→first-review→last-review time window, leakage yes/no, blinding_integrity.
  6. 结论 / Conclusion — 3–5 sentences, honestly recording any dimension the baseline won (hiding it defeats the cross-model adversarial check).

⑥ WIKI — the A/B verdict as schema-valid nodes + edges

Write the post-final A/B record per schemas/node_schema.json (§ below), append the edges, and print the stdout JSON. The reviewers (different model families) — never this orchestrator — produce the verdict; the orchestrator only seals, stages, unseals, and records (acceptance-gate: the loop drives, it cannot acquit).

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
61
Forks
4
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
comic-blind-comparison-review
Source
github.com/wanshuiyin/aris-movie-director