fortify

SkillMonitoring & ops

Systematic ablation study runner. After research:run finds improvements, fortify identifies component candidates from git diff + diary, creates isolated git worktrees per ablation (main repo never modified), runs metric+guard in each worktree, ranks component importance, and optionally generates reviewer Q&A calibrated to a target venue.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the fortify skill

What this skill tells your AI

The instructions your AI receives, as published by borda/ai-rig in plugins/cc_research/skills/fortify/SKILL.md and read by ahel’s review.

Ablation study runner — after /research:run finds improvements, fortify identifies which components contributed, generates ablation variants (remove one component at a time), runs each in isolated git worktrees (main repo never modified), ranks component importance, optionally generates reviewer Q&A calibrated to venue.

NOT for: initial optimization loop (use /research:run); methodology validation (use /research:judge); paper-vs-code consistency (use /research:verify); hypothesis generation (use research:scientist directly). Fortify runs ablation studies on completed runs only.

MAX_ABLATION_CANDIDATES:  8 (ceiling — scientist produces 3–8; --max-ablations caps further)
METRIC_TIMEOUT_MS:        360000 (6 min — deliberately larger than run's 120s/300s VERIFY_TIMEOUT_SEC, to accommodate slower ablation-variant runs)
GUARD_TIMEOUT_MS:         360000
GIT_OP_TIMEOUT_MS:        15000
SANITY_DIVERGENCE_PCT:    2.0 (full-variant vs best_metric mismatch threshold)
IMPORTANCE_CLASS_CRITICAL: 50.0 (% of full metric lost)
IMPORTANCE_CLASS_SIGNIFICANT: 10.0
FORTIFY_DIR_BASE:         .experiments
STATE_DIR_BASE:           .experiments/state
METRIC_CMD_SOURCE:        state.json .config.metric_cmd (campaign's real metric — no fail-open default; env METRIC_CMD overrides)
GUARD_CMD_SOURCE:         state.json .config.guard_cmd (campaign's real guard — no fail-open default; env GUARD_CMD overrides)

Environment overrides — set these before invoking the skill to override per-variant defaults:

  • METRIC_CMD — command run inside each variant worktree to measure the ablation metric (override; default sourced from source-run state.json .config.metric_cmd)
  • GUARD_CMD — command run inside each variant worktree to detect regressions (override; default sourced from source-run state.json .config.guard_cmd)
  • STATE_DIR_BASE — base directory for source-run state lookups (default: .experiments/state)
  • FORTIFY_DIR_BASE — base directory for per-run fortify artifacts (default: .experiments)

The constants block defaults above are YAML-only — bash blocks read environment variables (with ${VAR:-default} fallback) and never source values directly from YAML.

  • Key boundaries: end of F2 — ablation candidates identified by scientist; end of F4 — all ablation variant worktrees completed.
  • Preserve at F2: FORTIFY_DIR (TMPDIR key), RUN_ID (TMPDIR key), ablation-candidates.jsonl path, best_metric from source run.
  • Preserve at F4: FORTIFY_DIR, RUN_ID, results.jsonl path — ready for F5 importance ranking.
  • F4 loop is resume-safe: 4a-init's resume guard skips variants already terminal (non-timeout) in results.jsonl, so a mid-loop compaction resumes at the first pending variant — never re-runs a completed ablation; a timed-out variant still retries.
  • Clear at F1 start (stale prior run) and at start of F8 terminal summary.

Agent Resolution

Agent resolution: load and follow the protocol below. Contains foundry check + fallback table. If foundry not installed: use table to substitute each foundry:X with general-purpose. Agent this skill dispatches: research:scientist (same plugin — no fallback if research plugin installed).

# loads: compaction-contract.md
export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
_RESEARCH_SHARED=$(python "${CLAUDE_PLUGIN_ROOT:-plugins/cc_research}/bin/resolve_shared.py" 2>/dev/null)  # timeout: 5000
[ -z "$_RESEARCH_SHARED" ] && { echo "! Plugin path resolution failed — ensure research plugin installed and CLAUDE_PLUGIN_ROOT set, or invoke /research:fortify from project root."; exit 1; }
echo "$_RESEARCH_SHARED" > "${TMPDIR:-/tmp}/research-shared-${CSID}"  # cold resolve — every later site reads this sentinel instead of re-running python
cat "$_RESEARCH_SHARED/agent-resolution.md"
AgentFallback if absent
research:scientistgeneral-purpose (ablation candidate identification and reviewer Q&A quality reduced — ⚠ general-purpose agent may not emit the JSON envelope this skill parses; surface partial output and surface ⚠ in F7 report)

CRITICAL: Worktree-based isolation

Do NOT use git checkout -b <branch> for ablations — dirties main working tree, corrupts concurrent tool calls. Each ablation gets own git worktree under $FORTIFY_DIR/worktrees/<variant>, created from best_commit. Main working tree NEVER modified. Cleanup: git worktree remove --force per variant; git worktree prune on interrupt.

Fortify Mode (Steps F1–F8)

Triggered by fortify or fortify <run-id|program.md>.

Task tracking: create tasks for F1, F2, F3, F4, F5, F6, F7, F8 at start — before any tool calls.

Step F1: Locate source run, parse flags, and validate judge approval

Extract flags: --venue <VENUE>, --max-ablations <N>, --skip-run, --keep "<items>".

# --keep value (compaction-contract.md §keep)
KEEP_ITEMS=""
if [[ "$ARGUMENTS" =~ --keep[[:space:]]\"([^\"]+)\" ]]; then
    KEEP_ITEMS="${BASH_REMATCH[1]}"
fi
export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
# stale contract cleanup (compaction-contract.md §Lifecycle)
rm -f .temp/state/skill-contract.md  # timeout: 5000
echo "${KEEP_ITEMS:-}" > "${TMPDIR:-/tmp}/fortify-keep-items-${CSID}"  # for F2/F4 contract

# empty = no --venue → F6 skip rule fires; no default venue
VENUE=""
if [[ "$ARGUMENTS" =~ --venue[[:space:]]+([^[:space:]]+) ]]; then
    VENUE="${BASH_REMATCH[1]}"
fi
case "$VENUE" in
  ""|CVPR|NeurIPS|ICML|workshop) ;;
  *) echo "fortify: invalid --venue '$VENUE' — valid: CVPR, NeurIPS, ICML, workshop"; exit 2 ;;
esac
echo "$VENUE" > "${TMPDIR:-/tmp}/fortify-venue-${CSID}"  # for F6 (Check 41)

Unsupported flag check: load and follow the protocol below. Supported flags for this skill: --venue, --max-ablations, --skip-run, --keep.

# loads: unsupported-flag-protocol.md
export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
IFS= read -r _RESEARCH_SHARED < "${TMPDIR:-/tmp}/research-shared-${CSID}" 2>/dev/null || _RESEARCH_SHARED=""  # warm read (Check 41)
cat "$_RESEARCH_SHARED/unsupported-flag-protocol.md"

Input resolution (priority order):

  1. Explicit <run-id> argument → read $STATE_DIR_BASE/<run-id>/state.json
  2. Explicit <program.md> argument → scan $STATE_DIR_BASE/*/state.json for matching program_file, pick latest with status: completed or status: goal-achieved
  3. No argument → scan $STATE_DIR_BASE/, pick latest with status: completed or status: goal-achieved
  4. None found → stop:
    fortify: No completed run found. Run /research:run first.
    

Initialize base directory bash variables (constants YAML is not auto-exported; environment overrides honored):

STATE_DIR_BASE="${STATE_DIR_BASE:-.experiments/state}"
FORTIFY_DIR_BASE="${FORTIFY_DIR_BASE:-.experiments}"

METRIC_CMD/GUARD_CMD are NOT defaulted here — they are sourced from the source-run state.json after $RUN_ID resolves (see block below). No fail-open default: git diff --stat HEAD always exits 0, so a defaulted guard would be a no-op that never catches ablation regressions.

Assign $RUN_ID from input resolution above (must be set before guard block uses it):

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
STATE_DIR_BASE="${STATE_DIR_BASE:-.experiments/state}"  # default (Check 41)
_ARG1=$(echo "$ARGUMENTS" | awk '{print $1}')
if [ -n "$_ARG1" ] && [ "${_ARG1#-}" = "$_ARG1" ] && [ ! -f "$_ARG1" ] && [ -d "$STATE_DIR_BASE/$_ARG1" ]; then
  RUN_ID="$_ARG1"
elif [ -n "$_ARG1" ] && [ -f "$_ARG1" ]; then
  # program.md path  <!-- loads: find_run_id.py -->
  RUN_ID=$(python "${CLAUDE_PLUGIN_ROOT:-plugins/cc_research}/bin/find_run_id.py" "$STATE_DIR_BASE" --match-program "$_ARG1" 2>/dev/null)
else
  RUN_ID=$(python "${CLAUDE_PLUGIN_ROOT:-plugins/cc_research}/bin/find_run_id.py" "$STATE_DIR_BASE" 2>/dev/null)
fi
[ -z "$RUN_ID" ] && { echo "fortify: No completed run found. Run /research:run first."; exit 1; }
echo "$RUN_ID" > "${TMPDIR:-/tmp}/fortify-run-id-${CSID}"  # for later blocks (Check 41)

Source metric_cmd/guard_cmd from the campaign — env override first, then source-run state.json .config; fail closed if neither present (a defaulted git diff --stat HEAD guard always exits 0 → no-op that never catches regressions):

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
STATE_DIR_BASE="${STATE_DIR_BASE:-.experiments/state}"  # default (Check 41)
IFS= read -r RUN_ID < "${TMPDIR:-/tmp}/fortify-run-id-${CSID}" 2>/dev/null || RUN_ID=""  # reload (Check 41)
STATE_JSON="$STATE_DIR_BASE/$RUN_ID/state.json"
METRIC_CMD="${METRIC_CMD:-$(jq -r '.config.metric_cmd // empty' "$STATE_JSON" 2>/dev/null)}"
GUARD_CMD="${GUARD_CMD:-$(jq -r '.config.guard_cmd // empty' "$STATE_JSON" 2>/dev/null)}"
if [ -z "$METRIC_CMD" ] || [ -z "$GUARD_CMD" ]; then
  echo "fortify: BLOCKED — $STATE_JSON .config missing metric_cmd/guard_cmd; ablation needs the campaign's real metric and guard commands."
  echo "Re-run /research:run to regenerate state.json, or set METRIC_CMD/GUARD_CMD in the environment."
  exit 1
fi
# echo not printf — 4d/4e reload via `read -r VAR<f||VAR=""`; no trailing newline → read fails → || wipes value → false BLOCK
echo "$METRIC_CMD" > "${TMPDIR:-/tmp}/fortify-metric-cmd-${CSID}"  # for 4d (Check 41)
echo "$GUARD_CMD" > "${TMPDIR:-/tmp}/fortify-guard-cmd-${CSID}"    # for 4e (Check 41)

Guard: judge approval required. Judge skill writes verdict to .reports/research/judge-<branch>-<date>.md — scan for APPROVED verdict line:

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
STATE_DIR_BASE="${STATE_DIR_BASE:-.experiments/state}"  # default (Check 41)
IFS= read -r RUN_ID < "${TMPDIR:-/tmp}/fortify-run-id-${CSID}" 2>/dev/null || RUN_ID=""  # reload (Check 41)
JUDGE_VERDICT_FILE=$(ls -t .reports/research/judge-*.md 2>/dev/null | head -1)  # timeout: 5000
if [ -z "$JUDGE_VERDICT_FILE" ]; then
  echo "fortify: BLOCKED — no judge verdict found in .reports/research/."
  echo "Ablation studies require an approved baseline. Run: /research:judge <program.md>"
  exit 1
fi
JUDGE_VERDICT=$(grep -i '^[*]*[Vv]erdict[*]*:' "$JUDGE_VERDICT_FILE" | head -1 | sed 's/\*\*//g' | sed -E 's/.*[Vv]erdict[: ]+//' | sed 's/[[:space:]]*$//')  # trailing strip only — keep internal spaces ("NEEDS REVISION")

PROGRAM_FILE=$(grep -iE '^[*]*(Program(_file)?|Program file)[*]*:' "$JUDGE_VERDICT_FILE" | head -1 | sed 's/\*\*//g' | sed -E 's/.*:[[:space:]]*//' | sed 's/[[:space:]]*$//')
# F-02: empty PROGRAM_FILE — no program metadata, can't verify
if [ -z "$PROGRAM_FILE" ]; then
    echo "fortify: BLOCKED — judge verdict missing Program: field; cannot verify verdict applies to current experiment."
    echo "Re-run: /research:judge <program.md> to generate a fresh verdict with required metadata."
    exit 1
fi
# explicit state.json path — never CWD-relative
STATE_PROGRAM=$(jq -r '.program_file // ""' "$STATE_DIR_BASE/$RUN_ID/state.json" 2>/dev/null)
# realpath both — judge report may be relative, state.json absolute; raw compare would false-BLOCK
_PF_ABS=$(realpath "$PROGRAM_FILE" 2>/dev/null || echo "$PROGRAM_FILE")
_SP_ABS=$(realpath "$STATE_PROGRAM" 2>/dev/null || echo "$STATE_PROGRAM")
if [ -n "$STATE_PROGRAM" ] && [ -n "$PROGRAM_FILE" ] && [ "$_PF_ABS" != "$_SP_ABS" ]; then
    printf "! BLOCKED — judge verdict references program '%s' but current experiment is for '%s'\n" "$PROGRAM_FILE" "$STATE_PROGRAM"
    printf "Run: /research:judge %s\n" "$STATE_PROGRAM"
    exit 1
fi
if [ -n "$PROGRAM_FILE" ] && [ ! -f "$PROGRAM_FILE" ]; then
    printf "! BLOCKED — program file %s referenced by judge verdict not found on disk\n" "$PROGRAM_FILE"
    exit 1
fi
echo "$PROGRAM_FILE" > "${TMPDIR:-/tmp}/fortify-program-file-${CSID}"  # for F6 (Check 41)
echo "$JUDGE_VERDICT" > "${TMPDIR:-/tmp}/fortify-judge-verdict-${CSID}"  # for gate block; echo not printf — newline keeps read exit 0 (Check 41)

Verify JUDGE_VERDICT == "APPROVED". The program cross-match above guarantees the verdict was issued for the current experiment — fortify cannot ablate against a different program's verdict. Apply explicit bash gate — prose alone never halts execution:

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
# REJECTED default = fail closed; gate appends actual verdict itself, message never interpolates it
_GATE_MSG="fortify: BLOCKED — no APPROVED judge verdict found for this program; stopping before implementation (F1–F7).
Ablation studies require an approved baseline. Run: /research:judge <program.md>"
python "${CLAUDE_PLUGIN_ROOT:-plugins/cc_research}/bin/gate-on-sentinel.py" "${TMPDIR:-/tmp}/fortify-judge-verdict-${CSID}" APPROVED REJECTED "$_GATE_MSG" || exit 1

Note: do NOT infer from methodology.md alone — methodology_rating: sound is one input to verdict, not verdict itself. Only ## Verdict line in judge output file is authoritative.

Read from state.json: goal, best_metric, best_commit, config (including metric_cmd, guard_cmd, compute), program_file.

Also read baseline_commit — iteration 0 commit from experiments.jsonl (first line, status: "baseline", field "commit").

Pre-compute run directory (each in separate Bash call):

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
FORTIFY_DIR=$(python "${CLAUDE_PLUGIN_ROOT:-plugins/cc_research}/bin/make_run_dir.py" "fortify" ".experiments" 2>/dev/null)  # timeout: 5000
[ -z "$FORTIFY_DIR" ] && { echo "! make_run_dir.py failed — ensure research plugin installed"; exit 1; }
mkdir -p "$FORTIFY_DIR"
echo "$FORTIFY_DIR" > "${TMPDIR:-/tmp}/fortify-dir-${CSID}"  # for F2/F4-loop/F6; make_run_dir non-idempotent (Check 41)
WORKTREE_BASE="$FORTIFY_DIR/worktrees"
mkdir -p "$WORKTREE_BASE"
STATE_DIR="${STATE_DIR_BASE:-.experiments/state}/fortify-$(basename "$FORTIFY_DIR")"
mkdir -p "$STATE_DIR"

Step F2: Identify ablation candidates via scientist

Gather two inputs for scientist:

  1. Git diff: run git diff <baseline_commit>...<best_commit> --stat (summary) and full git diff <baseline_commit>...<best_commit>. If full diff exceeds ~200 lines, write to $FORTIFY_DIR/diff.txt via Write tool; otherwise inline in prompt.
  2. Experiment history: paths to experiments.jsonl and diary.md from source run directory.

Agent budget — each spawn costs ~120,851 tok of fixed overhead (~73 tool-calls' worth) plus ~12.0 s/call, so work under ~73 calls is cheaper done inline: spawn nothing. Keep each agent near ~55 tool-calls; past ~60 they stall without returning an envelope, forcing reconstruction from disk. Every spawn prompt must require an envelope even on exhaustion — partial: true plus what was finished.

Spawn research:scientist via Agent(subagent_type="research:scientist", prompt="...") with health monitoring (15-min cutoff, one 5-min extension — same pattern as judge J3).

Before building the prompt, substitute all bash variables into a single concrete string — never pass literal <FORTIFY_DIR> or <path> placeholders to the agent:

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
STATE_DIR_BASE="${STATE_DIR_BASE:-.experiments/state}"  # default (Check 41)
IFS= read -r RUN_ID < "${TMPDIR:-/tmp}/fortify-run-id-${CSID}" 2>/dev/null || RUN_ID=""  # reload (Check 41)
IFS= read -r FORTIFY_DIR < "${TMPDIR:-/tmp}/fortify-dir-${CSID}" 2>/dev/null || FORTIFY_DIR=""  # reload (Check 41)
EXPERIMENTS_PATH="$STATE_DIR_BASE/$RUN_ID/experiments.jsonl"
DIARY_PATH="$STATE_DIR_BASE/$RUN_ID/diary.md"
F2_PROMPT="Act as an ML ablation study designer.

Read:
- git diff at ${FORTIFY_DIR}/diff.txt (or inline if small)
- experiments.jsonl at ${EXPERIMENTS_PATH} (filter for entries with status: 'kept')
- diary.md at ${DIARY_PATH} (if exists)

Identify 3-8 distinct logical components that were changed during this run.
A component = a logically independent change that can be removed independently.

For each component produce one JSON line to ${FORTIFY_DIR}/ablation-candidates.jsonl:
{
  \"component_id\": <int>,
  \"name\": \"<descriptive name, e.g. 'learning rate warmup'>\",
  \"description\": \"<what it does and why it was introduced>\",
  \"files\": [\"<file:line range>\"],
  \"revert_commits\": [\"<commit SHA>\"],
  \"expected_importance\": \"HIGH|MEDIUM|LOW\"
}

Write your analysis to ${FORTIFY_DIR}/candidates-analysis.md.
Include ## Confidence block.
Return ONLY: {\"status\":\"done\",\"components\":N,\"file\":\"${FORTIFY_DIR}/ablation-candidates.jsonl\",\"confidence\":0.N}"

Pass $F2_PROMPT (fully expanded) as the prompt= argument to Agent(...).

Synchronous spawn note: the F2 scientist is spawned synchronously (not run_in_background=true), so CLAUDE.md §6 sentinel polling is unreachable mid-call. Timeout is handled post-hoc — after Agent() returns, check $FORTIFY_DIR/ablation-candidates.jsonl; if missing or empty, stop with "fortify: Scientist timed out. Check $FORTIFY_DIR/ for partial output." and surface with ⏱.

Read ablation-candidates.jsonl after scientist completes. If --max-ablations <M> specified and component count + 1 (for full variant) exceeds M: sort by expected_importance (HIGH first, then MEDIUM, then LOW), keep top M-1 components plus always include full sanity-check variant. Log dropped components: print a warning listing each dropped component by component_id and expected_importance so users can verify the scientist's importance estimates before proceeding. Include this list in the F7 report under ## Dropped Variants.

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
# boundary 1: F2 done, candidates ready (compaction-contract.md §Lifecycle)
IFS= read -r _RUN_ID < "${TMPDIR:-/tmp}/fortify-run-id-${CSID}" 2>/dev/null || _RUN_ID=""
IFS= read -r _FORTIFY_DIR < "${TMPDIR:-/tmp}/fortify-dir-${CSID}" 2>/dev/null || _FORTIFY_DIR=""
IFS= read -r _KEEP < "${TMPDIR:-/tmp}/fortify-keep-items-${CSID}" 2>/dev/null || _KEEP=""
_KEEP_APPEND=""; [ -n "$_KEEP" ] && _KEEP_APPEND="; user-keep: $_KEEP"
mkdir -p .temp/state  # timeout: 5000
{
    echo "## Active Skill Contract"
    echo "- skill: research:fortify · phase: ablation-execution (after F2 candidates identified)"
    echo "- run-dir: ${_FORTIFY_DIR}"
    echo "- preserve: run-id=${_RUN_ID}, fortify-dir=${_FORTIFY_DIR}, candidates=${_FORTIFY_DIR}/ablation-candidates.jsonl, variants=${_FORTIFY_DIR}/variants.jsonl${_KEEP_APPEND}"
    echo "- next: F3 generate variants → F4 run worktrees sequentially"
} > .temp/state/skill-contract.md  # timeout: 5000

--skip-run early exit: if --skip-run flag present, print candidate table (component_id, name, description, files, expected_importance) and exit. No ablation execution. Mark tasks F3, F4, F5, F6, F7 as skipped via TaskUpdate. Print all three lines (no .reports/research/fortify-*.md is written in --skip-run mode — only ablation-candidates.jsonl lives under $FORTIFY_DIR; surface $FORTIFY_DIR explicitly so the user can locate the candidate list):

fortify: --skip-run — <N> candidates identified.
Candidates artifact: $FORTIFY_DIR/ablation-candidates.jsonl
Next: /research:fortify without --skip-run to execute ablations

Jump to F8 (skip-run variant).

Step F3: Generate ablation variants

For each component from F2, one ablation variant: no-<component-name> (slugified — lowercase, spaces to hyphens). Plus one full variant (sanity check — should reproduce best_metric).

Write variant configs to $FORTIFY_DIR/variants.jsonl via Write tool — one JSON line per variant:

{"variant_name": "full", "component_removed": null, "revert_commits": [], "revert_strategy": "none"}
{"variant_name": "no-<name>", "component_removed": "<name>", "revert_commits": ["<sha1>", "<sha2>"], "revert_strategy": "git-revert"}

Step F4: Run ablation variants via worktrees

Run each variant sequentially — parallel worktrees would conflict.

Before loop — store original working directory and pre-create worktree-paths accumulator:

export CSID="${CLAUDE_CODE_SESSION_ID:-$PPID}"
ORIG_DIR="$(pwd)"  # timeout: 3000
echo "$ORIG_DIR" > "${TMPDIR:-/tmp}/fortify-orig-dir-${CSID}"  # for 4f cleanup cd (Check 41)
WORKTREE_PATHS_FILE=$(mktemp -t fortify-XXXX) || { echo "! BLOCKED — mktemp failed (tmpfs full or permission denied); cannot create cleanup accumulator. Aborting."; exit 1; }  # timeout: 3000
echo "$WORKTREE_PATHS_FILE" > "${TMPDIR:-/tmp}/fortify-paths-ptr-${CSID}"
echo 1 > "${TMPDIR:-/tmp}/fortify-variant-idx-${CSID}"  # 1-based cursor — loop's only identity across fences (Check 41)

STATE_DIR_BASE="${STATE_DIR_BASE:-.experiments/state}"  # default (Check 41)
IFS= read -r RUN_ID < "${TMPDIR:-/tmp}/fortify-run-id-${CSID}" 2>/dev/null || RUN_ID=""  # reload (Check 41)
best_commit=$(jq -r '.best_commit // empty' "$STATE_DIR_BASE/$RUN_ID/state.json" 2>/dev/null)
echo "$best_commit" > "${TMPDIR:-/tmp}/fortify-source-best-commit-${CSID}"  # raw value (Check 41)

# F-04: must be concrete SHA — branch tip would advance + pollute history on later revert
if ! best_commit_sha=$(git rev-parse --verify "$best_commit^{commit}" 2>/dev/null); then
    echo "! BLOCKED — best_commit '$best_commit' does not resolve to a commit. Re-run source experiment or correct state.json."
    exit 1
fi
best_commit="$best_commit_sha"  # worktree add always sees a SHA → detached HEAD
echo "$best_commit" > "${TMPDIR:-/tmp}/fortify-best-commit-${CSID}"  # for 4a worktree-add (Check 41)

On interrupt (user abort or unexpected error mid-loop): cd "$ORIG_DIR" first, then git worktree prune (timeout: 15000) to clean up partially created worktrees before exiting. Interrupt cleanup is manual by necessity — a shell trap cannot survive the Bash call that registers it (see the no-trap note below 4a), so the accumulator file is the durable record of what needs removing; re-run the post-loop sweep block below to drain it.

Loop control: run 4a-init through 4g-advance once per line of variants.jsonl, in file order. The iteration cursor lives in the fortify-variant-idx-${CSID} sentinel (initialized to 1 above, advanced at 4g-advance — or by the resume guard on skip) — never in a shell variable, which would die between blocks. Stop the loop when 4a-init reports the cursor is past the last line.

Each Bash call costs a ~12 s round-trip, so adjacent steps that share a first token and have no model decision between them are merged: 4a-init now carries the cursor read, the resume guard, and the cleanup pre-registration in one block, taking the loop from 13 calls per variant to 11. The remaining splits are load-bearing, not oversight — each is annotated at its own step (4b/4b-assert and the 4f pair keep cd out of compound commands; 4c's two calls keep git in first-token position; 4d and 4e each hold a timeout: 360000 and cannot share one call under the 600 s ceiling; 4g-advance must follow 4g's on-disk record). Do not "optimize" them back together.

4a-init. Read the current variant's spec by cursor + resume guard + pre-register cleanup path (one block — must run before any 4a/4b/4c/4d/4e block, which all read the name it persists):

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
27
Forks
4
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
fortify
Source
github.com/borda/ai-rig