challenge-resolve

SkillDev tools

Independently review and fix a scoped diff through bounded convergence rounds, with evidence-backed closure and explicit stop/recovery decisions. Use for requested adversarial review-and-fix loops, not a single read-only review.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the challenge-resolve skill

What this skill tells your AI

The instructions your AI receives, as published by borda/ai-rig in plugins/codex-rig/skills/challenge-resolve/SKILL.md and read by ahel’s review.

Before asking, read User Questions.

Challenge and Resolve

Read, apply ../../shared/adversarial-loop.md before dispatch or edits. That shared procedure owns the algorithm; this entrypoint owns Codex artifacts, the closing gate. Read evidence-contract.md before source capture or reviewer dispatch. Also read ../../shared/native-skill-contract.md and ../../shared/specialist-orchestration.md for authority, recurrence, reviewer admission, evidence limits.

Input Schema

{
  "goal": "required review-and-fix objective",
  "scope_files": ["bounded source paths relevant to the task; discover them when initially unknown"],
  "specification": "required behavior or acceptance criteria",
  "symptom": "reported problem, feature objective, or preventive review",
  "caller_run": "optional existing workflow run to resume",
  "done_when": "independent final review is clean and normal checks pass"
}

Use default max 3 review rounds, including initial W_0. For a problem or feature without named files, inspect the repository to identify and bound the relevant source before freezing a review; a pre-existing implementation can have an empty diff. In-scope feasible fixes are authorized by a review-and-fix request; a structural finding does not authorize scope expansion, public API changes, commits, installs, network access or publication. A caller's stricter scope or admission still applies. The retained request proves delivery of the parent's declared criteria to the challenger; report any uncertainty about interpreting the user's scope instead of claiming machine authentication of natural language.

Review each frozen scope with a fresh independent challenger process. Give it source, diff, criteria, and original finding counterexamples needed for verification; never give it resolver handoffs, explanations of changes, or parent closure claims. The main process owns old/new identity, severity, the weighted ledger, resolver assignments, and stop decisions. New challenge results use the same strict artifact checks under candidate and canonical filenames; archived JSON remains readable as data but cannot certify current completion.

Workflow

01: Establish scope and evidence ownership

Read ../../shared/helper-cli-contract.md; create a run with create_run.py --skill challenge-resolve. Record caller_run when supplied, without overwriting its artifacts. Retain baseline source and acceptance evidence. Write loop-report.md with Scope, Rounds, Findings, Remediation, and Verification sections. Identify implementation author and allowed reviewer route before dispatch.

Before any remediation, list every requested path, relevant unchanged consumer, and known interaction in loop-report.md under Scope; assign each to the first challenger scope or mark it unreviewed with a reason. Finish the first challenge over every admitted part before treating the gathered findings as the discovery baseline. A complete-file chunk plan proves byte delivery only; the parent must inspect each reviewer's stated coverage and keep an uninspected path or interaction visible. If full initial coverage is unavailable, stop an overall clean claim rather than spending repeated fix rounds on a partial baseline.

For each workflow branch in scope, record one source route before dispatch: entrypoint, producer, next ordinary consumer, distinct sibling path where one exists, and the acceptance test that exercises them. Include unchanged files on that route in supporting source; ask the challenger to state which routes it assessed. Missing route source or a reviewer-declared unassessed route blocks a complete discovery baseline. If a later challenger finds a defect already present in an earlier frozen source, mark the earlier route coverage gap, add the missed producer or consumer, and reassess that route before another fix round. A test that manually upgrades a producer's output cannot prove the ordinary producer-to-consumer path.

For a large scope, first assess the default Code Review native instruction-bounded route; the optional isolated Codex review route's 2 MiB transport limit is not a native-route limit. A size estimate alone cannot establish independence-unavailable: freeze the full source and diff, inspect actual native admission, then use the chunk path below if the full input cannot be admitted or meaningfully reviewed. Preserve the route's actual rejection evidence; never send truncated source or infer model capacity from a byte count.

Bounded chunk path

Use chunk_diff.py plan --repository <root> --out <coordinator-run>/chunks --goal <goal> --specification <criteria> --done-when <condition> [--scope-path <requested-path> ...] [--scope-mode changed|all] --budget-bytes <route-budget> after reading its --help. For remediation continuing an earlier authenticated run, also pass --caller-run <earlier-run> --prior-ledger <earlier-run>/loop-ledger.json --prior-coverage <coverage.json>; the coverage list assigns every final-round signature to source_paths, one review_kind (chunk, interaction, or group), and sorted chunk_ids. Retain the original validated reviewer receipt and select source paths that include the implementation, validator, next consumer, and relevant tests for each defect. The planner copies and hashes the earlier ledger and appends its canonical signature, tier, structural flag, evidence, and declared source-path block to the specification; use the resulting manifest request for the coordinator and every child. The checker rereads the declared caller run ledger and compares its exact bytes with the copy. The schema-1 manifest declares first-run or continuation origin; a declared continuation cannot omit prior ledger and coverage, and candidate and final coordinator results must copy the exact origin into metadata.challenge_origin. The copied ledger's structural validity does not reauthenticate historical reviewer provenance. Without an authenticated earlier ledger or defensible source assignment, stop the prior-finding closure claim. For a first review, pass the same exact nonblank request fields to coordinator and children. Select changed for a diff or all for a whole requested file or plugin scope; omitting --scope-path selects the repository. Choose a conservative budget for source JSON plus patch while reserving room for role card, instructions, and output; a plan measures input, not reviewer comprehension or host capacity. The helper inventories selected tracked, deleted, and non-ignored untracked files, writes exact source and supporting diff bytes per whole-file chunk, and records disjoint coverage in chunks.json. Never choose files or examples from a previous task. An oversized single file remains an explicit chunk; attempt a permitted route for it and continue independent chunks even if it needs separate remediation.

Run chunk_diff.py check --manifest <coordinator-run>/chunks/chunks.json before dispatch. Treat chunking as an internal division of the already authorized scope: no new scope approval is needed when the union equals the original request and no excluded file is added. Process chunks sequentially by default through separate ordinary challenge-resolve runs created under the coordinator's chunks/runs/ root. Gather the initial independent challenge for every admitted chunk and known interaction before changing source when the route and authority permit; then triage the union and assign fixes. Each run receives its chunk's complete source JSON and diff, relevant unchanged callers/consumers, the same specification in a schema-2 loop-evidence.request, and an independent Code Review challenger. The existing three-round cap, progress timing, action ledger, reviewer provenance, gates, and result validation apply to each run; chunk count does not increase any one run's round cap. Parent owns edits and pauses dispatch while source changes; replan, recheck, and independently review affected final chunks after every fix. Do not reuse stale child results.

After all child runs, assess every unordered pair of chunks for changed interfaces, producers/consumers, shared state, and other material interactions. Record reviewed for a material interaction or no-interaction with concrete source-based rationale. Also declare each known material interaction spanning three or more chunks as a group. Each pair decision and declared group needs an ordinary independent challenge-resolve run whose frozen scope contains every file in every member chunk, including files the parent considers unrelated. One run may assess several decisions and groups when it admits their complete union; separate runs may cover smaller unions. A child result, summary-only context, parent-only reconciliation, or one file per chunk cannot substitute. Send each assigned decision, its chunk IDs and exact rationale in the run's request.specification as Interaction assessment: followed by a newline and json.dumps(decisions, sort_keys=True, ensure_ascii=True), where decisions are ordered pair entries followed by group entries, each containing chunk_ids, decision, and evidence. Ask the challenger to examine whether the decisions are correct and whether additional multi-chunk groups are material; any identified omission is a finding and requires a new declared group and fresh review. The parent owns semantic judgment; no checker can prove an undeclared interaction absent. If complete input cannot be admitted or plausible higher-order interactions remain unassessed, keep overall clean blocked. Resolve findings through the parent, then replan and rerun affected current-source reviews.

Write chunks/results.json as {"schema_version":1,"chunks":[{"id":"chunk-001","result_path":"runs/challenge-resolve/<child-run>/result.json"}],"interactions":[{"chunk_ids":["chunk-001","chunk-002"],"decision":"reviewed","evidence":"specific interface and why these chunks interact","result_path":"runs/challenge-resolve/<interaction-run>/result.json"}],"groups":[{"chunk_ids":["chunk-001","chunk-002","chunk-003"],"evidence":"specific three-chunk contract","result_path":"runs/challenge-resolve/<group-run>/result.json"}]} with one ordered child entry per planned chunk, one ordered decision per chunk pair, and sorted unique group entries of three or more chunk IDs. Use groups: [] when no group is declared. A no-interaction pair retains its result_path and an evidence rationale so an independent challenger assesses it. The checker binds each child request to the exact coordinator goal, specification, and done_when; an interaction request keeps the same goal and done_when and appends exactly the canonical assessment block to the coordinator specification. It also binds each interaction run to the exact union of complete member chunks, then revalidates the ordinary clean result. Run chunk_diff.py check --manifest <coordinator-run>/chunks/chunks.json --results <coordinator-run>/chunks/results.json; it checks current source, disjoint inventory coverage, complete pair decisions, declared groups, request delivery, source/diff binding, clean loop status, and artifact validation. Set new coordinator candidate metadata to "challenge_scope":{"mode":"chunked","manifest_path":"chunks/chunks.json","results_path":"chunks/results.json"}; the final result validator reruns this check. A missing decision, stale or unadmitted assessment, or non-clean run keeps the overall request open. Reconcile interfaces and producer/consumer behavior across chunks with source context and relevant project gates; an omitted group remains a semantic limit even when the checker passes. Parallel independent read passes are allowed only when each run binds immutable source and the parent still serializes fixes and acceptance; sequential dispatch is the default.

For a stopped coordinator with at least one validated child or interaction result and pending or non-clean work, use the separate schema-1 partial chunks/results.json map with the same chunks, interactions, and groups lists. Include every completed run's result path; omit pending chunk entries, and use null result paths for declared pending interactions or groups. Run chunk_diff.py check-stopped --manifest <coordinator-run>/chunks/chunks.json --results <coordinator-run>/chunks/results.json --summary; copy its exact status: stopped summary, validated run score series and decisions, pending_chunks, pending_interactions, pending_groups, and pending_prior_signatures lists into metadata.chunked_review. An assigned earlier signature absent from an authenticated supplied reviewer result stays pending even when all supplied child runs are otherwise clean. The result validator reruns this stopped check and accepts only a failed coordinator and failed review gate. This path records a truthful partial stop; it never satisfies the schema-1 manifest clean coverage check or supplies a combined score. With no validated child or interaction result, retain the ordinary no-independent-review stop and do not invent a child-backed summary. Chunk manifest version 1 is required for source, complete, and stopped checks; changing the version cannot bypass continuation obligations.

In chunk mode, each child and interaction run produces its ordinary validated result.json and final.md; the coordinator reports those results and the coverage check without fabricating a combined W score or a full-scope reviewer receipt. Do not call the original request clean until every child, pair assessment, and declared group run is independently clean on current source, the parent has examined all interaction rationales and possible undeclared groups, the results check passes, cross-chunk behavior has evidence, and normal project gates pass. A passed checker establishes coverage of declared decisions and groups, not mathematical proof that no other group exists. The coordinator's summary names any remaining decision, group, or semantic coverage limit and its exact stop reason. No run result substitutes for a caller's separate completion gate.

When chunking resumes an earlier validated review loop, retain that loop's score series and finding signatures as historical evidence. Before dispatch, use the schema-1 plan's prior coverage assignment; inspect each assigned run's frozen primary and supporting source and run its no-model input preflight to confirm the declared paths and prior-evidence block are present. check --results rejects a clean coordinator unless the assigned authenticated current reviewer returns each exact prior signature with its original tier and structural flag as verified-fixed or rejected and its declared source paths were delivered. The declaration is supplied by the caller, so an undeclared or false caller history remains outside local artifact proof. A path list cannot prove all semantic dependencies were chosen; missing or uncertain source stays unassessed and blocks closure. Each chunk has its own local round index, so introduce every child progress table with its chunk ID and the earlier run's score series; never show an unqualified Iteration 1 that appears to reset the original loop. Keep a coordinator cross-run-history.md with one chronological row per validated original, child, and interaction review: identify the run and local round, source scope, retained ledger, six severity counts and weighted score as old + new, where old means the same underlying defect appeared in any earlier validated row across these runs. Preserve verified-fixed, rejected, and unassessed earlier signatures in that history even though they add zero to an inspected chunk's score. Do not convert an unassessed signature into an open finding merely because its validating module was omitted. This comparison describes finding recurrence across different scopes; it is not a combined score, a convergence ratio, or an overall verdict. Report both the local checker table and the cross-run comparison, with their different old definitions stated beside them. A rejected receipt stays advisory and never becomes a history row. A stopped parent loop retains its stop and prior-finding obligations until a complete current-source review of those obligations validates their disposition; new chunk baselines do not reset its three-round or recurrence history.

02: Run the shared bounded procedure

Use one serial challenger loop per scoped run, including each chunk or reviewed interaction:

  1. Freeze complete current source and diff. Start one fresh independent challenger subprocess with that scope, the specification, and every prior finding signature plus its original counterexample and evidence when this is a repeat round. Include the implementation, next consumer, validating helper, and relevant tests needed to judge each earlier finding; if they cannot fit or are unavailable, record that finding unassessed and stop its closure claim before dispatch. Ask the challenger for structured current findings and evidence on each original counterexample, without any resolver handoff or account of what changed. Keep its original structured response and reviewer provenance.
  2. Before any fix, score, or next challenge, triage each returned finding on separate axes: compare the failed invariant and proof with every signature in all earlier validated rounds and parent/child runs to mark new, carried, or reopened; classify its demonstrated impact as security, critical, high, medium, low, or nit; and for first-reported findings, classify the failing source as present earlier, introduced by remediation, scope-added, or undated with cited snapshot/coverage evidence. A changed title or file location is not a new defect, and severity does not determine identity. Retain the old signature for the same failed invariant; if its severity conflicts with the prior tier, stop ledger promotion until the conflict has evidence and is resolved under the stable-tier contract. If the reviewer lacked source needed to assess a prior signature, record unassessed in cross-run history and seek complete coverage rather than treating absence as closure or the old signature as a fresh finding. Record one schema-2 loop-evidence.rounds[].triage mapping per raw reviewer finding, with a reason and evidence for any corrected identity or severity. Keep the raw response intact; copy the mapped findings into loop-ledger.json, validate source and reviewer evidence, run adversarial_loop.py --ledger <run-directory>/loop-ledger.json --progress, and show its weighted old/new table. One reviewer avoids same-round classification conflicts; an invalid response or missing provenance stops the affected run with the returned findings still visible.
  3. Triage every finding against source. Record a refutation only with evidence; put every open signature and parent action in loop-actions.json, then validate it. Before fix dispatch, write a short per-signature closure plan under loop-report.md Remediation: the failed invariant, original counterexample, next ordinary consumer, all known sibling routes or write sites, and the source paths whose actual changes must be inspected. Delegate feasible fixes by disjoint finding domain, with one owner per file set. Give each solver this plan and a focused acceptance check that fails on the reviewed source. An action labeled fix is an assignment, never proof of completion. Parent owns integration, gates, and any shared-file conflict.
  4. Wait for every fix agent, then inspect the actual source diff and verify each attempted repair against its original counterexample and at least one distinct sibling route when one exists, through the next consumer. Record the failing-before and passing-after checks, or the exact missing capability and remaining finding, in loop-report.md under the signature. A passing unit test of only the producer, a child-reported changed-file list, or a solver summary cannot close a consumer-path finding. Freeze corrected source and the supporting modules used by those checks before another review; a solver's handoff and local tests leave fixes fixed-pending-verification until a fresh independent challenger checks the current source. Keep the solver report in the parent record, never in the next challenger input. If the invariant still fails, keep the finding open, investigate the incomplete fix, and do not dispatch a repeat challenge merely to rediscover it.
  5. Repeat from step 1 while the score converges and the three-round cap, authority, and recurrence rules permit. For each same-scope repeat review, report the prior round's open weighted score and the current open weight of all previously reported signatures, including reopened ones. If that old-finding residue exceeds half the prior open score, name the resolver shortfall and stop under the existing plateau/non-converging gate; new discoveries cannot conceal the shortfall. The existing total-score rule may also stop when old debt halves but new findings keep the total above half. A different-scope child has no comparable ratio: retain the parent stop, show its prior finding status, and wait for complete current-source coverage before claiming reduction. Never infer closure from an omitted signature alone.
  6. At clean, plateau, cap, stale-source, conflicting evidence, or unavailable-review stop, report only validated round scores and all remaining findings. For unexpected blockers, stop the affected route, give one recommended remediation action plus meaningful alternatives, and ask the user for direction before another affected dispatch or fix. Keep unrelated authorized work separate.

Read adversarial_loop.py --help before using it. Retain loop-ledger.json, loop-actions.json, round-<index>.diff, current snapshot current.diff, and each independent report. Show the progress table exactly once after each newly validated round, with Iteration | Critical | High | Medium | Low | Nits | Weighted score and literal old + new cells; never show a placeholder or repeat it in unrelated status updates. Keep the final Results table separate, with ranked severity counts per challenge. A completed reviewer response without a validated ledger is review completed, canonical round blocked, not an unqualified not-run; preserve its original findings and score each response descriptively while naming the missing gate. No paid retry follows from an older wave approval.

Collect scoped source snapshots with the existing collect_diff.py snapshot mode, retain loop-evidence.json per evidence-contract.md. Use existing Code Review routing, frozen contexts, specialist manifests for reviewer provenance; never manufacture a second runtime evidence format. Every participating reviewer receives the exact snapshot contents, diff, response contract, returns one structured response containing every finding, with only the route-required provenance header outside it. Preserve original reports and runtime evidence.

Before dispatch through the isolated Codex reviewer route, run validate_evidence.py --preflight-plan <frozen-plan.json> --request-evidence <run>/loop-evidence.json --supporting-source <frozen-supporting.json> when supporting paths exist; omit the last option otherwise. This no-model gate checks canonical current source and supporting bytes, evidence-declared source copies, the current scoped tracked patch, plan digests, and exact line-start source, diff, request, and supporting labels in the frozen challenger context. Then run the route's separate no-model host check. Stop on either failure; preserve the frozen plan and ask for a remediation or retry decision before a new paid attempt.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
28
Forks
4
Last commit
Sep 2026
Advanced
Item type
skill
Key
challenge-resolve
Source
github.com/borda/ai-rig