Verify Diff
SkillFiles & storageEmpirically verify that a code diff achieves its stated objective by dispatching a bias-free executor that actually runs auto-derived evaluation scenarios against the post-diff file. The executor returns a JSON verdict with `suggested_edits`, and this skill applies them, a single pass by default, or iteratively up to `Max iterations` when the caller raises it. Non-interactive, no user prompts. Use after applying an Edit when you need a dynamic cross-check that complements static reviewers like skill-review.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Verify Diff skill
What this skill tells your AI
The instructions your AI receives, as published by hiroro-work/claude-plugins in .claude/skills/verify-diff/SKILL.md and read by ahel’s review.
The convergence signal is the executor itself returning suggested_edits: [] on a pass verdict. By default the skill runs a single pass; a caller that raises Max iterations loops until that signal, max iterations is reached, or a safety rail trips.
It never prompts the user.
Invocation contract
The caller passes these fields in natural language (the skill extracts them from the invocation text):
Description(required for explicit-args mode; absent for auto-derive mode) — the original problem the diff is supposed to addressSuggested fix direction(required for explicit-args mode; absent for auto-derive mode) — how the diff was meant to be shapedTarget file(required for explicit-args mode; absent for auto-derive mode) — one relative path (single-file scope; multi-file diffs are out of scope in this mode)Base ref(optional, defaultHEAD) — git ref to diff against (both modes)Max iterations(optional, default1) — upper bound on the refinement loop (both modes; auto-derive applies the same upper bound to each per-skill loop). Default1is a single detect-and-apply pass — a caller that wants the applied fixes re-verified raises it explicitly
Mode determination
A field counts as provided iff the caller supplied a non-empty, non-whitespace value. Empty string and whitespace-only count as absent.
-
All three of
Description/Suggested fix direction/Target fileprovided → explicit-args mode (run## Workflow§ Step 1 – Step 5 unchanged). -
All three absent → auto-derive mode (run
## Auto-derive modeinstead). -
1 or 2 of the three provided (incomplete framing) → return early with the explicit-args mode Step 1 schema (does not enter auto-derive mode):
{"status": "skipped", "reason": "incomplete args", "iterations_used": 0, "applied_edits_count": 0, "unresolved_gaps": [], "reverted_paths": [], "objective_met": "unknown"}
The caller must not stage changes while this skill is running. The skill reads the working tree vs Base ref; staged content would mix into the diff and corrupt the verdict.
Auto-derive mode
Triggered when Description, Suggested fix direction, and Target file are all absent (see ## Invocation contract § Mode determination). The mode infers intent from the diff itself and verifies on a per-skill basis. The per-iteration loop semantics defined in ## Workflow § Step 3 are reused verbatim per skill, with the auto-derive-specific overrides spelled out in A2 § Per-iter loop semantics.
A1. Collect diff and group by skill
-
Run
git diff <Base ref>(Base refdefaults toHEAD). -
If the diff is empty, return early with the auto-derive aggregate shape (matches A3 schema):
{"mode": "auto-derive", "status": "skipped", "reason": "empty diff", "iterations_used_total": 0, "applied_edits_count_total": 0, "non_skill_files": [], "per_skill": {}}Auto-derive treats an empty diff as
skippedrather thanconflict(the explicit-args mode policy). -
Compute
affected_filesfrom the diff (git diff --name-only <Base ref>). -
Group each file by skill prefix:
skills/<name>/...→ skill<name>.claude/skills/<name>/...→ skill<name>- paths outside both prefixes →
non_skill_filesbucket (not dispatched; reported in the aggregate verdict) - Dedup: if the same skill
<name>appears under both prefixes (a file underskills/<name>/...and another under.claude/skills/<name>/...), merge both into the same skill key.fileskeeps every distinct path string that appeared in the diff (both prefixes if they were both listed).primary_fileselection (A2 § Primary file selection) prefers theskills/<name>/...form when both are present.
-
If
skill_groupsis empty (only non-skill files were changed), return early withstatus=skipped,reason="no skill files in diff". The aggregate verdict still records the skipped paths so the caller can see what was bypassed:{"mode": "auto-derive", "status": "skipped", "reason": "no skill files in diff", "iterations_used_total": 0, "applied_edits_count_total": 0, "non_skill_files": ["<path>", "..."], "per_skill": {}}
A2. Per-skill dispatch
Auto-derive pre-registers N skills × Max iterations task rows total, namespaced by <skill>. For each (skill, files) pair in skill_groups:
-
Pre-register iteration tasks —
TaskCreateone task per iteration named<skill> iter 1…<skill> iter <Max iterations>following the same protocol as## Workflow§ Step 3 pre-registration (markin_progressviaTaskUpdatebefore dispatch,completedafter parse + apply, append— skipped: <reason>on early exit). -
Primary file selection — pick
primary_file=<skill-dir>/SKILL.mdif it appears infiles; otherwise the first entry infilessorted lexicographically. This choice only labels the per-skill verdict'sprimary_filefield for human-readable anchoring and influences the dispatch payload's framing — it does not affect per-skill verdict semantics (status/applied_edits_count/objective_metetc.), which are emitted for the wholefilesset. -
Per-iter snapshot — on iter 1,
Readthe full current contents of every entry infiles(per-iter snapshot) and rungit diff <Base ref> -- <files>to capture the per-skill diff. The A1 (1) full-tree diff is used only for skill-prefix splitting in A1 and is not forwarded to the dispatch payload; the--- DIFF ---payload carries the per-skill diff only. Oni ≥ 2, only re-Readthe subset offileswhose path appeared in a successfully-appliedsuggested_editsentry during iteri-1(untouched files keep their iter-1 snapshot) and re-rungit diff <Base ref> -- <files>so the per-skill diff reflects edits that landed in prior iterations. -
Dispatch payload assembly — invoke the
Agenttool to dispatch a fresh executor. Assemble the dispatch prompt from the four sections below, each framed with a clear--- LABEL ---fence:--- DIFF ---: the per-skill multi-file diff captured inA2 § Per-iter snapshot--- AFFECTED FILES ---: each entry infilesas a### <path>sub-heading followed by that file's full current contents (publicity-review pattern)--- INFERENCE PROMPT ---: the body ofreferences/auto-derive-prompt.md§ Executor prompt injected verbatim, with<skill-name>substituted by the current skill's name. The reference holds the two-phase Phase 1 INFER INTENT (1–2 sentences) + Phase 2 VERIFY (scenarios + checklist) prompt--- RESPONSE FORMAT ---:references/auto-derive-prompt.md§ Response format injected verbatim — output schema (includinginferred_intentand per-entryfileforsuggested_edits) plus the "1–3 lines of surrounding context" convention
Agent unavailable fallback: detect availability and fall back per the canonical write-up in
rules-reviewSKILL.md§ 5. Review(the "Detecting Agent availability" / "Fallback when Agent is unavailable" paragraphs). Auto-derive specialization: when falling back, walk the executor prompt inline-sequentially against each per-skill group in the main thread and emit the same A3 aggregate JSON defined below so callers' parse path is identical. -
Schema differences from explicit-args mode (the response format adds two extensions over
## Workflow§ Step 3 (a)):inferred_intent: a new top-level required field (string, 1–2 sentences). Its value is captured from iter 1 only — seeA2 § Per-iter loop semantics(6.1) below.suggested_editsper-entry shape:{file, old_string, new_string, rationale}—fileis required and must equal one of the paths infiles. This is the publicity-review / skill-review multi-file pattern (different from explicit-args mode's{old_string, new_string, rationale}, which targets the singleTarget file).
-
Per-iter loop semantics — reuse
## Workflow§ Step 3 (b) Parse & apply and § Step 3 (c) safety rails verbatim, with these auto-derive-specific overrides:-
(b) sub-case 2 schema violation is extended: in addition to the existing checks, require non-empty string
inferred_intentat the top level and non-empty stringfileon everysuggested_editsentry. -
(b) sub-case 5 apply edits the file named by
suggested_edits[i].file(not the explicit-argsTarget file). Before eachEdit, verifysuggested_edits[i].file ∈ F(the per-skillfilesset); if not, record the path in anout_of_scopelist and skip the entry without callingEdit(no working-tree write occurs, so no revert is needed for that entry — see (c) Scope rail).old_stringmismatches on in-scope entries remain the no-op fallback.applied_edits_countincrements only forEditcalls that succeeded; record thefilepath of each successfully-applied edit into the per-iter setDconsumed by (c) Frontmatter rail. -
(c) Scope rail (per-skill, judgment annotation): after the apply phase, if
out_of_scopeis non-empty, the executor attempted to write paths the dispatch payload did not authorize — emit the per-skill verdict withstatus: "conflict",reason: "scope violation",reverted_paths: out_of_scope. Nogit checkout HEAD --runs — no offending write landed.reverted_pathsreports the rejected paths. Other per-skill loops continue. -
(c) Frontmatter rail: applies per edited file in this iter's apply phase (the set
Ddefined in (b) sub-case 5 — paths whoseEditcall returned success this iter). For each such file with a----delimited YAML frontmatter block, re-Readand parse it; on parse failure, rungit checkout HEAD -- <file>and emit the per-skill verdict withstatus: "conflict",reason: "frontmatter broken",reverted_paths=[<file>]. Files without frontmatter skip this rail (same as explicit-args mode). -
(6.1)
inferred_intentpersistence: captureinferred_intentfrom the iter 1 verdict into main-thread context and treat that value as fixed for the rest of the per-skill loop. iter 2+ verdicts may return a differentinferred_intentbecause the executor reruns Phase 1 each dispatch — do not overwrite the iter-1 value.inferred_intentwrite rules at per-skill loop exit (closed list, covers every termination path so the field is unambiguous regardless ofstatus):- iter 1 produced no parseable verdict (sub-case 1 / 2 / dispatch error) and the per-skill loop terminates immediately with
status=skipped→ writeinferred_intent: null(there is no iter-1 value to fix). - All other termination paths —
convergedat any iter,conflictfrom (c) Scope rail or Frontmatter rail at any iter,skipped (divergent gaps)at iter ≥ 2, max-iterunresolved— write the iter-1 captured value intoinferred_intentregardless ofstatus. The captured value is fixed at iter 1 and survives subsequent iter outcomes.
Divergence comparison (sub-case 4) uses the existing
(remaining_gaps, regressions)multiset pair only;inferred_intentis excluded. - iter 1 produced no parseable verdict (sub-case 1 / 2 / dispatch error) and the per-skill loop terminates immediately with
-
-
Per-skill verdict — assembled in main-thread context at per-skill loop exit (no JSON is emitted yet; A3 aggregates and emits a single block at the end):
{ "primary_file": "<path>", "files": ["<path>", ...], "inferred_intent": "<1-2 sentences>", // iter-1 value, fixed per (6.1) "status": "converged|unresolved|skipped|conflict", "iterations_used": N, "objective_met": "yes|partial|no|unknown", "applied_edits_count": N, "unresolved_gaps": ["..."], "reverted_paths": ["..."], "reason": "..." or null }- Field semantics for
unresolved_gaps/reverted_paths/reasonmirror## Workflow§ Step 5 — Emit structured summary (the explicit-args contract). Per-skillstatusis the 4-value enum (converged|unresolved|skipped|conflict); the newpartialvalue lives only at the top level (see A3). - Per-skill
skippedpaths: (b) sub-case 1 verdict missing/malformed, (b) sub-case 2 schema violation, (b) sub-case 4 divergent gaps, and dispatch error (theAgentcall itself errored / timed out / returned empty). These mirror the explicit-args mode'sskippedpaths. Because each skill is an independent executor, askippedper-skill verdict alongside otherconvergedper-skill verdicts is the typical case that triggers the top-levelpartialrule in A3.
- Field semantics for
A3. Aggregate per-skill verdicts
After every per-skill loop has finished, emit a single fenced JSON block at the end of the invocation matching this schema (this replaces ## Workflow § Step 5 — Emit structured summary for the auto-derive mode code path):
{
"mode": "auto-derive",
"status": "converged|partial|unresolved|skipped|conflict",
"iterations_used_total": 0,
"applied_edits_count_total": 0,
"non_skill_files": ["<path>", "..."],
"per_skill": {
"<skill-name>": { "...per-skill verdict object from `A2 § Per-skill verdict`..." }
},
"reason": null
}
iterations_used_total and applied_edits_count_total are sums across every per-skill verdict.
Top-level status rule (precedence, first-match-wins):
- any per-skill
status == conflict→conflict - else any per-skill
status == unresolved→unresolved - else every per-skill
status == converged→converged - else every per-skill
status == skipped→skipped - else (a mix of
convergedandskippedonly — neitherconflictnorunresolved) →partial
partial is auto-derive-only and is never emitted by the explicit-args mode contract.
Top-level reason rule:
status: "skipped"from A1 early-return: enum string ("empty diff"or"no skill files in diff").status: "skipped"from A3 aggregation (every per-skill isskipped): JSONnullat the top level. Each skill's individualreasonsurvives in its per-skill verdict object.- All other top-level statuses (
converged/partial/unresolved/conflict): JSONnull.
Workflow
Step 1 — Extract context
- Parse the five fields from the invocation text. If
Target fileis missing or empty, return early:{"status": "skipped", "reason": "missing target", "iterations_used": 0, "applied_edits_count": 0, "unresolved_gaps": [], "reverted_paths": [], "objective_met": "unknown"} - Run
git diff <Base ref> -- <Target file>. This captures working-tree-vs-base; no staging is assumed. - If the diff is empty, return early with
status=conflictandreason="empty diff".
Step 2 — Derive evaluation scenarios (inline, no dispatch)
Generate 1–2 evaluation scenarios plus a requirements checklist from the caller's Finding framing before the executor loop begins. This runs in the skill's main thread — Read only, no Agent dispatch.
Readthe first 100 lines ofTarget fileto pick up the frontmatter,description, and opening intent prose.- Produce at least one "median" scenario — the typical use case the
DescriptionandSuggested fix directionimply. If those two fields are too thin to anchor a median, fall back to the target file's own frontmatterdescription: "does the diff achieve what this file declares as its purpose?" This fallback always yields a scenario. - Add one "edge" scenario only if the Finding text explicitly names a boundary condition, a failure mode, or an input class the median does not cover. Otherwise stop at one scenario — do not pad.
- For each scenario, write a requirements checklist of 3–7 items. Frame each item as an intent / behavior / property the diff must achieve, not as a reference to specific text fragments or line numbers. When the Finding inherently involves a specific fragment (e.g. a rename), frame the item around the resulting behavior (e.g. "no occurrences of the old name remain in the specified scope") rather than requiring an exact substring match. At least one item must be tagged
[critical]. - Hold the scenarios and checklists in main-thread context for the duration of the run. Do not regenerate mid-run.
Step 3 — Iteration loop (i = 1 .. Max iterations)
Pre-register iteration tasks — before entering the loop, TaskCreate one task per iteration named iteration 1, iteration 2, ..., iteration <Max iterations>. Mark in_progress (via TaskUpdate) before each dispatch, completed after parse+apply (for a converged verdict, "apply" is a no-op — mark completed immediately after parsing the verdict). On early convergence (verdict matches the converged rule) or safety-rail triggered exit (skipped / conflict), mark remaining iteration tasks completed with note matching the exit reason (e.g. skipped: converged at iter 2). The "note" lives in the task's description field (the content field under the TodoWrite fallback) — append as — <reason>. Where the Task tools are unavailable (e.g. the VSCode extension, or Claude Code before v2.1.142), use the equivalent TodoWrite operations instead — the status values and pre-register semantics are identical; allowed-tools grants both. Pre-registration is load-bearing when a caller raises Max iterations: without it, the executor-driven loop tends to stop after the first iteration that looks acceptable, even when gaps remain that further iterations could close.
(a) Dispatch bias-free executor
At the start of every iteration, Read the full current contents of Target file so the snapshot reflects prior edits. On i ≥ 2, also re-run git diff <Base ref> -- <Target file> so the diff reflects edits that landed in prior iterations.
Invoke the Agent tool to dispatch a fresh executor. Assemble the dispatch prompt from the four sections below, each framed with a clear section label (e.g. --- TARGET FILE ---, --- UNIFIED DIFF ---, --- SCENARIOS ---, --- EXECUTOR PROMPT ---) so the executor can parse each payload unambiguously. Use the same order as the labels above:
- TARGET FILE:
Target file's full current contents (verbatim) - UNIFIED DIFF: the unified diff
- SCENARIOS: the scenarios and requirements checklists from Step 2 — Derive evaluation scenarios (verbatim, identical across all iterations)
- EXECUTOR PROMPT: the executor prompt and JSON schema below (verbatim)
Executor prompt (include verbatim in the executor prompt):
You are a fresh executor of the target file. You have not seen the original Finding framing — only the scenarios and checklist below. Actually execute each scenario against the target as written; do not merely read and judge. Produce scenario artifacts in your response body, then report unclear points, discretionary fills, and retries in natural language.
Judge whether the diff resolves the problem the scenarios encode AND follows the scenarios' intent. Check for regressions — changes that break behavior the original file relied on. Return
objective_met: "yes"only if there are no remaining gaps AND no regressions. Otherwise return"partial"(direction is right but gaps remain) or"no"(diff does not address the objective).Gate reachability rule (required): when your final verdict is
objective_met: "yes"ANDregressions: [], you must returnsuggested_edits: []. Do not emit speculative or nice-to-have edits on a pass — record any such observations inremaining_gapsinstead.suggested_editsis the convergence signal and must be empty when you are declaring the work done.If a scenario requires running a script or command,
cdinto a scratch/temp directory outside the skill tree under test before invoking it (or otherwise redirect its output there) — prefer the session scratchpad or the system temp directory. Never let output artifacts (e.g. captured stdout/stderr) land inside the directory being verified.
Response format (include verbatim in the executor prompt):
Write your reasoning and scenario execution in natural language, then end your response with a single fenced JSON block matching this schema:
```json { "objective_met": "yes|partial|no", "remaining_gaps": ["<short phrase>"], "regressions": ["<short phrase>"], "suggested_edits": [ {"old_string": "<unique snippet>", "new_string": "<replacement>", "rationale": "<why>"} ], "confidence": "high|medium|low" } ```
old_stringmust match exactly one location in the current file. Include 1–3 lines of surrounding context so the snippet is unique — short one-liners collide and cause the Edit to fail.
(b) Parse & apply — evaluate in this order, first match wins
- Verdict missing or malformed — no fenced JSON block found, or JSON parse fails → return
status=skipped,reason="verdict parse failure". - Schema violation —
objective_metis not one ofyes|partial|no, or required keys are missing → returnstatus=skipped,reason="verdict schema violation". - Converged —
objective_met == "yes"ANDregressionsis empty → exit loop withstatus=convergedand proceed directly to Step 5 — Emit structured summary. Ifsuggested_editsis nonempty, discard the edits (do not apply them). Safety rails (c) do not run (no edit was applied this iteration). - Divergence — only when
i >= 2: if bothremaining_gapsANDregressionscontain the same elements as the previous iteration's values (compare as multisets — sort each array textually before comparison so a reordered-but-identical report still counts as divergence), the loop is not making progress → returnstatus=skipped,reason="divergent gaps". - Otherwise — apply
suggested_editsin order:- Re-Read the target file before each Edit so
old_stringmatches current contents. - If an
old_stringis not found, skip that edit and continue with the next. The skip is a no-op fallback, not an error. - After the edits (applied or skipped), run the safety rails in (c), then continue to iteration
i + 1.
- Re-Read the target file before each Edit so
(c) Per-iteration safety rails — run only if at least one edit was applied
- Frontmatter integrity — Re-Read the file. If the file begins with a
----delimited YAML frontmatter block, parse it; if parsing fails:
Returngit checkout HEAD -- <Target file>status=conflict,reason="frontmatter broken",reverted_paths=[<Target file>]. If the file has no frontmatter block at all (e.g., a plain source file), skip this rail. - Scope — Run
git diff --name-only. If any returned path is not<Target file>:
Returngit checkout HEAD -- <each offending path>status=conflict,reason="scope violation",reverted_paths=[<each offending path>].
Step 4 — Max iterations reached without convergence
Set status=unresolved, unresolved_gaps = <last remaining_gaps>. applied_edits_count reflects edits that actually landed (not skipped).
Step 5 — Emit structured summary
Emit a single fenced JSON block at the end of the response, matching the schema for the mode that ran:
- explicit-args mode: emit the schema below (4-value
statusenum). Existing callers (e.g.dev-workflow-triage's (d) verdict parser) consume this contract. - auto-derive mode: emit the aggregate schema defined in
## Auto-derive mode§ A3 (5-valuestatusenum includingpartial, plus the per-skill nesting).
The explicit-args mode schema:
{
"status": "converged|unresolved|skipped|conflict",
"iterations_used": N,
"objective_met": "yes|partial|no|unknown",
"applied_edits_count": N,
"unresolved_gaps": ["..."],
"reverted_paths": ["..."],
"reason": "verdict parse failure|verdict schema violation|divergent gaps|frontmatter broken|scope violation|missing target|incomplete args|empty diff|dispatch error|null"
}
The |null token at the end of the reason enum means JSON null (not the string "null").
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 47
- Forks
- 3
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
verify-diff- Source
- github.com/hiroro-work/claude-plugins