Scenario Refine Loop
SkillAI & modelsLets your agent check generated images against their brief and keep revising them until they match.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Scenario Refine Loop skill
About this skill
Use when a Scenario output must be checked and improved rather than accepted first roll: iterating a generation until it matches the brief, fixing a batch that came back off-brief, retrying failed shots methodically, wiring an automated generate, review, revise loop, or deciding whether to re-prompt
What this skill tells your AI
The instructions your AI receives, as published by scenario-labs/skills in skills/scenario-refine-loop/SKILL.md and read by ahel’s review.
Overview
Agents fail generation QA in two symmetric ways: accepting the first roll, or rewording the whole prompt and re-rolling until the budget dies. Both skip the same two artifacts, a written rubric and a diagnosis. The loop that converges: rubric before generating, a small batch, a recorded verdict per asset, the cheapest targeted fix per failure, a hard round cap. Connection and the core loop: see the scenario skill. Critic tool contracts: scenario-asset-analysis. Baseline discipline: scenario-consistency. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.
Quick reference
| Step | Do |
|---|---|
| 1. Rubric | Before generating, turn the brief into pass/fail lines a viewer can check ("subject centered on a plain field"), never taste words |
| 2. Generate | The smallest batch that tests the recipe, at the cheapest size or quality tier the schema offers on which every rubric line can still be judged; dry_run when cost matters |
| 3. Critique | asset_analyze: up to 10 images per call, one instruction embedding the rubric and a fixed per-image output shape |
| 4. Fix | Route every fail line to the cheapest fix that addresses it (table below) |
| 5. Stop | A clean round ships; three rounds without one, or one line failing twice under different fixes, means report, not respin |
A tier change is a new generation, not the keeper enlarged: re-run only the keeper's recipe at delivery tier and re-critique it, or keep what passed and upscale it with a fidelity upscaler found by search target="models", filters={"tags": ["image-upscale"]}, public=true (fidelity versus creative picks and sizing in scenario-image-editing, video upscaling in scenario-video).
When the bar is the configured brand brief rather than a task rubric, and the team's Quality Gate add-on is enabled, critique images with asset_quality_gate_run instead: its reasons and suggestions feed the fix table directly (scenario-quality-gate; where the gate is missing it degrades to this asset_analyze path).
Fix routing, cheapest first:
| The verdict says | Fix |
|---|---|
| One local defect on a keeper | Masked inpaint of that region; on a schema with no mask field, an instruction edit of the keeper naming that one change (scenario-image) |
| A uniform finish off (grade, tint, crop) | A deterministic tool pass (scenario-image-editing), not a re-roll |
| Wrong content, composition, or rendered palette | Edit the delta clause, re-run from the approved baseline |
| Identity or style drift | Tighten the enumeration, add or re-role references (scenario-consistency) |
| Every line failing | Change the model: re-discover with recommend (capability-shaped; search is for a name, a private model, or a tag-filtered lane such as image-upscale), keep the prompt |
Change one variable per round. A round that swaps prompt, references, and model at once cannot attribute the improvement, so the next failure restarts from zero.
The two rules that keep the loop honest
- Verdicts cite the rubric, not taste. Instruct the critic to answer per image, in order:
<index>: pass|fail, <the failed line>. "Could be better" is not a verdict; a loop chasing better instead of the brief sands off exactly what made the direction distinctive and converges on generic output. - Regenerate from the baseline, never from the last attempt. Fixes re-run from the approved reference and prompt; chaining output to output compounds drift (
scenario-consistencyexplains why).
Worked example: four icons against a brief
- Rubric from the brief, five lines: single object, centered, plain field, palette #2A9D8F and #E9C46A only, no text.
- Generate four, one
model_runeach perscenario-image, thenjobs_wait. - One
asset_analyzecall with all four ids inimagesand the rubric-plus-shape instruction; answers land as text assets,asset_downloadthem to read the verdicts. - Two fail. Icon 2 is off palette, which is rendered content rather than a finish: edit the delta clause by pinning the hex codes in the prompt and re-run from the baseline. Icon 4 has one smeared edge: masked inpaint of that corner. Icons 1 and 3 ship untouched.
jobs_waitthe fix runs, then re-critique only the two new assets with the byte-identical instruction. Clean round: stop, file the keepers in a collection (scenario-asset-analysis).
Common mistakes
- Judging by glancing at
asset_displayin chat: unrecorded impressions do not accumulate; verdicts do. - Asking
asset_analyzeto improve or fix the image: it returns text only; every fix is a new run. - Re-rolling the whole batch because one item failed: route per item.
- Writing the rubric after seeing the batch: it inherits the batch's flaws as the standard.
- Running the loop uncapped: failed jobs are reimbursed, unsatisfying ones are not. Gate rounds on the costs the runs themselves report (
dry_runprices the next round ahead);usagetotals lag and answer the report after the run, not the mid-run gate. - Retrying a criterion a third time on the same model: two misses under two different fixes is evidence about the model, not bad luck.
Signals
- GitHub stars
- 681
- Forks
- 82
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
scenario-refine-loop- Source
- github.com/scenario-labs/skills