Lofn Image — Codex-backed vision pipeline

SkillMedia

Run the Lofn image/visual pipeline (steps 00–10) backed by Codex — contest-grade, render-ready image prompts (Flux noun-first by default, GPT-Image-2 directive mode optional). Use for images, pictures, artwork, portraits, visual concepts, or "make an image with the full pipeline". Expects a Phase-0/1 orchestrator packet from the `lofn` skill; if none exists, run `lofn` first. Do NOT use for music, video, story prose, or QA-only audits.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Lofn Image — Codex-backed vision pipeline skill

What this skill tells your AI

The instructions your AI receives, as published by localsymmetry/lofn in .agents/skills/lofn-image/SKILL.md and read by ahel’s review.

Produces render-ready image prompts at Lofn competition grade (the catalog behind 11 first-place finishes). Depth lives in skills/image/; this skill runs it with Codex as the engine (hybrid execution per .agents/skills/lofn/EXECUTION.md).

Before you start

  1. Confirm a Phase-0/1 packet exists (core_seed.md, 04_metaprompt.md, 05_pair_assignments.md, 06_vision_handoff.md with the ICB / Panel Ledger, filled CREATIVE_CONTEXT.md). No packet → run lofn first.
  2. Pick the renderer mode (the metaprompt/dispatch may set TARGET_RENDERER):
    • FLUX / unset (daily & community challenges) → read skills/image/renderer_flux_rules.md. Description-style, noun-first, 80–150 words.
    • GPT_I2 (PRO) → read skills/image/renderer_gpt_image2_rules.md. Five-slot directive, camera-spec, 250–400 words; the orchestrator also swaps in the Typography Structuralist / Physics Epistemologist / Storybook-Assassin panel slots (see skills/orchestration/SKILL.md).
  3. Visual quality bar + density checklist: vault/VISION_QA_DEPTH_AUDIT.md (the Visual Somatic Gate).

Execution (hybrid)

Coordinator 00–05 inline, then 6 pairs as parallel subagents for 06–10. Inject the full CREATIVE_CONTEXT.md everywhere. Default cardinality: 6 pairs × 4 = 24 prompts → rank → top picks.

Coordinator steps (inline)

StepFileArtifact
00skills/image/steps/00_Generate_Image_Aesthetics_And_Genres.mdstep00_aesthetics_and_genres.md
01skills/image/steps/01_Generate_Image_Essence_And_Facets.mdstep01_essence_and_facets.md
02skills/image/steps/02_Generate_Image_Concepts.mdstep02_concepts.md (12 concepts)
03skills/image/steps/03_Generate_Image_Artist_And_Critique.mdstep03_artist_and_critique.md
04skills/image/steps/04_Generate_Image_Medium.mdstep04_medium.md
05skills/image/steps/05_Generate_Image_Refine_Medium.mdstep05_refine_medium.md6 pairs

Per-pair steps (parallel subagents, one chain per pair)

StepFilePer-pair artifact
06skills/image/steps/06_Generate_Image_Facets.mdpair_{NN}_step06_facets.md
07skills/image/steps/07_Generate_Image_Aspects_Traits.mdpair_{NN}_step07_aspects_traits.md
08skills/image/steps/08_Generate_Image_Generation.mdpair_{NN}_step08_generation.md (4 prompts)
09skills/image/steps/09_Generate_Image_Artist_Refined.mdpair_{NN}_step09_artist_refined.md
10skills/image/steps/10_Generate_Image_Revision_Synthesis.mdpair_{NN}_step10_revision_synthesis.md

Describe-render self-check (one capped pass, reuses the existing max-3-attempt loop — EXECUTION.md §4). Before a pair returns its step-10 prompts, it predicts in 2–3 sentences what its Flux/GPT-Image-2 prompt would actually PRODUCE — the literal frame a renderer would emit from this exact text (the composition, the dominant material, the light, what reads at thumbnail size) — then diffs that predicted frame against the Golden Seed. Phrase it adversarially: "name the one way this would render generic" (the stock-portrait centering, the material named but not foregrounded, the emotion stated in words the renderer can't draw, the Storybook-Assassin softness creeping back in). If the predicted frame drifts from the seed or names a generic outcome, self-repair ONCE through the same repair loop, then move on. This is one inline pass by the pair itself — no dedicated render-verifier subagent, no recursion, no new tier. It governs fidelity only; the noun-first / word-count / banned-opener contract below is unchanged.

The image prompt contract (hard gate — non-waivable)

Lead with seed/scene, end with the checklist.

FLUX mode (default):

  • Noun-first, present-tense. Write as if captioning an image that already exists. ❌ Forbidden imperative openers: Create, Design, Make, Render, Generate, Depict, Show, Draw, Build, Produce. ✅ "An archive fairy kneels in layered photogravure darkness…"
  • 80–150 words. Dense, not long — weaker models lose coherence past ~150. Every word earns its place.
  • Medium/material named in the first third ("Corroded mirror glass and torn washi frame…"); name concrete materials (cracked egg tempera, indigo cloth, dull bronze) — these are the hooks Flux renders.
  • ⭐⛔ TEXTURE IS PREDICATED, NEVER APPOSED (hard gate, added 2026-08-15). Every named material must be attached to the subject or to what the subject is doing. A material may not stand as a free-floating caption on the medium. This is a rule about grammar, not quantity — our 1st-of-869 winner is 3.7× more texture-dense than the 2026-08-09 run, and capping texture would optimise away from it.
    • Apposed (inventory): "oil on coarse linen, ridged alla prima" · "Antique chromolithographic medical plate, 1024x1536 portrait orientation"
    • Predicated (event): "her vast vellum wings—fractured and trembling—display intricate fractal scorch marks" · "gold leaf veins flicker with kinetic energy"
    • Self-check: for each material, name the verb it hangs off. A material with no verb is a caption — rewrite it or cut it.
  • THE UNANSWERED QUESTION IS MANDATORY (image_narrative_p100_floor). ART_SOUL principle 4. The prompt must contain an arrested moment or an unresolved question — something about to happen, mid-transformation, or looked at beyond the frame. Measured median in the 2026-08-09 run was ZERO; in the winner it runs hotter than the texture (5.56 vs 9.26 material, ratio 1.25). "The silent question of how long such fragile beauty can survive."
  • LIGHT FROM THE SUBJECT (image_light_p100_floor). ART_SOUL principle 2 — light emanates, it does not merely illuminate. Name the source and put it inside or on the subject.
  • Emotion SHOWN, not named until a possible final reveal ("one eye carries borrowed strain", not "defiant tenderness").
  • No camera specs, no Kelvin numbers — use mood words ("cold institutional light", "face-dominant, soft background blur"). Palette as color words.
  • Storybook Assassin ban (all modes) — BANS THE HAZE ADJECTIVE, NOT THE LIGHT MECHANISM (scoped 2026-08-15). No "ethereal, dreamlike, whimsical, gentle light, soft glow, magical, delicate, floating" as atmosphere adjectives standing in for an image. ⛔ This is not a ban on emanation: read as a blanket lexical ban it deleted most of the working vocabulary of principle 2 and of the fairy register that won four first places, and light prevalence fell 79% → 58%. The test: does the phrase name a MECHANISM a renderer can draw? ❌ "an ethereal, magical glow" · ✅ "light coming off her hands through the vellum". No living-artist names — reference techniques ("like a hand-tinted ambrotype"), not artists. Hands mentioned simply.

GPT_I2 mode: five-slot directive, front-loaded camera spec (85mm, f/1.8), lighting source+angle+Kelvin+shadow behavior, background declared (VOID/STRUCTURED/MINIMAL), 3+ named devices with visible behavior; apply the mandatory anti-default substitutions ("warm rim light" → "hard axial light from below"; "centered" → "lower third with aggressive negative space"; etc.). 250–400 words.

Both modes: subject legible at first glance (thumbnail impact); material specificity; emotional legibility in pose/expression. Aspect ratio: upload challenges 3:4, TikTok/Stories 9:16, landscape 16:9 (default 9:16 unless told otherwise).

⭐ ONE SEAM — N = 1 (hard gate, COMPETITION_LEARNINGS L38)

A piece is made OF a material, never ABOUT materials. The medium is the SUBSTANCE, never the SUBJECT.

  • One dominant material = the world. At most one intruding material = the wound. One seam, placed at the junction that carries the meaning — and the subject must be DOING something across it.
  • Multi-style collision engines are RETIRED (six-ways-v2 / Faerie Convocation / Impossible Bouquet lineage: N art movements sharing one frame at named collision devices). They produce no hero, N spectacle layers, and seams that vanish at 128 px. A catalogue has no verb. Variety belongs ACROSS a set — one obscure medium per image, differentiated against the field — never inside one frame.
  • KEEP the historical-medium anchor. Without a named lineage the renderer defaults to "pretty AI portrait". The anchor is the good half; the collision is the bad half.
  • Pair self-check (state it in step 10): count the distinct material systems in the frame. If the count is >2, cut until it isn't. Name the single seam and what the subject does across it.

THE VINCENT TEST — post-render, pre-submission (COMPETITION_LEARNINGS L39)

NightCafe's auto-generated "Creation Summary by Vincent" has never seen the prompt — a free blind first-read. If its headline names an EVENT, the piece works; if it names a STYLE, TECHNIQUE or MEDIUM, the picture is about its own craft — repair or hold. This is THE BLIND RULE for images; it costs nothing, so it is never skipped.

⭐ THE SIX-WORD EVENT — pre-render Vincent proxy (step 10, added 2026-08-15)

Vincent is free but it fires after the credits are spent. The same failure is catchable in text for nothing.

At step 10, each pair states in ≤6 words what is HAPPENING in its frame — no medium, no period, no technique, no format. If the only honest answer names a medium, the piece has no event: repair once through the existing loop, then surface.

Measured on our own work 2026-08-14 — Anatomia Fairia, the best-performing recent creation (8 likes), returned the Vincent headline "Antique Fairy Anatomy Medical Plate Portrait": a period, a medium, a technique and a format, and not one thing happening. Its subject is seated — the least eventful verb available. Compare the 1st-of-869: "a rebellious fairy kneels amid charred manuscripts… wings fractured and trembling."

  • ❌ "Antique chromolithographic medical plate" → names a medium. FAIL.
  • ✅ "She reaches for an already-burnt page" → names an event. PASS.

QA & delivery

Run lofn-qa with the Visual Somatic Gate (vault/VISION_QA_DEPTH_AUDIT.md, 7-element density checklist) + the renderer pre-gen check from the rules file. Also run python3 scripts/measure_visual_drift.py --gate <run-dir> — the machine reader for the VISUAL RETURN FLOORS in vault/gates.yaml.

RANKING ORDER (re-ordered 2026-08-15). Rank by the event (what is happening) → thumbnail impact → emanating light → visual concreteness → material specificity → emotional legibility → model fit. Material specificity was 3rd and is now 5th, and the event did not appear in this list at all — which is how a run shipped 24 prompts at a median narrative of zero. Select top picks. Save each selected prompt per skills/lofn-core/OUTPUT.md (note image_model + aspect_ratio in frontmatter); INDEX last. Do not call render tools — emit paste-ready prompt text; the user renders (FAL Flux / GPT Image 2). Refinement (anatomy/hands) is a post-render inpainting pass on the 1–2 chosen images only.

Signals

GitHub stars
22
Forks
1
Last commit
Aug 2026
Advanced
Catalog kind
skill
Gateway key
lofn-image-localsymmetry
Source
github.com/localsymmetry/lofn