Qwen Image Edit Workflows

SkillMedia

Build Qwen Image Edit workflows covering model loading, conditioning, LoRAs, prompt patterns, and XY plot testing

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Qwen Image Edit Workflows skill

What this skill tells your AI

The instructions your AI receives, as published by artokun/comfyui-mcp in plugin/skills/qwen-image-edit/SKILL.md and read by ahel’s review.

Overview

Qwen Image Edit uses a vision-language model (Qwen2.5-VL) to edit images based on natural language instructions. The model "sees" the source image through CLIP conditioning and generates an edited version.

Models

Required Components

ComponentNodeModel NameNotes
UNETUNETLoaderqwen_image_edit_2511_bf16.safetensorsOfficial 2511 edit model (bf16)
CLIPCLIPLoader (type=qwen_image)qwen_2.5_vl_7b_fp8_scaled.safetensorsShared across all Qwen models
VAEVAELoaderqwen_image_vae.safetensorsQwen-specific VAE

Alternative UNET Models

ModelPathFocus
qwenImageEditRemix_v10qwenImageEditRemix_v10.safetensorsCommunity remix, general editing
qwenUltimateRealism_v11Qwen/imageized/qwenUltimateRealism_v11.safetensorsProduct photography, hyper-realistic
copaxTimelessQwen/realistic/copaxTimeless_qwenUltraRealistic.safetensorsUltra-realistic portraits
qwnImageEdit_v16Bf16Qwen/abliterated/qwnImageEdit_v16Bf16.safetensorsAbliterated (uncensored)

Conditioning Nodes

TextEncodeQwenImageEditPlusAdvance_lrzjason (Recommended)

From the qweneditutils custom node pack. The Advanced variant is preferred because it:

  • Outputs a LATENT directly (no need for separate EmptyLatentImage)
  • Has separate VL-resize and non-resize image slots for fine control
  • Supports target_size control for output resolution
  • Includes a pad/center/disabled crop method with pad_info output
Required Inputs:
  - clip: CLIP
  - prompt: STRING — natural language edit instruction

Optional Inputs:
  - vae: VAE — needed for image encoding and latent output
  - vl_resize_image1-3: IMAGE — images that get VL-resized (downscaled for vision encoder)
  - not_resize_image1-3: IMAGE — images kept at full resolution
  - target_size: [1024, 1344, 1536, 2048, 768, 512] (default 1024)
  - target_vl_size: [392, 384] (default 384)
  - upscale_method: [lanczos, bicubic, area]
  - crop_method: [pad, center, disabled]
  - instruction: STRING — system instruction template (has sensible default)

Outputs (10):
  [0] conditioning_with_full_ref: CONDITIONING — use as positive conditioning
  [1] latent: LATENT — auto-scaled latent, feed directly to KSampler
  [2] target_image1: IMAGE — processed target-size image
  [3] target_image2: IMAGE
  [4] target_image3: IMAGE
  [5] vl_resized_image1: IMAGE — VL-resized version
  [6] vl_resized_image2: IMAGE
  [7] vl_resized_image3: IMAGE
  [8] conditioning_with_first_ref: CONDITIONING — conditioning with only first ref
  [9] pad_info: ANY — padding info for later unpadding

Key advantage: Output [1] (latent) eliminates the need for a separate EmptyLatentImage or VAEEncode node. The Advanced node handles latent creation internally at the correct resolution.

Other Conditioning Variants

  • TextEncodeQwenImageEditPlus (Phr00t v2, built-in) is simpler: 4 image inputs, outputs only CONDITIONING. Requires separate EmptyLatentImage. Good for quick edits.
  • TextEncodeQwenImageEditPlus_lrzjason: 5 image inputs, resize toggles, but less control than Advance
  • TextEncodeQwenImageEditPlusPro_lrzjason: Per-image VL resize selection via vl_resize_indexs string, main_image_index control

Lightning LoRAs (Fast Generation)

4-Step Lightning (2511 Edit)

{
  "class_type": "LoraLoaderModelOnly",
  "inputs": {
    "model": ["<unet_node>", 0],
    "lora_name": "Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors",
    "strength_model": 1.0
  }
}

Settings: steps=4, cfg=1.0, sampler=euler, scheduler=simple, denoise=1.0

4-Step Lightning (General Qwen)

For non-edit models (txt2img, 2512):

  • Qwen-Image-Lightning-4steps-V1.0.safetensors (strength 1.0)

8-Step Lightning

  • Qwen-Image-Lightning-8steps-V1.0.safetensors, higher detail than 4-step

Sampler Settings

PresetStepsCFGSamplerSchedulerDenoiseLoRA
Lightning 4-step (2511 edit)41.0eulersimple1.02511-Lightning-4steps
Lightning 8-step81.0eulersimple1.0Lightning-8steps
Standard edit404.0eulersimple0.75none
Quality edit504.0eulersimple0.5-0.8none

The sub-1.0 denoise rows REQUIRE a VAEEncode latent. A denoise low enough to shorten the sampling schedule — which 0.5-0.8 certainly is — keeps part of the incoming latent, so that latent has to BE the source image. Wire latent_image from a VAEEncode of the source (or from a node that emits a source-derived latent, like TextEncodeQwenImageEditPlusAdvance_lrzjason output [1]). Pairing these rows with an EmptyLatentImage runs clean and returns a flat, near-uniform field — an empty latent has no source content to preserve. Feeding the reference through TextEncodeQwenImageEditPlus does not rescue it: that image rides on CONDITIONING, which steers denoising but never seeds the sampler's starting state.

Denoise for editing: Lower denoise = closer to source — provided the latent IS the source. 0.5-0.8 range for standard editing on a VAEEncode latent. Lightning uses 1.0 (model handles fidelity internally).

Resolutions

This table is for Qwen-Image TEXT-TO-IMAGE. Do not pick an edit graph's output size from it. An edit graph's geometry is decided by the SOURCE image, not by you — see "Resolution on an edit graph" below. Choosing 1104x1472 here for an edit was #2681.

Qwen-Image operates at ~1.6 megapixels natively:

AspectResolutionUse Case
Square1328x1328General
Portrait 3:41104x1472Portraits
Portrait 9:16928x1664Phone format
Landscape 4:31472x1104Landscape scenes
Landscape 16:91664x928Widescreen
Video-ready832x480For WAN 2.2 FLF pipeline

For video pipelines: Use 832x480 to match WAN 2.2's default resolution.

Resolution on an edit graph

TextEncodeQwenImageEdit and TextEncodeQwenImageEditPlus do not take a size. They scale every reference image to a hard-coded int(1024 * 1024) px — ~1.05 MP, at the source's own aspect ratio — VAE-encode it, and hand it to the model as a reference latent (comfy_extras/nodes_qwen.py).

The model then lays the reference tokens and the tokens it is generating on one shared, centred coordinate grid (comfy/ldm/qwen_image/model.py, process_img), so reference position (i, j) and output position (i, j) mean the same place only when the two grids are the same size. That is the whole reason the official templates run the one image through FluxKontextImageScale and then feed the sampler a VAEEncode of that scaled image — both branches then see the same pixels at the same ~1 MP scale and the same aspect. Every PREFERRED_KONTEXT_RESOLUTIONS entry is ~1.05 MP for the same reason.

It is agreement to within the encoder's round-to-8, not exact equality, and the difference is worth knowing precisely. FluxKontextImageScale snaps to a preferred pair; the encoder then renormalises that to its own 1,048,576 px budget. For 12 of the 18 preferred pairs the two land on the same latent grid. For the other 6 — 688x1504, 800x1328, 832x1248 and their landscape mirrors — the reference lands one latent row or column off: at 800x1328 the sampler's grid is 100x166 and the reference's is 99x165. ComfyUI's own bundled 2511 template does exactly this, so a sub-patch offset is evidently fine in practice. The failure this page is about is one of SCALE, not of rounding — an empty latent at 1104x1472 sits 1.24x away linearly, not one row.

So on an edit graph you do not choose a resolution — you inherit one:

  • Right: LoadImage -> FluxKontextImageScale -> (TextEncodeQwenImageEditPlus and VAEEncode) -> KSampler latent_image.
  • Wrong: an EmptyLatentImage at a size from the table above. Its dimensions are literals; the reference's are computed from the source when the graph runs. At 1104x1472 (1.63 MP) against a 1.05 MP reference the grids are 1.24x apart linearly, the model cannot copy detail across them, and it re-synthesises the subject instead — materials come back looking plastic/CGI and printed detail comes back as a generic shape (#2681). Nothing errors; the image just is not the edit you asked for.

create_workflow (action:"validate") now flags this pairing as edit_reference_empty_latent.

Prompt Patterns

Edit Instructions (Natural Language)

"Change the black cat into a cute girl with a black bodysuit and jeans"
"Make the sky a dramatic sunset with orange and purple clouds"
"Add a red sports car parked in front of the house"
"Remove the person on the left and fill with the background"

Multi-Angle LoRA (qwen-image-edit-2511-multiple-angles-lora)

Uses <sks> token with structured angle/distance prompts:

<sks> front view eye-level shot close-up
<sks> front-right quarter view low-angle shot medium shot
<sks> back view elevated shot wide shot

Template: <sks> {direction} view {angle} shot {distance}

Directions: front, front-right quarter, right side, back-right quarter, back, back-left quarter, left side, front-left quarter Angles: low-angle, eye-level, elevated, high-angle Distances: close-up, medium shot, wide shot

Negative Conditioning

Always use ConditioningZeroOut for negative conditioning with Qwen edit:

{
  "class_type": "ConditioningZeroOut",
  "inputs": { "conditioning": ["<positive_cond_node>", 0] }
}

Complete Workflow: Lightning Edit (Advanced Node)

Uses TextEncodeQwenImageEditPlusAdvance_lrzjason, which outputs the latent directly, so no EmptyLatentImage is needed.

{
  "1": { "class_type": "UNETLoader", "inputs": { "unet_name": "qwen_image_edit_2511_bf16.safetensors", "weight_dtype": "default" }},
  "2": { "class_type": "LoraLoaderModelOnly", "inputs": { "model": ["1", 0], "lora_name": "Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors", "strength_model": 1 }},
  "3": { "class_type": "CLIPLoader", "inputs": { "clip_name": "qwen_2.5_vl_7b_fp8_scaled.safetensors", "type": "qwen_image" }},
  "4": { "class_type": "VAELoader", "inputs": { "vae_name": "qwen_image_vae.safetensors" }},
  "5": { "class_type": "LoadImage", "inputs": { "image": "<source_image.png>" }},
  "6": { "class_type": "TextEncodeQwenImageEditPlusAdvance_lrzjason", "inputs": {
    "clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0],
    "vl_resize_image1": ["5", 0],
    "target_size": 1024, "target_vl_size": 384,
    "upscale_method": "lanczos", "crop_method": "pad"
  }},
  "7": { "class_type": "ConditioningZeroOut", "inputs": { "conditioning": ["6", 0] }},
  "8": { "class_type": "KSampler", "inputs": {
    "model": ["2", 0],
    "positive": ["6", 0],
    "negative": ["7", 0],
    "latent_image": ["6", 1],
    "seed": 42, "steps": 4, "cfg": 1, "sampler_name": "euler", "scheduler": "simple", "denoise": 1
  }},
  "9": { "class_type": "VAEDecode", "inputs": { "samples": ["8", 0], "vae": ["4", 0] }},
  "10": { "class_type": "SaveImage", "inputs": { "images": ["9", 0], "filename_prefix": "qwen_edit" }}
}

Key connections:

  • "latent_image": ["6", 1]: KSampler gets its latent directly from the Advanced node's output [1]
  • "positive": ["6", 0]: conditioning_with_full_ref from output [0]
  • "vl_resize_image1": ["5", 0]: source image goes into VL-resize slot (downscaled for vision encoder)

Simpler Alternative (Phr00t v2)

If qweneditutils custom node is unavailable, use the built-in TextEncodeQwenImageEditPlus with a separate EmptyLatentImage:

{
  "6": { "class_type": "TextEncodeQwenImageEditPlus", "inputs": {
    "clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0], "image1": ["5", 0]
  }},
  "8": { "class_type": "EmptyLatentImage", "inputs": { "width": 1024, "height": 1024, "batch_size": 1 }}
}

Replace node 6 and add node 8. KSampler latent_image connects to ["8", 0] instead of ["6", 1].

Do not leave that EmptyLatentImage wired to the sampler. It is shown above only because it is what the Plus encoder's own signature leaves you needing, and it fails in a different way on each side of denoise 1.0:

  • Below 1.0 it renders a flat, near-uniform field, always. The encoder emits CONDITIONING only, so the latent is genuinely empty, and a truncated sigma schedule exists to preserve an incoming latent that here has nothing in it (#2678).
  • At 1.0 it renders a plausible image that is not your source — unless its width and height happen to equal the geometry the encoder derived, which is ~1.05 MP at the source's aspect ratio. A 1024x1024 empty latent over an exactly-square source does line up, and is fine. 1104x1472 over that same source does not: the model aligns reference and output on one shared grid, cannot copy across grids that far apart, and re-synthesises instead (#2681). See "Resolution on an edit graph".

The second case is the trap, because it depends on a source you may not have looked at and it fails silently. A VAEEncode removes the coincidence — it cannot be the wrong size, because it is derived from the same pixels the encoder saw.

Feed latent_image from a VAEEncode of the same image you gave the encoder, scaled once up front so both branches see the same pixels at the same scale:

{
  "5b": { "class_type": "FluxKontextImageScale", "inputs": { "image": ["5", 0] }},
  "6":  { "class_type": "TextEncodeQwenImageEditPlus", "inputs": {
    "clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0], "image1": ["5b", 0]
  }},
  "8":  { "class_type": "VAEEncode", "inputs": { "pixels": ["5b", 0], "vae": ["4", 0] }}
}

KSampler latent_image connects to ["8", 0]. Any denoise is then meaningful: 1.0 for a full edit, sub-1.0 to stay closer to the source.

Basic Variant (Official ComfyUI Example)

The official "Qwen 2511 Edit Simple" example uses newer built-in nodes for model patching and image scaling:

Additional nodes in the official pipeline:

  • ModelSamplingAuraFlow (shift=3.1): Flow matching shift applied to the UNET. Used instead of ModelSamplingSD3.
  • CFGNorm (strength=1): Normalizes CFG guidance for more stable generation. Applied after ModelSamplingAuraFlow.
  • FluxKontextImageScale: Auto-scales input images to the correct resolution for Qwen. No manual size parameters needed.
  • FluxKontextMultiReferenceLatentMethod (method=index_timestep_zero): Applied to both positive and negative conditioning. Handles multi-reference latent indexing.
  • VAEEncode: Encodes the scaled image to latent (instead of EmptyLatentImage).

Official pipeline flow:

UNETLoader → [LoraLoaderModelOnly] → ModelSamplingAuraFlow (shift=3.1) → CFGNorm (strength=1) → MODEL
CLIPLoader (qwen_image) → CLIP
VAELoader → VAE

LoadImage → FluxKontextImageScale → scaled_image
  ├─ TextEncodeQwenImageEditPlus (positive) → FluxKontextMultiReferenceLatentMethod → positive CONDITIONING
  ├─ TextEncodeQwenImageEditPlus (negative, empty) → FluxKontextMultiReferenceLatentMethod → negative CONDITIONING
  └─ VAEEncode → LATENT

KSampler → VAEDecode → SaveImage

Official sampler settings:

VariantStepsCFGSamplerSchedulerDenoiseLoRA
Standard404.0eulersimple1.0none
Lightning41.0eulersimple1.02511-Lightning-4steps

Note: The FluxKontextMultiReferenceLatentMethod and FluxKontextImageScale nodes may not be needed when using Comfy's official model files directly, but may be required with community-repackaged models.

XY Plot Technique (from Widgets.json)

For batch-testing multiple edit variations, use the Easy Nodes XY Plot system:

  1. Text Multiline nodes define parameter lists (e.g., directions, angles)
  2. Split String breaks them into indexed options
  3. easy textIndexSwitch selects one at a time
  4. easy promptReplace substitutes {X}, {Y}, {Z} placeholders in the base prompt
  5. easy XYPlotAdvanced + easy XYInputs: PromptSR drives the sweep
  6. easy pipeIn bundles model/clip/vae/latent into a pipeline

This produces a grid image showing all combinations, useful for finding the best angle/distance/style for a given subject.

VRAM Considerations

  • Qwen 2511 edit bf16: ~10GB VRAM
  • CLIP (fp8): ~7GB VRAM
  • VAE: ~200MB
  • Total: ~17-18GB, fits comfortably on 24GB GPUs
  • Always clear_vram before loading if switching from another model family

Tips

  1. Upload source images first with upload_image (action:"image") before building the workflow
  2. Match output resolution to the next pipeline step (e.g., 832x480 for WAN FLF) — but only on a generation graph. On an edit graph the size is the source's; resize the RESULT afterwards instead of sampling at the size you want
  3. Lightning LoRA + denoise 1.0 works well. The model handles structure preservation through conditioning
  4. Take an edit graph's latent_image from a VAEEncode of the source, not from an EmptyLatentImage — at sub-1.0 denoise the empty latent decodes to a flat, near-uniform field (#2678), and at denoise 1.0 it is right only if its literal size happens to equal the geometry the encoder derived from the source, which is exactly the coincidence a VAEEncode removes (#2681). Neither failure errors. create_workflow (action:"validate") flags both pairings (partial_denoise_empty_latent, edit_reference_empty_latent)
  5. The lrzjason Pro variant is best for multi-image compositions where you need fine control over which images get VL-resized
  6. Use get_workflow (action:"analyze") to understand any saved Qwen edit workflow before modifying or executing it. It returns a structured summary, not raw JSON. Only use get_workflow when you need the actual JSON for enqueue_workflow or create_workflow (action:"modify").

Sources

  • Official: none found.
  • Empirical: sampler values, wiring, and prompt notes from working graphs in packs/ and observed renders; not a vendor prompting guide.

Signals

GitHub stars
739
Forks
120
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
qwen-image-edit
Source
github.com/artokun/comfyui-mcp