Skill: DiT 量化后推理验证(可选)

SkillAI & models

After W8A8 dynamic quantization, uses MindIE-SD smoke tests to verify that DiT quantized weights can be loaded (optional skill). Difference from `quant-tuning-evaluate` (DiT extension section): this skill only performs a lightweight "runs through + weights usable" verification (4, 16 prompts, results

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Skill: DiT 量化后推理验证(可选) skill

What this skill tells your AI

The instructions your AI receives, as published by kali20gakki/msagent in skills/quantizer/quant-tuning-infer-dit/SKILL.md and read by ahel’s review.

适用与不适用

  • 适用:已经 quant-tuning-quantize-dit 产出量化权重,需验证"能跑通 + 权重可用"。
  • 不适用:
    • 精度指标评估(FID/CLIP-score 等;本 skill 不计算)
    • 服务化推理(不引入 vLLM-Ascend / Tritonserver 等后端)
    • LLM/VLM 推理验证(走对应 LLM/VLM 子 skill)

协作关系

quantization-accuracy-tuning-orchestrator (workflow)
        │  按 model_family 路由到 dit 分支
        ▼ 调用(可选)
quant-tuning-infer-dit (tool)
        │
        ▼ 后端
  MindIE-SD
        │
        ▼ 输出
  {workdir}/infer_outputs/

输入参数

参数类型必填说明
quantized_pathstring✅quant-tuning-quantize-dit 产出路径(MindIE-SD 格式;对应其 output 字段 quantized_path)
prompt_liststring[]✅文本 prompt 列表;建议 4~16 条
num_inference_stepsint字段名 / 默认值按目标 DiT 推理仓差异较大;详见 inference_config_field_map.md §2 默认值表。
guidance_scalefloat同上;FLUX 字段同名 guidance_scale,HunyuanVideo / Wan 系字段名不同,详见字段表 §1。
output_dirstring✅推理产物目录;推荐 {workdir}/infer_outputs/
inference_repostring✅推理仓路径(pipeline 所在)
model_familystring✅dit

推理后端

后端说明
MindIE-SDNPU 上跑 MindIE-SD;适用 HunyuanVideo、Wan2.2-T2V/I2V/TI2V、FLUX.1-dev(已迁移)、Wan2.1、Sana、SD3、SD1.5/SDXL 等
  • 不依赖 vLLM-Ascend:MindIE-SD 单次调用即可
  • 不在本阶段做服务化:单次调用,不引入额外后端

工作流

┌───────────────────────────┐
│ 1. 入参校验                │
│ - quantized_path          │
│ - prompt_list 非空         │
└──────────┬────────────────┘
           ▼
┌───────────────────────────┐
│ 2. 加载 MindIE-SD 后端     │
└──────────┬────────────────┘
           ▼
┌───────────────────────────┐
│ 3. 加载量化 DiT pipeline  │
│ - 替换 transformer 为      │
│   量化版本                 │
└──────────┬────────────────┘
           ▼
┌───────────────────────────┐
│ 4. 逐 prompt 生成          │
│ - num_inference_steps     │
│ - guidance_scale          │
└──────────┬────────────────┘
           ▼
┌───────────────────────────┐
│ 5. 保存到 output_dir      │
│ - {idx:04d}.png           │
│ - 记录耗时                 │
└──────────┬────────────────┘
           ▼
┌───────────────────────────┐
│ 6. 汇总                    │
│ status=ok                 │
│ 成功数/总请求数             │
│ 平均推理耗时                │
└───────────────────────────┘

输出结果

成功

{
  "protocol": "msagent.subagent_io",
  "subagent_type": "quant-tuning-infer-dit",
  "status": "ok",
  "output": {
    "status": "ok",
    "backend": "mindie-sd",
    "output_dir": "<workdir>/infer_outputs/",
    "success_count": 8,
    "total_count": 8,
    "avg_inference_sec": 12.3,
    "generated_files": [
      "<workdir>/infer_outputs/0000.png",
      "<workdir>/infer_outputs/0001.png"
    ],
    "commands": [
      {"name": "load_pipeline", "command": "MindIE-SD pipeline 加载"},
      {"name": "inference", "command": "pipeline(prompt=...).images[0]"}
    ]
  }
}

失败

立即中止,回传:

{
  "protocol": "msagent.subagent_io",
  "subagent_type": "quant-tuning-infer-dit",
  "status": "failed",
  "error": {
    "code": "INFERENCE_ERROR",
    "message": "OOM / NaN / pipeline 加载失败摘要",
    "suggestion": "OOM / NaN → 减小 num_inference_steps 或换设备;pipeline 加载失败 → 检查 quantized_path 是否 MindIE-SD 格式;报 stderr 摘要,等待 orchestrator 决策"
  }
}

错误处理

错误类型处理
quantized_path 不存在立即中止,提示用户先跑 quant-tuning-quantize-dit
MindIE-SD pipeline 加载失败检查 quantized_path 中是否含 MindIE-SD 格式
OOM / NaN减小 num_inference_steps 或换设备;本 skill 不重试
prompt list 为空立即中止

约束

  • MindIE-SD only:禁止引入 vLLM-Ascend / Tritonserver 等后端
  • 不做精度指标评估:本 skill 仅验证"能跑通 + 权重可用"
  • 不引入服务化:单次调用,避免重复拉起服务
  • 不修改 LLM/VLM skill 既有行为字段

常见错误

错误原因解决
transformer not found in <quantized_path>量化权重目录结构与 diffusers 不一致检查 quant-tuning-quantize-dit save_path
NaN detected in latents量化精度不足回退到 quant-tuning-quantize-dit 检查 exclude 列表
OOM设备内存不足减小 num_inference_steps 或换设备

检查清单

参考

Signals

GitHub stars
31
Forks
8
Last commit
Sep 2026
Advanced
Item type
skill
Key
quant-tuning-infer-dit
Source
github.com/kali20gakki/msagent