Speculative Decoding Draft-Model Training

SkillMonitoring & ops

Guides your agent through training and debugging a speculative decoding draft model step by step.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Speculative Decoding Draft-Model Training skill

About this skill

Train, debug, and validate a speculative decoding draft model (EAGLE3, DFlash, DSpark, Domino) through the ModelOpt launcher pipeline. Use when the user wants to add a new model to a draft-training pipeline, asks why a pipeline run failed, wants experiment logs reviewed, or wants to check whether a

What this skill tells your AI

The instructions your AI receives, as published by nvidia/model-optimizer in plugins/modelopt/skills/speculative-decoding/SKILL.md and read by ahel’s review.

Everything needed to take a target model from "no draft head" to "validated acceptance rate" lives in this directory. Work through the stages below in order for a new model; jump straight to a stage when you already know which one you need.

Two axes: stage and algorithm

The pipeline is the same shape for every draft-model algorithm — synthesize data, dump base-model hidden states, train the draft, benchmark acceptance rate. What changes between algorithms is which training script and recipe run, which knobs matter, and which failures are typical.

So this skill is split along those two axes, and you almost always read one file from each:

AxisDirectoryWhat it holds
Stagereferences/stages/The procedure — algorithm-independent
Algorithmreferences/algorithms/The data sheet — scripts, recipe, knobs, thresholds, known failures

Read the stage file for what to do, and the algorithm sheet for the values to plug in. When a stage file says "see the algorithm sheet", it means the section of references/algorithms/<algorithm>.md with the matching heading.

Stages

StageReferenceUse when
1. Configurereferences/stages/configure.mdAdding a model that has no pipeline YAML yet
2. Review logsreferences/stages/review-logs.mdA run finished (or died) and you want a pass/fail summary
3. Triagereferences/stages/triage.mdA task failed and you need root cause plus a fix
4. Validatereferences/stages/validate.mdAll tasks passed and you need to confirm the acceptance rate gate

Review-logs and triage overlap by design: review-logs is the fast sweep across all tasks, triage is the deep dive into one failing task. Start with review-logs unless the user already knows which task broke.

Algorithms

All recipes live in modelopt_recipes/general/speculative_decoding/<algorithm>.yaml.

AlgorithmSheetFamily
EAGLE3references/algorithms/eagle3.mdAutoregressive draft head
DFlashreferences/algorithms/dflash.mdBlock diffusion
DSparkreferences/algorithms/dspark.mdDFlash backbone + Markov head + optional confidence head
Dominoreferences/algorithms/domino.mdDFlash backbone + GRU causal correction head

DSpark and Domino are DFlash variants, not separate pipelines: same recipe_type: speculative_dflash, same training script, same dflash.* config namespace, selected by dflash_architecture_config.projector_type. Read references/algorithms/dflash.md first, then the variant's sheet for the delta.

If the user's algorithm has no sheet yet, the stage procedures still apply — derive the missing values from an existing launcher example for that algorithm (tools/launcher/examples/*/*/hf_*_<algorithm>.yaml) and its recipe, then write the sheet as you go. references/algorithms/README.md defines what a sheet must contain.

End-to-end: a new model

  1. Confirm the algorithm and find the closest existing launcher example.

  2. Configure — write the pipeline YAML (references/stages/configure.md).

  3. Preview with --dryrun, then submit:

    cd tools/launcher
    uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes
    
  4. Register the job and set up monitoring per the monitor skill.

  5. Review logs when it finishes (references/stages/review-logs.md).

  6. Triage anything that failed (references/stages/triage.md), fix, re-run only the failed tasks onward via pipeline.task_N.skip=true.

  7. Validate once all tasks pass (references/stages/validate.md).

Model-support gaps that need code changes land in modelopt/torch/speculative/ and require a separate ModelOpt PR — the pipeline YAML alone cannot fix an unrecognized architecture.

Signals

GitHub stars
5k
Forks
698
Last commit
Sep 2026
Advanced
Item type
skill
Key
speculative-decoding-nvidia
Source
github.com/nvidia/model-optimizer