VeOmni Patchgen Modeling Protocol
SkillAI & modelsAuthor or refresh a VeOmni model's patchgen-generated modeling under generated/, GPU and/or NPU config, dense or MoE, text / VLM / Omni. Covers the patchgen decorators, sharing patches across sibling models via name_map, MoE fused-expert weight loading, Ulysses SP in multimodal forwards, __init__.py registration, running codegen, and the test cases. This is the modeling step of adding a new model, not only of refreshing an existing one. Trigger: 'add patchgen for a model', 'write a patch_gen_config', 'regenerate the generated modeling', 'add NPU patchgen', 'port a model to patchgen', 'transformers v5 migration'. Never hand-edit anything under generated/.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the VeOmni Patchgen Modeling Protocol skill
What this skill tells your AI
The instructions your AI receives, as published by bytedance-seed/veomni in .agents/skills/veomni-patchgen-model/SKILL.md and read by ahel’s review.
Purpose: add or refresh a model's patchgen-generated modeling under
veomni/models/transformers/<model>/generated/. VeOmni pins
transformers==5.16.1 and ships patchgen-generated modeling for every
supported transformers-family model. The non-transformers architectures
(flux, movqgan, wan) have no generated/ directory and are out of scope.
References (read first, load on demand):
docs/design/patchgen.md— patchgen DSL, CLI, CI drift checkdocs/transformers_v5/transformers_v5_moe_weight_loading.md— MoE fused-expert layout + runtime converterdocs/transformers_v5/veomni_flash_attention_kernel_adapter.md— FA custom-name adapterdocs/transformers_v5/testing_new_model.md— test case SOP for a new model
What to read for your model
This file is the spine: it applies to every model. The category-specific
material lives in references/ — load only what your model needs.
| Your model | Also read |
|---|---|
| Any model, before Phase 1 | references/model-examples.md — pick the closest existing model and mirror it |
| Has routed experts (MoE) | references/moe.md — Phase 2 expert patches, Phase 3 checkpoint converter, MoE pitfalls |
| Has a vision / audio / speech tower (VLM or Omni) | references/multimodal.md — SP-aware multimodal forward, metadata precompute, dummy_forward, subtree pruning, VLM/Omni pitfalls |
| Text-only and dense | neither — the spine plus the examples file is the whole protocol |
A text-only dense GPU model therefore reads this file plus the examples, and skips about 380 lines of MoE and multimodal material. A VLM+MoE model reads everything. Read the spine first either way; the reference files add to it and never replace a phase.
Phase 0: Environment + Reference Setup
0.1 Verify transformers venv
Patchgen runs against transformers==5.16.1. Before touching code:
source .venv/bin/activate
python -c "import transformers; print(transformers.__version__)"
If not 5.16.1, re-sync the default env:
uv sync --frozen --extra gpu --group dev
source .venv/bin/activate
0.2 (Strongly recommended) Drop HF reference source into .agents_workspace/
.agents_workspace/ is gitignored. Keeping the upstream HF source next to your
patchgen config is the single biggest accelerator for catching subtle
signature/contract drift while iterating.
Use the pinned version as the directory name so several pins can coexist:
PIN=$(python -c "import transformers; print(transformers.__version__)")
mkdir -p ".agents_workspace/hf_reference/<m>/v${PIN}"
curl -fsSL -o ".agents_workspace/hf_reference/<m>/v${PIN}/modeling_<m>.py" \
"https://github.com/huggingface/transformers/raw/v${PIN}/src/transformers/models/<m>/modeling_<m>.py"
-f matters: without it a missing tag or renamed module returns 404 and curl
writes the error page into modeling_<m>.py with exit status 0, so you would
diff against an HTML page and not notice.
For VLMs also grab processing_<m>.py / image_processing_<m>.py /
configuration_<m>.py if you expect processor-side or config-shape work.
If you are refreshing an existing patchgen-generated file across a
transformers minor bump (the pin the generated file was produced against → the
new pin), pull both versions side-by-side and diff to spot contract drift —
substitute the
<old_ver> / <new_ver> tags with the actual versions you are migrating
between:
mkdir -p .agents_workspace/hf_reference/<m>/{old,new}
curl -fsSL -o .agents_workspace/hf_reference/<m>/old/modeling_<m>.py \
"https://github.com/huggingface/transformers/raw/<old_ver>/src/transformers/models/<m>/modeling_<m>.py"
curl -fsSL -o .agents_workspace/hf_reference/<m>/new/modeling_<m>.py \
"https://github.com/huggingface/transformers/raw/<new_ver>/src/transformers/models/<m>/modeling_<m>.py"
diff -u .agents_workspace/hf_reference/<m>/{old,new}/modeling_<m>.py | less
0.3 For a pin bump: survey signature drift across all configs first
patchgen <config> (without --dry-run) runs the generated file through ruff,
so a patch body referencing a symbol upstream no longer defines fails loudly
with F821 / F811. That catches removed names. It does not catch a
patch whose target still exists but whose signature changed — the patch
keeps applying and silently runs against the wrong contract.
Before touching any config, index both upstream versions with ast and compare
the parameter lists of every target named in each config's override_method /
replace_class / replace_function call. Targets missing from both versions
are VeOmni-added methods (patchgen uses override_method to inject them) and
should be filtered out, or they drown the real findings.
docs/transformers_v5/upgrade_5_9_to_5_16.md records what that survey turned up
for the 5.9 → 5.16 bump and how each class of breakage was resolved — read it
before starting a new bump, the categories repeat.
Note that --dry-run returns before the ruff step, so it reports success on
files that cannot even import. Never use it as the pass/fail signal.
Things to watch for in upstream contracts:
@can_return_tuple,@capture_outputs,@merge_with_config_defaults,@auto_docstringdecorators → affect behavior of youroverride_method. When youoverride_methodon a@auto_docstring-decorated method, every parameter you declare in the new signature must also appear in the patched docstring'sArgs:block — otherwiseauto_docstringwill emit warnings at import time about "undocumented parameter". For Omni-style overrides that add params likeaudio_feature_lengths,feature_lens,aftercnn_lens,rope_deltas,image_grid_thw,video_grid_thw, etc., copy the upstream docstring and append minimal one-line entries for every new param.attention_maskmay be a dict — HF v5 routinely passesattention_mask={"full_attention": <tensor>, ...}keyed by attention type. Any patched forward that forwardsattention_masktocompute_3d_position_ids/get_rope_index/ other tensor-expecting helpers must defensively unwrapattention_mask.get("full_attention", None)when it's a dict.
VLM and Omni models have four more upstream contracts to check before writing
any patch — placeholder masks, get_{image,video}_features return shapes, the
packed position-ids layout and mrope shape collapse. See
references/multimodal.md, "Phase 0: upstream contracts".
Keep this directory around through commit; delete it after the PR merges (it's already gitignored so it won't leak into the repo).
Before You Start: Create a Plan
Track the phases with whatever todo/plan tool the running agent provides. Suggested plan:
Phase 0: Verify venv + drop HF reference files -> in_progress
Phase 1: Scope & audit upstream surface -> pending
Phase 2: Draft <model>_gpu_patch_gen_config.py -> pending
Phase 3: (MoE only) Add checkpoint converter -> pending
Phase 4: Wire __init__.py to expose generated classes -> pending
Phase 5: Run patchgen + verify diff -> pending
Phase 6: Add test cases -> pending
Phase 7: Run tests (single-GPU + e2e) -> pending
Phase 8: Docs + commit; review before PR or substantive update -> pending
Drop phases that don't apply (e.g. Phase 3 for non-MoE models).
Phase 1: Scope & Audit
Input: model name <M> (e.g. qwen3_5, glm_moe_dsa).
Operations:
- Locate
veomni/models/transformers/<M>/. If the directory does not exist yet you are being called as the modeling step of/veomni-new-model: create it, and read that skill's Phase 1 first so the category (text / VLM / Omni, dense / MoE, GPU-only or GPU+NPU) is already decided when you get here. - If a patchgen-generated file already exists under
veomni/models/transformers/<M>/generated/you are refreshing an existing config (e.g. picking up upstream changes, adding NPU sibling, fixing a bug). Otherwise you are writing the first config for this model. Either way, the rest of this protocol applies identically. - Decide backend coverage:
- GPU only → one
<m>_gpu_patch_gen_config.py+ onegenerated/patched_modeling_<m>_gpu.py. - GPU + NPU → add sibling
<m>_npu_patch_gen_config.pythat writesgenerated/patched_modeling_<m>_npu.py; mirror theglm_moe_dsaorqwen3_vllayout.
- GPU only → one
- Check model category. Each entry below names the closest existing model;
references/model-examples.mdsays what to copy out of it, file by file.- Text-only LLM → reference
qwen3/(orllama/for the minimal example) - MoE → reference
qwen3_moe/(plus converter work in Phase 3) - VLM (non-MoE) → reference
qwen3_vl/ - VLM + MoE → reference
qwen3_vl_moe/(multimodal forward + SP scatter, ViT dummy forward, Flash-attn kwargs popping,get_position_id_func) - Omni (non-MoE thinker + speech subtree to exclude) → reference
qwen2_5_omni/(audio/vision SP + dummy_forward, talker/token2wav/BigVGAN exclusion,log_probs/entropyoutput dataclass, no parallel_plan/converter) - Omni MoE → reference
qwen3_omni_moe/
- Text-only LLM → reference
- Check upstream source (
from transformers.models.<m> import modeling_<m>). Confirm class/function names still exist; MoE expert layouts especially diverge between sibling models — seedocs/transformers_v5/transformers_v5_moe_weight_loading.md. - Note related configs/loaders to preserve:
MODELING_REGISTRY,MODEL_CONFIG_REGISTRYinveomni/models/loader.py; any auto-config registrations. - Look for a sibling model you can borrow patches from: e.g. qwen3_5_moe
reuses GatedDeltaNet/ViT patches from
qwen3_5via direct import +name_map={"Qwen3_5": "Qwen3_5Moe"}. Prefer reuse over copy-paste when the upstream classes are structural duplicates with only a name-prefix difference. - Compare upstream and VeOmni parameter keys, including constructor overrides and nested modules. For any mismatch, follow the user-decision rule in veomni-new-model before choosing a model rename or checkpoint conversion. This also applies to refreshes and dependency upgrades. A resolution already authorized in the current task does not require another confirmation.
Validation: you have a concrete list of patches to apply, the reference model directory to mirror, and the backend/category decision pinned down.
Phase 2: Draft <M>_gpu_patch_gen_config.py
Create veomni/models/transformers/<M>/<M>_gpu_patch_gen_config.py at the model root.
Skeleton (mirror qwen3_gpu_patch_gen_config.py):
from veomni.patchgen.patch_spec import PatchConfig, create_patch_from_external
config = PatchConfig(
source_module="transformers.models.<m>.modeling_<m>",
target_file="patched_modeling_<m>_gpu.py",
description="<M> with LigerKernel GPU replacements + VeOmni SP/fused-loss patches",
)
Patch primitives:
| Effect | patchgen decorator / API |
|---|---|
| Replace whole class (RMSNorm, MLP, Experts) | @config.replace_class("<Class>") or create_patch_from_external(...) for liger |
| Replace module-level function (rotary, loss) | @config.replace_function("<name>") |
| Override a single method (Attention.forward, Model.forward, ForCausalLM.forward) | @config.override_method("<Class>.<method>") |
Add attribute / extra super().__init__() wiring | @config.modify_init("<Class>") |
| Reuse patch from a sibling config (name-prefix difference) | config.override_method("<NewClass>.<m>", replacement=<imported_fn>, name_map={"OldPrefix": "NewPrefix"}) — non-decorator form. Caveat: name_map only rewrites symbol names at the AST level; it does NOT align field sets between sibling output dataclasses (e.g. dense ModelOutputWithPast vs MoE ModelOutputWithPast with extra router_logits). Any <OldClass>Output(...) constructor call in the body gets its name rewritten but keeps the original arg list, silently dropping MoE-only fields. Clone the body when return dataclasses differ. |
| Supporting import needed in generated file | config.add_import("<module>", names=[...]) (or alias=..., is_from_import=False) |
| Remove an upstream import the generated file should NOT keep | config.drop_import_names("<symbol>", ...) |
| Inject raw code (try/except import fallback, helper fn used by patched code) near top of generated file | config.add_post_import_block("""...""") |
| Remove unused class from output | config.exclude_from_output("<Class>") |
| Inherit an entire sibling GPU config into an NPU config (reuse helpers / imports / post-import blocks; only override device-specific kernels) | config.helpers.extend(gpu_config.helpers) + config.post_import_blocks.extend(gpu_config.post_import_blocks) + config.additional_imports.extend(gpu_config.additional_imports) + import each <fn>_patched and re-register via config.override_method(...). See qwen3_vl_npu_patch_gen_config.py |
Cross-config reuse pattern (qwen3_5_moe reusing qwen3_5):
from veomni.models.transformers.qwen3_5.qwen3_5_gpu_patch_gen_config import (
qwen3_5_gated_deltanet_forward_patched,
qwen3_5_vision_model_forward,
# ...
)
_NAME_MAP = {"Qwen3_5": "Qwen3_5Moe"}
config.override_method(
"Qwen3_5MoeGatedDeltaNet.forward",
replacement=qwen3_5_gated_deltanet_forward_patched,
name_map=_NAME_MAP,
description="...",
)
name_map rewrites symbol references inside the replacement body so the shared
function transparently targets the correct class namespace. Use it to avoid
duplicating ~hundreds of lines per sibling model.
Common v5 patch set (steal from qwen3):
create_patch_from_external→LigerRMSNormreplacing<M>RMSNorm(for models with a "1 + weight" centered RMSNorm formulation — e.g. Qwen3Next variants — useLigerRMSNormForQwen3Nextinstead; check the upstream RMSNorm definition).create_patch_from_external→LigerSwiGLUMLPreplacing<M>MLP.@config.replace_function("apply_rotary_pos_emb")→liger_rotary_pos_emb. Exception: do NOT replace rotary when the model uses partial rotary (partial_rotary_factor < 1.0) ormrope_interleaved=True— liger applies RoPE to the full head_dim and produces NaN. Qwen3_5Moe explicitly skips this; leave an inline comment in the patchgen config when you do.@config.override_method("<M>Model.forward")→ keep SP-friendly shape handling.@config.override_method("<M>ForCausalLM.forward")(orForConditionalGeneration.forwardfor VLM) → fused cross-entropy path viaself.loss_function(logits=logits, labels=labels, vocab_size=..., hidden_states=..., weights=self.lm_head.weight, **kwargs). Note VLM top-level models useconfig.text_config.vocab_size, notconfig.vocab_size.- DecoderLayer varlen metadata — if the model has linear-attention / Mamba /
GatedDeltaNet layers, override
<M>DecoderLayer.forwardto passcu_seq_lens_qthrough (see qwen3_5_moe), and import cu-free FLA impls viaadd_post_import_blockwith a try/except fallback.
MoE models add three more patches here — expert replacement, _moe_implementation
propagation and the expert parallel plan. See references/moe.md, "Phase 2 additions".
VLM / Omni models add the SP-aware multimodal forward and the metadata
precompute contract, and Omni models also prune the speech subtree. See
references/multimodal.md.
Flash attention: VeOmni custom names
(veomni_flash_attention_{2,3,4}_with_sp) are handled globally by
transformers.integrations.hub_kernels.load_and_register_attn_kernel adapter —
no per-model patching needed. Just keep attn_implementation names unchanged
in configs. See
docs/transformers_v5/veomni_flash_attention_kernel_adapter.md.
Patch comment style:
Every decorated patch function / replaced class must be preceded by a
numbered header block enumerating what changed and why, and every modified
region inside the body must be bracketed by inline # --- Patch.N ---
markers that correspond to the header numbers. The comments survive into the
generated patched_modeling_*.py, giving reviewers a self-documenting diff
against the upstream HF source.
# Patch: <Class>.<method>
# 1. <what changed> — <why>
# 2. <next change> — <why>
@config.override_method("<Class>.<method>", description="...")
def <name>_patched(self, ...):
...
# --- Patch.1 ---
<modified region>
# --- Patch.1 ---
...
# --- Patch.2 ---
<other modified region>
# --- Patch.2 ---
Guidelines:
- Header numbering is local to the function; reuse the same number for all inline markers that belong to the same logical change.
- For removed/replaced upstream lines, keep the original as a commented
line inside the
# --- Patch.N ---block (seeqwen2_5_vl_gpu_patch_gen_config.py's vision-attentionmax_seqlenpatch) so the diff against HF is self-documenting. - Mention upstream-contract subtleties explicitly (e.g.
BaseModelOutputWithPoolingreturn type,pooler_outputtuple-of-tensors) — these are the most common source of regressions when HF bumps minor versions.
Regen command (put at top of file as docstring, mirror qwen3):
patchgen \
veomni.models.transformers.<m>.<m>_gpu_patch_gen_config \
-o veomni/models/transformers/<m>/generated --diff
Validation: file is syntactically valid (import it: python -c "import veomni.models.transformers.<m>.<m>_gpu_patch_gen_config") and every behaviour
identified in Phase 1 has a corresponding decorator here.
Phase 3: MoE Checkpoint Tensor Converter (MoE models only)
Skip this phase entirely for dense models.
Verify the HF checkpoint layout empirically for every MoE model. Add a runtime
converter only when that layout differs from the v5 fused-expert layout;
direct v5-compatible checkpoints need no converter. The verification procedure,
converter templates, and round-trip safety checks are in references/moe.md,
"Phase 3: checkpoint tensor converter".
Phase 4: Wire __init__.py
Pick one of three patterns based on Phase 1's backend + capability decision.
Pattern A — text LLM / dense (qwen3 style):
from ...loader import MODELING_REGISTRY
@MODELING_REGISTRY.register("<m>")
def register_<m>_modeling(architecture: str):
from .generated.patched_modeling_<m>_gpu import (
<M>ForCausalLM,
<M>Model,
)
if "ForCausalLM" in architecture:
return <M>ForCausalLM
return <M>Model
Pattern B — MoE with a converter selected in Phase 3 (qwen3_moe style): same as A, plus register the converter on each generated model class. Use Pattern A when the verified checkpoint layout needs no conversion:
from .checkpoint_tensor_converter import create_<m>_checkpoint_tensor_converter
for model_cls in (<M>ForCausalLM, <M>Model, ...):
model_cls._create_checkpoint_tensor_converter = staticmethod(
create_<m>_checkpoint_tensor_converter
)
staticmethod(...) is required — the loader calls it as
model._create_checkpoint_tensor_converter(model).
Pattern C — GPU + NPU sibling (glm_moe_dsa / qwen3_vl style): branch on
IS_NPU_AVAILABLE between the two generated modules:
from ....utils.device import IS_NPU_AVAILABLE
from ...loader import MODELING_REGISTRY
@MODELING_REGISTRY.register("<m>")
def register_<m>_modeling(architecture: str):
if IS_NPU_AVAILABLE:
from .generated.patched_modeling_<m>_npu import <M>ForCausalLM, <M>Model
else:
from .generated.patched_modeling_<m>_gpu import <M>ForCausalLM, <M>Model
if "ForCausalLM" in architecture:
return <M>ForCausalLM
return <M>Model
Rules:
- All logic lives in the patchgen config + generated file. Do not create
hand-written
modeling_<m>.py/gpu_patch.py/npu_patch.py— those files have been retired across the codebase. - For NPU (Pattern C): write a separate
<m>_npu_patch_gen_config.py— do not toggle GPU vs NPU kernels inside a single config via runtimeifs.
Phase 5: Run Patchgen + Verify Diff
- Regenerate.
make patchgen(patchgen --all --diff) rebuilds every model's generated file, which is the safe default because it cannot leave a GPU/NPU sibling behind. Target a single module only when you want a fast loop:patchgen \ veomni.models.transformers.<m>.<m>_gpu_patch_gen_config \ -o veomni/models/transformers/<m>/generated --diff -v - Inspect
generated/patched_modeling_<m>_gpu.py:- Header lists every patch you defined under "Patches applied".
- Patched classes/methods carry the
# [PATCHED ...]markers. - Relative imports (
from ...activations) rewritten to absolute (from transformers.activations).
- Inspect
generated/patched_modeling_<m>_gpu.diff— every hunk must correspond to an intentional patch. Unexpected hunks (e.g. whitespace, unrelated classes) indicate a misconfigured patchgen config. make quality/ruff formaton the generated file (patchgen pipeline runs ruff, but double-check).- Check CI drift guard:
Must exit 0.patchgen --check--fixoverwrites checked-in files if drift is intentional. - If
make style/ruff --fixauto-removed unused imports from the generated*.py(this happens when patchgen pulls an import from HF source that the patched version doesn't use, e.g.torch_compilable_checkin transformers v5.2), the sibling*.difffile becomes stale against the post-fix*.py. Re-sync with:
Do NOT manually re-runpatchgen --check --fixpatchgen(without--check) to "fix" it — that would re-introduce the unused imports and you'd ping-pong between ruff and patchgen.patchgen --check --fixwrites the diff against the post-style-fix.py, which is what CI expects.
Never edit generated/*.py by hand — always go back to the patchgen config
and regenerate. This is a hard rule called out in AGENTS.md.
Phase 6: Add Test Cases
Follow docs/transformers_v5/testing_new_model.md. Every file below is already
enumerated in a CI workflow — the unit-test ones, or gpu_e2e_test.yml /
npu_e2e_test.yml for the e2e tables — so appending a case needs no workflow
change. That is exactly why this phase extends tables instead of adding files.
tests/models/test_model_registry.py and
tests/models/test_models_logits_equal_v5.py are part of the minimum too:
the first proves the registry returns the generated class, the second that it
is numerically equal to upstream.
If you think you need a new test file, read .agents/knowledge/testing.md first.
Minimum coverage:
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 2k
- Forks
- 272
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
veomni-patchgen-model- Source
- github.com/bytedance-seed/veomni