quark-torch-file2file-quantization

SkillDocs & knowledge

Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. Use when the user wants to run file2file quantization, adapt a new safetensors checkpoint without loading the full model, register an external LLMTemplate, inspect sharded checkpoint naming, generate wrapper or conversion scripts, or validate low-memory sharded quantization outputs. Trigger for "run file2file quantization", "quantize without loading the model", "large safetensors low-memory quantization", "file2file for DeepSeek/Qwen/MoE", "safetensors naming incompatible", "register LLMTemplate externally".

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the quark-torch-file2file-quantization skill

What this skill tells your AI

The instructions your AI receives, as published by amd/quark in .claude/skills/quark-torch-file2file-quantization/SKILL.md and read by ahel’s review.

Read and follow the instructions in .claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization/SKILL.md.

Signals

GitHub stars
166
Forks
33
Last commit
Aug 2026
Advanced
Catalog kind
skill
Gateway key
quark-torch-file2file-quantization
Source
github.com/amd/quark