Quark Torch Quant-Perf
SkillSearchLets your agent run and manage model quantization and speed-testing workflows for PyTorch and HuggingFace models.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Quark Torch Quant-Perf skill
About this skill
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. Use whenever the user asks for mixed-precision search, Quark MXFP4/FP8/PTPC-FP8 quantization with accuracy and performance validation, vLLM throughput, TraceLens bottleneck analysis
What this skill tells your AI
The instructions your AI receives, as published by amd/quark in skills/quark-torch-quant-perf/SKILL.md and read by ahel’s review.
Purpose
Translate a user's PyTorch or HuggingFace quantization request into the current
quark-quant-perf CLI, run or resume its managed Orchestrator, monitor the
session to a terminal state, and report the generated artifacts. Preserve the
user's accuracy, workload, source, and explicit performance requirements
throughout.
Use the managed pipeline:
input and runtime validation
→ baseline health
→ quantize or mixed-precision search
→ real quantized accuracy gate
→ [measure or optimize] baseline and quantized throughput
→ [optimize, when the target is missed] conditional TraceLens / vendor tuning /
GEAK → candidate validation and retention
→ FINAL reports
Performance modes skip inapplicable optional stages; they do not create
separate FINAL stages. The Orchestrator enters FINAL after the selected stages
complete or a terminal failure is recorded. The report subcommand may later
regenerate terminal artifacts without rerunning GPU work.
Do not replace a stage with a custom evaluation, benchmark, or proxy metric.
Inputs
- Model path or HuggingFace model ID.
- Accuracy-gap requirement and optional performance intent.
- GPU type, GPU index, TP, ISL, OSL, and concurrency when specified.
- Fixed quantization intent, or a mixed-precision search space.
- Framework and kernel source policy: explicit, auto, or readonly.
- A discoverable Quark checkout containing
quark-torch-ptqfor fixed direct PTQ;QUARK_ROOTmay select one explicitly. - Existing session directory when resuming.
- Optional TraceLens architecture JSON and backend overrides.
Only --model is required by the CLI. Preserve CLI defaults for omitted
options rather than inventing values.
Before constructing a command:
- Run
quark-quant-perf --helpfor the installed CLI options and defaults. - Read CLI mapping for intent-to-option mapping.
- Read session lifecycle before resume, monitoring, diagnosis, or final reporting.
- Inspect the session's
state.jsonandprogress.jsonwhen resuming or reporting.
Treat the installed CLI help as authoritative if a local reference differs.
Outputs: session_report.md
The managed FINAL stage produces:
<session>/reports/final.json
<session>/reports/final.md
<session>/session_breakdown.json
<session>/session_report.md
Treat session_breakdown.json as the complete structured fact source and
session_report.md as the detailed human-readable result. A terminal failure
may still have complete reports and useful quantized artifacts.
Interaction Flow
-
Intake
- Resolve local paths to absolute paths.
- Preserve the requested model, accuracy gap, performance mode, target gain, workload, GPU, tensor parallelism, search modes, and repository policy.
- Treat explicitly named quantized precision modes as a closed set, separate
from any runtime-required
nativefallback. Do not infer compound modes from their components. Normalize aliases one-to-one; for example, usemxfp4_fp8only when the user explicitly requests that mode, W4A8, or MXFP4 weights with FP8 activations. - If the user does not name precision modes, omit
--layer-precision-candidatesand preserve the GPU-specific automatic search space. - If the user does not request throughput or performance optimization,
leave performance at the default
off. - Inspect existing session state before creating a duplicate run.
-
Route
- Omit
--quant-strategyfor mixed-precision search. - Use
--quant-strategy "<normalized intent>"for one fixed PTQ recipe. - For fixed PTQ, use the automatically discovered Quark checkout when
available. Set
QUARK_ROOTonly when discovery fails or the user selects a specific checkout. - When writable framework or kernel repositories are supplied, use
--workspace-source explicitwith those repositories. - When repositories are not supplied, use
--workspace-source autoso runtime repair and PerfOpt can discover or materialize writable sources. - Use
--workspace-source readonlyonly when the user explicitly requests evaluation or inspection without source modification. - Treat fresh or no-history execution as an experience-store policy, not a
source policy. It never implies
readonly; keepautounless the user separately prohibits source modification. - Before launch, if the command contains
--workspace-source readonly, identify the user's explicit no-source-modification request. If there is none, remove the option and use the defaultauto. - Apply the source policy to every run. It governs runtime activation and
eligible repair work even when performance mode is
offormeasure; PerfOpt source modification remains specific tooptimize.
- Omit
-
Plan
- Maintain the native task plan for every multi-stage run. Use
update_planin Codex andTaskCreate/TaskUpdatein Claude Code. - Before launch, show every applicable top-level task derived from the
resolved run intent. Always include validation, baseline health,
quantization or search, the real accuracy gate, and FINAL reporting. Add
throughput for
measureoroptimize, and add the PerfOpt decision plus conditional bottleneck, optimization, and retention tasks foroptimize. Do not use one universal task list for every performance mode. - Keep exactly one plan step in progress.
- Mark an applicable conditional task as completed with the skip reason when
the Orchestrator does not need it. Add candidate, retry, repair, and kernel
subtasks dynamically when command output,
state.json,progress.json, or an artifact shows that they exist. - On resume or context compaction, reconstruct the complete run-specific plan from the persisted run specification and current session evidence rather than conversation memory.
- Maintain the native task plan for every multi-stage run. Use
-
Execute
- Launch through
quark-quant-perf; do not call internal stage functions as a replacement pipeline. - Keep long runs in a managed terminal or tool session that can be polled;
do not detach them with
nohupor shell&. - Keep the exact command, session directory, and process handle visible.
- Do not pause, signal, or terminate an active process without explicit user approval unless immediate system safety or data loss is at risk.
- Launch through
-
Monitor and summarize
- Poll process output while the run is active. When no new output arrives,
inspect
state.jsonandprogress.jsonabout every 30 seconds. Do not redraw an unchanged plan, but always refresh it before a status response or wait. - Use
progress.jsonfor the currentstageandstage_detail, and usestate.jsonfor completed, failed, resumed, and terminal facts. - During mixed-precision search, show the persisted
total_configs_evaluated/total_configs_availablecounts. Also account forcandidate_cursor,candidate_queue,partial_timeout, andtermination_reasonwhen describing export fallback or a salvaged search. - Keep all applicable top-level tasks visible throughout the run. Enrich the active task and add retry, repair, candidate, and kernel subtasks when their corresponding attempts, measurements, journeys, or artifacts appear.
- Use terminal output to enrich the current task, not as the sole evidence for completion.
- Distinguish search accuracy from the authoritative real accuracy gate.
- Distinguish local kernel speedup from final retained end-to-end gain.
- Continue through FINAL even when accuracy or performance targets fail.
- Report the terminal status, selected quantization, accuracy, throughput, retained patches, rejected attempts, and artifact paths.
- Poll process output while the run is active. When no new output arrives,
inspect
Command Rules
Use current canonical option names:
--max-search-candidates
--search-timeout
--layer-precision-candidates
--kv-cache-precision-candidates
--search-gsm8k-num-samples
--geak-direction-budget
Do not emit removed historical names:
--max-rounds
--layer-mode
--kv-cache-mode
--geak-budget
Convert percentages to ratios:
- 3% maximum accuracy drop →
--accuracy-gap 0.03 - 35% target speedup →
--target-gain 1.35 - 2x throughput target →
--target-gain 2.0
Performance intent:
- No performance request → omit both options; the effective mode is
off. - Throughput measurement only →
--performance-mode measure. - Optimization target →
--target-gain MULTIPLIER; this implies--performance-mode optimize. - Explicit optimize without a target uses the CLI default target of
1.2.
Precision and backend intent:
- No requested layer modes → omit
--layer-precision-candidatesand use the GPU-specific defaults. - Explicit layer modes → pass only the named quantized modes;
nativeremains an implicit fallback. - No requested backend → omit the MXFP4/W4A8 backend flags and preserve their CLI or environment defaults.
- Explicit backend → pass only the corresponding backend option; backend selection does not add a precision mode to the search space.
Use --tracelens-gpu-arch-json when the installed TraceLens package lacks the
requested GPU architecture data. Do not copy architecture data into
site-packages.
Use --retry-accuracy-gate when the real accuracy gate must be reopened. Use
--retry-perfopt only when reusable accuracy and quant-only throughput
evidence still match the current runtime fingerprint.
Monitoring Rules
- A status request, parameter correction, turn interruption, or new message is not permission to stop an active task.
- Do not infer current state from old warnings in logs.
- Use
state.json, terminal command output, and FINAL reports for conclusions. - A local GEAK result is not a final performance result.
- A patch is retained only after the managed correctness, accuracy, and end-to-end performance gates accept it.
- Preserve unrelated user changes and dirty worktrees.
Recovery
- Reuse the same session when the user asks to continue or resume.
- Diagnose the failing boundary before changing code or configuration.
- Use
--recheck-baselineonly to bypass an exact cached baseline failure. - Use
--retry-accuracy-gateafter a failed accuracy stage or after a runtime change invalidates the saved accuracy fingerprint. - Use
--retry-perfoptafterperf_failedonly when upstream evidence remains reusable. - Do not silently change TP, model, ISL, OSL, concurrency, quantization modes, KV-cache modes, or benchmark implementation during recovery.
- If source paths or runtime identity change, expect fingerprints to invalidate cached measurements and rerun the required managed gates.
- Preserve failure logs and candidate evidence even when the candidate source change is reverted.
Signals
- GitHub stars
- 174
- Forks
- 35
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
quark-torch-quant-perf- Source
- github.com/amd/quark