Quark Torch Quant-Perf

SkillSearch

Lets your agent run and manage model quantization and speed-testing workflows for PyTorch and HuggingFace models.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Quark Torch Quant-Perf skill

About this skill

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. Use whenever the user asks for mixed-precision search, Quark MXFP4/FP8/PTPC-FP8 quantization with accuracy and performance validation, vLLM throughput, TraceLens bottleneck analysis

What this skill tells your AI

The instructions your AI receives, as published by amd/quark in skills/quark-torch-quant-perf/SKILL.md and read by ahel’s review.

Purpose

Translate a user's PyTorch or HuggingFace quantization request into the current quark-quant-perf CLI, run or resume its managed Orchestrator, monitor the session to a terminal state, and report the generated artifacts. Preserve the user's accuracy, workload, source, and explicit performance requirements throughout.

Use the managed pipeline:

input and runtime validation
→ baseline health
→ quantize or mixed-precision search
→ real quantized accuracy gate
→ [measure or optimize] baseline and quantized throughput
→ [optimize, when the target is missed] conditional TraceLens / vendor tuning /
  GEAK → candidate validation and retention
→ FINAL reports

Performance modes skip inapplicable optional stages; they do not create separate FINAL stages. The Orchestrator enters FINAL after the selected stages complete or a terminal failure is recorded. The report subcommand may later regenerate terminal artifacts without rerunning GPU work.

Do not replace a stage with a custom evaluation, benchmark, or proxy metric.

Inputs

  • Model path or HuggingFace model ID.
  • Accuracy-gap requirement and optional performance intent.
  • GPU type, GPU index, TP, ISL, OSL, and concurrency when specified.
  • Fixed quantization intent, or a mixed-precision search space.
  • Framework and kernel source policy: explicit, auto, or readonly.
  • A discoverable Quark checkout containing quark-torch-ptq for fixed direct PTQ; QUARK_ROOT may select one explicitly.
  • Existing session directory when resuming.
  • Optional TraceLens architecture JSON and backend overrides.

Only --model is required by the CLI. Preserve CLI defaults for omitted options rather than inventing values.

Before constructing a command:

  1. Run quark-quant-perf --help for the installed CLI options and defaults.
  2. Read CLI mapping for intent-to-option mapping.
  3. Read session lifecycle before resume, monitoring, diagnosis, or final reporting.
  4. Inspect the session's state.json and progress.json when resuming or reporting.

Treat the installed CLI help as authoritative if a local reference differs.

Outputs: session_report.md

The managed FINAL stage produces:

<session>/reports/final.json
<session>/reports/final.md
<session>/session_breakdown.json
<session>/session_report.md

Treat session_breakdown.json as the complete structured fact source and session_report.md as the detailed human-readable result. A terminal failure may still have complete reports and useful quantized artifacts.

Interaction Flow

  1. Intake

    • Resolve local paths to absolute paths.
    • Preserve the requested model, accuracy gap, performance mode, target gain, workload, GPU, tensor parallelism, search modes, and repository policy.
    • Treat explicitly named quantized precision modes as a closed set, separate from any runtime-required native fallback. Do not infer compound modes from their components. Normalize aliases one-to-one; for example, use mxfp4_fp8 only when the user explicitly requests that mode, W4A8, or MXFP4 weights with FP8 activations.
    • If the user does not name precision modes, omit --layer-precision-candidates and preserve the GPU-specific automatic search space.
    • If the user does not request throughput or performance optimization, leave performance at the default off.
    • Inspect existing session state before creating a duplicate run.
  2. Route

    • Omit --quant-strategy for mixed-precision search.
    • Use --quant-strategy "<normalized intent>" for one fixed PTQ recipe.
    • For fixed PTQ, use the automatically discovered Quark checkout when available. Set QUARK_ROOT only when discovery fails or the user selects a specific checkout.
    • When writable framework or kernel repositories are supplied, use --workspace-source explicit with those repositories.
    • When repositories are not supplied, use --workspace-source auto so runtime repair and PerfOpt can discover or materialize writable sources.
    • Use --workspace-source readonly only when the user explicitly requests evaluation or inspection without source modification.
    • Treat fresh or no-history execution as an experience-store policy, not a source policy. It never implies readonly; keep auto unless the user separately prohibits source modification.
    • Before launch, if the command contains --workspace-source readonly, identify the user's explicit no-source-modification request. If there is none, remove the option and use the default auto.
    • Apply the source policy to every run. It governs runtime activation and eligible repair work even when performance mode is off or measure; PerfOpt source modification remains specific to optimize.
  3. Plan

    • Maintain the native task plan for every multi-stage run. Use update_plan in Codex and TaskCreate / TaskUpdate in Claude Code.
    • Before launch, show every applicable top-level task derived from the resolved run intent. Always include validation, baseline health, quantization or search, the real accuracy gate, and FINAL reporting. Add throughput for measure or optimize, and add the PerfOpt decision plus conditional bottleneck, optimization, and retention tasks for optimize. Do not use one universal task list for every performance mode.
    • Keep exactly one plan step in progress.
    • Mark an applicable conditional task as completed with the skip reason when the Orchestrator does not need it. Add candidate, retry, repair, and kernel subtasks dynamically when command output, state.json, progress.json, or an artifact shows that they exist.
    • On resume or context compaction, reconstruct the complete run-specific plan from the persisted run specification and current session evidence rather than conversation memory.
  4. Execute

    • Launch through quark-quant-perf; do not call internal stage functions as a replacement pipeline.
    • Keep long runs in a managed terminal or tool session that can be polled; do not detach them with nohup or shell &.
    • Keep the exact command, session directory, and process handle visible.
    • Do not pause, signal, or terminate an active process without explicit user approval unless immediate system safety or data loss is at risk.
  5. Monitor and summarize

    • Poll process output while the run is active. When no new output arrives, inspect state.json and progress.json about every 30 seconds. Do not redraw an unchanged plan, but always refresh it before a status response or wait.
    • Use progress.json for the current stage and stage_detail, and use state.json for completed, failed, resumed, and terminal facts.
    • During mixed-precision search, show the persisted total_configs_evaluated / total_configs_available counts. Also account for candidate_cursor, candidate_queue, partial_timeout, and termination_reason when describing export fallback or a salvaged search.
    • Keep all applicable top-level tasks visible throughout the run. Enrich the active task and add retry, repair, candidate, and kernel subtasks when their corresponding attempts, measurements, journeys, or artifacts appear.
    • Use terminal output to enrich the current task, not as the sole evidence for completion.
    • Distinguish search accuracy from the authoritative real accuracy gate.
    • Distinguish local kernel speedup from final retained end-to-end gain.
    • Continue through FINAL even when accuracy or performance targets fail.
    • Report the terminal status, selected quantization, accuracy, throughput, retained patches, rejected attempts, and artifact paths.

Command Rules

Use current canonical option names:

--max-search-candidates
--search-timeout
--layer-precision-candidates
--kv-cache-precision-candidates
--search-gsm8k-num-samples
--geak-direction-budget

Do not emit removed historical names:

--max-rounds
--layer-mode
--kv-cache-mode
--geak-budget

Convert percentages to ratios:

  • 3% maximum accuracy drop → --accuracy-gap 0.03
  • 35% target speedup → --target-gain 1.35
  • 2x throughput target → --target-gain 2.0

Performance intent:

  • No performance request → omit both options; the effective mode is off.
  • Throughput measurement only → --performance-mode measure.
  • Optimization target → --target-gain MULTIPLIER; this implies --performance-mode optimize.
  • Explicit optimize without a target uses the CLI default target of 1.2.

Precision and backend intent:

  • No requested layer modes → omit --layer-precision-candidates and use the GPU-specific defaults.
  • Explicit layer modes → pass only the named quantized modes; native remains an implicit fallback.
  • No requested backend → omit the MXFP4/W4A8 backend flags and preserve their CLI or environment defaults.
  • Explicit backend → pass only the corresponding backend option; backend selection does not add a precision mode to the search space.

Use --tracelens-gpu-arch-json when the installed TraceLens package lacks the requested GPU architecture data. Do not copy architecture data into site-packages.

Use --retry-accuracy-gate when the real accuracy gate must be reopened. Use --retry-perfopt only when reusable accuracy and quant-only throughput evidence still match the current runtime fingerprint.

Monitoring Rules

  • A status request, parameter correction, turn interruption, or new message is not permission to stop an active task.
  • Do not infer current state from old warnings in logs.
  • Use state.json, terminal command output, and FINAL reports for conclusions.
  • A local GEAK result is not a final performance result.
  • A patch is retained only after the managed correctness, accuracy, and end-to-end performance gates accept it.
  • Preserve unrelated user changes and dirty worktrees.

Recovery

  • Reuse the same session when the user asks to continue or resume.
  • Diagnose the failing boundary before changing code or configuration.
  • Use --recheck-baseline only to bypass an exact cached baseline failure.
  • Use --retry-accuracy-gate after a failed accuracy stage or after a runtime change invalidates the saved accuracy fingerprint.
  • Use --retry-perfopt after perf_failed only when upstream evidence remains reusable.
  • Do not silently change TP, model, ISL, OSL, concurrency, quantization modes, KV-cache modes, or benchmark implementation during recovery.
  • If source paths or runtime identity change, expect fingerprints to invalidate cached measurements and rerun the required managed gates.
  • Preserve failure logs and candidate evidence even when the candidate source change is reverted.

Signals

GitHub stars
174
Forks
35
Last commit
Sep 2026
Advanced
Item type
skill
Key
quark-torch-quant-perf
Source
github.com/amd/quark