quark-torch-llm-eval

SkillCloud & infra

End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — docker container setup OR host (no-docker) runtime, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. Use when the user wants to evaluate, benchmark, or compare an LLM's accuracy. Trigger for "evaluate this model", "run gsm8k/mmlu/mmlu_pro/aime/gpqa/hellaswag/arc", "test accuracy", "measure perplexity", "compare quantized model accuracy", "does this mxfp4 model lose accuracy". For evaluating Quark Agent Skills themselves, use quark-torch-eval-runner instead.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the quark-torch-llm-eval skill

What this skill tells your AI

The instructions your AI receives, as published by amd/quark in .claude/skills/quark-torch-llm-eval/SKILL.md and read by ahel’s review.

Read and follow the instructions in .claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval/SKILL.md.

Signals

GitHub stars
166
Forks
33
Last commit
Aug 2026
Advanced
Catalog kind
skill
Gateway key
quark-torch-llm-eval
Source
github.com/amd/quark