xpu-discover

SkillDev tools

Inventory Intel GPUs (Arc, Arc Pro, Data Center GPU Max) on a Linux host. Detect devices, check driver health, list processes using each XPU, run a quick diagnostic, and read live utilisation.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the xpu-discover skill

What this skill tells your AI

The instructions your AI receives, as published by intel/skills in skills/xpu-discover/SKILL.md and read by ahel’s review.

xpu-smi is Intel's nvidia-smi. Sees only Intel GPUs (Arc, Arc Pro, Battlemage, Flex, Max).

Quickstart

Run in order. If step 1 is empty, stop.

xpu-smi discovery                 # 1. inventory
xpu-smi diag --precheck           # 2. driver/firmware health
xpu-smi ps                        # 3. processes using each GPU
xpu-smi diag -d 0 -l 1            # 4. quick functional test
xpu-smi stats -d 0                # 5. utilisation snapshot
xpu-smi dump -d 0 -m 0,5,18 -i 1  # 6. live CSV stream (Ctrl-C)

All commands accept -j for JSON output (use when parsing).

CUDA -> Intel cheat sheet

CUDAIntel
nvidia-smixpu-smi discovery
nvidia-smi -Lxpu-smi discovery -j
nvidia-smi pmon -c 1xpu-smi ps
nvidia-smi dmonxpu-smi dump -d <id> -m 0,5,18 -i 1
nvidia-smi --query-gpu=...xpu-smi stats -d <id> -j
nvidia-smi topo -mxpu-smi topology -m
CUDA_VISIBLE_DEVICES=0ZE_AFFINITY_MASK=0
cuda-memcheckxpu-smi diag -d 0 -l 1

CUDA refugee footgun: CUDA_VISIBLE_DEVICES=99 silently hides all GPUs; ZE_AFFINITY_MASK=99 crashes the Level Zero loader with an assertion. Always check xpu-smi discovery for valid IDs (start at 0) before setting the mask.

What each subcommand returns

discovery — inventory

One stanza per Intel GPU. Key fields:

  • Device ID — small integer, used as -d and as ZE_AFFINITY_MASK value.

  • PCI BDF Address — stable across reboots (e.g. 0000:36:00.0).

  • DRM Device — /dev/dri/card0, used in --device for Docker.

  • Device Name — Battlemage shows Intel(R) Graphics [0xe2XX] rather than the marketing name; driver quirk, not a problem. Map the PCI device ID in brackets to the product SKU:

    PCI device IDProduct SKUConfirmed
    0xe20bArc B580yes (lspci on hardware)
    0xe211Arc Pro B60yes (pci.ids)
    0xe220Arc Pro B50yes (pci.ids)
    0xe221Arc Pro B65yes (pci.ids)
    0xe223Arc Pro B70yes (lspci on hardware)

    Full table provided above. Cross-check with lspci -d 8086: -nn (prints [8086:XXXX]).

Empty output -> kernel didn't enumerate any Intel GPU. See "Troubleshooting".

diag --precheck — driver health

Scans journalctl for known Intel-GPU error categories (GuC/HuC firmware, IOMMU, PCIe, DRM, i915/Xe, Level Zero init).

  • All Pass -> proceed.
  • Any Critical -> see "Error pattern routing" below.

Scope with --since today / --since yesterday / --listtypes (show every error category).

ps — what's using each GPU

Lists processes holding Level Zero handles + shared/device memory in MiB. Desktop processes (plasmashell, xauth_*) are normal on a workstation; only worry about a stale model server still holding memory.

diag -d <id> -l <level> — does it compute

LevelTimeImpact
-l 1secondssafe on a busy box
-l 2mediumimpacts other workloads
-l 3minutesimpacts performance

Reading on a workstation: the Software Permission sub-test fails when other processes already hold the device, so the overall line says Fail even when Computation Check: Pass. Read per-sub-test rows. For a clean run, log out of the GUI session and run from TTY, or run inside a privileged container.

Pick individual tests with --singletest:

IDTestWhen
1Computationdoes it compute
2Memory Errorsuspected ECC / bit-flip
3Memory Bandwidthsanity-check HBM/GDDR
4Media Codecvideo pipelines
5PCIe Bandwidthsuspected slot/cable issue
6Powerthermal/TDP investigation
7Computation functionalquick sanity (lighter than 1)
8Media Codec functionallighter media check
9Xe Link Throughputmulti-GPU peer link
10Xe Link all-to-allmulti-GPU; needs -d -1

Example: xpu-smi diag -d 0 --singletest 1,3 -j.

stats -d <id> — utilisation snapshot

Many fields show N/A on consumer Battlemage drivers (Arc Pro B70 included) — counter-wiring limitation, not a bug. Memory (column 5) and power (column 18) generally work. For utilisation while running, prefer xpu-smi dump.

dump -d <id> -m <metrics> -i <interval> — live CSV stream

Useful metric IDs for LLM serving:

  • 0 GPU utilisation (%)
  • 5 GPU memory used (MiB)
  • 18 GPU power (W)
xpu-smi dump -d 0 -m 0,5,18 -i 1 > xpu.csv &
# ... run your model ...
kill %1

Privilege note: metric 0 reads MEI telemetry, restricted on consumer parts. Without sudo, that column shows N/A. Metrics 5 and 18 work unprivileged.

topology -m — multi-GPU connectivity

Matrix of Xe Link / PCIe switch / hostbridge between XPU pairs. Only useful on multi-XPU systems.

Error pattern routing

When diag --precheck flags a critical error:

CategoryCause + fix
Level Zero Init ErrorDriver/userspace mismatch. Confirm xe (Battlemage) or i915 (older) modules loaded: lsmod | grep -E 'i915|xe'. Reload or reboot.
GuC / HuC Not RunningMissing firmware blob. Check dmesg | grep -i 'GuC|HuC'; install linux-firmware.
IOMMU CatastrophicKernel cmdline. On consumer boards: intel_iommu=on iommu=pt.
PCIe ErrorReseat card / check slot; re-run xpu-smi diag -d <id> --singletest 5.
DRM ErrorStuck context from a crashed desktop session. Logout/login (or reboot) clears it.
i915 Not Loaded on BattlemageBattlemage uses xe, not i915. Confirm modinfo xe; precheck error is misleading on this generation.

Troubleshooting

SymptomFix
xpu-smi: command not foundInstall Intel level-zero packages (Ubuntu / RHEL / Arch all package xpu-smi). Binary lands at /usr/bin/xpu-smi.
discovery empty but card present(1) Kernel didn't bind: lspci -k -s <bdf> should show Kernel driver in use:. (2) Bound to vfio: lsmod | grep vfio. (3) Inside container without /dev/dri: add --device /dev/dri --privileged.
diag fails inside containerContainer needs --privileged for diag ioctls beyond the standard render-node interface.
Two XPUs present, one visibleprintenv ZE_AFFINITY_MASK; unset for full inventory, re-export for workloads.

Env vars

VariablePurpose
ZE_AFFINITY_MASKWhich XPU(s) a process sees (0, 0,1, ...). Invalid IDs crash the L0 loader (unlike CUDA's silent-hide).
ZE_FLAT_DEVICE_HIERARCHYFLAT exposes tiles as separate root devices; COMPOSITE (default) groups under one. Battlemage is single-tile, doesn't matter.
ZE_ENABLE_VALIDATION_LAYER=1L0 loader prints API misuse — useful when something silently returns wrong device count.

References

Signals

GitHub stars
21
Forks
9
Last commit
Sep 2026
Advanced
Item type
skill
Key
xpu-discover
Source
github.com/intel/skills