xpu-container-run

SkillCloud & infra

Launch a Docker container with Intel GPU access on Linux. Encodes the correct combination of `--device /dev/dri`, render-group access, `--ipc=host`, `ZE_AFFINITY_MASK` pinning, Hugging Face cache mount, and `--entrypoint /bin/bash` for interactive use. Use when running any Intel-XPU container (vLLM-XPU, sglang-xpu, torch-XPU, llama.cpp SYCL, etc.) and the device must be visible inside. The CUDA analogue is `docker run --gpus all`, Intel has no `--gpus` flag, you pass the Direct Rendering Manager (DRM) nodes directly.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the xpu-container-run skill

What this skill tells your AI

The instructions your AI receives, as published by intel/skills in skills/xpu-container-run/SKILL.md and read by ahel’s review.

Intel GPUs do not plug into Docker via --gpus all. There is no nvidia-container-toolkit equivalent. Pass the kernel's Direct Rendering Manager (DRM) character devices into the container and grant the right group ownership.

CUDA → Intel cheat sheet

CUDAIntel
docker run --gpus all--device /dev/dri --group-add "$(getent group render | cut -d: -f3)"
docker run --gpus '"device=0"'-e ZE_AFFINITY_MASK=0
--ipc=hostsame
--shm-size=16gsame (alternative to --ipc=host)
--runtime nvidianothing — xe/i915 is in-kernel
nvidia-smi inside containerxpu-smi discovery

No "Intel container toolkit" needed; passing the DRM nodes is enough.

Image source

Comes from the runner skill:

  • vLLM serving → vllm-xpu-run (vllm/vllm-openai-xpu:latest)
  • SGLang → sglang-xpu-run (built from upstream docker/xpu.Dockerfile)
  • PyTorch / Transformers → torch-xpu-run

<image> below is whichever you picked.

Quickstart — interactive shell, one GPU

Confirm the image name with the user before running — this binds host GPU devices into the container.

docker run --rm -it \
    --device /dev/dri \
    --group-add "$(getent group render | cut -d: -f3)" \
    --ipc=host \
    -e ZE_AFFINITY_MASK=0 \
    -e HF_TOKEN="$HF_TOKEN" \
    -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
    --entrypoint /bin/bash \
    <image>
FlagWhy
--device /dev/driPass every Intel GPU's DRM nodes. Use --device /dev/dri/renderD128 for just the first GPU's render node (least privilege).
--group-add "$(getent group render | cut -d: -f3)"Joins the container user to the host's render group by GID (not name) so it works in images where a render group with a different GID — or no render group at all — exists. Required when nodes are mode 0660/0640. Skip causes EACCES on Level Zero init.
--ipc=hostvLLM and torch.distributed use /dev/shm and POSIX semaphores. --shm-size=16g is a private-IPC alternative.
-e ZE_AFFINITY_MASK=0Pin to GPU 0. See xpu-discover for IDs. Always set explicitly.
-v ~/.cache/huggingface:...Share the host model cache; avoid re-download.
--entrypoint /bin/bashOverride server-image autostart for interactive use.

When --privileged is needed

Exception, not rule. Required only for:

  • unitrace / VTune collectors that read PMU MSRs.
  • GPU firmware updates (xpu-smi updatefw).
  • xpu-smi diag --singletest 5 (PCIe bandwidth) and similar low-level diag tests.

For running models and most profiling, --device /dev/dri is enough. Add --privileged only when you hit a specific permission failure pointing at it.

Server mode, multi-GPU, --net=host

See references/server-and-multi-gpu.md for daemon-style server launches, one-process-per-GPU vs single-process TP / PP layouts, the oneCCL CCL_ZE_IPC_EXCHANGE=pidfd setting for multi-XPU TP, and when --net=host is actually needed.

Verifying the container sees the GPU

xpu-smi discovery
SymptomCauseFix
xpu-smi: command not foundimage lacks xpu-smiuse a different image or skip this check
empty discovery tableno /dev/dri passedadd --device /dev/dri
Level Zero init failed / EACCESuser not in render groupadd `--group-add "$(getent group render
wrong GPU countZE_AFFINITY_MASK inherited from hostpass mask explicitly with -e
diag works on host, fails in containercontainer not privilegedadd --privileged, or skip diag inside container

Common errors

  • failed to create shim task: permission denied → container runtime can't open /dev/dri/card0. Add --privileged or check host file mode.
  • LIBZE_LOADER: Failed to load level-zero loader → image missing libze1 / intel-level-zero-gpu. Use a different image.
  • RuntimeError: Cannot find any XPU devices (PyTorch) → Level Zero loaded but no device visible. Re-check ZE_AFFINITY_MASK and run xpu-smi discovery in the container.
  • bus error early in vLLM/PyTorch startup → shared memory too small. Use --ipc=host or raise --shm-size.

Env vars

VariablePurpose
ZE_AFFINITY_MASKWhich XPU(s) visible.
HF_TOKENHugging Face auth.
HF_HOMEOverride in-container HF cache path.
HUGGINGFACE_HUB_CACHEOlder alias; some images still use it.
OMP_NUM_THREADSCap CPU threads; 1 for single-process serving.
CCL_ZE_IPC_EXCHANGE=pidfdMulti-GPU-friendly oneCCL IPC mechanism.
IGC_EnableAluBinding=1Battlemage matmul-codegen hint; bench both.
ONEAPI_DEVICE_SELECTOR=level_zero:0Belt-and-suspenders pin alongside ZE_AFFINITY_MASK.

References

Signals

GitHub stars
21
Forks
9
Last commit
Sep 2026
Advanced
Item type
skill
Key
xpu-container-run
Source
github.com/intel/skills