vllm-xpu-run

SkillMedia

Serve a Hugging Face safetensors model on an Intel GPU with upstream vLLM-XPU's OpenAI-compatible API, or check whether a model or architecture is currently documented on XPU. Covers live support lookup, image choice, container launch, known serve-flag requirements, model-impl fallback, and attention/quant compatibility. Use to launch /v1/chat/completions or /v1/completions, troubleshoot a launch, or check model support. Not for choosing the best quantization, KV dtype, DP/TP layout, context, or concurrency (use model-config-recommend).

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the vllm-xpu-run skill

What this skill tells your AI

The instructions your AI receives, as published by intel/skills in skills/vllm-xpu-run/SKILL.md and read by ahel’s review.

Use the official upstream vllm/vllm-openai-xpu:latest image.

Upstream vLLM has a first-class XPU backend. The CLI is identical to the CUDA build (vllm serve <model>); device is detected from torch.xpu.is_available(). There is no --device xpu flag.

The official image already sets ENTRYPOINT ["vllm", "serve"]. Pass the model id and serve flags directly after the image name. Do not append another vllm serve: that produces vllm serve vllm serve <model> and the container exits with code 2.

Current supported models and architectures

When the user asks which models vLLM supports on Intel XPU, fetch the current upstream page at request time:

https://docs.vllm.ai/en/stable/models/hardware_supported_models/xpu/

Do not answer from memory and do not copy a static model list into this skill. Report both the explicitly listed Model rows and the Architecture column, because the recommended model table is not an exhaustive checkpoint allowlist.

For a specific unlisted Hugging Face model, read its current config.json and compare every value in architectures with the live page's Architecture column. Report the evidence precisely:

  • Exact model row → explicitly documented on the fetched page.
  • Architecture match only → the architecture is documented on XPU, but this exact checkpoint is not explicitly validated by the page; perform a generation smoke test before claiming full support.
  • Neither matches → not documented by the current XPU page; this is not proof of impossibility.

Include the source URL and retrieval date in the answer. If the page cannot be fetched, report that failure and offer to retry rather than substituting a remembered list. Do not infer XPU support merely from general vLLM, CUDA, or Transformers support.

CUDA → XPU cheat sheet

CUDA conventionXPU convention
vllm/vllm-openai:latestvllm/vllm-openai-xpu:latest
--gpus all--device /dev/dri + -v /dev/dri/by-path:/dev/dri/by-path:ro + --ipc=host
--dtype auto--dtype bfloat16 (explicit)
CUDA graphs default--enforce-eager
--tensor-parallel-size Nsame; pin N XPUs in ZE_AFFINITY_MASK

Quickstart (single GPU)

Confirm the target image tag and model with the user before running the docker run command — container launches bind host devices and download multi-GB weights.

Pull the official XPU image from https://hub.docker.com/r/vllm/vllm-openai-xpu. The examples use :latest; pin an immutable digest for reproducible deployments. Generated plans should call scripts/emit_launch.sh instead of copying the template manually — that keeps image policy, proxy env propagation, quantization flags, and multi-XPU topology in one place.

docker run -d --name vllm-xpu \
    --device /dev/dri \
    -v /dev/dri/by-path:/dev/dri/by-path:ro \
    --group-add "$(getent group render | cut -d: -f3)" \
    --ipc=host \
    -e ZE_AFFINITY_MASK=0 \
    -e VLLM_WORKER_MULTIPROC_METHOD=spawn \
    -e HTTP_PROXY -e HTTPS_PROXY -e NO_PROXY \
    -e http_proxy -e https_proxy -e no_proxy \
    -e HF_TOKEN="$HF_TOKEN" \
    -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
    -p 8000:8000 \
    vllm/vllm-openai-xpu:latest \
    Qwen/Qwen2.5-1.5B-Instruct \
        --dtype bfloat16 \
        --enforce-eager \
        --block-size=64 \
        --max-model-len 4096 \
        --gpu-memory-utilization 0.85

Wait for Application startup complete, then test:

curl -s http://localhost:8000/v1/chat/completions \
    -H 'Content-Type: application/json' \
    -d '{"model":"Qwen/Qwen2.5-1.5B-Instruct",
         "messages":[{"role":"user","content":"Hi."}],"max_tokens":32}'

Return the response content or a short summary to the user. A running HTTP server without a successful generation is not a validated deployment.

Cleanup: docker stop vllm-xpu && docker rm vllm-xpu.

Drop --rm on first launches so logs survive a crashed init.

To serve on a remote Intel GPU host over ssh (the local machine → remote-box workflow), see references/remote-deploy.md.

Flag rationales

FlagWhy
--dtype bfloat16Battlemage runs bf16 better than fp16; auto may pick fp16 from the checkpoint config. Set for unquantized serving only — combining with --quantization produces conflicts.
--enforce-eagerConservative default. Graph capture / torch.compile on XPU is experimental. After a stable eager run, drop it and re-bench; keep only if TPOT/TTFT improve and content stays correct.
--block-size=64Validated default for the XPU paged-attention path on Battlemage. Bench higher (128, 256) once 64 is correct.
--gpu-memory-utilization 0.85Default 0.92 fails on workstations with active GUI sessions. Drop to 0.70 with browsers open; raise to 0.92 on headless servers.
-v /dev/dri/by-path:/dev/dri/by-path:roSome oneCCL/device-discovery configurations scan the host's /dev/dri/by-path symlinks even for a single-GPU launch. Include this read-only mount to support those configurations; if it is omitted, affected images can abort at engine initialization with opendir failed: could not open device directory.
--max-model-len 4096KV cache is allocated up-front. Start at 4096, raise in 2× steps until OOM, back off one step.
--trust-remote-codeSecurity opt-in — permits arbitrary Python from the model repo to run in your engine. Set only when you trust the publisher.
--disable-sliding-windowWorkaround when SWA produces incorrect output for a specific (model, image) combination. Don't apply blindly — disabling SWA on a model designed for it inflates KV memory.
--model-impl transformersFallback for Model architectures ['<X>'] are not supported. Slower but correct. Upgrade transformers in the container if that also fails.

For pooling / embedding / reranker, serve with --dtype bfloat16 or --quantization fp8.

Env vars

vLLM's env-var surface is image-version-specific. After launch, verify there are no silent rejects:

docker logs <name> 2>&1 | grep -i "Unknown vLLM environment"   # must be empty
VariablePurpose
ZE_AFFINITY_MASKWhich XPU(s) the server sees.
HF_TOKENHF auth.
TRITON_CACHE_DIRPersist compiled XPU Triton kernels.
CCL_ZE_IPC_EXCHANGE=pidfdoneCCL IPC over Docker PID namespace (multi-GPU).
VLLM_LOGGING_LEVEL=DEBUGVerbose engine logs.
VLLM_WORKER_MULTIPROC_METHOD=spawnRequired — fork deadlocks oneCCL init on XPU.
VLLM_ALLOW_LONG_MAX_MODEL_LEN=1Set when extending RoPE past the model card's default.
VLLM_MLA_DISABLE=1Workaround for MLA models (DeepSeek-V2/V3, MiniMax-Text-01); no-op otherwise.

Quantization, attention backend selection

See references/quantization.md for the full (quant × KV dtype × attention backend) pairing table and live-kernel verification, plus FP8 / AWQ / GPTQ / MXFP4 / AutoRound specifics.

Multi-GPU, tuning, speculative decoding

See references/multi-gpu-and-tuning.md for tensor parallel, oneCCL / XCCL collective env vars, --block-size / --max-num-batched-tokens sweeps, speculative decoding (EAGLE3 / MTP / n-gram), and legacy env-var aliases for older images.

Attention backend gotcha (read before quantising)

--kv-cache-dtype fp8 and the W4A8 quant kernels (AWQ / GPTQ / MXFP4) require --attention-backend TRITON_ATTN. The default FA-XPU backend does not implement fp8 KV — vLLM exits with NotImplementedError at engine init. Full pairing table in references/quantization.md.

Common errors

  • Model architectures ['<X>'] are not supported → add --model-impl transformers. If still failing, upgrade transformers in the container or use a newer image.
  • RuntimeError: Cannot find any XPU devices → container missing GPU access; verify with xpu-smi discovery inside the container.
  • Free memory on device xpu:0 ... is less than desired GPU memory utilization → drop --gpu-memory-utilization to 0.85 or 0.70.
  • Other OOM at engine init → --max-model-len too large; halve.
  • OOM after a few requests → cap --max-num-seqs 16.
  • Server hangs at Detected platform: xpu → oneCCL init. Check --ipc=host; for multi-GPU, check both XPUs in ZE_AFFINITY_MASK.
  • Crash at init with oneCCL: ze_fd_manager.cpp ... init_device_fds: opendir failed: could not open device directory → /dev/dri/by-path not visible in the container. Add -v /dev/dri/by-path:/dev/dri/by-path:ro. Fires on single-GPU too (the worker all_reduces at init). Setting CCL_ZE_IPC_EXCHANGE=pidfd alone does not fix it — the drmfd fallback still scans by-path.
  • tensor parallel size N is not allowed → ZE_AFFINITY_MASK has fewer than N XPUs.
  • HTTP 400 "model not found" → model field in JSON must match /v1/models exactly.
  • Gibberish output → dtype mismatch. Force --dtype bfloat16. Last resort: --override-attention-dtype float32.
  • Triton compile error on first request → set TRITON_CACHE_DIR to a mounted volume so the next run starts hot.

Verifying device placement

xpu-smi dump -d 0 -m 5,18 -i 1 | head -5

Memory should sit at gigabytes once the engine is ready. <100 MiB while the server reports ready means the model loaded on CPU.

What this skill does NOT cover

  • Choosing the best quantization, KV dtype, DP/TP layout, context, or concurrency → model-config-recommend. Questions such as "How should I configure vLLM on my Arc cards?" belong there even when the user also names a model.
  • SGLang serving → sglang-xpu-run.
  • Pure PyTorch / Transformers → torch-xpu-run.
  • Throughput / TTFT / TPOT measurement → vllm-xpu-bench.
  • Profile-level slowness → vllm-xpu-profile.
  • SYCL kernel fixes — out of scope.
  • intel/llm-scaler-vllm images — out of scope.

References

Signals

GitHub stars
21
Forks
9
Last commit
Sep 2026
Advanced
Item type
skill
Key
vllm-xpu-run
Source
github.com/intel/skills