GR00T

SkillCloud & infra

Use when working on NVIDIA GR00T deployment, model download, finetuning, evaluation, serving, inference, conversion, status checks, validation, routing, or CUDA alignment.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the GR00T skill

What this skill tells your AI

The instructions your AI receives, as published by nebius/nebius-physical-ai in skills/tools/groot/SKILL.md and read by ahel’s review.

When To Use

Use this skill for NVIDIA GR00T robot foundation model workbench changes, especially NGC/Hugging Face model handling, embodiment tags, checkpoint conversion, serving, inference, and validation.

Procedure

  1. Start with the current command surface:

    npa workbench groot --help
    
  2. Use download to stage model artifacts and credentials, finetune for training, eval for offline scoring, serve and infer for runtime calls, and convert when transforming checkpoints for downstream use. For declarative N1.7 training, use workflows/testing/groot-1-7-finetune.yaml; its gpu_count is a positive single-node world size that controls both the scheduler allocation and the real trainer.

  3. Use status, system-info, and list for operational checks. Keep ensure-ingress, register-byovm, reload-env, and cleanup-partial scoped to setup and recovery flows.

  4. Preserve credential redaction for NGC, Hugging Face, S3, and SSH values.

  5. Before deploy provisions or updates anything, validate actual Hugging Face access to both the selected GR00T checkpoint and its runtime-fetched nvidia/Cosmos-Reason2-2B dependency. The operator's HF token and upstream permissions are the only local gate for gated weights; do not add a manual acceptance flag or a model-check bypass.

Three-Tier Contract

  • CLI: list, deploy, download, finetune, eval, serve, infer, convert, status, and system-info are the main user commands.
  • SDK/API: keep model, checkpoint, and storage path normalization in shared helpers so service and CLI routes do not diverge.
  • YAML: groot-1-7-finetune.yaml calls workbench.groot.finetune in the stage's own image. Keep model/source pins, checkpoint S3 URIs, run ID, and GPU count explicit; counts above one must use the upstream torchrun path.

Routing And Validation

  • GR00T does not require RT cores for the standard model paths.
  • Route throughput-heavy training/eval to H100/H200 unless a command or image specifically requires another target.
  • CUDA 13 alignment is vendor-paced on NVIDIA x86_64 CUDA 13 and is not a Nebius infrastructure blocker.

Runtime Isaac bootstrap (the container ships no Isaac Sim)

Before a GR00T Isaac simulation run, load skills/atomic/third-party-eula-preflight/SKILL.md. Standard inference and fine-tuning do not trigger this preflight.

The npa-groot image contains no NVIDIA Isaac Sim or Isaac Lab code. It used to bake Omniverse Kit, which made it non-redistributable; Isaac is now downloaded on first use of /isaac-sim/python.sh from https://pypi.nvidia.com, into a cache volume, under the operator's own EULA acceptance. Full rationale: docs/workbench/container-packaging.md and skills/atomic/solution-licensing/SKILL.md.

What this changes in practice:

  • Only Isaac simulation defaults acceptance. NPA defaults ACCEPT_EULA=Y on that path. Empty, N, NO, 0, FALSE, or --no-accept-eula exits 78 before download. Y, YES, 1, and TRUE are accepted case-insensitively; other values are invalid. The launcher derives OMNI_KIT_ACCEPT_EULA=YES internally; do not expose duplicate user plumbing. Keep PRIVACY_CONSENT and telemetry off. Standard GR00T inference and fine-tuning do not require Isaac acceptance.
  • Reach Isaac through /isaac-sim/python.sh (the value of ISAAC_LAB_PYTHON). That is the bootstrap shim, and it is what every SkyPilot template, the sim2real engine and the workbench CLI already use. A bare python3 is the system interpreter and will not find Isaac.
  • Never invoke the shim from a Dockerfile RUN. It would download and bake ~4.5 GB of Isaac into a layer. Build-time work uses the image's own venv python.
  • Budget the first start. Measured on RTX PRO 6000: 111 s cold, 32 ms warm, 10.04 GiB of cache. Pre-warm a shared volume once per node/PVC with npa/docker/workbench/common/warm-isaac-cache.yaml, then run workload pods with NPA_ISAAC_CACHE_READONLY=1. Otherwise every pod pays it, and 8 GPU pods on a node download ~36 GB.
  • isaac-bootstrap status reports what is cached without needing acceptance or network; isaac-bootstrap verify additionally launches Isaac Sim headless (needs a GPU).
  • No NGC credentials are needed to build or run this image.

GR00T inference and fine-tuning run in GROOT_VENV (python -m npa.smoke.test_groot_functional covers inference) and need no Isaac or EULA acceptance, so they pay no first-run download. Only the Isaac Lab simulation paths do.

Gotchas

  • NGC credentials are required for gated NGC model refs. Do not print token values in diagnostics.
  • Managed VM deploy defaults to in-place updates for existing aliases. Terraform plans that would destroy or replace critical infrastructure are blocked unless the operator passes --replace and confirms with --yes.
  • BYOVM deploys record endpoint_strategy: public or endpoint_strategy: ssh_fallback in ~/.npa/config.yaml; live commands honor that strategy and can self-heal blocked public endpoints through a transient SSH-local route.
  • Known issue: output truncation at high step counts must be validated with artifacts, not subjective evaluation.
  • Multi-GPU fine-tuning defaults to NCCL's native transport selection. When a clean two-rank collective proves that both P2P and SHM are unsafe on a single-node host, use the workflow's nccl_transport=socket compatibility fallback. It disables both transports, so do not use it speculatively on healthy high-bandwidth hosts.

Verify

npa/.venv/bin/python -m pytest npa/tests/guardrails/test_skills_index.py -q
npa/.venv/bin/python -m pytest npa/tests/orchestration/npa_workflow/test_groot_finetune_workflow.py -q

The skill smoke invokes current GR00T training help and parses the reference N1.7 workflow in addition to the download/convert command checks.

Signals

GitHub stars
28
Forks
15
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
groot
Source
github.com/nebius/nebius-physical-ai