Install GPU Operator in host-owned mode
The host owns NVIDIA kernel modules, GSP firmware, userspace and Container Toolkit. GPU Operator 26.7.0 owns discovery, validation, CDI and the device plugin only. Never enable its driver or RDMA-driver containers on the custom-fork lane.
Choose the sole NFD owner
Section titled “Choose the sole NFD owner”NFD (Node Feature Discovery) labels hardware features on Nodes.
- Fresh cluster from this tutorial: GPU Operator owns bundled NFD; use
nfd.enabled=true. - Existing cluster that already has NFD: set
nfd.enabled=falseand verify the existing NFD is healthy.
Network Operator must later use nfd.enabled=false in both cases. Two NFD installations compete over labels and lifecycle; stop if more than one is installed.
kubectl get deploy,ds --all-namespaces -l app.kubernetes.io/name=node-feature-discoveryInstall the chart
Section titled “Install the chart”Run on: workstation with cluster-admin credentials.
Writes: gpu-operator namespace, CRDs and operator operands.
Expected: operator operands settle Running/Completed without a driver DaemonSet.
Stop: if host driver/runtime checks in the k3s guide failed, or the chart tries to install a driver/toolkit.
helm repo add nvidia https://helm.ngc.nvidia.com/nvidiahelm repo updatehelm upgrade --install gpu-operator nvidia/gpu-operator \ --namespace gpu-operator \ --create-namespace \ --version 26.7.0 \ --set driver.enabled=false \ --set toolkit.enabled=false \ --set nfd.enabled=true \ --set cdi.enabled=trueFor the existing-NFD lane change only the final NFD value:
helm upgrade --install gpu-operator nvidia/gpu-operator \ --namespace gpu-operator --create-namespace --version 26.7.0 \ --set driver.enabled=false --set toolkit.enabled=false \ --set nfd.enabled=false --set cdi.enabled=trueConfirm ownership from rendered values, not pod names alone:
helm get values gpu-operator -n gpu-operator -o yamlkubectl -n gpu-operator get pods -o widekubectl -n gpu-operator wait --for=condition=Available deployment --all --timeout=10mkubectl -n gpu-operator get daemonsetExpected values are exactly driver.enabled: false and toolkit.enabled: false. No operator driver DaemonSet may run. Diagnose operands that are CrashLoopBackOff before continuing; common boundaries are missing host modules, missing CDI/runtime integration or unsatisfied NFD labels.
Require one allocatable GPU on each DGX Spark
Section titled “Require one allocatable GPU on each DGX Spark”kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpuFor the two fresh GB10 nodes, the GPU column must be 1. Do not copy higher allocator values from another site: an extended-resource count is an allocator unit, not a physical-GPU count. Experienced operators may use a different parameterized resource key/count only when their installed device plugin defines it.
Run the intermediate CUDA proof
Section titled “Run the intermediate CUDA proof”Run on: workstation.
Creates: one temporary pod using one physical GPU and runtimeClassName: nvidia.
Expected: CUDA allocates memory and synchronizes.
Stop: if scheduling, CDI injection, allocation or synchronization fails. nvidia-smi visibility alone is insufficient.
cat <<'EOF' | kubectl apply -f -apiVersion: v1kind: Podmetadata: name: cuda-allocation-smoke namespace: defaultspec: restartPolicy: Never runtimeClassName: nvidia containers: - name: cuda image: nvcr.io/nvidia/k8s/cuda-sample:vectoradd-cuda12.5.0 resources: limits: nvidia.com/gpu: 1EOFkubectl wait pod/cuda-allocation-smoke --for=jsonpath='{.status.phase}'=Succeeded --timeout=10mkubectl logs cuda-allocation-smokekubectl delete pod cuda-allocation-smokeExpected logs end with Test PASSED. This proves one container can execute CUDA; it does not prove RDMA, TP=2 collectives or the recipe image. Continue with Multus and RDMA.
See NVIDIA’s GPU Operator getting started guide, platform support and pre-installed driver/toolkit guidance.