Add Multus, Whereabouts and RDMA
Cilium must already pass its connectivity test. Multus is a CNI multiplexer: it adds a secondary macvlan interface to a pod after Cilium creates the primary interface. Whereabouts allocates secondary IPs. NVIDIA Network Operator 26.7.0 is used only to advertise a narrow RDMA shared-device resource backed by inbox mlx5; it must not install OFED, NFD, Multus or other CNI components.
Discover the k3s CNI paths
Section titled “Discover the k3s CNI paths”Run on: each GPU host, root/read-only.
Expected: all nodes report the same k3s CNI binary and configuration locations.
Stop: if paths differ; do not substitute ordinary /opt/cni/bin or /etc/cni/net.d mounts.
CNI_BIN=/var/lib/rancher/k3s/data/cniCNI_CONF=/var/lib/rancher/k3s/agent/etc/cni/net.dtest -d "$CNI_BIN"test -d "$CNI_CONF"printf 'CNI_BIN=%s\nCNI_CONF=%s\n' "$CNI_BIN" "$CNI_CONF"test -x "$CNI_BIN/macvlan"The selected k3s release exposes a stable /var/lib/rancher/k3s/data/cni link. Follow k3s’ Multus guidance, not a generic distribution’s CNI paths.
Install Multus thick plugin
Section titled “Install Multus thick plugin”Use the pinned v4.2.2 thick-plugin manifest, rewritten locally so every CNI host mount uses the discovered k3s paths:
curl --fail --location --output /tmp/multus-thick-v4.2.2.yaml \ https://raw.githubusercontent.com/k8snetworkplumbingwg/multus-cni/v4.2.2/deployments/multus-daemonset-thick-plugin.ymlpython3 - /tmp/multus-thick-v4.2.2.yaml <<'PY'from pathlib import Pathimport sys
path = Path(sys.argv[1])text = path.read_text()text = text.replace('/etc/cni/net.d', '/var/lib/rancher/k3s/agent/etc/cni/net.d')text = text.replace('/opt/cni/bin', '/var/lib/rancher/k3s/data/cni')path.write_text(text)PYgrep -nE '/etc/cni/net.d|/opt/cni/bin' /tmp/multus-thick-v4.2.2.yaml && exit 1 || truekubectl apply -f /tmp/multus-thick-v4.2.2.yamlkubectl -n kube-system rollout status daemonset/kube-multus-ds --timeout=5mInspect /tmp/multus-thick-v4.2.2.yaml before applying. Both installer variables and host mounts must resolve to the same k3s paths. Stop if Multus replaces the Cilium primary configuration instead of creating a Multus wrapper with Cilium as delegate.
Install Whereabouts 0.9.2
Section titled “Install Whereabouts 0.9.2”curl --fail --location --output /tmp/whereabouts-v0.9.2.yaml \ https://raw.githubusercontent.com/k8snetworkplumbingwg/whereabouts/v0.9.2/doc/crds/daemonset-install.yamlpython3 - /tmp/whereabouts-v0.9.2.yaml <<'PY'from pathlib import Pathimport sys
path = Path(sys.argv[1])text = path.read_text()text = text.replace('/etc/cni/net.d', '/var/lib/rancher/k3s/agent/etc/cni/net.d')text = text.replace('/opt/cni/bin', '/var/lib/rancher/k3s/data/cni')path.write_text(text)PYgrep -nE '/etc/cni/net.d|/opt/cni/bin' /tmp/whereabouts-v0.9.2.yaml && exit 1 || truekubectl apply -f /tmp/whereabouts-v0.9.2.yamlkubectl -n kube-system rollout status daemonset/whereabouts --timeout=5mkubectl get crd ippools.whereabouts.cni.cncf.io overlappingrangeipreservations.whereabouts.cni.cncf.ioBefore applying, inspect the rewritten DaemonSet’s CNI config/binary hostPath mounts. Do not apply it if a generic host path remains.
Install Network Operator as allocator only
Section titled “Install Network Operator as allocator only”Network Operator’s bundled NFD, Multus, CNI and OFED management are disabled. The host’s pinned open driver and inbox mlx5 remain authoritative.
helm repo add nvidia https://helm.ngc.nvidia.com/nvidiahelm repo updatehelm upgrade --install network-operator nvidia/network-operator \ --namespace network-operator \ --create-namespace \ --version 26.7.0 \ --set nfd.enabled=false \ --set multus.enabled=false \ --set sriovNetworkOperator.enabled=false \ --set ofedDriver.deploy=falsekubectl -n network-operator wait --for=condition=Available deployment --all --timeout=10mhelm get values network-operator -n network-operator -o yamlStop if Network Operator creates an OFED driver pod, a second NFD, or another Multus DaemonSet.
Define the secondary network
Section titled “Define the secondary network”Read real values chosen with the network operator; no example fabric values are safe defaults:
read -r -p 'Physical RoCE interface: ' ROCE_INTERFACEread -r -p 'Secondary subnet CIDR: ' ROCE_SUBNETread -r -p 'First allocatable IP: ' ROCE_RANGE_STARTread -r -p 'Last allocatable IP: ' ROCE_RANGE_ENDread -r -p 'Fabric MTU: ' ROCE_MTUread -r -p 'Reserved CIDR or IP: ' ROCE_EXCLUDEThe subnet must not overlap host, pod or service networks. Reserve host and infrastructure addresses through Whereabouts exclusions.
cat <<EOF | kubectl apply -f -apiVersion: k8s.cni.cncf.io/v1kind: NetworkAttachmentDefinitionmetadata: name: roce-net namespace: vllmspec: config: |- { "cniVersion": "0.3.1", "type": "macvlan", "master": "${ROCE_INTERFACE}", "mode": "bridge", "mtu": ${ROCE_MTU}, "ipam": { "type": "whereabouts", "range": "${ROCE_SUBNET}", "range_start": "${ROCE_RANGE_START}", "range_end": "${ROCE_RANGE_END}", "exclude": ["${ROCE_EXCLUDE}"] } }EOFThe documented fabric is L2-only, so the manifest omits gateway. If your routed secondary network requires one, add the real address selected by the network operator rather than inventing it.
Advertise one RDMA unit per node
Section titled “Advertise one RDMA unit per node”Identify the runtime network values
Section titled “Identify the runtime network values”The manifest generator needs the resource key created by the policy below, the HCA names selected by that policy, and the GID index for the RoCE address family and VLAN. Inspect them on each GPU host:
ibdev2netdevshow_gidsrdma link showUse the same verified GID index on both ranks. Stop if the selected HCA does not map to the intended secondary-network interface or if the two hosts expose different mappings.
Find the actual netdevice names and PCI driver on both hosts:
rdma link showibdev2netdevethtool -i "$ROCE_INTERFACE"Create a narrow shared-device policy matching only the intended physical interface. Do not use rdmaMax: 63; production allocator counts are not physical-device counts. This fresh lane needs one allocatable unit for its one rank per node.
cat <<EOF | kubectl apply -f -apiVersion: mellanox.com/v1alpha1kind: NicClusterPolicymetadata: name: nic-cluster-policyspec: rdmaSharedDevicePlugin: config: | { "configList": [{ "resourceName": "roce", "rdmaHcaMax": 1, "selectors": { "ifNames": ["${ROCE_INTERFACE}"], "drivers": ["mlx5_core"] } }] }EOFkubectl get nicclusterpolicy nic-cluster-policy -o yamlkubectl get nodes -o custom-columns=NAME:.metadata.name,RDMA:.status.allocatable.rdma\.com/roceExpected: each GB10 agent advertises at least one rdma.com/roce. That is an allocator unit granting access to the selected RDMA device, not a count of HCAs or ports.
Prove attachment, then GPU RDMA
Section titled “Prove attachment, then GPU RDMA”Create two diagnostic pods, one pinned to each GPU node, with:
- annotation
k8s.v1.cni.cncf.io/networks: vllm/roce-net; - request and limit
rdma.com/roce: 1; runtimeClassName: nvidiafor the later GPU test.
Set DIAG_POD to each resulting pod in turn:
read -r -p 'Diagnostic pod name: ' DIAG_PODkubectl -n vllm get pod "$DIAG_POD" -o jsonpath='{.metadata.annotations.k8s\.v1\.cni\.cncf\.io/network-status}'kubectl -n vllm exec "$DIAG_POD" -- ip -brief addresskubectl -n vllm exec "$DIAG_POD" -- rdma link showThe network-status annotation must show one Cilium primary interface and one distinct Whereabouts address on the secondary network. Re-run a primary-network DNS/API connectivity check after Multus installation. A secondary-interface ping proves only IP attachment.
Before applying the model LWS, perform NVIDIA’s supported GPUDirect RDMA verification between the two diagnostic pods using the exact flags documented for the installed perftest build. Run the receiving command on one pod and the matching CUDA-buffer client on the other; require successful transfer and reported bandwidth. Stop if the test falls back to host memory, cannot open the mlx5 device, or reports GID/MTU/address-family mismatch.
Record the selected NetworkAttachmentDefinition name, HCA, GID index, socket/Gloo interface and extended resource key for the recipe builder.