Deploy the DeepSeek TP=2 group
This deployment uses two nodes, one physical GB10 and one vLLM rank per node, with tensor parallel size 2. The model snapshot is pinned to 6821d6ad3681a4b137b066b76094fa82ebd0a380. The default image is the exact digest accepted live on 14 September 2026.
The rank-zero pod owns the HTTP API on port 8888. The rank-one pod is headless and participates in GPU collectives. LWS controls group lifecycle and leader-first startup; it does not schedule inference requests. Cilium carries Kubernetes and control traffic, while the separately attached Multus network carries the configured collective path. The figure and the rest of that reasoning are in traffic paths in a two-node TP=2 group.
Render a complete artifact
Section titled “Render a complete artifact”Storage paths
Section titled “Storage paths”model_storage_path is a durable root shared by model downloads. The default /var/lib/models is safe because model-sync keeps snapshots revision-specific and publishes this recipe below /models/deepseek-v4-flash-vision in the container. jit_storage_path must be model-specific because generated kernels and compiler artifacts depend on the model and engine configuration. The default is /var/cache/vllm/deepseek-v4-flash-vision. Create both paths on each GPU node and ensure the container runtime can write to them.
Node placement
Section titled “Node placement”List labels before choosing a selector or topology contract:
kubectl get nodes --show-labelskubectl get nodes -L kubernetes.io/arch,nvidia.com/gpu.productnode_selector uses key=value entries and must select both GB10 nodes. topology_key is a label key whose values distinguish the eligible failure domains. topology_values is the comma-separated set accepted for this group. Add deliberate site labels if the standard labels do not express the pair. Do not use hostnames as a topology scheme unless pinning the deployment to those exact machines is intentional.
Runtime network interfaces
Section titled “Runtime network interfaces”Kubernetes pods use eth0 for Gloo coordination and NCCL socket bootstrap, so the LWS generator fixes both variables to that primary pod interface. NCCL_IB_HCA separately selects the RDMA devices used for GPU collectives. Its rocep1s0f1,roceP2p1s0f1 preset is the pair used by this DeepSeek deployment, not a hardware guarantee: another workload in the same cluster exposes a differently named second device. Use the Multus and RDMA guide to inspect both hosts and type over the preset when their enumeration differs.
Docker host values
Section titled “Docker host values”Docker uses host networking. Set leader_ip and worker_ip to stable host addresses that reach each other. Set the socket and Gloo interfaces to real host interfaces, and the HCA list to the host RDMA devices mounted into each container. Kubernetes NetworkAttachmentDefinition names and extended-resource keys do not apply to Docker.
Open the recipe builder, select Kubernetes / LeaderWorkerSet, and supply:
- namespace and workload name;
- absolute model and JIT host paths;
- Hugging Face Secret name and key;
- NetworkAttachmentDefinition, HCA, GID index and control interface;
- exact node selector, topology key and allowed topology values;
- the RDMA extended-resource key and allocator-unit count.
Each rank requests one nvidia.com/gpu; the fresh GPU Operator setup guarantees that contract, so it is not a form field. The generator defaults the secondary attachment to roce-net, the RDMA resource to rdma.com/roce, its allocator count to one, and the GID index to 3. These match the setup guides but remain editable because network policy and GID tables are site-specific. Do not copy resource keys or high allocator counts from another site unless your own device plugins expose that contract.
The builder withholds copy/download actions when required values are missing or invalid. It never asks for a token. Keep Validated recipe image selected for the accepted digest; selecting the resolved latest image is explicitly unvalidated for this recipe.
Create namespace and credential
Section titled “Create namespace and credential”Run on: workstation.
Writes: namespace and one Secret.
Expected: Secret contains the local environment value without printing it.
Stop: if HF_TOKEN is empty; never place it in rendered YAML.
: "${HF_TOKEN:?export HF_TOKEN locally first}"read -r -p 'Recipe namespace: ' RECIPE_NAMESPACEread -r -p 'Hugging Face Secret name: ' HF_SECRETread -r -p 'Secret data key: ' HF_SECRET_KEYkubectl create namespace "$RECIPE_NAMESPACE" --dry-run=client -o yaml | kubectl apply -f -kubectl -n "$RECIPE_NAMESPACE" create secret generic "$HF_SECRET" \ --from-literal="$HF_SECRET_KEY=$HF_TOKEN" \ --dry-run=client -o yaml | kubectl apply -f -Inspect and apply the download
Section titled “Inspect and apply the download”Download the generated YAML. The builder has already rejected missing fields and private references; perform a local Kubernetes parse before apply:
kubectl apply --dry-run=client -f deepseek-v4-flash-vision-lws.yamlkubectl apply -f deepseek-v4-flash-vision-lws.yamlThe generated pods select runtimeClassName: nvidia, mount persistent /models and /cache, use 64 GiB memory-backed shared memory, and run vllm-image model-sync. Only the worker waits for the leader rendezvous. There is no liveness probe: long GPU kernels must not trigger a destructive group restart.
Observe startup
Section titled “Observe startup”read -r -p 'Recipe workload name: ' RECIPE_NAMEkubectl -n "$RECIPE_NAMESPACE" get leaderworkerset,pod -wkubectl -n "$RECIPE_NAMESPACE" describe leaderworkerset "$RECIPE_NAME"kubectl -n "$RECIPE_NAMESPACE" get pods -l leaderworkerset.sigs.k8s.io/name="$RECIPE_NAME" -o wideExpected sequence:
- both nodes validate/download the identical pinned snapshot;
- rank zero opens rendezvous port 25000;
- rank one resolves the LWS leader address and connects;
- the two model containers become Ready only after the rank-zero
/v1/modelsAPI responds and the worker engine child is alive.
Stop if ranks land on the same topology value, image digests differ, rank one becomes an API Service endpoint, or a helper exits with configuration/DNS/port/model errors. Startup and readiness use the same HTTP health basis with different kubelet budgets; readiness is not a stronger semantic model check.
Direct rank-zero acceptance
Section titled “Direct rank-zero acceptance”The generated rank-zero Service is debugging access. Port-forward it from the workstation:
kubectl -n "$RECIPE_NAMESPACE" port-forward service/"$RECIPE_NAME-serve" 8888:8888In another terminal, first verify the served name:
curl --fail --silent http://127.0.0.1:8888/v1/modelsThen make three real requests against the recipe’s served model name (deepseek-v4-flash-vision for that recipe), which is what --served-model-name publishes and what /v1/models returns: a text completion, a tool-call request with an explicit tool schema, and a vision request whose message includes one valid image. Require valid OpenAI-compatible JSON and the expected content/tool result for each.
This direct Service checks the rank-zero API and both-rank engine group. It bypasses EPP selection. Continue with request routing to configure the Gateway path.
Operational stop
Section titled “Operational stop”A healthy distributed group may spend a long time in GPU kernels. Never force-delete its pods or force-kill GB10 processes merely because a request or shutdown is slow. Use ordinary Kubernetes deletion and allow the generated 120-second termination grace. If maintenance is required, drain requests, remove the complete group gracefully, verify no GPU holders, then proceed.