Skip to content

Deploy the DeepSeek TP=2 group

This deployment uses two nodes, one physical GB10 and one vLLM rank per node, with tensor parallel size 2. The model snapshot is pinned to 6821d6ad3681a4b137b066b76094fa82ebd0a380. The default image is the exact digest accepted live on 14 September 2026.

The rank-zero pod owns the HTTP API on port 8888. The rank-one pod is headless and participates in GPU collectives. LWS controls group lifecycle and leader-first startup; it does not schedule inference requests. Cilium carries Kubernetes and control traffic, while the separately attached Multus network carries the configured collective path. The figure and the rest of that reasoning are in traffic paths in a two-node TP=2 group.

model_storage_path is a durable root shared by model downloads. The default /var/lib/models is safe because model-sync keeps snapshots revision-specific and publishes this recipe below /models/deepseek-v4-flash-vision in the container. jit_storage_path must be model-specific because generated kernels and compiler artifacts depend on the model and engine configuration. The default is /var/cache/vllm/deepseek-v4-flash-vision. Create both paths on each GPU node and ensure the container runtime can write to them.

List labels before choosing a selector or topology contract:

Terminal window
kubectl get nodes --show-labels
kubectl get nodes -L kubernetes.io/arch,nvidia.com/gpu.product

node_selector uses key=value entries and must select both GB10 nodes. topology_key is a label key whose values distinguish the eligible failure domains. topology_values is the comma-separated set accepted for this group. Add deliberate site labels if the standard labels do not express the pair. Do not use hostnames as a topology scheme unless pinning the deployment to those exact machines is intentional.

Kubernetes pods use eth0 for Gloo coordination and NCCL socket bootstrap, so the LWS generator fixes both variables to that primary pod interface. NCCL_IB_HCA separately selects the RDMA devices used for GPU collectives. Its rocep1s0f1,roceP2p1s0f1 preset is the pair used by this DeepSeek deployment, not a hardware guarantee: another workload in the same cluster exposes a differently named second device. Use the Multus and RDMA guide to inspect both hosts and type over the preset when their enumeration differs.

Docker uses host networking. Set leader_ip and worker_ip to stable host addresses that reach each other. Set the socket and Gloo interfaces to real host interfaces, and the HCA list to the host RDMA devices mounted into each container. Kubernetes NetworkAttachmentDefinition names and extended-resource keys do not apply to Docker.

Open the recipe builder, select Kubernetes / LeaderWorkerSet, and supply:

  • namespace and workload name;
  • absolute model and JIT host paths;
  • Hugging Face Secret name and key;
  • NetworkAttachmentDefinition, HCA, GID index and control interface;
  • exact node selector, topology key and allowed topology values;
  • the RDMA extended-resource key and allocator-unit count.

Each rank requests one nvidia.com/gpu; the fresh GPU Operator setup guarantees that contract, so it is not a form field. The generator defaults the secondary attachment to roce-net, the RDMA resource to rdma.com/roce, its allocator count to one, and the GID index to 3. These match the setup guides but remain editable because network policy and GID tables are site-specific. Do not copy resource keys or high allocator counts from another site unless your own device plugins expose that contract.

The builder withholds copy/download actions when required values are missing or invalid. It never asks for a token. Keep Validated recipe image selected for the accepted digest; selecting the resolved latest image is explicitly unvalidated for this recipe.

Run on: workstation.
Writes: namespace and one Secret.
Expected: Secret contains the local environment value without printing it.
Stop: if HF_TOKEN is empty; never place it in rendered YAML.

Terminal window
: "${HF_TOKEN:?export HF_TOKEN locally first}"
read -r -p 'Recipe namespace: ' RECIPE_NAMESPACE
read -r -p 'Hugging Face Secret name: ' HF_SECRET
read -r -p 'Secret data key: ' HF_SECRET_KEY
kubectl create namespace "$RECIPE_NAMESPACE" --dry-run=client -o yaml | kubectl apply -f -
kubectl -n "$RECIPE_NAMESPACE" create secret generic "$HF_SECRET" \
--from-literal="$HF_SECRET_KEY=$HF_TOKEN" \
--dry-run=client -o yaml | kubectl apply -f -

Download the generated YAML. The builder has already rejected missing fields and private references; perform a local Kubernetes parse before apply:

Terminal window
kubectl apply --dry-run=client -f deepseek-v4-flash-vision-lws.yaml
kubectl apply -f deepseek-v4-flash-vision-lws.yaml

The generated pods select runtimeClassName: nvidia, mount persistent /models and /cache, use 64 GiB memory-backed shared memory, and run vllm-image model-sync. Only the worker waits for the leader rendezvous. There is no liveness probe: long GPU kernels must not trigger a destructive group restart.

Terminal window
read -r -p 'Recipe workload name: ' RECIPE_NAME
kubectl -n "$RECIPE_NAMESPACE" get leaderworkerset,pod -w
kubectl -n "$RECIPE_NAMESPACE" describe leaderworkerset "$RECIPE_NAME"
kubectl -n "$RECIPE_NAMESPACE" get pods -l leaderworkerset.sigs.k8s.io/name="$RECIPE_NAME" -o wide

Expected sequence:

  1. both nodes validate/download the identical pinned snapshot;
  2. rank zero opens rendezvous port 25000;
  3. rank one resolves the LWS leader address and connects;
  4. the two model containers become Ready only after the rank-zero /v1/models API responds and the worker engine child is alive.

Stop if ranks land on the same topology value, image digests differ, rank one becomes an API Service endpoint, or a helper exits with configuration/DNS/port/model errors. Startup and readiness use the same HTTP health basis with different kubelet budgets; readiness is not a stronger semantic model check.

The generated rank-zero Service is debugging access. Port-forward it from the workstation:

Terminal window
kubectl -n "$RECIPE_NAMESPACE" port-forward service/"$RECIPE_NAME-serve" 8888:8888

In another terminal, first verify the served name:

Terminal window
curl --fail --silent http://127.0.0.1:8888/v1/models

Then make three real requests against the recipe’s served model name (deepseek-v4-flash-vision for that recipe), which is what --served-model-name publishes and what /v1/models returns: a text completion, a tool-call request with an explicit tool schema, and a vision request whose message includes one valid image. Require valid OpenAI-compatible JSON and the expected content/tool result for each.

This direct Service checks the rank-zero API and both-rank engine group. It bypasses EPP selection. Continue with request routing to configure the Gateway path.

A healthy distributed group may spend a long time in GPU kernels. Never force-delete its pods or force-kill GB10 processes merely because a request or shutdown is slow. Use ordinary Kubernetes deletion and allow the generated 120-second termination grace. If maintenance is required, drain requests, remove the complete group gracefully, verify no GPU holders, then proceed.