Two-node tensor-parallel deployment A client sends HTTP inference requests through a Kubernetes Service to the rank-zero vLLM pod, while rank zero and rank one exchange tensor-parallel collectives over a separate RoCE network. DEPLOYMENT / TESTED LWS SHAPE TP=2 deployment on two DGX Spark nodes LeaderWorkerSet fixes placement and startup; vLLM owns the TP=2 execution group. GB10 NODE A / WORKER 0 GB10 NODE B / WORKER 1 HTTP :8888 ONLY RANK 0 RDMA COLLECTIVE RDMA COLLECTIVE CLIENT Inference client text / tool / vision SERVICE Rank-zero serving Service selector: worker-index=0 / :8888 POD Leader pod / rank 0 vLLM API :8888 / one GB10 Model snapshot PINNED REVISION POD Worker pod / rank 1 headless / one GB10 Model snapshot PINNED REVISION NODE-LOCAL /models + /cache NODE-LOCAL /models + /cache LEGEND API REQUEST TP COLLECTIVE STARTUP CONTROL PHYSICAL NODE BOUNDARY

The API path ends at rank zero

The Service selects only worker-index=0. Rank zero receives the OpenAI-compatible request on port 8888 and coordinates inference; rank one is not an API endpoint.

Collectives use a second path

Both ranks participate in TP=2. Tensor data crosses the secondary RoCE attachment, separately from the primary-network HTTP request and LWS startup rendezvous.