Skip to content

Build and deploy on NVIDIA DGX Spark

Latest publicationDigest resolved during the trusted site buildInspect release data

This project builds local-inference-lab/vLLM and its source dependencies for NVIDIA DGX Spark. It publishes an ARM64 OCI image and generates Kubernetes and Docker configuration for the two-node TP=2 deployments I run on two DGX Spark systems. It is not a builder or recipe catalog for upstream vllm-project/vllm releases.

Kubernetes and Docker

Generate Kubernetes and Docker configuration

Each page uses the DGX Spark image built from the pinned local-inference-lab/vLLM revision. It keeps the tested image separate from the newest publication, then renders Kubernetes and Docker artifacts from one authored definition. Site-specific network, storage, and resource values remain yours to supply.

Running on LWS

DeepSeek V4 Flash Vision

One million token context, tensor parallel size two, one DGX Spark system per node, and the exact digest accepted live.

TP 22 nodesARM64
Open recipe →

Running on LWS

Qwen3.8 Flash Next

NVFP4 weights, a 1,048,576-token context from a 4x YaRN override, MTP speculation at depth three, MXFP8 draft experts on B12X kernels, and PLE tables in page-locked host memory.

TP 22 nodesARM64
Open recipe →

k3s setup guide

Install the Kubernetes dependencies

A linear route through the 64 KiB kernel driver, Cilium, GPU and RDMA dependencies, LWS, and inference-aware routing.

CiliumMultusllm-d
Open setup guide →
  • local-inference-lab/vllm#800 adds bounded shared-memory broadcast waits and gates the adaptive reader and writer waits behind VLLM_EXPERIMENTAL_SHM_BROADCAST_ADAPTIVE_SPIN=1. The Qwen3.8 Flash Next recipe enables it with a 50 ms maximum grace.
  • local-inference-lab/b12x#384 keeps prepared launchers alive through the PCIe and RoCE launcher factories and runs collective priming behind a deadline-enforced barrier, so later preparation waits until a timed-out callback exits.
  • local-inference-lab/vllm#955 routes ModelOpt MXFP8 MoE experts, such as the Qwen3.8 step-5500 MTP draft, to native B12X kernels instead of Marlin, preparing them at their sharded width.
  • local-inference-lab/b12x#449 adds native block-scaled MXFP8 W8A8 grouped MoE, including 32-aligned intermediate sizes such as the 320-wide per-rank experts at TP=2. vllm#955 depends on it.