Running on LWS
DeepSeek V4 Flash Vision
One million token context, tensor parallel size two, one DGX Spark system per node, and the exact digest accepted live.
Open recipe →This project builds local-inference-lab/vLLM and its source dependencies for NVIDIA DGX Spark. It publishes an ARM64 OCI image and generates Kubernetes and Docker configuration for the two-node TP=2 deployments I run on two DGX Spark systems. It is not a builder or recipe catalog for upstream vllm-project/vllm releases.
Kubernetes and Docker
Each page uses the DGX Spark image built from the pinned local-inference-lab/vLLM revision. It keeps the tested image separate from the newest publication, then renders Kubernetes and Docker artifacts from one authored definition. Site-specific network, storage, and resource values remain yours to supply.
Running on LWS
One million token context, tensor parallel size two, one DGX Spark system per node, and the exact digest accepted live.
Open recipe →Running on LWS
NVFP4 weights, a 1,048,576-token context from a 4x YaRN override, MTP speculation at depth three, MXFP8 draft experts on B12X kernels, and PLE tables in page-locked host memory.
Open recipe →k3s setup guide
A linear route through the 64 KiB kernel driver, Cilium, GPU and RDMA dependencies, LWS, and inference-aware routing.
Open setup guide →VLLM_EXPERIMENTAL_SHM_BROADCAST_ADAPTIVE_SPIN=1. The Qwen3.8 Flash Next recipe enables it with a 50 ms maximum grace.