Qwen3.8 Flash Next on two DGX Spark systems
Runtime recipe · Qwen3.8 Flash Next
This is the public, parameterized form of the two-node NVIDIA DGX Spark deployment I run for Qwen3.8-Flash-Next-NVFP4. Its image contains the pinned local-inference-lab/vLLM fork rather than an arbitrary upstream vLLM release. The recipe fixes the model revision, engine flags, environment, topology, and the image the configuration was accepted on. You supply the details that belong to your cluster or Docker hosts.
- Model
- local-inference-lab/Qwen3.8-Flash-Next-NVFP4
- Model revision
- 60215d26cf5e42c2db6128774032d57fc62678da
- Topology
- 2 nodes · TP=2 · one DGX Spark per rank
- Context
- 1,048,576 tokens
What is fixed
Section titled “What is fixed”The recipe preserves the live runtime command and tuning accepted on 2026-09-30: B12X GDN prefill and decode, the B12X linear and MoE backends, buffered InstantTensor loading, NVFP4 weights through modelopt_mixed, MTP speculation at depth 3, FP8 KV cache, full and piecewise CUDA graphs capped at a 32-token capture size, GPU utilization 0.74, eight sequences, an 8,192-token batch budget, decode-aware prefill scheduling, and KV-event publishing from rank zero. It also selects the tested A16 dense-activation path with quantized B12X MXFP8 activations, one-shot RoCE all-reduce limits, the experimental adaptive shared-memory broadcast spin (VLLM_EXPERIMENTAL_SHM_BROADCAST_ADAPTIVE_SPIN=1, maximum grace 50 ms), and the image’s AOT artifacts. The first boot with a 900 ms maximum grace deadlocked one group in B12X preparation; 50 ms booted both groups. One failure does not isolate the spin setting as the cause.
Two things changed at this cut, and both trace back to the checkpoint.
The weights moved. The accepted revision is 60215d26cf5e: QAD step 5500 with PLE step 1000 (branch qad-step5500-ple1000), pinned by commit so a branch move cannot silently change what is served. Its MTP draft experts are MXFP8 - E4M3 weights with UE8M0 scale factors - while the target experts stay NVFP4. That requires native B12X MXFP8 MoE; without it --moe-backend b12x rejects the draft at kernel selection. The live rank-0 log shows what was chosen: Using 'B12X' NvFp4 MoE backend, Using 'B12X_MXFP8' MxFp8 MoE backend (user-requested), and nothing on Marlin. At TP=2 the 640-wide MTP experts shard to 320 per rank, and B12X runs them at 320: UE8M0 scales cover 32-wide K blocks, and 32-aligned intermediate sizes use separate up and gate descriptors, so no padding is prepared.
The memory budget moved with them. Step 5500 loads 42.77 GiB per rank against 40.32 GiB at step 4000, and at 0.76 the startup profile pass froze two GB10 hosts with no kernel log. 0.74 hands back 2.48 GiB of the node’s 123.79 GiB and restores the absolute headroom step 4000 had. Measured on the live ranks after the change: one group reports 24.78 GiB of KV cache memory, 3,414,080 tokens, 3.26x concurrency for a full-window request; the other reports 25.29 GiB, 3,483,537 tokens, 3.32x. That spread is node memory state at boot, not a configuration difference.
The image those ranks run is published by this repository as v20261001.1: vLLM 0a6739845a32 on dev/rv-mxfp8-mtp and B12X 965e748f74fe on feat/mxfp8-moe, digest sha256:43aaf6d1b5e0…dae6cbde. The RecipeBuilder offers it as the validated image.
Two settings deserve attention because they are not interchangeable with defaults.
The PLE tables live in page-locked host memory. --additional-config '{"ple_table_memory":"ram"}' copies the roughly 26.8 GiB of n-gram tables (13.41 GiB per rank at TP=2) from the checkpoint’s safetensors once at boot. Disk mode — the io_uring DiskTable reading rows by byte offset on every forward pass — was retired on 2026-09-22 after the recurring asynchronous cudaErrorIllegalAddress class surfaced through its synchronize point on prefix-cache-resumed requests; that fault spans both PLE modes and both image cuts tested, so it is tracked as an open engine issue rather than a mode defect. The loader must still be instanttensor, and the published model directory must stay local and readable for every load and restart.
Two environment controls need care. Never set CUDA_LAUNCH_BLOCKING: b12x does not function with synchronous kernel launches. VLLM_QWEN3_8_PREFILL_COALESCE is opt-in and functional only from the current cycle pin, which wires the prefill checkpoint blocks into the NVIDIA GDN decoder; on earlier cuts that already carry the coalesce feature, setting it aborts boot with “all mamba groups must share cache scheduling parameters”.
The context comes from a YaRN override. The checkpoint declares max_position_embeddings 262,144 with rope_type default, so the 1,048,576-token window is a 4x override applied through --hf-overrides. The override must repeat the scaled length in max_position_embeddings as well as carrying the rope parameters: _get_and_verify_max_len multiplies the derived length by the rope factor only for rope types outside a list that includes yarn, because Transformers’ own _compute_yarn_parameters treats max_position_embeddings as already scaled. Omit that key and the engine stops at ModelConfig with “User-specified max_model_len (1048576) is greater than the derived max_model_len”. VLLM_ALLOW_LONG_MAX_MODEL_LEN is deliberately unset: it would downgrade that error to a warning instead of declaring the length the rope code assumes. The override also keeps mrope_section and mrope_interleaved, because get_rope() only preserves M-RoPE, which the vision position ids need, when mrope_section is present. Serving this at 1M tokens also requires vLLM 5a70cefff7e0 or later: earlier revisions did not pass the override into the MTP draft model, so the draft stayed at 262,144 and rejected a full-length prefill.
This is the shape I run twice. The live LeaderWorkerSet carries replicas: 2, so two independent TP=2 groups share one InferencePool and the endpoint picker chooses between two rank-zero APIs. The generated manifest stays at one group - that is the smallest deployment this recipe can claim to have validated end to end - but it emits the flags prefix-aware routing needs regardless: --enable-scale-out and --kv-events-config on rank zero only, and the kv-events 5556 / kv-replay 5559 container ports on both. How inference request routing works covers what the router does with those events and why publishing from the headless rank would silently break prefix accounting.
This is the LWS configuration I run. The Docker target uses the same engine definition, but I have not run it on two separate hosts.
The deployment benchmark guide explains how to run the benchmark against another deployment.
Measured performance
Section titled “Measured performance”Method: llm_decode_bench.py from local-inference-lab/llm-inference-bench 0.7.5, against the group’s rank-zero endpoint on the model port directly, not through the Gateway. Thirty-second duration-mode cells, one sample per cell, temperature 1, maximum 8,192 output tokens, integrated prefill scouts, loop detection on. The measurement ran from a separate x86-64 workstation, so no hardware panel was collected, and burst/end-to-end cells were not requested.
Conditions: the run was taken on one TP=2 group confirmed idle before and after (zero running, zero waiting); the other group was left serving. Configuration: GitOps revision 818be2ad4, image version above. The padded-draft columns are the same benchmark on the other group, running the previous image (vLLM b607be360037, B12X b600bf26c8e4, MTP experts padded 320 to 384) without adaptive spin; both images and both node pairs differ between the columns.
Capacity: every cell held its concurrency with nothing queued - the C=8 cells report eight running requests and zero waiting. The recipe sets –max-num-seqs 8, so C=8 exactly fills the scheduler’s sequence slots. Aggregate figures are duration-weighted per-request decode, not total server throughput, because concurrent generation intervals overlap.
Speculative normalisation: steps per second is decode throughput divided by the accept length the engine measured for that cell (accepted draft tokens plus the bonus token, against a nominal four tokens per step). It removes acceptance-rate variation, which moves between runs at temperature 1: aggregate decode at C=1 with empty context read 45.1 tok/s at accept length 1.93 on the padded-draft run and 52.5 at 2.24 here, while steps per second read 23.4 on both.
Run-to-run spread: engine steps per second agree with the padded-draft run within three percent in every decode cell, and prefill agrees within one percent up to 64K. That matches the kernel measurement: unpadded MTP experts save nothing at the one to four tokens the draft processes per step. The 128K prefill cell varies between runs - 2,416, 1,873 and 2,369 tok/s across three single samples - so read it as a range, not a quotable number.
Not measured: prefix-cache reuse through the router, gateway-path overhead, mixed workloads, and any figure for a cluster other than the one described above. These are observations of this stack on GB10, not hardware guarantees.
Requests enter at rank zero and collectives cross the secondary network; the shape is the same for every recipe here and is described in traffic paths in a two-node TP=2 group.
Manifest generator
Choose a target and supply the details that belong to your hosts or cluster. Values stay in this page and are never sent to a service.
Loading recipe definition...
No manifests generated
Complete the required fields, then select Generate manifests.
Before applying output
Section titled “Before applying output”- Each GPU node needs at least 130 GiB free on the model path. The checkpoint’s tensor payload is about 99 GiB; ram mode reads it once at boot, so the published root must stay local and readable for every load and restart.
- For Kubernetes, complete the cluster prerequisite checklist and confirm the NVIDIA runtime, LWS controller, CNI, Multus, and RDMA allocator are ready.
- For Docker, synchronize the model on both hosts before starting either serving container. Start rank zero first, then let the worker’s rendezvous check succeed before its headless server starts.
- Stop a GB10 serving container gracefully with a 120-second timeout. Long kernel execution is not grounds for force-killing a GPU process.
- A digest-qualified latest image is newer, not equivalent to the validated recipe image. Selecting it changes the validation status shown by the generator.