Latest image publication
Every image on this page packages a pinned revision of local-inference-lab/vLLM and its DGX Spark dependency set. These are not generic upstream vLLM images.
Release channel
The trusted Pages build resolves the public repository’s latest tag to a multiarchitecture index digest, finds the matching immutable publication tag, and records the resolution time. Resolution fails unless the index carries exactly the linux/arm64 and linux/amd64 manifests. The recipe generator consumes that static record and always emits a digest-qualified reference.
The selected recipe’s validated digest remains the default. Choosing Latest published image displays an explicit warning because publication does not prove that the new image completed the recipe’s two-node runtime acceptance.
Published metadata
Resolved immutable image
- Immutable tag
- Loading...
- Digest reference
- Loading...
- Resolved at
- Loading...
- Recipe status
- Not inherited
What each published lineage carries
Section titled “What each published lineage carries”A publication pins one vLLM revision and one b12x revision together. The pairing matters: the fork’s vLLM branches import source-owned b12x modules, so a stale b12x pin fails at model load rather than at build time.
Qwen3.8 Flash Next lineage
Section titled “Qwen3.8 Flash Next lineage”sha256:6391d8c2f20c3fd476606f28a45a92d6cbc0ae20486ef84769fc759046f244d3 · 16 September 2026
- vLLM
5a70cefff7e011d909cb21a576333652ac2c05d9, b12x759e98370230e8488ac27482acf16feacb6a10dd. - The vLLM revision passes the target model’s positional overrides into the MTP draft model. Before it, the draft resolved 262,144 positions from the checkpoint while the target resolved 1,048,576, and a full-length prefill was rejected.
- Restores the per-rank provider census and the symmetric-preparation guard, so a rank that fails to publish its providers raises a structured error instead of stalling the group.
- Builder side: console scripts are now materialized into
/opt/venv/bin, which is what makesninjapresent for runtime C++ extension builds; the image contract and publisher accept a pinned commit rather than a branch name. - Pull requests: vllm-multiarch-oci#25 for the builder and recipe, local-inference-lab/vllm#777 for the MTP override fix upstream.
Qwen3.8 Flash Next with MXFP8 MTP experts
Section titled “Qwen3.8 Flash Next with MXFP8 MTP experts”sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde · 1 October 2026 · v20261001.1
- vLLM
0a6739845a3249b07a30ad9b1e720e5e3fb6236f(dev/rv-mxfp8-mtp), b12x965e748f74fe9dac0a35c497a0947c5b108617c5(feat/mxfp8-moe). First public release whose source pair reachesmain. - Native block-scaled MXFP8 W8A8 MoE on both architectures: Qwen3.8 Flash Next step 5500’s MTP draft experts are MXFP8 and run on
B12X_MXFP8at their unpadded 320-wide TP=2 shard; the NVFP4 target experts run onB12X. Nothing selects Marlin. - Runtime dependencies move to current CUDA-era pins (nvidia-cutlass-dsl 4.8.0, tilelang 0.1.15, cuda-python 13.4.1, nixl, numba, PyNvVideoCodec, fastokens); locks target manylinux_2_34.
- Serving acceptance and the idle benchmark are recorded in the release notes and the Qwen recipe, which pins this digest as its validated image.
- Pull requests: vllm-multiarch-oci#36 for the source pair, dependency bump and docs, #37 for the per-architecture base-creation label the first release attempt tripped.
- Current state: both open recipes’
validation.imagedigests are public releases again.
DeepSeek V4 Flash Vision lineage
Section titled “DeepSeek V4 Flash Vision lineage”sha256:42d5d72bd79a371c3b5a756fe2a1814f3c9393dd9a73d65bcefa01eb8f28d726
This is the image the DeepSeek recipe was accepted on. It remains the recipe’s default; a newer publication is a different lineage, not an upgrade of this one.
Architecture support
Section titled “Architecture support”Every publication is one OCI index with a native linux/arm64 and a native linux/amd64 manifest. Both manifests are built and contract-tested on every publication pipeline run; the publisher rejects an index that is missing either architecture.
- ARM64 (GB10, DGX Spark) is the production architecture. The two-node recipes and their runtime acceptance evidence all run here.
- x86-64 (RTX 5090, SM120) is built from the same pins and passes the same image contract. A single-GPU serving smoke measured 46.5 tokens/s single-stream on an RTX 5090 under WSL2. Two x86-64 specifics from that smoke: invoke the entrypoint through the
vllm servesubcommand, because a bare positional model path fails in this image, and set a small--max-num-seqs(8) on a 32 GiB card, because the default 256 exceeds the Mamba cache budget and leaves KV space for roughly 52K tokens at 0.90 utilization. Distributed topologies on x86-64 are untested.