Included upstream changes
Published images are built from pinned fork revisions of local-inference-lab/vLLM and local-inference-lab/B12X, not from their upstream branches. A revision alone does not say which proposed changes it carries, so this page lists them.
A change reaches the build in one of two ways. It is either already merged into the pinned fork revision, or it is applied to the pinned source at build time from a patch in this repository. Both are recorded below, and the same list is written into the image’s org.opencontainers.image.description label.
Pinned at 0a6739845a32, branched from local-inference-lab/vllm integration/karmic-kraken-beta.
- Native MXFP8 MTP through the ModelOpt B12X backend — no upstream pull request
- Merged into the pinned revision (4 commits).
- User-directed feature-branch pin. Preserves E4M3 weights and UE8M0 scales through the paired B12X W8A8 recipe and prepares the step5500 MTP experts at their unpadded TP=2 intermediate (320). Bounded SM120 numerical, ModelOpt lifecycle (I=640 and I=320) and graph tests pass. No upstream PR is open.
- Partial port of vLLM PR 779: Qwen HC ownership and checkpoint coalescing — no upstream pull request
- Merged into the pinned revision (9 commits).
- Carries the HC/coalescing runtime port to Qwen4Exp, not the complete upstream PR or an upstream beta merge. HC and coalescing remain opt-in. GPU serving was not exercised in this rebase.
- local-inference-lab/vllm#800 — Bounded shared-memory broadcast waits
- Merged into the pinned revision (1 commits).
- Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute — no upstream pull request
- Applied at build time by
third_party/vllm_flash_attn_cute_namespace.patch. - Packaging repair for this builder’s symlink install path, which skips the import rewrite the upstream CMake copy performs. It carries no upstream change.
- Applied at build time by
Pinned at 965e748f74fe, branched from local-inference-lab/b12x integration/karmic-kraken-beta.
- Native block-scaled MXFP8 W8A8 MoE execution — no upstream pull request
- Merged into the pinned revision (3 commits).
- User-directed feature-branch pin. Implements lossless E4M3/UE8M0 K32 preparation and grouped expert execution, including 32-aligned intermediate sizes through separate up/gate descriptors. Numerical (I=640/320/96), CUDA graph replay and bounded memcheck pass on SM120 with sm_120f artifacts. SM120 MoE kernel time at I=320 unpadded vs I=384 padded: equal at 1-4 tokens, 6-10% lower at 16-64 tokens. No SM121 runtime or serving-performance claim. No upstream PR is open.
- PLE checkpoint exports, prepared launcher closures and collective entry barriers — no upstream pull request
- Merged into the pinned revision (11 commits).
- Fork-only changes remain outside upstream beta, including the PLE export_checkpoint API from B12X PR 386 required by Qwen coalescing.
Keeping this list honest
Section titled “Keeping this list honest”scripts/vllmb12x-included-changes.py --check fails when this page, the README table, or the image description no longer matches the ledger, and when the recorded patch series no longer matches the patches Bazel applies.
scripts/vllmb12x-included-changes.py --verify --clone-root <directory> goes further and checks the ledger against the fork history: a pull request is integrated by replaying its commits, so the command asserts that every commit of every listed pull request is present in the pinned revision, and that the recorded commit count still matches.