VLLMB12X runtime configuration
Generated by scripts/vllmb12x-runtime-config.py. Do not edit by hand.
Compared Sources
Section titled “Compared Sources”- Stock vLLM merge-base:
85a78f57a0e4d825f19b9cff243068d9a3aac7b2. - Locked randomvariable/vLLM integration before profile patches:
5dd5bd5dde76fd1ab4c0458ff1f72f85fa86e812. - Locked B12X integration:
9135f93ae533ac23db1c66aa8e1279ebb990c338. - The build applies
third_party/vllm_flash_attn_cute_namespace.patchto the locked vLLM tree.
This inventory contains runtime controls added or whose default or accepted values changed relative to the stock vLLM merge-base, plus every B12X environment control consumed by the locked B12X source. Descriptions come from source help text, field docstrings, source comments, or the checked-in description mixin when the source has no prose. Static inspection records each environment value as a string because the source performs its own parsing. unset means the cited read has no source default. source expression identifies an indirect control name or nonliteral source default.
For regular vLLM configuration, see the vLLM configuration reference.
Controls
Section titled “Controls”| Setting | Description | Accepted / source |
|---|---|---|
B12X_AUTOTUNEenvironmentdefault: '1' | Enable B12X kernel autotuning during preparation. | environment stringb12x/preparation/session.py:291 |
B12X_AUTOTUNE_EXHAUSTIVEenvironmentdefault: '0' | Extend B12X autotuning to the exhaustive search space. | environment stringb12x/preparation/_efficiency.py:6 |
B12X_COLLECTIVE_BARRIER_TIMEOUTenvironmentdefault: '120' | Maximum preparation time in seconds for cross-rank collective readiness. | environment stringb12x/preparation/session.py:121 |
B12X_COMPILE_CACHE_DIRenvironmentdefault: unset | Enable or size B12X generated-kernel compilation caches. | environment stringb12x/_lib/compiler.py:681 |
B12X_COMPILE_DISK_CACHEenvironmentdefault: '1' | Enable or size B12X generated-kernel compilation caches. | environment stringb12x/_lib/compiler.py:666 |
B12X_COMPILE_MEMORY_CACHEenvironmentdefault: '1' | Enable or size B12X generated-kernel compilation caches. | environment stringb12x/_lib/compiler.py:653 |
B12X_COMPILE_MEMORY_CACHE_SIZEenvironmentdefault: '1024' | Enable or size B12X generated-kernel compilation caches. | environment stringb12x/_lib/compiler.py:658 |
B12X_COMPILE_SPEC_MEMOenvironmentdefault: '1' | Enable or size B12X generated-kernel compilation caches. | environment stringb12x/_lib/compiler.py:126 |
B12X_COMPILE_WORKERSenvironmentdefault: '8' | Enable or size B12X generated-kernel compilation caches. | environment stringb12x/preparation/session.py:268 |
B12X_DENSE_ATOM_24environmentdefault: '0' | Select a B12X dense GEMM specialization or its workspace behavior. | environment stringb12x/_lib/dense_gemm.py:150 |
B12X_DENSE_FUSED_QUANTenvironmentdefault: '0' | MX-FP6 fused activation quantization: fuse BF16 activation quantization into the GEMM's DMA producer prologue, eliminating the separate quant kernel and the HBM round-trip for activation codes+scales. Currently m=1 only (decode hot path). The GEMM's producer warp does a full-row amax scan, derives gs/alpha, then quantizes each K-tile's 32-element blocks directly into sA/sSFA smem. Distinct from the MXFP8 fused_quant_a machinery. | environment stringb12x/_lib/dense_gemm.py:167 |
B12X_DENSE_PERSISTENT_SCRATCHenvironmentdefault: '1' | Select a B12X dense GEMM specialization or its workspace behavior. | environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:56 |
B12X_DENSE_PER_ROW_GSenvironmentdefault: '1' | Persistent stream-local decode-quantization workspaces. The quantizer writes them before the GEMM consumes them on the same stream; GEMM outputs are not pooled because they escape the linear call. | environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:52 |
B12X_DENSE_PER_ROW_IN_KERNELenvironmentdefault: '1' | The fallback host chain and in-kernel per-row scaling are bit-identical. | environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:60 |
B12X_DENSE_SPLITK_TURBOenvironmentdefault: '1' | Select a B12X dense GEMM specialization or its workspace behavior. | environment stringb12x/_lib/dense_gemm.py:145 |
B12X_DIRECT_CUTE_OPTIONSenvironmentdefault: '' | Pass direct CuteDSL compiler options to the B12X MoE path. | environment stringb12x/moe/fused_moe/_impl.py:10000 |
B12X_DISABLE_BF16_GEMVenvironmentdefault: '' | Disable the B12X BF16 GEMV fast path when set. | environment stringb12x/gemm/bf16_gemv/api.py:29 |
B12X_DISK_BACKENDenvironmentdefault: 'io_uring' | Select the disk-backed PLE row-staging backend: io_uring or GDS. | environment stringb12x/sequence/_shared/disk_table.py:166 |
B12X_DSA_VALIDATE_PAGE_IDSenvironmentdefault: '0' | Validate DSA page identifiers in the B12X indexer path. | environment stringb12x/attention/dsa_indexer/_impl.py:56 |
B12X_DYNAMIC_SWAP_ABenvironmentdefault: unset | Select the dynamic A/B swap policy for B12X fused MoE kernels. | environment stringb12x/moe/fused_moe/_impl.py:2151 |
B12X_DYNAMIC_TILE_MNenvironmentdefault: unset | Select the dynamic M/N tile policy for B12X fused MoE kernels. | environment stringb12x/moe/fused_moe/_impl.py:1691 |
B12X_ENABLE_FP6_MICROenvironmentdefault: '0' | Enable the B12X FP6 micro-kernel path. | environment stringb12x/moe/_shared/kernels/micro.py:127 |
B12X_FP6_ACT_FMT_OVERRIDESenvironmentdefault: '' | Override FP6 activation formats selected by the B12X checkpoint. | environment stringb12x/quantization/mxfp6/fp6_checkpoint.py:248 |
B12X_FUSED_INDEXERenvironmentdefault: '1' | Enable the fused B12X sparse indexer path. | environment stringb12x/attention/dsa_indexer/_tuning.py:133 |
B12X_INDEXER_DIRECT_Kenvironmentdefault: '1' | Triage kill-switch: B12X_INDEXER_DIRECT_K=0 forces every variant back to the staged pipeline (which also restores the v2 cache keys), so serving can A/B the direct-L2 score in place without a checkout change. | environment stringb12x/attention/dsa_indexer/fused_indexer.py:114 |
B12X_LOG_CUTE_COMPILESenvironmentdefault: '' | Enable B12X CuteDSL compilation progress logging. | environment stringb12x/_lib/compiler.py:691 |
B12X_LOG_CUTE_COMPILE_ARGSenvironmentdefault: '' | Log arguments passed to B12X CuteDSL compilation. | environment stringb12x/_lib/compiler.py:1217 |
B12X_LOG_CUTE_COMPILE_STACKenvironmentdefault: '' | Log the Python stack for B12X CuteDSL compilation. | environment stringb12x/_lib/compiler.py:711 |
B12X_LOG_CUTE_COMPILE_STACK_DEPTHenvironmentdefault: '' | Set the maximum stack depth in B12X CuteDSL compile logs. | environment stringb12x/_lib/compiler.py:722 |
B12X_MHC_PDLenvironmentdefault: '0' | Enable programmatic dependent launch for MHC normalization kernels. | environment stringb12x/norm/mhc/_kernels.py:158 |
B12X_MHC_PREFILL_FINALIZE_THREADSenvironmentdefault: '256' | Set the CUDA thread count for the MHC prefill finalize kernel. | environment stringb12x/norm/mhc/_kernels.py:305 |
B12X_MHC_PREFILL_GRAM_THREADSenvironmentdefault: os.getenv('B12X_MHC_PREFILL_THREADS', '1024') | Set the CUDA thread count for the MHC Gram-trick prefill path. | environment stringb12x/norm/mhc/_kernels.py:166 |
B12X_MHC_PREFILL_TF32_TMA_CHUNK_MIN_TOKENSenvironmentdefault: '4096' | Set the token threshold that selects the MHC TF32 TMA chunk specialization. | environment stringb12x/norm/mhc/_kernels.py:227 |
B12X_MHC_PREFILL_TF32_TMA_CHUNK_M_WARPSenvironmentdefault: '12' | Set the M-dimension compute warp count for the MHC TF32 TMA chunk specialization. | environment stringb12x/norm/mhc/_kernels.py:237 |
B12X_MHC_PREFILL_TF32_TMA_CHUNK_N_WARPSenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TF32_TMA_N_WARPS', '1') | Set the N-dimension compute warp count for the MHC TF32 TMA chunk specialization. | environment stringb12x/norm/mhc/_kernels.py:243 |
B12X_MHC_PREFILL_TF32_TMA_CHUNK_STAGESenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TF32_TMA_STAGES', '2') | Set the asynchronous pipeline stage count for the MHC TF32 TMA chunk specialization. | environment stringb12x/norm/mhc/_kernels.py:279 |
B12X_MHC_PREFILL_TF32_TMA_CHUNK_TILE_Kenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TF32_TMA_TILE_K', '64') | Set the K tile dimension for the MHC TF32 TMA chunk specialization. | environment stringb12x/norm/mhc/_kernels.py:273 |
B12X_MHC_PREFILL_TF32_TMA_CHUNK_TILE_Menvironmentdefault: '192' | Set the M tile dimension for the MHC TF32 TMA chunk specialization. | environment stringb12x/norm/mhc/_kernels.py:255 |
B12X_MHC_PREFILL_TF32_TMA_CHUNK_TILE_Nenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TF32_TMA_TILE_N', '24') | Set the N tile dimension for the MHC TF32 TMA chunk specialization. | environment stringb12x/norm/mhc/_kernels.py:261 |
B12X_MHC_PREFILL_TF32_TMA_LONG_MIN_TOKENSenvironmentdefault: '8192' | Set the token threshold that selects the MHC TF32 TMA long-context specialization. | environment stringb12x/norm/mhc/_kernels.py:230 |
B12X_MHC_PREFILL_TF32_TMA_LONG_M_WARPSenvironmentdefault: '8' | Set the M-dimension compute warp count for the MHC TF32 TMA long-context specialization. | environment stringb12x/norm/mhc/_kernels.py:287 |
B12X_MHC_PREFILL_TF32_TMA_LONG_N_WARPSenvironmentdefault: '1' | Set the N-dimension compute warp count for the MHC TF32 TMA long-context specialization. | environment stringb12x/norm/mhc/_kernels.py:290 |
B12X_MHC_PREFILL_TF32_TMA_LONG_STAGESenvironmentdefault: '2' | Set the asynchronous pipeline stage count for the MHC TF32 TMA long-context specialization. | environment stringb12x/norm/mhc/_kernels.py:302 |
B12X_MHC_PREFILL_TF32_TMA_LONG_TILE_Kenvironmentdefault: '64' | Set the K tile dimension for the MHC TF32 TMA long-context specialization. | environment stringb12x/norm/mhc/_kernels.py:299 |
B12X_MHC_PREFILL_TF32_TMA_LONG_TILE_Menvironmentdefault: '128' | Set the M tile dimension for the MHC TF32 TMA long-context specialization. | environment stringb12x/norm/mhc/_kernels.py:293 |
B12X_MHC_PREFILL_TF32_TMA_LONG_TILE_Nenvironmentdefault: '24' | Set the N tile dimension for the MHC TF32 TMA long-context specialization. | environment stringb12x/norm/mhc/_kernels.py:296 |
B12X_MHC_PREFILL_TF32_TMA_M_WARPSenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TF32_TMA_WARPS', os.getenv('B12X_MHC_PREFILL_TMA_WARPS', '1')) | Set the M-dimension compute warp count for the MHC TF32 TMA prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:192 |
B12X_MHC_PREFILL_TF32_TMA_N_WARPSenvironmentdefault: '1' | Set the N-dimension compute warp count for the MHC TF32 TMA prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:201 |
B12X_MHC_PREFILL_TF32_TMA_STAGESenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TMA_STAGES', '1') | Set the asynchronous pipeline stage count for the MHC TF32 TMA prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:221 |
B12X_MHC_PREFILL_TF32_TMA_TILE_Kenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TMA_TILE_K', '256') | Set the K tile dimension for the MHC TF32 TMA prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:215 |
B12X_MHC_PREFILL_TF32_TMA_TILE_Menvironmentdefault: os.getenv('B12X_MHC_PREFILL_TMA_TILE_M', '16') | Set the M tile dimension for the MHC TF32 TMA prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:206 |
B12X_MHC_PREFILL_TF32_TMA_TILE_Nenvironmentdefault: '8' | Set the N tile dimension for the MHC TF32 TMA prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:212 |
B12X_MHC_PREFILL_TF32_TMA_WARPSenvironmentdefault: os.getenv('B12X_MHC_PREFILL_TMA_WARPS', '1') | Set the compute warp count for the MHC TF32 TMA prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:194 |
B12X_MHC_PREFILL_THREADSenvironmentdefault: '512' | Set the CUDA thread count for the MHC prefill kernel. | environment stringb12x/norm/mhc/_kernels.py:164 |
B12X_MHC_PREFILL_TMA_STAGESenvironmentdefault: '3' | Set the asynchronous pipeline stage count for the MHC prefill TMA kernel. | environment stringb12x/norm/mhc/_kernels.py:190 |
B12X_MHC_PREFILL_TMA_TILE_Kenvironmentdefault: '64' | Set the K tile dimension for the MHC prefill TMA kernel. | environment stringb12x/norm/mhc/_kernels.py:188 |
B12X_MHC_PREFILL_TMA_TILE_Menvironmentdefault: '128' | Set the M tile dimension for the MHC prefill TMA kernel. | environment stringb12x/norm/mhc/_kernels.py:182 |
B12X_MHC_PREFILL_TMA_TILE_Nenvironmentdefault: '16' | Set the N tile dimension for the MHC prefill TMA kernel. | environment stringb12x/norm/mhc/_kernels.py:185 |
B12X_MHC_PREFILL_TMA_WARPSenvironmentdefault: '8' | Set the compute warp count for the MHC prefill TMA kernel. | environment stringb12x/norm/mhc/_kernels.py:178 |
B12X_MICRO_DYNAMIC_CUTOVER_PAIRSenvironmentdefault: unset | Set the active expert-pair count at which micro-MoE changes to its dynamic execution path. | environment stringb12x/moe/fused_moe/_impl.py:2225 |
B12X_MICRO_REUSE_COMPILEDenvironmentdefault: '1' | Compile planning produces deferred programs that must never be memoized or returned in place of a compiled kernel. | environment stringb12x/moe/fused_moe/_impl.py:9980 |
B12X_MICRO_SHARE_INPUT_ACROSS_EXPERTSenvironmentdefault: '1' | Enable reuse of one micro-MoE input load across the selected experts. | environment stringb12x/moe/fused_moe/_impl.py:13318 |
B12X_MLA_SM120_PREFILL_MGenvironmentdefault: '1' | Enable the SM120 multi-group MLA prefill kernel. | environment stringb12x/attention/_shared/mla/prefill.py:398 |
B12X_MLA_SM120_PREFILL_PACK_HILO_ROWSenvironmentdefault: '1' | Pack high- and low-position MLA prefill rows into one SM120 kernel input. | environment stringb12x/attention/_shared/mla/prefill_mg.py:3948 |
B12X_MOE_TILE_MNenvironmentdefault: unset | Override the MxN MMA tile shape used by the B12X fused MoE kernel. | environment stringb12x/moe/fused_moe/_impl.py:9893 |
B12X_MSA_DECODE_PAGEMAXenvironmentdefault: '1' | Enable the page-maximum reduction used by the MSA decode indexer. | environment stringb12x/attention/dsa_indexer/_impl.py:885 |
B12X_NVFP4_SPLIT_DECODEenvironmentdefault: '0' | Split NVFP4 decode work into the materialized decode path instead of the monolithic path. | environment stringb12x/moe/fused_moe/_impl.py:12116 |
B12X_PACKED_B_EXPAND_AHEADenvironmentdefault: '1' | Expand-ahead for packed-B: at k_block 0 the MMA warps wait for stage s+1 and expand it in place, overlapping the expansion with ALL of stage s's MMA work instead of putting it on the critical path at the stage boundary. It requires at least four pipeline stages of producer slack. | environment stringb12x/_lib/dense_gemm.py:157 |
B12X_PACKED_B_MIN_Nenvironmentdefault: '12288' | Minimum out_features for which the GEMM streams the 3:4-packed weight directly (b_packed=True) instead of a cached 1-byte/code expansion. Default 12288 enables shapes with enough N tiles for the packed path. | environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:125 |
B12X_PAGED_EXTEND_QWEN_FP8_PV_REPACKenvironmentdefault: unset | Enable Qwen FP8 value-projection repacking during paged attention extension. | environment stringb12x/attention/paged/forward_extend_generic.py:3525 |
B12X_PAGED_KV_TMA_PLANE_SWIZZLEenvironmentdefault: '' | Set the TMA plane swizzle used to load paged KV data. | environment stringb12x/attention/paged/_forward.py:215 |
B12X_PAGED_MSAenvironmentdefault: unset | Enable MSA block-sparse paged attention. | environment stringb12x/attention/paged/_scratch.py:127 |
B12X_PAGED_MSA_UNION_PREFILLenvironmentdefault: '1' | Enable the MSA union-prefill path for paged attention. | environment stringb12x/attention/paged/_scratch.py:136 |
B12X_PCIE_ALLREDUCE_ALGORITHMenvironmentdefault: 'auto' | Select the B12X PCIe all-reduce algorithm: automatic, hierarchical, island reduce-scatter, or one-shot. | environment stringb12x/comm/pcie/pcie_allreduce.py:58 |
B12X_PCIE_DMA_A2A_CHUNKSenvironmentdefault: '0' | Override the number of chunks used by the PCIe DMA all-to-all path. | environment stringb12x/comm/pcie/pcie_dma.py:258 |
B12X_PCIE_DMA_FP8environmentdefault: '0' | Enable the FP8 PCIe DMA all-reduce path. | environment stringb12x/comm/pcie/pcie_dma.py:101 |
B12X_PCIE_DMA_PIECESenvironmentdefault: '0' | Override the number of pieces used by the PCIe DMA all-reduce path. | environment stringb12x/comm/pcie/pcie_dma.py:257 |
B12X_PCIE_FUSED_CTAS_PER_ROWenvironmentdefault: '0' | Recover the CTA count from the already-derived single-CTA/config formula without changing the launch contract. | environment stringb12x/comm/pcie/pcie_oneshot.py:1977 |
B12X_PCIE_FUSED_THREADSenvironmentdefault: '256' | Set the CUDA thread count for the fused PCIe collective kernel. | environment stringb12x/comm/pcie/pcie_oneshot.py:1655 |
B12X_PCIE_HIERARCHICAL_BF16X2environmentdefault: '1' | Enable vectorized BF16x2 loads in the hierarchical PCIe all-reduce kernel. | environment stringb12x/comm/pcie/pcie_hierarchical.py:67 |
B12X_PCIE_HIERARCHICAL_BF16X2_MAX_ELEMENTSenvironmentdefault: '7168' | Set the largest reduction size eligible for vectorized BF16x2 hierarchical PCIe all-reduce. | environment stringb12x/comm/pcie/pcie_hierarchical.py:74 |
B12X_PCIE_HIERARCHICAL_DEFERRED_CONSUMPTIONenvironmentdefault: '0' | Defer consumption of hierarchical PCIe all-reduce partial results. | environment stringb12x/comm/pcie/pcie_hierarchical.py:158 |
B12X_PCIE_HIERARCHICAL_DOUBLE_BUFFERenvironmentdefault: '0' | Enable double buffering of staging memory in hierarchical PCIe all-reduce. | environment stringb12x/comm/pcie/pcie_hierarchical.py:155 |
B12X_PCIE_HIERARCHICAL_NANOSLEEP_CYCLESenvironmentdefault: '24' | Set spin-wait nanosleep cycles for hierarchical PCIe all-reduce. | environment stringb12x/comm/pcie/pcie_hierarchical.py:36 |
B12X_PCIE_HIERARCHICAL_THREADSenvironmentdefault: '224' | Set the CUDA thread count for hierarchical PCIe all-reduce. | environment stringb12x/comm/pcie/pcie_hierarchical.py:51 |
B12X_PCIE_ISLAND_RS_NANOSLEEP_CYCLESenvironmentdefault: '24' | Set spin-wait nanosleep cycles for island reduce-scatter PCIe all-reduce. | environment stringb12x/comm/pcie/pcie_island_rs.py:50 |
B12X_PCIE_ISLAND_RS_THREADSenvironmentdefault: '512' | Default CUDA thread count for the TP16 equal-quarter launch. | environment stringb12x/comm/pcie/pcie_island_rs.py:40 |
B12X_PCIE_ONESHOT_BLOCK_LIMITenvironmentdefault: '8' | Set the maximum number of CUDA blocks used by one-shot PCIe all-reduce. | environment stringb12x/comm/pcie/pcie_oneshot.py:1526 |
B12X_PCIE_ONESHOT_PUSHenvironmentdefault: '0' | Enable the one-shot PCIe push transport, which writes each rank's input into every peer's eager slot. | environment stringb12x/comm/pcie/pcie_oneshot.py:128 |
B12X_PCIE_ONESHOT_THREADSenvironmentdefault: '256' | Set the CUDA thread count for one-shot PCIe all-reduce. | environment stringb12x/comm/pcie/pcie_oneshot.py:1524 |
B12X_PCIE_TEST_VISIBLE_SM_COUNTenvironmentdefault: '0' | Override the visible SM count for PCIe collective topology validation. | environment stringb12x/comm/pcie/pcie_oneshot.py:704 |
B12X_PCIE_TP2_PLAIN_REMOTE_PUSHenvironmentdefault: unset | Enable the qualified plain TP2 remote-write PCIe transport. | environment stringb12x/comm/pcie/pcie_oneshot.py:145 |
B12X_PCIE_TP2_REMOTE_PUSHenvironmentdefault: '0' | Enable the qualified fused TP2 remote-write PCIe transport. | environment stringb12x/comm/pcie/pcie_oneshot.py:134 |
B12X_PCIE_TP4_REMOTE_PUSHenvironmentdefault: '0' | Enable the qualified fused TP4 remote-write PCIe transport. | environment stringb12x/comm/pcie/pcie_oneshot.py:154 |
B12X_PCIE_TP8_OWNER_REDUCEenvironmentdefault: '1' | Enable the topology-scoped fused TP8 owner-reduce PCIe transport. | environment stringb12x/comm/pcie/pcie_oneshot.py:160 |
B12X_PCIE_VOCAB_ARGMAX_NANOSLEEP_CYCLESenvironmentdefault: '24' | Set spin-wait nanosleep cycles for PCIe vocabulary argmax reduction. | environment stringb12x/comm/pcie/pcie_vocab_argmax.py:71 |
B12X_PREPARATION_TRACE_DIRenvironmentdefault: unset | Write B12X preparation timing traces to this directory. | environment stringb12x/preparation/_timing.py:14 |
B12X_PROBE_COMMANDenvironmentdefault: '' | Configure B12X PCIe overlap probe diagnostics. | environment stringb12x/comm/pcie/overlap_probe.py:1190 |
B12X_PROBE_SOURCE_REVISIONenvironmentdefault: '' | Configure B12X PCIe overlap probe diagnostics. | environment stringb12x/comm/pcie/overlap_probe.py:1191 |
B12X_PROBE_SOURCE_WORKTREEenvironmentdefault: '' | Configure B12X PCIe overlap probe diagnostics. | environment stringb12x/comm/pcie/overlap_probe.py:1192 |
B12X_ROCE_CACHE_DIRenvironmentdefault: unset | Directory for the compiled B12X RoCE proxy and its source-hash cache. | environment stringb12x/comm/roce/_proxy.py:44 |
B12X_SQG_XOR_CHEB_T12_SMEMenvironmentdefault: '1' | Enable shared-memory storage for the SQG XOR-Chebyshev T12 kernel path. | environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:172 |
B12X_TIMINGenvironmentdefault: '0' | Enable B12X kernel timing output. | environment stringb12x/_lib/dense_gemm.py:137 |
B12X_TIMING_THRESHOLD_MSenvironmentdefault: os.getenv('VLLM_B12X_TIMING_THRESHOLD_MS', '0') | Set the B12X kernel timing output threshold in milliseconds. | environment stringb12x/_lib/dense_gemm.py:140 |
B12X_TUNING_CACHE_VERSIONenvironmentdefault: '1' | Version the B12X preparation tuning-cache format. | environment stringb12x/preparation/_cache.py:28 |
B12X_VALIDATE_PAGED_INDEXER_CUDA_VALUESenvironmentdefault: '0' | Enable B12X validation checks for runtime values. | environment stringb12x/attention/dsa_indexer/paged.py:400 |
B12X_W4A16_SMALL_M_DIRECTenvironmentdefault: '1' | Use the direct small-M W4A16 MoE kernel instead of its fallback path. | environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:9093 |
B12X_W4A16_SMALL_M_HOST_BARRIER_RESETenvironmentdefault: '1' | Reset the small-M W4A16 MoE host barrier between launches. | environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:9132 |
B12X_W4A16_SMALL_M_SPLITKenvironmentdefault: '0' | Enable split-K decomposition for the small-M W4A16 MoE kernel. | environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:162 |
B12X_W4A8_TINY_DECODEenvironmentdefault: '1' | Default on; B12X_W4A8_TINY_DECODE=0 is the kill switch. | environment stringb12x/moe/fused_moe/_impl.py:12227 |
B12X_WO_B_FUSED_TILEenvironmentdefault: '' | Environment variable read by the cited source. | environment stringb12x/gemm/_shared/wo_mxfp8.py:2808 |
B12X_WO_QUANT_CHUNKS_PER_PROGRAMenvironmentdefault: '16' | Environment variable read by the cited source. | environment stringb12x/gemm/_shared/wo_mxfp8.py:49 |
VLLM_B12X_TIMINGenvironmentdefault: '0' | Configure a local-inference-lab vLLM B12X execution path. | environment stringb12x/_lib/dense_gemm.py:137 |
VLLM_B12X_TIMING_THRESHOLD_MSenvironmentdefault: '0' | Configure a local-inference-lab vLLM B12X execution path. | environment stringb12x/_lib/dense_gemm.py:142 |
VLLM_EXPERIMENTAL_SHM_BROADCAST_ADAPTIVE_SPINenvironmentdefault: "0" | Experimental opt-in for adaptive shm_broadcast reader and writer waits. The default retains the fixed one-second reader grace and writer yields. | environment stringvllm/envs.py:768 |
VLLM_QWEN3_8_HC_PREFILL_MODEenvironmentdefault: 'off' | HC prefill row ownership: off, control, or shard. Opt-in pending serving qualification; ranks must agree at startup. | environment stringvllm/envs.py:1754 |
VLLM_QWEN3_8_PREFILL_COALESCEenvironmentdefault: "0" | Select a Qwen3.8 Flash Next execution path. | environment stringvllm/envs.py:1757 |
VLLM_SHM_BROADCAST_ADAPTIVE_ALPHAenvironmentdefault: "0.25" | EMA coefficient for observed inter-read intervals, in (0, 1]: higher adapts to cadence changes faster but tracks noise. | environment stringvllm/envs.py:799 |
VLLM_SHM_BROADCAST_ADAPTIVE_BUDGET_MSenvironmentdefault: "1.0" | Tunables for the experimental adaptive shm_broadcast reader spin grace. Readers busy-loop for a grace period after the last read before parking on the poller; the adaptive policy derives that grace from an EMA of observed inter-read intervals (T_ema): grace = clamp(B * B / T_ema, MIN_GRACE, MAX_GRACE) B (BUDGET) is the pivot: at T_ema == B the grace equals B; faster traffic spins toward MAX_GRACE (arrivals near-certain within the grace), slower traffic parks within MIN_GRACE. The grace is monotonically decreasing in T_ema, so slower traffic never increases spin, and it is bounded regardless of estimate staleness. Passing a float to SpinCondition's busy_loop_s overrides the policy with a fixed grace. All grace values are in milliseconds. Time-scale pivot B: grace == B when T_ema == B. | environment stringvllm/envs.py:784 |
VLLM_SHM_BROADCAST_ADAPTIVE_MAX_GRACE_MSenvironmentdefault: "2.0" | Upper bound on the spin grace: the most time a reader busy-loops even under the fastest observed traffic. | environment stringvllm/envs.py:794 |
VLLM_SHM_BROADCAST_ADAPTIVE_MIN_GRACE_MSenvironmentdefault: "0.05" | Lower bound on the spin grace: an idle reader parks within this window instead of spinning. | environment stringvllm/envs.py:789 |
VLLM_SHM_BROADCAST_WRITE_PARK_MAX_MSenvironmentdefault: "1.0" | Ceiling for the writer's park step once its own grace expires: the writer sleeps in doubling steps from 50us up to this bound while waiting for the slowest reader to release a block. Raising it trades write latency for fewer wakeups on a writer blocked by slow readers. | environment stringvllm/envs.py:806 |
Semantic Extensions
Section titled “Semantic Extensions”--hf-overridesretains dictionary YaRN and rope values when vLLM builds an in-model MTP draftModelConfig(local-inference-lab/vLLM #777).VLLM_QWEN3_8_PREFILL_COALESCErequires B12X #386’s prepared PLE checkpoint export and vLLM 5dd5bd5dde76 or later, which wires the prefill checkpoint blocks into the NVIDIA GDN decoder. On earlier cuts that carry the coalesce feature from 61f93c53ee, setting it aborts boot with “all mamba groups must share cache scheduling parameters”.- Never set
CUDA_LAUNCH_BLOCKING. b12x does not function with synchronous kernel launches, so the variable is not a usable debug lever on this stack. - The adaptive shared-memory controls alter wait behavior only and are excluded from vLLM compile factors. They do not select kernels or invalidate compiled graphs.
Generator Identity
Section titled “Generator Identity”16e5efc2ab3647ea305bef521710ab1344365132ceff76a827c90cdc168e29b2