Skip to content

VLLMB12X runtime configuration

Generated by scripts/vllmb12x-runtime-config.py. Do not edit by hand.

  • Stock vLLM merge-base: 85a78f57a0e4d825f19b9cff243068d9a3aac7b2.
  • Locked randomvariable/vLLM integration before profile patches: 5dd5bd5dde76fd1ab4c0458ff1f72f85fa86e812.
  • Locked B12X integration: 9135f93ae533ac23db1c66aa8e1279ebb990c338.
  • The build applies third_party/vllm_flash_attn_cute_namespace.patch to the locked vLLM tree.

This inventory contains runtime controls added or whose default or accepted values changed relative to the stock vLLM merge-base, plus every B12X environment control consumed by the locked B12X source. Descriptions come from source help text, field docstrings, source comments, or the checked-in description mixin when the source has no prose. Static inspection records each environment value as a string because the source performs its own parsing. unset means the cited read has no source default. source expression identifies an indirect control name or nonliteral source default. For regular vLLM configuration, see the vLLM configuration reference.

SettingDescriptionAccepted / source
B12X_AUTOTUNEenvironment
default: '1'
Enable B12X kernel autotuning during preparation.environment stringb12x/preparation/session.py:291
B12X_AUTOTUNE_EXHAUSTIVEenvironment
default: '0'
Extend B12X autotuning to the exhaustive search space.environment stringb12x/preparation/_efficiency.py:6
B12X_COLLECTIVE_BARRIER_TIMEOUTenvironment
default: '120'
Maximum preparation time in seconds for cross-rank collective readiness.environment stringb12x/preparation/session.py:121
B12X_COMPILE_CACHE_DIRenvironment
default: unset
Enable or size B12X generated-kernel compilation caches.environment stringb12x/_lib/compiler.py:681
B12X_COMPILE_DISK_CACHEenvironment
default: '1'
Enable or size B12X generated-kernel compilation caches.environment stringb12x/_lib/compiler.py:666
B12X_COMPILE_MEMORY_CACHEenvironment
default: '1'
Enable or size B12X generated-kernel compilation caches.environment stringb12x/_lib/compiler.py:653
B12X_COMPILE_MEMORY_CACHE_SIZEenvironment
default: '1024'
Enable or size B12X generated-kernel compilation caches.environment stringb12x/_lib/compiler.py:658
B12X_COMPILE_SPEC_MEMOenvironment
default: '1'
Enable or size B12X generated-kernel compilation caches.environment stringb12x/_lib/compiler.py:126
B12X_COMPILE_WORKERSenvironment
default: '8'
Enable or size B12X generated-kernel compilation caches.environment stringb12x/preparation/session.py:268
B12X_DENSE_ATOM_24environment
default: '0'
Select a B12X dense GEMM specialization or its workspace behavior.environment stringb12x/_lib/dense_gemm.py:150
B12X_DENSE_FUSED_QUANTenvironment
default: '0'
MX-FP6 fused activation quantization: fuse BF16 activation quantization into the GEMM's DMA producer prologue, eliminating the separate quant kernel and the HBM round-trip for activation codes+scales. Currently m=1 only (decode hot path). The GEMM's producer warp does a full-row amax scan, derives gs/alpha, then quantizes each K-tile's 32-element blocks directly into sA/sSFA smem. Distinct from the MXFP8 fused_quant_a machinery.environment stringb12x/_lib/dense_gemm.py:167
B12X_DENSE_PERSISTENT_SCRATCHenvironment
default: '1'
Select a B12X dense GEMM specialization or its workspace behavior.environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:56
B12X_DENSE_PER_ROW_GSenvironment
default: '1'
Persistent stream-local decode-quantization workspaces. The quantizer writes them before the GEMM consumes them on the same stream; GEMM outputs are not pooled because they escape the linear call.environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:52
B12X_DENSE_PER_ROW_IN_KERNELenvironment
default: '1'
The fallback host chain and in-kernel per-row scaling are bit-identical.environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:60
B12X_DENSE_SPLITK_TURBOenvironment
default: '1'
Select a B12X dense GEMM specialization or its workspace behavior.environment stringb12x/_lib/dense_gemm.py:145
B12X_DIRECT_CUTE_OPTIONSenvironment
default: ''
Pass direct CuteDSL compiler options to the B12X MoE path.environment stringb12x/moe/fused_moe/_impl.py:10000
B12X_DISABLE_BF16_GEMVenvironment
default: ''
Disable the B12X BF16 GEMV fast path when set.environment stringb12x/gemm/bf16_gemv/api.py:29
B12X_DISK_BACKENDenvironment
default: 'io_uring'
Select the disk-backed PLE row-staging backend: io_uring or GDS.environment stringb12x/sequence/_shared/disk_table.py:166
B12X_DSA_VALIDATE_PAGE_IDSenvironment
default: '0'
Validate DSA page identifiers in the B12X indexer path.environment stringb12x/attention/dsa_indexer/_impl.py:56
B12X_DYNAMIC_SWAP_ABenvironment
default: unset
Select the dynamic A/B swap policy for B12X fused MoE kernels.environment stringb12x/moe/fused_moe/_impl.py:2151
B12X_DYNAMIC_TILE_MNenvironment
default: unset
Select the dynamic M/N tile policy for B12X fused MoE kernels.environment stringb12x/moe/fused_moe/_impl.py:1691
B12X_ENABLE_FP6_MICROenvironment
default: '0'
Enable the B12X FP6 micro-kernel path.environment stringb12x/moe/_shared/kernels/micro.py:127
B12X_FP6_ACT_FMT_OVERRIDESenvironment
default: ''
Override FP6 activation formats selected by the B12X checkpoint.environment stringb12x/quantization/mxfp6/fp6_checkpoint.py:248
B12X_FUSED_INDEXERenvironment
default: '1'
Enable the fused B12X sparse indexer path.environment stringb12x/attention/dsa_indexer/_tuning.py:133
B12X_INDEXER_DIRECT_Kenvironment
default: '1'
Triage kill-switch: B12X_INDEXER_DIRECT_K=0 forces every variant back to the staged pipeline (which also restores the v2 cache keys), so serving can A/B the direct-L2 score in place without a checkout change.environment stringb12x/attention/dsa_indexer/fused_indexer.py:114
B12X_LOG_CUTE_COMPILESenvironment
default: ''
Enable B12X CuteDSL compilation progress logging.environment stringb12x/_lib/compiler.py:691
B12X_LOG_CUTE_COMPILE_ARGSenvironment
default: ''
Log arguments passed to B12X CuteDSL compilation.environment stringb12x/_lib/compiler.py:1217
B12X_LOG_CUTE_COMPILE_STACKenvironment
default: ''
Log the Python stack for B12X CuteDSL compilation.environment stringb12x/_lib/compiler.py:711
B12X_LOG_CUTE_COMPILE_STACK_DEPTHenvironment
default: ''
Set the maximum stack depth in B12X CuteDSL compile logs.environment stringb12x/_lib/compiler.py:722
B12X_MHC_PDLenvironment
default: '0'
Enable programmatic dependent launch for MHC normalization kernels.environment stringb12x/norm/mhc/_kernels.py:158
B12X_MHC_PREFILL_FINALIZE_THREADSenvironment
default: '256'
Set the CUDA thread count for the MHC prefill finalize kernel.environment stringb12x/norm/mhc/_kernels.py:305
B12X_MHC_PREFILL_GRAM_THREADSenvironment
default: os.getenv('B12X_MHC_PREFILL_THREADS', '1024')
Set the CUDA thread count for the MHC Gram-trick prefill path.environment stringb12x/norm/mhc/_kernels.py:166
B12X_MHC_PREFILL_TF32_TMA_CHUNK_MIN_TOKENSenvironment
default: '4096'
Set the token threshold that selects the MHC TF32 TMA chunk specialization.environment stringb12x/norm/mhc/_kernels.py:227
B12X_MHC_PREFILL_TF32_TMA_CHUNK_M_WARPSenvironment
default: '12'
Set the M-dimension compute warp count for the MHC TF32 TMA chunk specialization.environment stringb12x/norm/mhc/_kernels.py:237
B12X_MHC_PREFILL_TF32_TMA_CHUNK_N_WARPSenvironment
default: os.getenv('B12X_MHC_PREFILL_TF32_TMA_N_WARPS', '1')
Set the N-dimension compute warp count for the MHC TF32 TMA chunk specialization.environment stringb12x/norm/mhc/_kernels.py:243
B12X_MHC_PREFILL_TF32_TMA_CHUNK_STAGESenvironment
default: os.getenv('B12X_MHC_PREFILL_TF32_TMA_STAGES', '2')
Set the asynchronous pipeline stage count for the MHC TF32 TMA chunk specialization.environment stringb12x/norm/mhc/_kernels.py:279
B12X_MHC_PREFILL_TF32_TMA_CHUNK_TILE_Kenvironment
default: os.getenv('B12X_MHC_PREFILL_TF32_TMA_TILE_K', '64')
Set the K tile dimension for the MHC TF32 TMA chunk specialization.environment stringb12x/norm/mhc/_kernels.py:273
B12X_MHC_PREFILL_TF32_TMA_CHUNK_TILE_Menvironment
default: '192'
Set the M tile dimension for the MHC TF32 TMA chunk specialization.environment stringb12x/norm/mhc/_kernels.py:255
B12X_MHC_PREFILL_TF32_TMA_CHUNK_TILE_Nenvironment
default: os.getenv('B12X_MHC_PREFILL_TF32_TMA_TILE_N', '24')
Set the N tile dimension for the MHC TF32 TMA chunk specialization.environment stringb12x/norm/mhc/_kernels.py:261
B12X_MHC_PREFILL_TF32_TMA_LONG_MIN_TOKENSenvironment
default: '8192'
Set the token threshold that selects the MHC TF32 TMA long-context specialization.environment stringb12x/norm/mhc/_kernels.py:230
B12X_MHC_PREFILL_TF32_TMA_LONG_M_WARPSenvironment
default: '8'
Set the M-dimension compute warp count for the MHC TF32 TMA long-context specialization.environment stringb12x/norm/mhc/_kernels.py:287
B12X_MHC_PREFILL_TF32_TMA_LONG_N_WARPSenvironment
default: '1'
Set the N-dimension compute warp count for the MHC TF32 TMA long-context specialization.environment stringb12x/norm/mhc/_kernels.py:290
B12X_MHC_PREFILL_TF32_TMA_LONG_STAGESenvironment
default: '2'
Set the asynchronous pipeline stage count for the MHC TF32 TMA long-context specialization.environment stringb12x/norm/mhc/_kernels.py:302
B12X_MHC_PREFILL_TF32_TMA_LONG_TILE_Kenvironment
default: '64'
Set the K tile dimension for the MHC TF32 TMA long-context specialization.environment stringb12x/norm/mhc/_kernels.py:299
B12X_MHC_PREFILL_TF32_TMA_LONG_TILE_Menvironment
default: '128'
Set the M tile dimension for the MHC TF32 TMA long-context specialization.environment stringb12x/norm/mhc/_kernels.py:293
B12X_MHC_PREFILL_TF32_TMA_LONG_TILE_Nenvironment
default: '24'
Set the N tile dimension for the MHC TF32 TMA long-context specialization.environment stringb12x/norm/mhc/_kernels.py:296
B12X_MHC_PREFILL_TF32_TMA_M_WARPSenvironment
default: os.getenv('B12X_MHC_PREFILL_TF32_TMA_WARPS', os.getenv('B12X_MHC_PREFILL_TMA_WARPS', '1'))
Set the M-dimension compute warp count for the MHC TF32 TMA prefill kernel.environment stringb12x/norm/mhc/_kernels.py:192
B12X_MHC_PREFILL_TF32_TMA_N_WARPSenvironment
default: '1'
Set the N-dimension compute warp count for the MHC TF32 TMA prefill kernel.environment stringb12x/norm/mhc/_kernels.py:201
B12X_MHC_PREFILL_TF32_TMA_STAGESenvironment
default: os.getenv('B12X_MHC_PREFILL_TMA_STAGES', '1')
Set the asynchronous pipeline stage count for the MHC TF32 TMA prefill kernel.environment stringb12x/norm/mhc/_kernels.py:221
B12X_MHC_PREFILL_TF32_TMA_TILE_Kenvironment
default: os.getenv('B12X_MHC_PREFILL_TMA_TILE_K', '256')
Set the K tile dimension for the MHC TF32 TMA prefill kernel.environment stringb12x/norm/mhc/_kernels.py:215
B12X_MHC_PREFILL_TF32_TMA_TILE_Menvironment
default: os.getenv('B12X_MHC_PREFILL_TMA_TILE_M', '16')
Set the M tile dimension for the MHC TF32 TMA prefill kernel.environment stringb12x/norm/mhc/_kernels.py:206
B12X_MHC_PREFILL_TF32_TMA_TILE_Nenvironment
default: '8'
Set the N tile dimension for the MHC TF32 TMA prefill kernel.environment stringb12x/norm/mhc/_kernels.py:212
B12X_MHC_PREFILL_TF32_TMA_WARPSenvironment
default: os.getenv('B12X_MHC_PREFILL_TMA_WARPS', '1')
Set the compute warp count for the MHC TF32 TMA prefill kernel.environment stringb12x/norm/mhc/_kernels.py:194
B12X_MHC_PREFILL_THREADSenvironment
default: '512'
Set the CUDA thread count for the MHC prefill kernel.environment stringb12x/norm/mhc/_kernels.py:164
B12X_MHC_PREFILL_TMA_STAGESenvironment
default: '3'
Set the asynchronous pipeline stage count for the MHC prefill TMA kernel.environment stringb12x/norm/mhc/_kernels.py:190
B12X_MHC_PREFILL_TMA_TILE_Kenvironment
default: '64'
Set the K tile dimension for the MHC prefill TMA kernel.environment stringb12x/norm/mhc/_kernels.py:188
B12X_MHC_PREFILL_TMA_TILE_Menvironment
default: '128'
Set the M tile dimension for the MHC prefill TMA kernel.environment stringb12x/norm/mhc/_kernels.py:182
B12X_MHC_PREFILL_TMA_TILE_Nenvironment
default: '16'
Set the N tile dimension for the MHC prefill TMA kernel.environment stringb12x/norm/mhc/_kernels.py:185
B12X_MHC_PREFILL_TMA_WARPSenvironment
default: '8'
Set the compute warp count for the MHC prefill TMA kernel.environment stringb12x/norm/mhc/_kernels.py:178
B12X_MICRO_DYNAMIC_CUTOVER_PAIRSenvironment
default: unset
Set the active expert-pair count at which micro-MoE changes to its dynamic execution path.environment stringb12x/moe/fused_moe/_impl.py:2225
B12X_MICRO_REUSE_COMPILEDenvironment
default: '1'
Compile planning produces deferred programs that must never be memoized or returned in place of a compiled kernel.environment stringb12x/moe/fused_moe/_impl.py:9980
B12X_MICRO_SHARE_INPUT_ACROSS_EXPERTSenvironment
default: '1'
Enable reuse of one micro-MoE input load across the selected experts.environment stringb12x/moe/fused_moe/_impl.py:13318
B12X_MLA_SM120_PREFILL_MGenvironment
default: '1'
Enable the SM120 multi-group MLA prefill kernel.environment stringb12x/attention/_shared/mla/prefill.py:398
B12X_MLA_SM120_PREFILL_PACK_HILO_ROWSenvironment
default: '1'
Pack high- and low-position MLA prefill rows into one SM120 kernel input.environment stringb12x/attention/_shared/mla/prefill_mg.py:3948
B12X_MOE_TILE_MNenvironment
default: unset
Override the MxN MMA tile shape used by the B12X fused MoE kernel.environment stringb12x/moe/fused_moe/_impl.py:9893
B12X_MSA_DECODE_PAGEMAXenvironment
default: '1'
Enable the page-maximum reduction used by the MSA decode indexer.environment stringb12x/attention/dsa_indexer/_impl.py:885
B12X_NVFP4_SPLIT_DECODEenvironment
default: '0'
Split NVFP4 decode work into the materialized decode path instead of the monolithic path.environment stringb12x/moe/fused_moe/_impl.py:12116
B12X_PACKED_B_EXPAND_AHEADenvironment
default: '1'
Expand-ahead for packed-B: at k_block 0 the MMA warps wait for stage s+1 and expand it in place, overlapping the expansion with ALL of stage s's MMA work instead of putting it on the critical path at the stage boundary. It requires at least four pipeline stages of producer slack.environment stringb12x/_lib/dense_gemm.py:157
B12X_PACKED_B_MIN_Nenvironment
default: '12288'
Minimum out_features for which the GEMM streams the 3:4-packed weight directly (b_packed=True) instead of a cached 1-byte/code expansion. Default 12288 enables shapes with enough N tiles for the packed path.environment stringb12x/quantization/mxfp6/fp6_dense_weights.py:125
B12X_PAGED_EXTEND_QWEN_FP8_PV_REPACKenvironment
default: unset
Enable Qwen FP8 value-projection repacking during paged attention extension.environment stringb12x/attention/paged/forward_extend_generic.py:3525
B12X_PAGED_KV_TMA_PLANE_SWIZZLEenvironment
default: ''
Set the TMA plane swizzle used to load paged KV data.environment stringb12x/attention/paged/_forward.py:215
B12X_PAGED_MSAenvironment
default: unset
Enable MSA block-sparse paged attention.environment stringb12x/attention/paged/_scratch.py:127
B12X_PAGED_MSA_UNION_PREFILLenvironment
default: '1'
Enable the MSA union-prefill path for paged attention.environment stringb12x/attention/paged/_scratch.py:136
B12X_PCIE_ALLREDUCE_ALGORITHMenvironment
default: 'auto'
Select the B12X PCIe all-reduce algorithm: automatic, hierarchical, island reduce-scatter, or one-shot.environment stringb12x/comm/pcie/pcie_allreduce.py:58
B12X_PCIE_DMA_A2A_CHUNKSenvironment
default: '0'
Override the number of chunks used by the PCIe DMA all-to-all path.environment stringb12x/comm/pcie/pcie_dma.py:258
B12X_PCIE_DMA_FP8environment
default: '0'
Enable the FP8 PCIe DMA all-reduce path.environment stringb12x/comm/pcie/pcie_dma.py:101
B12X_PCIE_DMA_PIECESenvironment
default: '0'
Override the number of pieces used by the PCIe DMA all-reduce path.environment stringb12x/comm/pcie/pcie_dma.py:257
B12X_PCIE_FUSED_CTAS_PER_ROWenvironment
default: '0'
Recover the CTA count from the already-derived single-CTA/config formula without changing the launch contract.environment stringb12x/comm/pcie/pcie_oneshot.py:1977
B12X_PCIE_FUSED_THREADSenvironment
default: '256'
Set the CUDA thread count for the fused PCIe collective kernel.environment stringb12x/comm/pcie/pcie_oneshot.py:1655
B12X_PCIE_HIERARCHICAL_BF16X2environment
default: '1'
Enable vectorized BF16x2 loads in the hierarchical PCIe all-reduce kernel.environment stringb12x/comm/pcie/pcie_hierarchical.py:67
B12X_PCIE_HIERARCHICAL_BF16X2_MAX_ELEMENTSenvironment
default: '7168'
Set the largest reduction size eligible for vectorized BF16x2 hierarchical PCIe all-reduce.environment stringb12x/comm/pcie/pcie_hierarchical.py:74
B12X_PCIE_HIERARCHICAL_DEFERRED_CONSUMPTIONenvironment
default: '0'
Defer consumption of hierarchical PCIe all-reduce partial results.environment stringb12x/comm/pcie/pcie_hierarchical.py:158
B12X_PCIE_HIERARCHICAL_DOUBLE_BUFFERenvironment
default: '0'
Enable double buffering of staging memory in hierarchical PCIe all-reduce.environment stringb12x/comm/pcie/pcie_hierarchical.py:155
B12X_PCIE_HIERARCHICAL_NANOSLEEP_CYCLESenvironment
default: '24'
Set spin-wait nanosleep cycles for hierarchical PCIe all-reduce.environment stringb12x/comm/pcie/pcie_hierarchical.py:36
B12X_PCIE_HIERARCHICAL_THREADSenvironment
default: '224'
Set the CUDA thread count for hierarchical PCIe all-reduce.environment stringb12x/comm/pcie/pcie_hierarchical.py:51
B12X_PCIE_ISLAND_RS_NANOSLEEP_CYCLESenvironment
default: '24'
Set spin-wait nanosleep cycles for island reduce-scatter PCIe all-reduce.environment stringb12x/comm/pcie/pcie_island_rs.py:50
B12X_PCIE_ISLAND_RS_THREADSenvironment
default: '512'
Default CUDA thread count for the TP16 equal-quarter launch.environment stringb12x/comm/pcie/pcie_island_rs.py:40
B12X_PCIE_ONESHOT_BLOCK_LIMITenvironment
default: '8'
Set the maximum number of CUDA blocks used by one-shot PCIe all-reduce.environment stringb12x/comm/pcie/pcie_oneshot.py:1526
B12X_PCIE_ONESHOT_PUSHenvironment
default: '0'
Enable the one-shot PCIe push transport, which writes each rank's input into every peer's eager slot.environment stringb12x/comm/pcie/pcie_oneshot.py:128
B12X_PCIE_ONESHOT_THREADSenvironment
default: '256'
Set the CUDA thread count for one-shot PCIe all-reduce.environment stringb12x/comm/pcie/pcie_oneshot.py:1524
B12X_PCIE_TEST_VISIBLE_SM_COUNTenvironment
default: '0'
Override the visible SM count for PCIe collective topology validation.environment stringb12x/comm/pcie/pcie_oneshot.py:704
B12X_PCIE_TP2_PLAIN_REMOTE_PUSHenvironment
default: unset
Enable the qualified plain TP2 remote-write PCIe transport.environment stringb12x/comm/pcie/pcie_oneshot.py:145
B12X_PCIE_TP2_REMOTE_PUSHenvironment
default: '0'
Enable the qualified fused TP2 remote-write PCIe transport.environment stringb12x/comm/pcie/pcie_oneshot.py:134
B12X_PCIE_TP4_REMOTE_PUSHenvironment
default: '0'
Enable the qualified fused TP4 remote-write PCIe transport.environment stringb12x/comm/pcie/pcie_oneshot.py:154
B12X_PCIE_TP8_OWNER_REDUCEenvironment
default: '1'
Enable the topology-scoped fused TP8 owner-reduce PCIe transport.environment stringb12x/comm/pcie/pcie_oneshot.py:160
B12X_PCIE_VOCAB_ARGMAX_NANOSLEEP_CYCLESenvironment
default: '24'
Set spin-wait nanosleep cycles for PCIe vocabulary argmax reduction.environment stringb12x/comm/pcie/pcie_vocab_argmax.py:71
B12X_PREPARATION_TRACE_DIRenvironment
default: unset
Write B12X preparation timing traces to this directory.environment stringb12x/preparation/_timing.py:14
B12X_PROBE_COMMANDenvironment
default: ''
Configure B12X PCIe overlap probe diagnostics.environment stringb12x/comm/pcie/overlap_probe.py:1190
B12X_PROBE_SOURCE_REVISIONenvironment
default: ''
Configure B12X PCIe overlap probe diagnostics.environment stringb12x/comm/pcie/overlap_probe.py:1191
B12X_PROBE_SOURCE_WORKTREEenvironment
default: ''
Configure B12X PCIe overlap probe diagnostics.environment stringb12x/comm/pcie/overlap_probe.py:1192
B12X_ROCE_CACHE_DIRenvironment
default: unset
Directory for the compiled B12X RoCE proxy and its source-hash cache.environment stringb12x/comm/roce/_proxy.py:44
B12X_SQG_XOR_CHEB_T12_SMEMenvironment
default: '1'
Enable shared-memory storage for the SQG XOR-Chebyshev T12 kernel path.environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:172
B12X_TIMINGenvironment
default: '0'
Enable B12X kernel timing output.environment stringb12x/_lib/dense_gemm.py:137
B12X_TIMING_THRESHOLD_MSenvironment
default: os.getenv('VLLM_B12X_TIMING_THRESHOLD_MS', '0')
Set the B12X kernel timing output threshold in milliseconds.environment stringb12x/_lib/dense_gemm.py:140
B12X_TUNING_CACHE_VERSIONenvironment
default: '1'
Version the B12X preparation tuning-cache format.environment stringb12x/preparation/_cache.py:28
B12X_VALIDATE_PAGED_INDEXER_CUDA_VALUESenvironment
default: '0'
Enable B12X validation checks for runtime values.environment stringb12x/attention/dsa_indexer/paged.py:400
B12X_W4A16_SMALL_M_DIRECTenvironment
default: '1'
Use the direct small-M W4A16 MoE kernel instead of its fallback path.environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:9093
B12X_W4A16_SMALL_M_HOST_BARRIER_RESETenvironment
default: '1'
Reset the small-M W4A16 MoE host barrier between launches.environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:9132
B12X_W4A16_SMALL_M_SPLITKenvironment
default: '0'
Enable split-K decomposition for the small-M W4A16 MoE kernel.environment stringb12x/moe/_shared/kernels/w4a16/kernel.py:162
B12X_W4A8_TINY_DECODEenvironment
default: '1'
Default on; B12X_W4A8_TINY_DECODE=0 is the kill switch.environment stringb12x/moe/fused_moe/_impl.py:12227
B12X_WO_B_FUSED_TILEenvironment
default: ''
Environment variable read by the cited source.environment stringb12x/gemm/_shared/wo_mxfp8.py:2808
B12X_WO_QUANT_CHUNKS_PER_PROGRAMenvironment
default: '16'
Environment variable read by the cited source.environment stringb12x/gemm/_shared/wo_mxfp8.py:49
VLLM_B12X_TIMINGenvironment
default: '0'
Configure a local-inference-lab vLLM B12X execution path.environment stringb12x/_lib/dense_gemm.py:137
VLLM_B12X_TIMING_THRESHOLD_MSenvironment
default: '0'
Configure a local-inference-lab vLLM B12X execution path.environment stringb12x/_lib/dense_gemm.py:142
VLLM_EXPERIMENTAL_SHM_BROADCAST_ADAPTIVE_SPINenvironment
default: "0"
Experimental opt-in for adaptive shm_broadcast reader and writer waits. The default retains the fixed one-second reader grace and writer yields.environment stringvllm/envs.py:768
VLLM_QWEN3_8_HC_PREFILL_MODEenvironment
default: 'off'
HC prefill row ownership: off, control, or shard. Opt-in pending serving qualification; ranks must agree at startup.environment stringvllm/envs.py:1754
VLLM_QWEN3_8_PREFILL_COALESCEenvironment
default: "0"
Select a Qwen3.8 Flash Next execution path.environment stringvllm/envs.py:1757
VLLM_SHM_BROADCAST_ADAPTIVE_ALPHAenvironment
default: "0.25"
EMA coefficient for observed inter-read intervals, in (0, 1]: higher adapts to cadence changes faster but tracks noise.environment stringvllm/envs.py:799
VLLM_SHM_BROADCAST_ADAPTIVE_BUDGET_MSenvironment
default: "1.0"
Tunables for the experimental adaptive shm_broadcast reader spin grace. Readers busy-loop for a grace period after the last read before parking on the poller; the adaptive policy derives that grace from an EMA of observed inter-read intervals (T_ema): grace = clamp(B * B / T_ema, MIN_GRACE, MAX_GRACE) B (BUDGET) is the pivot: at T_ema == B the grace equals B; faster traffic spins toward MAX_GRACE (arrivals near-certain within the grace), slower traffic parks within MIN_GRACE. The grace is monotonically decreasing in T_ema, so slower traffic never increases spin, and it is bounded regardless of estimate staleness. Passing a float to SpinCondition's busy_loop_s overrides the policy with a fixed grace. All grace values are in milliseconds. Time-scale pivot B: grace == B when T_ema == B.environment stringvllm/envs.py:784
VLLM_SHM_BROADCAST_ADAPTIVE_MAX_GRACE_MSenvironment
default: "2.0"
Upper bound on the spin grace: the most time a reader busy-loops even under the fastest observed traffic.environment stringvllm/envs.py:794
VLLM_SHM_BROADCAST_ADAPTIVE_MIN_GRACE_MSenvironment
default: "0.05"
Lower bound on the spin grace: an idle reader parks within this window instead of spinning.environment stringvllm/envs.py:789
VLLM_SHM_BROADCAST_WRITE_PARK_MAX_MSenvironment
default: "1.0"
Ceiling for the writer's park step once its own grace expires: the writer sleeps in doubling steps from 50us up to this bound while waiting for the slowest reader to release a block. Raising it trades write latency for fewer wakeups on a writer blocked by slow readers.environment stringvllm/envs.py:806
  • --hf-overrides retains dictionary YaRN and rope values when vLLM builds an in-model MTP draft ModelConfig (local-inference-lab/vLLM #777).
  • VLLM_QWEN3_8_PREFILL_COALESCE requires B12X #386’s prepared PLE checkpoint export and vLLM 5dd5bd5dde76 or later, which wires the prefill checkpoint blocks into the NVIDIA GDN decoder. On earlier cuts that carry the coalesce feature from 61f93c53ee, setting it aborts boot with “all mamba groups must share cache scheduling parameters”.
  • Never set CUDA_LAUNCH_BLOCKING. b12x does not function with synchronous kernel launches, so the variable is not a usable debug lever on this stack.
  • The adaptive shared-memory controls alter wait behavior only and are excluded from vLLM compile factors. They do not select kernels or invalidate compiled graphs.

16e5efc2ab3647ea305bef521710ab1344365132ceff76a827c90cdc168e29b2