chore(deps): update dependency flashinfer-python to v0.7.0 - #123
Closed
renovate[bot] wants to merge 1 commit into
Closed
renovate[bot] wants to merge 1 commit into
renovate[bot] wants to merge 1 commit into
Conversation
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
6 times, most recently
from
June 9, 2026 06:35
79bf3af to
e47106d
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
5 times, most recently
from
June 18, 2026 08:28
b7642ac to
e5caed6
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
2 times, most recently
from
June 22, 2026 08:24
8712535 to
4846306
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
2 times, most recently
from
July 2, 2026 12:20
bd9a2fa to
423e158
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
8 times, most recently
from
July 8, 2026 15:27
4234234 to
420b89e
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
5 times, most recently
from
July 14, 2026 08:35
ea10f4f to
6731996
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
3 times, most recently
from
August 7, 2026 23:40
015d3a2 to
9d22f6b
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
from
August 8, 2026 10:50
9d22f6b to
007fdc9
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
from
August 16, 2026 00:12
007fdc9 to
4197286
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
3 times, most recently
from
August 25, 2026 13:43
dcd1cb0 to
0cc8a5f
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
from
August 30, 2026 08:12
0cc8a5f to
ac43add
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
3 times, most recently
from
September 5, 2026 15:51
ca992bb to
658a9c9
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
7 times, most recently
from
September 14, 2026 13:17
fb81258 to
f802883
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
2 times, most recently
from
September 21, 2026 08:39
1a85f15 to
8b950c0
Compare
renovate
Bot
force-pushed
the
renovate/flashinfer-python-0.x
branch
from
September 22, 2026 13:31
8b950c0 to
30dbed6
Compare
Contributor
Author
Renovate Ignore NotificationBecause you closed this PR without merging, Renovate will ignore this update ( If you accidentally closed this PR, or if you changed your mind: rename this PR to get a fresh replacement PR. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR contains the following updates:
==0.6.16.post3→==0.7.0==0.6.18→==0.7.0Release Notes
flashinfer-ai/flashinfer (flashinfer-python)
v0.7.0Compare Source
These highlights are also published at flashinfer.ai/releases.
v0.7.0 Highlights
FlashInfer 0.7.0 makes the unified mixture-of-experts (MoE) API official, brings TRT-LLM Gen MoE kernels into readable Python source with PrimTS, and introduces Autotuner v2 and a formal experimental-API policy. It also expands sparse and linear attention, expert-parallel serving, and diffusion workloads across Blackwell GPUs.
Read the v0.7 overview and the accompanying deep dives on MegaMoE, Autotuner v2, and the experimental path.
Unified MoE API is official
MoELayeris now an official FlashInfer API, with the lower-level kernel entry points supported alongside it.QuantConfiggives weights, activations, and output explicitQuantFormatfields:QuantConfig(weight=QuantFormat.MXFP4)selects MXFP4 weights with BF16 activations, while addingactivation=QuantFormat.MXFP8selects W4A8. Backend coverage expands with quantization-specific CUTLASS runners, cuTile BF16 and NVFP4 MoE, and CuTe-DSL BF16 MoE on Hopper.TRT-LLM Gen MoE kernels as Python source with PrimTS
The experimental PrimTS MoE backend exposes kernels built with the CUTLASS DSL Primitives and Task Scheduling APIs, making the expert GEMMs available as readable Python source. It supports BF16, per-tensor and block-scaled FP8, and NVFP4/MXFP4 combinations, while reusing TRT-LLM Gen routing and finalization. Accuracy qualification targets B200; B300 qualification is still pending, and supported configurations have backend-specific restrictions. PrimTS attention also gains paged block-sparse attention, variable-window attention, and a unified
plan()/run()contract for reusable wrappers.Autotuner v2 tunes the way you serve
autotune_v2()lets applications measure candidates in eager or CUDA-graph execution and persist the results in an environment-specific cache managed by FlashInfer. Atomic cache entries support concurrent rank writes, andautotune_v2_reload()lets homogeneous ranks converge on shared results. In the reported vLLM validation, Qwen3-8B-FP8 at TP2 on B200 reduced its tuning window from 128 seconds on a cold start to 1 second on restart with the same cache; total startup was 410.7 seconds and 165.3 seconds, respectively. Applications opt in through the new API;autotune()remains available.Experimental APIs and backends have an explicit opt-in path
Experimental APIs are marked with
@flashinfer_experimental_api, and experimental backends have a dedicatedflashinfer.experimentalnamespace. Calling an experimental API or explicitly selecting a marked backend emits a warning; automatic selection includes marked experimental backends only whenFLASHINFER_ALLOW_EXPERIMENTAL_AUTO_BACKENDS=1is set. The policy defines admission and graduation criteria and keeps experimental implementations JIT-only, outside prebuilt packages.Expert-parallel MoE expands across Blackwell
moe_epadds an unquantized BF16 MegaMoE backend and a W4A8 split backend on B200, with MXFP8-packed dispatch for the latter. RTX PRO 6000 and DGX Spark gain an MXFP8 MegaMoE backend with functional correctness validated and performance tuning ongoing. The NCCL-EP split path supports CUDA-graph capture through a reusable handle with per-stepupdate(), validated through vLLM on four B200 GPUs.DeepSeek-V4 and MiniMax-M3 sparse attention
Blackwell gains paged FP8/MXFP4 indexer logits and more top-K choices, including a CUB backend with variable-length support. DeepSeek-V4 Flash sparse MLA supports an NVFP4 KV cache on SM120/SM121 through
kv_cache_format="nvfp4". MiniMax-M3 sparse attention gains source-distributed CAKE-generated kernels on SM100/SM103; the B200 kernel benchmark reports a 2.45× geometric-mean speedup over the MiniMax baseline across 11 comparable prefill, decode, speculative, and boundary cases.Linear attention for Kimi K3, Qwen 3.6, and Nemotron-H
Kimi K3 gains a speculative-verification entry point compatible with vLLM's recurrent verifier, plus CuTe-DSL recurrent prefill on B200/B300 and SM120. The experimental
RecurrentKDAPrefillWrappersupports planning and graph-safe prefix checkpoints. Qwen 3.6 gains experimental fused GDN decode steps on SM120 that combine projection, convolution, gating, and recurrent state updates. Nemotron-H gains source-built Mamba SSD-combined and selective-state-update backends on B200/GB300.Diffusion and video attention on Blackwell
Video Sparse Attention gains generated SM100/SM103 kernels, and the block-64 path adds a native CuTe-DSL implementation with Sage FP8 support. SM120 gains a generated Sage block-sparse backend and an optimized NVFP4 attention path for Cosmos workloads. MiniMax-H3 gains a prepared MXFP8 pre-attention pipeline on B200/B300, while PrimTS block-sparse attention adds proxy compensation for Sol-Attn workloads.
Communication for PCIe and NVLink deployments
PcieIpcAllReduceWorkspaceadds an intra-node CUDA-IPC all-reduce for two, four, or eight ranks on PCIe machines without NVLink. Blackwell fused all-gather matmul gains abackend="cake"option and a prepared callable for packed-QKV workloads, with tensor parallelism up to eight GPUs.Notice: packaging and dependencies
flashinfer-jit-cacheis now a small shim that depends on architecture-specific provider wheels. The usual installation command installs the provider set for the selected CUDA and CPU platform; to reduce image size, install only the provider needed for your GPU. If you mirror or vendor wheels, include the provider wheels as well as the shim, and keep their CUDA-specific versions aligned.nvidia-cudnn-frontend>=1.25.0>=1.29.0apache-tvm-ffi>=0.1.6,!=0.1.8,!=0.1.8.post0,<0.2>=0.1.11,<0.2nccl4py>=0.3.1nccl4py>=0.4.1andnccl-extensions>=0.1.0nvidia-cutlass-dslthrough[cu12]/[cu13]>=4.6.2a0>=4.7.0a0The base
nvidia-cutlass-dslrequirement remains>=4.6.2a0; the CUDA extras have the higher floor. Environments pinned to CUTLASS DSL 4.6.2 need compatible dependency pins before installing those extras.Notice: small-batch FP8 groupwise GEMM
Notice: API removals and behavior changes
Update callers of the removed APIs before upgrading:
comm.trtllm_custom_all_reducecomm.trtllm_allreduce_fusioncomm.trtllm_create_ipc_workspace_for_all_reducecomm.trtllm_create_ipc_workspace_for_all_reduce_fusionBatchDecodeMlaWithPagedKVCacheWrappermla.BatchMLAPagedAttentionWrapperend_forward()methods on decode, prefill, sparse, cascade, and POD wrappersfused_moe.QuantVariant,QuantConfig(variant=...),QuantConfig.from_variantQuantConfig(weight=..., activation=...)usingQuantFormatmamba.checkpointing_ssureplaces the per-stepold_x,old_B,old_dt,old_cumAdt, andcache_buf_idxarguments with ring-buffer caches (x_cache,B_cache,dt_cache,ring_start). Update cache allocation and argument binding; the caller-providedouttensor remains required.For the TRT-LLM backend of
BatchPrefillWithRaggedKVCacheWrapper,skip_all_rows_active_checkdefaults toTrue. Calls on the default fast path must have positive query and KV lengths for every row; passFalseto request device-derived checking when CPU length mirrors are absent.gated_delta_rule_mtpstill resolves an omitteddisable_state_updatetoTruein 0.7.0. Passdisable_state_update=TrueorFalseexplicitly to select the intended behavior and suppress the stale warning announcing a change in this release.What's Changed
DeviceBatchedTopKtop-k backend with variable-length support by @NaderAlAwar in #4442radix_filter) by @dhiraj113 in #4621@flashinfer_experimental_api, andflashinfer.experimentalnamespace by @bkryu in #4880New Contributors
Configuration
📅 Schedule: (UTC)
🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.
♻ Rebasing: Whenever PR is behind base branch, or you tick the rebase/retry checkbox.
🔕 Ignore: Close this PR and you won't be reminded about these updates again.
This PR was generated by Mend Renovate. View the repository job log.