Skip to content

chore(deps): update dependency flashinfer-python to v0.7.0 - #123

Closed
renovate[bot] wants to merge 1 commit into
mainfrom
renovate/flashinfer-python-0.x
Closed

renovate[bot] wants to merge 1 commit into
mainfrom
renovate/flashinfer-python-0.x

Conversation

@renovate

@renovate renovate Bot commented Jun 3, 2026 •

Copy link
Copy Markdown
Contributor

ℹ️ Note

This PR body was truncated due to platform limits.

This PR contains the following updates:

Package Change Age Confidence
flashinfer-python ==0.6.16.post3 → ==0.7.0 age confidence
flashinfer-python ==0.6.18 → ==0.7.0 age confidence

Release Notes

flashinfer-ai/flashinfer (flashinfer-python)

v0.7.0

Compare Source

These highlights are also published at flashinfer.ai/releases.

v0.7.0 Highlights

FlashInfer 0.7.0 makes the unified mixture-of-experts (MoE) API official, brings TRT-LLM Gen MoE kernels into readable Python source with PrimTS, and introduces Autotuner v2 and a formal experimental-API policy. It also expands sparse and linear attention, expert-parallel serving, and diffusion workloads across Blackwell GPUs.

Read the v0.7 overview and the accompanying deep dives on MegaMoE, Autotuner v2, and the experimental path.

Unified MoE API is official

MoELayer is now an official FlashInfer API, with the lower-level kernel entry points supported alongside it. QuantConfig gives weights, activations, and output explicit QuantFormat fields: QuantConfig(weight=QuantFormat.MXFP4) selects MXFP4 weights with BF16 activations, while adding activation=QuantFormat.MXFP8 selects W4A8. Backend coverage expands with quantization-specific CUTLASS runners, cuTile BF16 and NVFP4 MoE, and CuTe-DSL BF16 MoE on Hopper.

TRT-LLM Gen MoE kernels as Python source with PrimTS

The experimental PrimTS MoE backend exposes kernels built with the CUTLASS DSL Primitives and Task Scheduling APIs, making the expert GEMMs available as readable Python source. It supports BF16, per-tensor and block-scaled FP8, and NVFP4/MXFP4 combinations, while reusing TRT-LLM Gen routing and finalization. Accuracy qualification targets B200; B300 qualification is still pending, and supported configurations have backend-specific restrictions. PrimTS attention also gains paged block-sparse attention, variable-window attention, and a unified plan()/run() contract for reusable wrappers.

Autotuner v2 tunes the way you serve

autotune_v2() lets applications measure candidates in eager or CUDA-graph execution and persist the results in an environment-specific cache managed by FlashInfer. Atomic cache entries support concurrent rank writes, and autotune_v2_reload() lets homogeneous ranks converge on shared results. In the reported vLLM validation, Qwen3-8B-FP8 at TP2 on B200 reduced its tuning window from 128 seconds on a cold start to 1 second on restart with the same cache; total startup was 410.7 seconds and 165.3 seconds, respectively. Applications opt in through the new API; autotune() remains available.

Experimental APIs and backends have an explicit opt-in path

Experimental APIs are marked with @flashinfer_experimental_api, and experimental backends have a dedicated flashinfer.experimental namespace. Calling an experimental API or explicitly selecting a marked backend emits a warning; automatic selection includes marked experimental backends only when FLASHINFER_ALLOW_EXPERIMENTAL_AUTO_BACKENDS=1 is set. The policy defines admission and graduation criteria and keeps experimental implementations JIT-only, outside prebuilt packages.

Expert-parallel MoE expands across Blackwell

moe_ep adds an unquantized BF16 MegaMoE backend and a W4A8 split backend on B200, with MXFP8-packed dispatch for the latter. RTX PRO 6000 and DGX Spark gain an MXFP8 MegaMoE backend with functional correctness validated and performance tuning ongoing. The NCCL-EP split path supports CUDA-graph capture through a reusable handle with per-step update(), validated through vLLM on four B200 GPUs.

DeepSeek-V4 and MiniMax-M3 sparse attention

Blackwell gains paged FP8/MXFP4 indexer logits and more top-K choices, including a CUB backend with variable-length support. DeepSeek-V4 Flash sparse MLA supports an NVFP4 KV cache on SM120/SM121 through kv_cache_format="nvfp4". MiniMax-M3 sparse attention gains source-distributed CAKE-generated kernels on SM100/SM103; the B200 kernel benchmark reports a 2.45× geometric-mean speedup over the MiniMax baseline across 11 comparable prefill, decode, speculative, and boundary cases.

Linear attention for Kimi K3, Qwen 3.6, and Nemotron-H

Kimi K3 gains a speculative-verification entry point compatible with vLLM's recurrent verifier, plus CuTe-DSL recurrent prefill on B200/B300 and SM120. The experimental RecurrentKDAPrefillWrapper supports planning and graph-safe prefix checkpoints. Qwen 3.6 gains experimental fused GDN decode steps on SM120 that combine projection, convolution, gating, and recurrent state updates. Nemotron-H gains source-built Mamba SSD-combined and selective-state-update backends on B200/GB300.

Diffusion and video attention on Blackwell

Video Sparse Attention gains generated SM100/SM103 kernels, and the block-64 path adds a native CuTe-DSL implementation with Sage FP8 support. SM120 gains a generated Sage block-sparse backend and an optimized NVFP4 attention path for Cosmos workloads. MiniMax-H3 gains a prepared MXFP8 pre-attention pipeline on B200/B300, while PrimTS block-sparse attention adds proxy compensation for Sol-Attn workloads.

Communication for PCIe and NVLink deployments

PcieIpcAllReduceWorkspace adds an intra-node CUDA-IPC all-reduce for two, four, or eight ranks on PCIe machines without NVLink. Blackwell fused all-gather matmul gains a backend="cake" option and a prepared callable for packed-QKV workloads, with tensor parallelism up to eight GPUs.

Notice: packaging and dependencies

flashinfer-jit-cache is now a small shim that depends on architecture-specific provider wheels. The usual installation command installs the provider set for the selected CUDA and CPU platform; to reduce image size, install only the provider needed for your GPU. If you mirror or vendor wheels, include the provider wheels as well as the shim, and keep their CUDA-specific versions aligned.

Dependency 0.6.18.post1 0.7.0
nvidia-cudnn-frontend >=1.25.0 >=1.29.0
apache-tvm-ffi >=0.1.6,!=0.1.8,!=0.1.8.post0,<0.2 >=0.1.11,<0.2
NCCL-EP Python packages nccl4py>=0.3.1 nccl4py>=0.4.1 and nccl-extensions>=0.1.0
nvidia-cutlass-dsl through [cu12] / [cu13] >=4.6.2a0 >=4.7.0a0

The base nvidia-cutlass-dsl requirement remains >=4.6.2a0; the CUDA extras have the higher floor. Environments pinned to CUTLASS DSL 4.6.2 need compatible dependency pins before installing those extras.

Notice: small-batch FP8 groupwise GEMM

⚠️ The CUTLASS gemm_fp8_nt_groupwise path on SM100/SM103 can intermittently produce incorrect output for M <= 32 with scale_granularity_mnk=(1, 128, 128). This pre-existing issue remains in 0.7.0; see #​4396 for the investigation. Validate affected workloads before deployment. A different backend requires its own supported shapes and scale layout; there is no documented switch to disable only this CUTLASS fast path.

Notice: API removals and behavior changes

Update callers of the removed APIs before upgrading:

Removed Replacement
comm.trtllm_custom_all_reduce comm.trtllm_allreduce_fusion
comm.trtllm_create_ipc_workspace_for_all_reduce comm.trtllm_create_ipc_workspace_for_all_reduce_fusion
BatchDecodeMlaWithPagedKVCacheWrapper mla.BatchMLAPagedAttentionWrapper
No-op end_forward() methods on decode, prefill, sparse, cascade, and POD wrappers Remove the call
fused_moe.QuantVariant, QuantConfig(variant=...), QuantConfig.from_variant QuantConfig(weight=..., activation=...) using QuantFormat

mamba.checkpointing_ssu replaces the per-step old_x, old_B, old_dt, old_cumAdt, and cache_buf_idx arguments with ring-buffer caches (x_cache, B_cache, dt_cache, ring_start). Update cache allocation and argument binding; the caller-provided out tensor remains required.

For the TRT-LLM backend of BatchPrefillWithRaggedKVCacheWrapper, skip_all_rows_active_check defaults to True. Calls on the default fast path must have positive query and KV lengths for every row; pass False to request device-derived checking when CPU length mirrors are absent.

gated_delta_rule_mtp still resolves an omitted disable_state_update to True in 0.7.0. Pass disable_state_update=True or False explicitly to select the intended behavior and suppress the stale warning announcing a change in this release.

What's Changed
New Contributors

❗ Important

✂ PR body was truncated to here.


Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

♻ Rebasing: Whenever PR is behind base branch, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about these updates again.


  • If you want to rebase/retry this PR, check this box

This PR was generated by Mend Renovate. View the repository job log.

@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 6 times, most recently from 79bf3af to e47106d Compare June 9, 2026 06:35
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 5 times, most recently from b7642ac to e5caed6 Compare June 18, 2026 08:28
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 2 times, most recently from 8712535 to 4846306 Compare June 22, 2026 08:24
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.12 chore(deps): update dependency flashinfer-python to v0.6.13 Jun 25, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 2 times, most recently from bd9a2fa to 423e158 Compare July 2, 2026 12:20
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.13 chore(deps): update dependency flashinfer-python to v0.6.14 Jul 2, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 8 times, most recently from 4234234 to 420b89e Compare July 8, 2026 15:27
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 5 times, most recently from ea10f4f to 6731996 Compare July 14, 2026 08:35
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.16 chore(deps): update dependency flashinfer-python to v0.6.16.post1 Aug 3, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 3 times, most recently from 015d3a2 to 9d22f6b Compare August 7, 2026 23:40
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.16.post1 chore(deps): update dependency flashinfer-python to v0.6.16.post2 Aug 7, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch from 9d22f6b to 007fdc9 Compare August 8, 2026 10:50
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.16.post2 chore(deps): update dependency flashinfer-python to v0.6.16.post3 Aug 8, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch from 007fdc9 to 4197286 Compare August 16, 2026 00:12
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.16.post3 chore(deps): update dependency flashinfer-python to v0.6.17 Aug 16, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 3 times, most recently from dcd1cb0 to 0cc8a5f Compare August 25, 2026 13:43
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch from 0cc8a5f to ac43add Compare August 30, 2026 08:12
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.17 chore(deps): update dependency flashinfer-python to v0.6.18 Aug 30, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 3 times, most recently from ca992bb to 658a9c9 Compare September 5, 2026 15:51
@renovate renovate Bot changed the title chore(deps): update dependency flashinfer-python to v0.6.18 chore(deps): update dependency flashinfer-python to v0.6.18.post1 Sep 5, 2026
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 7 times, most recently from fb81258 to f802883 Compare September 14, 2026 13:17
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch 2 times, most recently from 1a85f15 to 8b950c0 Compare September 21, 2026 08:39
@renovate
renovate Bot force-pushed the renovate/flashinfer-python-0.x branch from 8b950c0 to 30dbed6 Compare September 22, 2026 13:31
@renovate

renovate Bot commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

Renovate Ignore Notification

Because you closed this PR without merging, Renovate will ignore this update (==0.7.0). You will get a PR once a newer version is released. To ignore this dependency forever, add it to the ignoreDeps array of your Renovate config.

If you accidentally closed this PR, or if you changed your mind: rename this PR to get a fresh replacement PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant