Skip to content

[perf] MiniMax-H3 LoRA training on a B200 node: 6.32 s to 2.53 s at 8 GPUs, mostly from the attention kernels - #1680

Draft
TarzanZhao wants to merge 3 commits into
modelscope:mainfrom
TarzanZhao:perf/minimax-h3-lora-training
Draft

[perf] MiniMax-H3 LoRA training on a B200 node: 6.32 s to 2.53 s at 8 GPUs, mostly from the attention kernels#1680
TarzanZhao wants to merge 3 commits into
modelscope:mainfrom
TarzanZhao:perf/minimax-h3-lora-training

MiniMax-H3: selective activation checkpointing, attention output kept

21badf8
Select commit
Loading
Failed to load commit list.

Workflow runs completed with no jobs