diff --git a/configs/nvidia-master.yaml b/configs/nvidia-master.yaml index ee6770e3e0..a8c2378c28 100644 --- a/configs/nvidia-master.yaml +++ b/configs/nvidia-master.yaml @@ -8105,6 +8105,12 @@ dsv41flash-fp4-gb200-vllm-agentic-dspark: search-space: # Engram weights use UVA DRAM; the KV cache stays GPU-resident. - { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] } + # TP2 halves the GPU count per replica. Weights rise to ~145 GiB on each + # 256 GiB GPU with the Engram tables in pinned host DRAM, so the arm + # takes the same caps as the B200 TP2 arm: batched tokens 4096 (the + # indexer's logits buffer is 32 GiB at the upstream 16384) and graph + # capture stopped at 512, leaving ~49 GiB of KV per GPU. + - { tp: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] } # SGLang arm for DeepSeek-V4.1-Flash AgentX on GB200, from the SGLang cookbook # (https://lmsysorg.mintlify.app/cookbook/autoregressive/DeepSeek/DeepSeek-V4_1). diff --git a/perf-changelog.yaml b/perf-changelog.yaml index 62f4d853ec..55326b2d45 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -8492,3 +8492,14 @@ - "Floor the TP2 --max-num-seqs at 16 (it stays 2x the concurrency above that, capped at 256): in run 35320655804 c1, c2 and c4 never started because FlashInfer's autotune dummy run batches max-num-seqs requests through the DSpark draft head and with 2-8 requests selected an invalid MXFP8 split-K tactic ((128, 8), (1, 1), True, False, 4), while c8 with 16 seqs and every larger point served" - "将 TP2 的 --max-num-seqs 下限设为 16(高于该值时仍为并发数的 2 倍,上限 256):运行 35320655804 中 c1、c2、c4 未能启动,原因是 FlashInfer 的 autotune dummy run 会让 max-num-seqs 个请求经过 DSpark draft head,在 2-8 个请求时选中了无效的 MXFP8 split-K tactic ((128, 8), (1, 1), True, False, 4),而 c8(16 seqs)及更大的点均正常服务" pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3216 + +- config-keys: + - dsv41flash-fp4-gb200-vllm-agentic-dspark + scenario-type: + - agentic-coding + description: + - "Add a TP2 arm to the GB200 vLLM DeepSeek-V4.1-Flash AgentX recipe alongside the existing TP4 arm, keeping the Engram tables in pinned host DRAM via --engram-config cpu_offload; weights rise to ~145 GiB on each 256 GiB GPU" + - "Apply the B200 TP2 caps in dsv41flash_fp4_vllm_mtp.sh to every TP2 arm: --max-num-batched-tokens 4096 (the sparse-attention indexer's [batched-tokens, 1M] fp8 buffer is 32 GiB at the upstream 16384), --max-num-seqs at twice the concurrency (16-256) and CUDA graph capture stopped at 512, which leaves ~49 GiB of KV per GPU; TP4 and TP8 arms keep the upstream defaults" + - "为 GB200 vLLM DeepSeek-V4.1-Flash AgentX 配方在现有 TP4 臂旁新增 TP2 臂,Engram 表继续通过 --engram-config cpu_offload 放在固定页主机 DRAM;每张 256 GiB GPU 的权重升至约 145 GiB" + - "在 dsv41flash_fp4_vllm_mtp.sh 中将 B200 TP2 的上限推广到所有 TP2 臂:--max-num-batched-tokens 4096(上游 16384 时 indexer 的 [batched-tokens, 1M] fp8 缓冲区达 32 GiB)、--max-num-seqs 为并发的两倍(16-256)、CUDA graph 捕获上限 512,为每张 GPU 留出约 49 GiB KV;TP4 与 TP8 臂沿用上游默认值" + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3320