From 833de7aaf214d91dfd3eca4e13c7f50a0ed76074 Mon Sep 17 00:00:00 2001 From: Wenyao Gao Date: Sat, 19 Sep 2026 23:59:08 -0700 Subject: [PATCH 1/3] chore: refresh the complete GB300 Kimi-K3 AgentX curve MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The 6 published GB300 Kimi-K3 disaggregated points come from the 2026-09-14 sweep, whose recipes did not yet declare the dcgm-power provider, so the job logs report USES_DCGM_POWER=0 and the points carry power_valid=0 with no power values. #3047 enabled the provider on those recipes, but its run was purged, so nothing power-enabled has been published for them. AgentX collapses spec_method, disagg and offload_mode into the curve scope, so a hardware resolves to exactly one curve and a partial selection would replace it rather than complete it. The entry selects both GB300 config keys and re-measures the full 11-point curve. 中文:已发布的 6 个 GB300 Kimi-K3 分离部署点来自 2026-09-14 的 sweep, 当时配方尚未声明 dcgm-power provider,作业日志显示 USES_DCGM_POWER=0, 这些点因此是 power_valid=0 且没有功耗数值。#3047 已为这些配方启用该 provider,但其运行被清除,所以至今没有带功耗的已发布数据。AgentX 把 spec_method、disagg 与 offload_mode 折叠进 curve scope,每个硬件只解析出 一条曲线,只选一部分会替换而不是补全曲线。本条目选中两个 GB300 config key,一次重测完整的 11 点曲线。 --- perf-changelog.yaml | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/perf-changelog.yaml b/perf-changelog.yaml index c5b4309f5f..5c6883278e 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -8455,3 +8455,13 @@ description: - "Update B200 vLLM AgentX to DSpark6 and a new image with TP8 and DEP8 configurations." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3274 + +- config-keys: + - kimik3-fp4-gb300-dynamo-vllm-agentic-dspark-mooncake-dcp8-agg + - kimik3-fp4-gb300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg + scenario-type: + - agentic-coding + description: + - "Refresh the complete GB300 Kimi-K3 AgentX curve so every published point carries measured power. The 6 published disaggregated points come from a 2026-09-14 sweep whose recipes did not yet declare the dcgm-power provider, so they carry no power values, and AgentX resolves one curve per hardware, so only a sweep selecting both config keys can replace them together." + - "重测 GB300 Kimi-K3 AgentX 整条曲线,使每个已发布的点都带实测功耗。已发布的 6 个分离部署点来自 2026-09-14 的 sweep,当时配方尚未声明 dcgm-power provider,因此没有功耗数值;而 AgentX 每个硬件只解析出一条曲线,因此只有同时选中两个 config key 的 sweep 才能整体替换它们。" + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3315 From ac57adaad178124ad016d12c74c8e63b39766229 Mon Sep 17 00:00:00 2001 From: Wenyao Gao Date: Sun, 20 Sep 2026 12:22:30 -0700 Subject: [PATCH 2/3] fix: migrate the GB300 Kimi-K3 recipes to the current telemetry schema MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit srtctl rejected the four disagg recipes with `Invalid config ... {'telemetry': {'provider': ['Unknown field.'], 'default_frequency': ['Unknown field.']}}` before submitting anything. When utils/srt-slurm moved to the upstream pin, `provider` and `default_frequency` were retired for `dcgm_exporter` and `collect_interval_ms`; 140 recipes migrated and six Kimi-K3 ones did not. That is why the disaggregated half of the GB300 curve failed in `Launch multi-node job script` while the aggregated half, already migrated, passed. default_frequency was a period in seconds, so 1.0 becomes collect_interval_ms: 1000, matching agg-dcp8-dspark4-mooncake.yaml in the same directory. 中文:srtctl 以 `Invalid config ... {'telemetry': {'provider': ['Unknown field.'], 'default_frequency': ['Unknown field.']}}` 拒绝了四个 分离部署配方,尚未提交任何任务。utils/srt-slurm 切到上游 pin 时,provider 与 default_frequency 已被 dcgm_exporter 和 collect_interval_ms 取代;140 个 配方完成了迁移,六个 Kimi-K3 配方没有。这正是 GB300 曲线中分离部署的一半在 `Launch multi-node job script` 阶段失败、而已迁移的聚合部署一半通过的原因。 default_frequency 的单位是秒,因此 1.0 对应 collect_interval_ms: 1000,与同 目录下的 agg-dcp8-dspark4-mooncake.yaml 一致。 --- .../agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake.yaml | 3 +-- .../agentx/disagg-1p2d-dcp8-dcp8-dspark4-mooncake.yaml | 3 +-- .../agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake.yaml | 3 +-- .../agentx/disagg-1p3d-dcp8-dcp8-dspark7-mooncake.yaml | 3 +-- perf-changelog.yaml | 2 ++ 5 files changed, 6 insertions(+), 8 deletions(-) diff --git a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake.yaml b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake.yaml index dd9a30f069..388929f9fb 100644 --- a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake.yaml +++ b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake.yaml @@ -174,8 +174,7 @@ srun_options: telemetry: enabled: true - provider: dcgm-power - default_frequency: 1.0 + collect_interval_ms: 1000 storage_subdir: power required: true startup_timeout_seconds: 120 diff --git a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p2d-dcp8-dcp8-dspark4-mooncake.yaml b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p2d-dcp8-dcp8-dspark4-mooncake.yaml index 6ea49208c6..38548cbc16 100644 --- a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p2d-dcp8-dcp8-dspark4-mooncake.yaml +++ b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p2d-dcp8-dcp8-dspark4-mooncake.yaml @@ -174,8 +174,7 @@ srun_options: telemetry: enabled: true - provider: dcgm-power - default_frequency: 1.0 + collect_interval_ms: 1000 storage_subdir: power required: true startup_timeout_seconds: 120 diff --git a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake.yaml b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake.yaml index 4bf4ecc656..c0a99ae534 100644 --- a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake.yaml +++ b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake.yaml @@ -174,8 +174,7 @@ srun_options: telemetry: enabled: true - provider: dcgm-power - default_frequency: 1.0 + collect_interval_ms: 1000 storage_subdir: power required: true startup_timeout_seconds: 120 diff --git a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark7-mooncake.yaml b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark7-mooncake.yaml index 0cd2a16883..5467b1df25 100644 --- a/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark7-mooncake.yaml +++ b/benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/gb300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark7-mooncake.yaml @@ -174,8 +174,7 @@ srun_options: telemetry: enabled: true - provider: dcgm-power - default_frequency: 1.0 + collect_interval_ms: 1000 storage_subdir: power required: true startup_timeout_seconds: 120 diff --git a/perf-changelog.yaml b/perf-changelog.yaml index 5c6883278e..536ba821c9 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -8463,5 +8463,7 @@ - agentic-coding description: - "Refresh the complete GB300 Kimi-K3 AgentX curve so every published point carries measured power. The 6 published disaggregated points come from a 2026-09-14 sweep whose recipes did not yet declare the dcgm-power provider, so they carry no power values, and AgentX resolves one curve per hardware, so only a sweep selecting both config keys can replace them together." + - "Migrate the four disagg Kimi-K3 recipes to the current srtctl telemetry schema: provider and default_frequency were retired for dcgm_exporter and collect_interval_ms when the submodule moved to the upstream pin, and srtctl rejected them as Unknown field before submitting anything." - "重测 GB300 Kimi-K3 AgentX 整条曲线,使每个已发布的点都带实测功耗。已发布的 6 个分离部署点来自 2026-09-14 的 sweep,当时配方尚未声明 dcgm-power provider,因此没有功耗数值;而 AgentX 每个硬件只解析出一条曲线,因此只有同时选中两个 config key 的 sweep 才能整体替换它们。" + - "将四个分离部署的 Kimi-K3 配方迁移到当前的 srtctl telemetry schema:submodule 切到上游 pin 时,provider 与 default_frequency 已被 dcgm_exporter 和 collect_interval_ms 取代,srtctl 在提交作业前就以 Unknown field 拒绝了它们。" pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3315 From ec8f11ad19d36a29805261f36cf0bb65de45d716 Mon Sep 17 00:00:00 2001 From: Wenyao Gao Date: Mon, 21 Sep 2026 00:00:23 -0700 Subject: [PATCH 3/3] docs: trim GB300 Kimi-K3 changelog MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 精简 GB300 Kimi-K3 更新日志,保留完整曲线功耗重测与 telemetry schema 迁移说明。 [skip-sweep] Wording only; benchmark configuration and selection are unchanged. --- perf-changelog.yaml | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/perf-changelog.yaml b/perf-changelog.yaml index 536ba821c9..9f2c9de535 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -8462,8 +8462,8 @@ scenario-type: - agentic-coding description: - - "Refresh the complete GB300 Kimi-K3 AgentX curve so every published point carries measured power. The 6 published disaggregated points come from a 2026-09-14 sweep whose recipes did not yet declare the dcgm-power provider, so they carry no power values, and AgentX resolves one curve per hardware, so only a sweep selecting both config keys can replace them together." - - "Migrate the four disagg Kimi-K3 recipes to the current srtctl telemetry schema: provider and default_frequency were retired for dcgm_exporter and collect_interval_ms when the submodule moved to the upstream pin, and srtctl rejected them as Unknown field before submitting anything." - - "重测 GB300 Kimi-K3 AgentX 整条曲线,使每个已发布的点都带实测功耗。已发布的 6 个分离部署点来自 2026-09-14 的 sweep,当时配方尚未声明 dcgm-power provider,因此没有功耗数值;而 AgentX 每个硬件只解析出一条曲线,因此只有同时选中两个 config key 的 sweep 才能整体替换它们。" - - "将四个分离部署的 Kimi-K3 配方迁移到当前的 srtctl telemetry schema:submodule 切到上游 pin 时,provider 与 default_frequency 已被 dcgm_exporter 和 collect_interval_ms 取代,srtctl 在提交作业前就以 Unknown field 拒绝了它们。" + - "Refresh the full GB300 Kimi-K3 AgentX curve with measured power across aggregated and disaggregated deployments." + - "Update four disaggregated recipes to use srtctl's dcgm_exporter and collect_interval_ms telemetry fields." + - "重测 GB300 Kimi-K3 AgentX 完整曲线,为聚合和分离部署补齐实测功耗。" + - "将四个分离部署配方的 telemetry 字段更新为 srtctl 的 dcgm_exporter 和 collect_interval_ms。" pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3315