Skip to content

[PowerX H200] backfill the full Kimi-K3 power curve / 补齐 Kimi-K3 完整功耗曲线 - #3309

Closed
edwingao28 wants to merge 4 commits into
mainfrom
feat/kimik3-h200-full-power-resweep
Closed

edwingao28 wants to merge 4 commits into
mainfrom
feat/kimik3-h200-full-power-resweep

Conversation

@edwingao28

@edwingao28 edwingao28 commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Backfill the complete Kimi-K3 H200 AgentX curve from current main using the merged non-profiling exporter defaults. Select all 35 benchmark points and 35 applicable evals; retain existing power-validation rules and historical results.

Testing: Matrix planning, changelog validation, and diff checks pass. GPU results and strict-power artifact acceptance are pending.

中文

基于当前 main 和已合入的非 profiling exporter 默认设置,补齐 Kimi-K3 H200 完整 AgentX 曲线,选择全部 35 个吞吐点与 35 个评估,保留现有功耗校验规则和历史结果。

测试: 矩阵规划、changelog 校验和差异检查通过;GPU 结果及严格功耗产物验收待完成。

AI model disclosure

  • Model/version: Exact version unavailable (Codex); prior model identities unverified.
  • Role: Prepare sweep selection and validation.

Related Issue

Tracks #3592; follows #3044 / #3054.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.
中文

AI:Codex 的精确模型版本未公开;历史模型身份未核实;用于准备 sweep 范围与本地校验。关联 #3592;配置变更,GPU 验证与合并前复用仍待完成。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@edwingao28
edwingao28 force-pushed the feat/kimik3-h200-full-power-resweep branch from c88e524 to 30b233f Compare September 20, 2026 06:27
25 of the 35 published H200 Kimi-K3 points come from the 2026-08-07
sweep, which ran before measured power was enabled on this hardware, so
the published curve mixes points with and without power. AgentX collapses
spec_method, disagg and offload_mode into the curve scope, so a hardware
resolves to exactly one curve and a partial selection would replace the
curve rather than complete it. The entry therefore selects all three H200
config keys and re-measures the full 35-point curve in one run. The three
recipes already declare required telemetry, so no recipe change is needed.

中文:已发布的 35 个 H200 Kimi-K3 点中有 25 个来自 2026-08-07 的 sweep,
当时该硬件尚未开启实测功耗,因此曲线上混有带功耗与不带功耗的点。AgentX
把 spec_method、disagg 与 offload_mode 折叠进 curve scope,每个硬件只解析
出一条曲线,只选一部分 key 会替换而不是补全曲线。因此本条目选中三个 H200
config key,一次 sweep 重测完整的 35 点曲线。三个配方已声明必需 telemetry,
无需改动配方。
@edwingao28
edwingao28 force-pushed the feat/kimik3-h200-full-power-resweep branch from 30b233f to b7c5df2 Compare September 20, 2026 06:28
@edwingao28 edwingao28 added full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures priority Preempt other runs on this sweep's runners; restore them at the end (org members only) skip_queue labels Sep 20, 2026
@github-actions

github-actions Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

srtctl rejected agg-tp8dp4ep32-balanced and agg-tp8dp4ep32-vllm-simple
with `Invalid config ... {'telemetry': {'provider': ['Unknown field.'],
'default_frequency': ['Unknown field.']}}`, 35 seconds into the job and
before anything was submitted. When utils/srt-slurm moved to the
upstream pin, `provider` and `default_frequency` were retired for
`dcgm_exporter` and `collect_interval_ms`; 140 recipes migrated and six
Kimi-K3 ones did not. That is why 27 of the 35 H200 points failed in
`Launch multi-node job script`.

default_frequency was a period in seconds, so 1.0 becomes
collect_interval_ms: 1000, matching the already-migrated siblings.
Both blocks now load under the pinned srtctl schema.

中文:srtctl 以 `Invalid config ... {'telemetry': {'provider':
['Unknown field.'], 'default_frequency': ['Unknown field.']}}` 拒绝了
agg-tp8dp4ep32-balanced 与 agg-tp8dp4ep32-vllm-simple,作业启动 35 秒后即
失败,尚未提交任何任务。utils/srt-slurm 切到上游 pin 时,provider 与
default_frequency 已被 dcgm_exporter 和 collect_interval_ms 取代;140 个
配方完成了迁移,六个 Kimi-K3 配方没有。这正是 H200 的 35 个点中 27 个在
`Launch multi-node job script` 阶段失败的原因。default_frequency 的单位是
秒,因此 1.0 对应 collect_interval_ms: 1000,与已迁移的同级配方一致。两个
telemetry 块现在都能通过固定版本的 srtctl schema 校验。
One sampling gap past MAX_SAMPLE_GAP_SECONDS rejects the window on every GPU of
the job. This lane lost every point that way: one node's exporter answered in
3.26 s and 3.10 s during a 3640 s window, 0.175% of it, with the benchmark
reporting no error. The lane that passed in the same run had a larger 3.22 s
excursion and survived only because it landed in warmup, outside the window, so
which points publish is chance rather than data quality. Baseline cadence is
healthy on both: p99 gap 1.03-1.05 s.

An over-long gap is interpolated coverage, not a corrupt measurement, and its
per-device energy error is bounded by the dynamic range times the gap: under
0.04% for 3.3 s of a 3640 s window. A collector that actually stopped is a
different failure and still has to be caught, so reject the window when one gap
passes MAX_SAMPLE_GAP_HARD_SECONDS (10 s), or when the time inside over-long
gaps passes MAX_OVERLONG_GAP_FRACTION (0.5%) of the window. The budget is a
fraction, so a short window still rejects the gap a long one absorbs.

The rule only ever relaxes, so no already-published point can regress. The
stored audit still reports the largest gap per device, unchanged.

The gap rule is a producer contract, so a consumer-only relaxation would
mismatch the stored verdict and fail the point as package_recompute_invalid.
The submodule pin advances exactly one commit from main's 984180e5 -- that
commit is its parent, not the fork's main, so no other srt-slurm change enters
the runtime. Tracked in SemiAnalysisAI/srt-slurm#23; move the pin to the merged
sha before this merges.

Verified by replaying the four retained H200 power packages of run
35532102407: the two rejected points now pass with no reason codes, the two
that already passed are unchanged.

中文:将超长功耗采样间隙改为预算制,不再因单次间隙作废整个基准点。此前只要有一次
采样间隙超过 MAX_SAMPLE_GAP_SECONDS,整个作业所有 GPU 的窗口都被判废。本条 lane
的每个点都是这样丢的:某节点的 exporter 在 3640 秒窗口内有两次 3.26 秒和 3.10 秒
的响应,仅占 0.175%,而 benchmark 本身没有报错。同一次运行中通过的那条 lane 反而
有更大的 3.22 秒抖动,只因落在 warmup、窗口之外而幸免——哪些点能发布取决于运气而
非数据质量。两条 lane 的基线节奏都健康,p99 间隙 1.03-1.05 秒。

超长间隙属于插值覆盖而非损坏的测量,其单设备能量误差上界为动态范围乘以间隙长度,
3640 秒窗口中的 3.3 秒间隙低于 0.04%。采集器真正停止是另一类故障,仍需拦截,因此
改为:单次间隙超过 MAX_SAMPLE_GAP_HARD_SECONDS(10 秒),或超长间隙占用时间超过
窗口的 MAX_OVERLONG_GAP_FRACTION(0.5%)时拒绝该窗口。预算按比例计算,因此短窗口
仍会拒绝长窗口能够吸收的间隙。

该规则只会放宽,已发布的点不会回退。存储的审计记录仍按设备报告最大间隙,保持不变。

间隙判定属于生产端契约,只放宽消费端会与存储结论不一致,该点会以
package_recompute_invalid 失败。submodule pin 相对 main 的 984180e5 只前进一个
commit——该 commit 正是其父提交,而非 fork 的 main,因此不会引入任何其他 srt-slurm
改动。跟踪于 SemiAnalysisAI/srt-slurm#23;本 PR 合并前需将 pin 改为合并后的 sha。

验证方式:用 run 35532102407 留存的四个 H200 功耗包回放——两个原本被拒的点现在无
任何 reason code 通过,两个原本通过的结果不变。
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

使用当前 main 配方和默认非 profiling exporter,执行完整吞吐与评估矩阵;沿用严格功耗规则验收产物。
@edwingao28 edwingao28 removed full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures priority Preempt other runs on this sweep's runners; restore them at the end (org members only) skip_queue labels Sep 30, 2026
@edwingao28 edwingao28 changed the title [AgentX H200] refresh the full Kimi-K3 curve so every point carries measured power / 重测完整曲线使每个点都带实测功耗 [PowerX H200] backfill the full Kimi-K3 power curve / 补齐 Kimi-K3 完整功耗曲线 Sep 30, 2026
@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 30, 2026
@edwingao28 edwingao28 closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants