Skip to content

fix(trainer-rank): account for recomputed mixer memory and head stages - #963

Draft
bradhilton wants to merge 52 commits into
mainfrom
dalinar/recompute-mixer-floor
Draft

bradhilton wants to merge 52 commits into
mainfrom
dalinar/recompute-mixer-floor

Conversation

@bradhilton

@bradhilton bradhilton commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Agent: Dalinar

Full one-layer recompute keeps a layer’s mixer live beside its MoE stage. Price that overlap and previously hidden residual, norm, GDN, shared-expert, routing and first-call TE workspaces. Where every MoE layer’s recompute is covered, charge one incoming gradient instead of one per boundary. HybridEP prices one dispatched copy on the EP group’s balanced rows when EP and CP groups coincide.

Price eligible warmed CP2 target-only fused head statistics as a separate backward stage; unsupported heads retain the existing floor. EP routing allowances become 1.4/1.6/2.0 for EP2/4/8, based on measured routing. Reports freeze these facts for CPU replay. This branch includes #1002; review that first. #978 adds per-rank layouts and #981 adds observations.

Brad’s decision remains required for strict fused statistics: a wave admitted using the staged head price raises if its fused kernel fails instead of using a wider FP32 fallback. Its CP peer can wait at the next collective, as with a rank-local OOM.

Recorded trainer-rank/replay CPU tests and local H200 runs passed. On the full stack with #1002, Qwen3.6-35B-A3B CP2 single-sequence cold errors were +0.9–2.0% at EP1 and +3.0–13.4% at EP2; warm errors were +2.8–24.9%. No measured case was under. Staged-head evidence covers 29 CP2 EP1/EP2 waves. These measurements do not establish a general peak bound.

EP2 still exceeds the 10% accuracy target, largely from the routing allowance. Routing measurements cover sampled workloads only; EP4/EP8 have no memory runs, and a small EP8 sample approached its 2.0 allowance. PyTorch peaks omit the HybridEP buffer and other external memory: one cold EP2 case grew 2.76 GB outside PyTorch while the estimate included only its 1.05 GB HybridEP buffer. A warm routing discount is separate work.

@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 06:42 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 07:09 — with GitHub Actions Failure
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 07:20 — with GitHub Actions Active
@bradhilton
bradhilton force-pushed the dalinar/recompute-mixer-floor branch from 1ffff5e to 4b3d1e0 Compare September 25, 2026 15:47
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 15:47 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 15:55 — with GitHub Actions Failure
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 15:59 — with GitHub Actions Active
bradhilton and others added 8 commits September 25, 2026 17:17
…kpoint floor

Full one-layer recompute replays a layer with gradients, so its mixer's
saved activations stay live beside that layer's MoE stage. The checkpoint
floor priced boundaries and the MoE stage only, which left context-parallel
runs short: Qwen3.6-35B-A3B at CP2 peaked 9-11 GB above the floor on the
most loaded rank. Price the larger of the model's attention and GDN mixers
per recomputed row, with context-parallel stage buffers and GDN exchange
copies, from allocator traces at CP1 and CP2.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Price GDN from its saved tensors (norm output, q/k with fp32 l2norm copies,
v, z, segment-layout tensors, gated norm and the chunk decay matrix) instead
of a ratio fit, and its context-parallel exchanges from hidden and value
widths rather than the key width. Divide CP attention extras by TP like the
retained widths, and say CP above 2 reuses the CP2 allowance.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The l2-normalized q and k are expanded to the value heads before they are
saved, so their width follows value_heads * key_head_dim, not twice the key
width. Qwen3.6 is unchanged; geometries with more value than key heads were
under-priced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The EP1 all-to-all holds its permuted copy and the exchanged rows at the
expert stage; HybridEP permutes while it dispatches and returns one tensor.
A Qwen3.6 CP2/EP2 allocator trace holds exactly one routed H-wide input
beside the FC1 and FC2 stage tensors (9,728 features per routed row), where
the planner charged two (11,776).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…pendent

Backward recomputes the last layer first, so the checkpoint floor's peak meets
every saved boundary but only the one incoming gradient. Where the MoE stage
is priced, charge that gradient instead of one per boundary (39 hidden rows
per token too many at 40 layers), and price what the old allowance was
silently covering, all from Qwen3.6-35B-A3B allocator traces:

- the recomputed layer's residual and pre-MLP norm output (2H per row);
- GDN's sixth value-width tensor (the projected q/k/v includes v);
- the shared expert's saved FC1 gate/up and GLU outputs;
- router scores and map plus the dispatcher's row-id map (EP1) or probability
  copy and handle (HybridEP);
- TE's cuBLAS workspaces, as growth until its GEMMs allocate them.

Without a priced MoE stage the per-boundary allowance stays: it also covers
dense MLP and other recompute work the floor does not price.

The EP>1 routed-row allowance becomes EP-dependent (1.4, 1.6, 2.0 at EP2, 4,
8), from pretrained Qwen3.6 routing of 3.5M tokens of retail agent
trajectories (worst layer 1.21, 1.41, 1.63) and one production EP2 run (1.35).
Routed rows are no longer rounded up to whole rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton force-pushed the dalinar/recompute-mixer-floor branch from 2b54d0c to c59a692 Compare September 25, 2026 18:50
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 18:51 — with GitHub Actions Failure
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 19:03 — with GitHub Actions Active
HybridEP dispatches the whole EP group's rows. When that group is this
rank's CP group, a balanced rank receives the group's rows over EP, not
the busiest CP rank's share: a CP2/EP2 real-data trace put 52,480 rows on
one rank while each layer dispatched exactly 8 x 96,794 pairs across both.
Price only the routed part (and its converted stages) on that share; the
shared expert, mixer and boundaries stay on local rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 20:12 — with GitHub Actions Failure
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 20:14 — with GitHub Actions Failure
Charge only the one incoming gradient when every decoder layer is a priced
MoE layer that encloses its FC1 stage, for each gradient group's slot. A
positive FC2-only coefficient, dense layers or a slot that reprices to zero
keep one gradient per boundary, which also covers unpriced recompute work.

Count an empty CP rank's padding row in the EP group's total: dispatch runs
at least one row per rank, and that row is routed too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 25, 2026 20:48 — with GitHub Actions Active
A layer counts as enclosed only if its FC1 converted stages are priced too,
unless FC1 has no adapter or the selected slot has no FC1 tensors. A slot
with FC1 adapters but no FC2 adapter prices FC2 rows from the original
metadata yet skips the whole converted-stage block, so it now keeps one
gradient per boundary. A slot's walk must enclose as many layers as the
constructor's, which already matched every decoder layer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton had a problem deploying to trainer-rank-gpu-validation September 25, 2026 21:08 — with GitHub Actions Error
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton and others added 13 commits September 29, 2026 03:51
Grouped planner replay (#1028) recomputes each subforward's cost from
primitive runtime facts, with group indices standing in for slot refs. The
adapter-gradient floor reads the gradient slot's unallocated LoRA gradients
from the live model, so replay could neither resolve the slot nor reproduce
the term. Capture each gradient group's slot kind and name and its pending
gradient bytes per decoder layer with the selection (runtime facts version
2), and have ReplayRank answer the floor's slot and pending-gradient readers
from those facts. The capture's stock-estimator check now covers the floor's
readers too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replay sizes per-layer boundary tuples from the report's recorded
num_layers, which only capture bounded; refuse counts outside capture's
1024-layer limit before any estimator runs. A slot without a kind (a
megatron-less reference) has no pending gradients, so reject kindless facts
that claim some. Test forged adapter facts, the layer bound, and an
instance-overridden pending-gradient reader.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…layout

Port #963 (through 35850f7) onto ART's extracted planner modules: the
recomputed layer's attention/GDN activations, routing state and TE
workspaces in the checkpoint floor; routed rows per group; and the head
backward priced as its own stage where traced. The planner-miss replay
freezes the new model and process readers (facts version 3) and records
the plan's head staging and routed rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Validate the recorded geometry and rank dimensions the recomputed-mixer
arithmetic reads; accept any live Triton threshold and mirror live's
empty head for non-positive rows; read the recompute and head-stage
readers only with a checkpointed decoder, as live pricing does; check
the producers of recorded staging and routed rows; and refuse a staged
head that the recorded topology and head facts cannot produce.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Update the fact-budget test's forged MoE terms to the shared-coefficient
shape.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 29, 2026 06:22 — with GitHub Actions Active
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton
bradhilton deployed to trainer-rank-gpu-validation September 29, 2026 06:46 — with GitHub Actions Active
bradhilton added a commit that referenced this pull request Sep 29, 2026
Per-rank CP layouts: the head stage takes the largest rank's adapter term
over its own boundaries, sharing the per-rank walk with the decoder stage's
extra.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bradhilton added a commit that referenced this pull request Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@bradhilton bradhilton changed the title Price the recomputed layer's attention or GDN activations in the checkpoint floor fix(trainer-rank): account for recomputed mixer memory and head stages Oct 2, 2026

This branch was successfully deployed

1 active deployment
trainer-rank-gpu-validation — 1763c67d Deployed Sep 29, 2026 by bradhilton via Run on 2x H200 #975
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant