Add affine AutoEP checkpoint placement - #8544
Conversation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
|
Hi @jinyouzhi thank you for your PR. This PR overall looks good to me, I have two questions:
|
|
Thank you for comments. |
…e-autoep-checkpoint-placement Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> # Conflicts: # deepspeed/module_inject/auto_ep_layer.py
Persist versioned per-parameter affine maps in AutoEP checkpoint metadata and use them for conversion and restore, while retaining placement descriptors for provenance and legacy checkpoints. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
|
I completed an end-to-end loss-resume validation using the DeepSpeedExamples script (deepspeedai/DeepSpeedExamples#1014 ) The experiment trains an uninterrupted 100-step baseline, saves a native checkpoint at step 50, then resumes both directly from the native checkpoint and after ds_to_universal.py conversion. The converter completed the AutoEP ZeRO-3 expert-state consolidation successfully, and all runs had finite loss values. For every resumed step (51–100), the native-resume and Universal-checkpoint-resume losses are exactly identical. Their shared difference from the uninterrupted baseline is max_abs_error=0.1374 and mean_abs_error=0.01052 ; because the native and Universal paths match at every step, this difference is not introduced by the Universal/AffineIR conversion. I attached the loss plot below.
|
Use explicit placement rank IDs for local expert lookup and reject checkpoints whose persisted affine map conflicts with placement provenance. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
| @@ -288,19 +336,27 @@ def consolidate_autoep_zero12_expert_states(temp_dir, output_dir, expert_param_i | |||
|
|
|||
There was a problem hiding this comment.
is consolidate_autoep_zero12_expert_states still needed?
There was a problem hiding this comment.
also if it is no longer called, does it mean this functionality is lost? Is it intentionally not called?
Consolidate ZeRO FP32 master expert weights and optimizer states after per-expert model files, and verify native and Universal loss-resume parity for ZeRO stages 1 and 2. Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
|
fix ZeRO1/2 issues and also have a test with ZeRO-2 @delock
|
Remove the unreachable expp_rank optimizer consolidation helper; ZeRO-1/2 optimizer state is handled by the active shard merger. Keep expert-file FP32 fallbacks in the temporary directory and let the merger write each final tensor once. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
Head branch was pushed to by a user without write access
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
1a6df6a to
87a00c9
Compare
Restore now takes this rank's shard from the map the layer published, which is the same piece list conversion uses run in the opposite direction. That completes the symmetry §7 of `affine_ir_spec.md` describes and that conversion has had since deepspeedai#8385: until now the two sides described the same layout in two different ways, and could disagree. Follows deepspeedai#8519 (layers emit the map) and deepspeedai#8575 (GPTBigCode and Yuan describe themselves). ### This one is not additive The map lives on the parameter, built by this job's layer. It is `M_t` — it describes the topology being restored into, regardless of how the source checkpoint was converted. So the new path is taken for **every** AutoTP restore, not only for checkpoints that were converted through a map. A parameter whose layer published no map falls back to the existing category keys unchanged. That is the correct reading of the spec rather than a shortcut: §6.1 keeps `M_t` out of the file precisely because the restoring job derives it from its own layers. But it does mean the blast radius here is wider than the three PRs before it, so the equivalence is worth checking rather than assuming. I ran `tests/unit/checkpoint/` and `tests/unit/runtime/tensor_parallel/` (385 tests) against this branch and against upstream master in a separate worktree. The failing set is **identical by name**, not merely the same count — 75 failed, 187 passed either way, all of them FusedAdam JIT, fp16-on-CPU, or the pre-existing `st_st_size` typo in `test_convert_checkpoint.py`. ### The tests are built to discriminate `test_restore_prefers_the_map_over_the_category_keys` gives one parameter deliberately contradictory geometry: the category keys describe a plain dim-0 split, so rank 0 would take the first half, while the map hands rank 0 the second half. Which half rank 0 receives says which side of the metadata restore actually consulted. This matters because the old path covers the same layouts correctly. Without a discriminating case the map branch could be deleted and every existing test would stay green. I checked: removing the branch fails that test, while `test_restore_falls_back_when_no_map_was_published` keeps passing. ### Validation 112 passed on CPU/gloo (`DS_ACCELERATOR=cpu LOCAL_SIZE=4`): the affine suite, the resume matrix, the coverage test, `tests/unit/runtime/tensor_parallel/`, and deepspeedai#8544's AutoEP affine suite, which this now sits on top of after the rebase. ### Scale powers Restore resolves a partition once per state file, so each state takes its own power of a piece's scale. The power table lives in affine.py as SCALE_POWER_BY_STATE — ds_to_universal.py had its own copy since deepspeedai#8519, so conversion and restore now read one definition rather than two that could drift. This also reaches the refusal built for scaled optimizer states, which only triggers on a non-parameter power. No layer emits a scaled map today, so it is ahead of the path rather than behind it. Related: deepspeedai#8252, deepspeedai#8230. cc @delock --------- Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com>


Follow up #8385
This pull request introduces support for flexible expert placement in DeepSpeed's AutoEP (Automatic Expert Placement) system, enabling non-uniform, non-contiguous, and replicated expert layouts. The main changes add a versioned expert placement descriptor, validation logic, and integration into checkpoint consolidation and metadata validation. This lays the groundwork for more advanced expert scheduling and model parallelism strategies.
The most important changes are:
AutoEP Expert Placement Descriptor and Affine Map Lowering
autoep_affine.pythat defines the expert placement descriptor, validation, legacy uniform descriptor synthesis, and lowering to affine maps for sharded tensor reconstruction. This enables flexible, versioned expert placement beyond the legacy uniform contiguous layout.Integration into Checkpoint Consolidation and Metadata
autoep_universal.pyto:num_local_experts * ep_size == num_expertswhen a placement is provided.Metadata Validation Enhancements
autoep_zero3_metadata.pyto:Documentation Updates
affine_ir_spec.md) to document the new AutoEP placement descriptor, its semantics, and its integration into the IR and runtime, clarifying the distinction between placement provenance and scheduling.Bugfixes and Robustness
affine.pyby skipping empty piece lists during tensor rebuilding, preventing errors when a rank has no assigned pieces.Related: #8252, #8230.