Skip to content

Add NoRender locomotion training and batched state access - #721

Draft
acrlw wants to merge 3 commits into
mainfrom
feat/locomotion-norender
Draft

acrlw wants to merge 3 commits into
mainfrom
feat/locomotion-norender

Conversation

@acrlw

@acrlw acrlw commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Description

G1 training without image observations previously used Hybrid and initialized backgrounds, lights, and materials. Joint/root state access, partial resets, and contact-history processing also submitted data repeatedly during rollout. This change integrates NoRender and DexSim's batched state APIs, reduces data movement in the training loop, and corrects throughput accounting after checkpoint restoration. Policy evaluation selects headless execution or the native Viewer through the renderer configuration.

1. NoRender training and policy evaluation

  • Add no-render to RenderCfg, the simulation CLI, and eval-policy. SimulationManager skips background, light, visual-material initialization, and Newton render-state publication in this mode while retaining the physical ground.
  • Require headless=True for NoRender and report configuration errors for explicit render synchronization or native camera creation. Both G1 Default/Newton environment and PPO configurations select NoRender.
  • Default Viewer evaluation to Hybrid when the saved training configuration uses NoRender; preserve an explicitly selected native renderer. Reject --viewer --renderer no-render before simulation creation.
  • Correct camera-extrinsic device placement and partial-environment indexing. Recording frames own contiguous memory so subsequent renders cannot overwrite submitted frames.

Primary implementation: embodichain/lab/sim/sim_manager.py, embodichain/lab/sim/cfg/simulation.py, and embodichain/learning/rl/policy_evaluation/cli.py.

2. Batched state access and partial resets

  • Add ArticulationData.fetch_state() to read joint positions/velocities and root poses/velocities together. Returned tensors use reusable buffers; retaining a historical state requires an explicit clone().
  • Add Articulation.set_state() to submit root state, joint state, control targets, and dynamics clearing together. Validate shapes before writing, clamp joint positions to their limits, and preserve unselected environment rows during partial resets.
  • Use DexSim's fetch_state, apply_state, and fetch_joint_properties in the Scene adapter, with centralized environment/joint selection and pose-layout conversion. Read joint limits and drive properties through the batch interface.
  • Combine locomotion reset writes and reset each sensor once per environment reset. Reuse static joint selections and observation indices after initialization.

Primary implementation: embodichain/lab/sim/objects/articulation.py, embodichain/lab/sim/objects/backends/scene.py, and embodichain_tasks/embodichain_tasks/locomotion/velocity/_embodichain.py.

3. Rollout and PPO data handling

  • Capture contact scattering and history updates in a CUDA Graph; recapture when buffers or history configuration change. Reuse registered Warp/Torch streams and retain event ordering between caller and sampling streams.
  • Reuse the base actor/critic state within one observation calculation while generating noise separately. Use countdowns for periodic events and velocity disturbances to reduce GPU index reads before an event is due.
  • Shallow-clone the TensorDict for GAE's added fields while sharing rollout tensors. Focused tests verify that input fields and values remain unchanged and policy parameters update.
  • Construct W&B scalar logs only when enabled. Record the starting step and time for each train() call, calculate throughput from newly collected transitions, and retain cumulative checkpoint counters.
  • Set G1 Newton solver limits to 20/50; use the native scene-node capacity default for G1 on both backends so NoRender creation stays within the supported range. Update the configuration tests. Add optional Newton contact priority, preserving the source value by default.

4. Regression coverage and local paired validation

The branch is based on main@d959eeed and retains main's action-contract interfaces. Follow-up commit 1630f972 synchronizes the Scene batch test doubles with the new state/property APIs, supplies the camera test's renderer configuration, accepts NoRender in packaged deployment checks, and classifies use_native_mesh_loader as an asset-loader option rather than a native shape property. It also adds regressions for per-environment batch properties, reusable state buffers, selected writes, negative batch status, set_state validation/clamping, and fresh-process NoRender lifecycle on both backends.

Validation on 2026-09-30 used this PR worktree and a freshly rebuilt local DexSim integration worktree at a6945b3e6. Core code loaded directly from this PR worktree; bundled tasks were installed from the same source. The benchmark and CUDA training tests were reused unchanged from #710, with task configuration resolved from this PR.

Check Result
State/reset, contact CPU paths, events, PPO, task configs, simulation configuration, evaluation CLI 434 passed; 2 CUDA cases excluded
Fresh-process NoRender world/URDF, stepping, selected reset, camera rejection, normal exit 2 passed: Default and Newton
G1/Go2 CUDA smoke and G1 checkpoint resume 6 passed across Default and Newton
Newton/MJWarp CUDA PPO with 4096 environments Finite rollout/losses and parameter update; checkpoint saved and loaded by independent evaluation processes
Hybrid camera Valid 160×120 RGB, pixel standard deviation 19.59
Headless policy evaluation 2 completed episodes, exit code 0
Native Viewer from NoRender training config Hybrid selected; 16 control steps / 64 physics steps, exit code 0
Black 26.3.1 7 follow-up Python files and repository-wide check passed
Public API documentation 2,322/2,322 exports documented
Project context Map check passed; existing capacity-default, renderer and state-buffer contracts remain accurate

CI follow-up on the current main merge

CI run 36660823947 tested the synthetic merge with main@7e56e321, rather than the PR branch alone. It reported 57 failures and 4 collection errors: 55 failures and all collection errors came from main's planner integration calling deepcopy without importing it; the other two failures came from the published DexSim package lacking the batch methods and NewtonCollisionDesc.priority. Follow-up commit 1282a68c supplies the standard-library import.

On that exact merge plus the import fix, 446 focused checks passed (1 deselected), covering the affected task/Gym configuration and package-resource paths, together with the two engine contracts using the local paired build. A broader local fast run yielded 5,663 passed, 3 failed, 7 skipped and 2 collection errors: local DexUni configuration differences, missing Gradio, and a packaging subprocess selecting an old binding. The packaging runtime source was corrected and its unchanged test passed in the focused run. This is not a claim that the full local suite or published-package CI is green.

The 4096-environment workload uses the repository G1 asset, NoRender, solver limits 20/50, and physics CUDA Graphs. Each update collects 24 control steps with four physics substeps per step. Three warmup updates precede five measured updates; PPO uses five epochs and four minibatches. Hardware/software: RTX 5090 D v2, Torch 2.10.0+cu128, Newton 1.6.0, MuJoCo/MJWarp 3.12.0, Warp 1.17.0.

Training SPS = 491,520 / Σ(rollout + PPO); rollout SPS = 491,520 / Σrollout. CUDA is synchronized at phase boundaries; creation and checkpoint operations are outside those throughput intervals. RAM and process VRAM are the maximum phase-boundary samples from scene creation through training. This short run verifies execution and throughput, not policy convergence.

Training SPS Rollout SPS Rollout(ms/update) PPO(ms/update) Per-update SPS CV Creation(s) RAM(MiB) Process VRAM(MiB)
205,991 247,540 397.12 80.10 2.95% 12.75 4,065.25 3,896

The benchmark verifies that the physics CUDA Graph is captured after warmup and rejects nonfinite values or MJWarp overflow. The process returned 0, but teardown still reported an entity-count difference of 32; one-environment headless evaluation reported a difference of 1. These warnings remain an engine cleanup acceptance item. The local run does not establish that teardown is warning-free.

5. Dependencies and merge readiness

  • Requires a DexSim package exposing Renderer.NORENDER, ArticulationBatch.fetch_state/apply_state/fetch_joint_properties, and NewtonCollisionDesc.priority.
  • CI still installs the published dexsim_engine==0.5.0 dependency. The local build retains the version string 0.5.0 but is identified by commit a6945b3e6; it does not replace the package installed by CI.
  • Keep this PR in Draft until those APIs are available in a published DexSim release, update the dependency pin, and rerun CI and paired acceptance using that release. Review the remaining teardown count warnings before merge.
  • Documentation and agent context cover renderer selection, batched-state buffer lifetime, contact history, and resumed-training SPS accounting. The follow-up tests/configuration fix requires no additional context change: the existing context already states that omitted capacity retains the native default.

Type of change

  • Enhancement: batched state access and training-loop data handling.
  • New feature: NoRender training and evaluation configuration.
  • Bug fix: resumed-training SPS, partial camera initialization, recording-frame ownership, and G1 NoRender scene capacity.
  • Documentation update.

Checklist

  • Formatted changed Python files and passed the repository-wide Black check.
  • Reviewed affected documentation and agent context; updated the API/contracts and explained why the follow-up needs no context edit.
  • Public API changes pass python docs/scripts/check_api_docs.py.
  • Added and synchronized relevant regression tests; completed local paired runtime validation.
  • Update the published DexSim dependency and pass CI / paired acceptance with that release.
  • Review the remaining teardown entity-count warnings.

@acrlw acrlw added enhancement New feature or request dexsim Things related to dexsim rl Features related to reinforcement learning labels Sep 29, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dexsim Things related to dexsim enhancement New feature or request rl Features related to reinforcement learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant