Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
G1 training without image observations previously used Hybrid and initialized backgrounds, lights, and materials. Joint/root state access, partial resets, and contact-history processing also submitted data repeatedly during rollout. This change integrates NoRender and DexSim's batched state APIs, reduces data movement in the training loop, and corrects throughput accounting after checkpoint restoration. Policy evaluation selects headless execution or the native Viewer through the renderer configuration.
1. NoRender training and policy evaluation
no-rendertoRenderCfg, the simulation CLI, andeval-policy.SimulationManagerskips background, light, visual-material initialization, and Newton render-state publication in this mode while retaining the physical ground.headless=Truefor NoRender and report configuration errors for explicit render synchronization or native camera creation. Both G1 Default/Newton environment and PPO configurations select NoRender.--viewer --renderer no-renderbefore simulation creation.Primary implementation:
embodichain/lab/sim/sim_manager.py,embodichain/lab/sim/cfg/simulation.py, andembodichain/learning/rl/policy_evaluation/cli.py.2. Batched state access and partial resets
ArticulationData.fetch_state()to read joint positions/velocities and root poses/velocities together. Returned tensors use reusable buffers; retaining a historical state requires an explicitclone().Articulation.set_state()to submit root state, joint state, control targets, and dynamics clearing together. Validate shapes before writing, clamp joint positions to their limits, and preserve unselected environment rows during partial resets.fetch_state,apply_state, andfetch_joint_propertiesin the Scene adapter, with centralized environment/joint selection and pose-layout conversion. Read joint limits and drive properties through the batch interface.Primary implementation:
embodichain/lab/sim/objects/articulation.py,embodichain/lab/sim/objects/backends/scene.py, andembodichain_tasks/embodichain_tasks/locomotion/velocity/_embodichain.py.3. Rollout and PPO data handling
train()call, calculate throughput from newly collected transitions, and retain cumulative checkpoint counters.20/50; use the native scene-node capacity default for G1 on both backends so NoRender creation stays within the supported range. Update the configuration tests. Add optional Newton contactpriority, preserving the source value by default.4. Regression coverage and local paired validation
The branch is based on
main@d959eeedand retains main's action-contract interfaces. Follow-up commit1630f972synchronizes the Scene batch test doubles with the new state/property APIs, supplies the camera test's renderer configuration, accepts NoRender in packaged deployment checks, and classifiesuse_native_mesh_loaderas an asset-loader option rather than a native shape property. It also adds regressions for per-environment batch properties, reusable state buffers, selected writes, negative batch status,set_statevalidation/clamping, and fresh-process NoRender lifecycle on both backends.Validation on 2026-09-30 used this PR worktree and a freshly rebuilt local DexSim integration worktree at
a6945b3e6. Core code loaded directly from this PR worktree; bundled tasks were installed from the same source. The benchmark and CUDA training tests were reused unchanged from #710, with task configuration resolved from this PR.CI follow-up on the current main merge
CI run 36660823947 tested the synthetic merge with
main@7e56e321, rather than the PR branch alone. It reported 57 failures and 4 collection errors: 55 failures and all collection errors came from main's planner integration callingdeepcopywithout importing it; the other two failures came from the published DexSim package lacking the batch methods andNewtonCollisionDesc.priority. Follow-up commit1282a68csupplies the standard-library import.On that exact merge plus the import fix, 446 focused checks passed (1 deselected), covering the affected task/Gym configuration and package-resource paths, together with the two engine contracts using the local paired build. A broader local fast run yielded 5,663 passed, 3 failed, 7 skipped and 2 collection errors: local DexUni configuration differences, missing Gradio, and a packaging subprocess selecting an old binding. The packaging runtime source was corrected and its unchanged test passed in the focused run. This is not a claim that the full local suite or published-package CI is green.
The 4096-environment workload uses the repository G1 asset, NoRender, solver limits
20/50, and physics CUDA Graphs. Each update collects 24 control steps with four physics substeps per step. Three warmup updates precede five measured updates; PPO uses five epochs and four minibatches. Hardware/software: RTX 5090 D v2, Torch 2.10.0+cu128, Newton 1.6.0, MuJoCo/MJWarp 3.12.0, Warp 1.17.0.Training SPS =
491,520 / Σ(rollout + PPO); rollout SPS =491,520 / Σrollout. CUDA is synchronized at phase boundaries; creation and checkpoint operations are outside those throughput intervals. RAM and process VRAM are the maximum phase-boundary samples from scene creation through training. This short run verifies execution and throughput, not policy convergence.The benchmark verifies that the physics CUDA Graph is captured after warmup and rejects nonfinite values or MJWarp overflow. The process returned 0, but teardown still reported an entity-count difference of 32; one-environment headless evaluation reported a difference of 1. These warnings remain an engine cleanup acceptance item. The local run does not establish that teardown is warning-free.
5. Dependencies and merge readiness
Renderer.NORENDER,ArticulationBatch.fetch_state/apply_state/fetch_joint_properties, andNewtonCollisionDesc.priority.dexsim_engine==0.5.0dependency. The local build retains the version string0.5.0but is identified by commita6945b3e6; it does not replace the package installed by CI.Type of change
Checklist
python docs/scripts/check_api_docs.py.