Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions agent_context/topics/rl-learning/rl-learning.md
Original file line number Diff line number Diff line change
Expand Up @@ -168,6 +168,10 @@ consume observations and write action, log-probability, entropy, and value
fields needed by their algorithm. Differentiable policies must expose
graph-preserving action sampling.

Trainer throughput counts transitions collected during the current `train()`
call. Restored `global_step` remains a cumulative checkpoint counter and is
excluded from that call's SPS numerator.

For actor-critic policies, `policy.obs_groups.actor` and `.critic` select ordered
observation groups. The collector and standard buffer preserve separate
`critic_obs` when configured; evaluation applies the same selection. The PPO
Expand All @@ -184,6 +188,11 @@ DexSim's Motion Policy Evaluator, preserving the original task's reset, step,
observation and action path. Supplying an Environment means the adapter owns
its camera lifecycle; Kit's default flat-ground camera is not applied to it.

Headless evaluation accepts `--renderer no-render`. When a saved training
configuration uses NoRender, Viewer evaluation defaults to Hybrid; an explicit
native renderer overrides that choice. `--viewer --renderer no-render` is
rejected before creating the simulation.

Tasks opt in through `PolicyViewerCameraCfg` and
`get_policy_viewer_target_pose()` (world XYZ + XYZW). The six bundled flat
velocity tasks own their presets; `policy_evaluation/_viewer_camera.py` owns
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,12 @@ full-batch tensors when partial writes must preserve other rows/DOFs. Avoid
DexSim's host-materialized selected-DOF path. Pose conversions follow the
[public quaternion contract](simulation-system.md#quaternion-and-pose-convention).

`ArticulationData.fetch_state()` reads joint position/velocity and root
pose/velocities together into existing data buffers. The Scene view reuses
DexSim's batch fetch and converts the root-pose layout once. Returned tensors
are borrowed buffers, not snapshots; clone values that must survive later
reads. Fetch after reset or direct writes rather than caching across them.

`Articulation.set_root_velocity()` writes selected world-frame linear and
angular velocities as `(N, 6)` rows. The Scene adapter validates the complete
input before writing one selected batch; other environment rows are preserved.
Expand Down
9 changes: 9 additions & 0 deletions agent_context/topics/simulation-system/rendering.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,15 @@
Read this for physics/render synchronization, native-window/DLSS configuration
and readiness reporting. Return to the [simulation overview](simulation-system.md).

## NoRender initialization

`RenderCfg(renderer="no-render")` selects `Renderer.NORENDER` and requires
headless mode. The manager skips background, light and visual-material setup,
keeps physical ground, and skips Newton render-state publication. DexSim uses
its existing package and device-free NoRender engine. Native camera and window
operations require a native renderer. Checkpoint evaluation with `--viewer`
defaults to Hybrid when the saved training configuration uses NoRender.

## Rendering does not advance physics

`SimulationManager.render_frame()` owns a read-only consumption phase after
Expand Down
22 changes: 22 additions & 0 deletions docs/source/api_reference/embodichain/embodichain.lab.sim.cfg.rst
Original file line number Diff line number Diff line change
Expand Up @@ -128,3 +128,25 @@ window remains a consumer regardless of this automatic policy.
physics_cfg_for_backend
physics_backend_from_cfg
validate_physics_cfg

Rigid-body property module
--------------------------

The rigid-body configuration types are also available from
``embodichain.lab.sim.cfg.rigid``. They describe mass, collision properties,
materials, and backend-specific overrides. ``NewtonCollisionPropertiesCfg``
accepts an optional ``priority`` for MuJoCo contact-parameter selection;
``None`` preserves the source value.

.. currentmodule:: embodichain.lab.sim.cfg.rigid

.. autosummary::

MassPropertiesCfg
DefaultRigidBodyPropertiesCfg
CollisionPropertiesCfg
DefaultCollisionPropertiesCfg
NewtonCollisionPropertiesCfg
RigidBodyMaterialCfg
NewtonRigidBodyMaterialCfg
RigidBodyPhysicsCfg
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,12 @@ Rigid Object Group
Articulation
------------

``robot.body_data.fetch_state()`` reads joint position, joint velocity, root
pose and root velocities together through DexSim's batch state interface.
Each call refreshes the values after stepping or state writes. The returned
tensors reuse the data object's buffers; clone tensors when keeping a past
snapshot. Individual data properties remain available for single-field reads.

.. autoclass:: Articulation
:members:
:inherited-members:
Expand Down
20 changes: 20 additions & 0 deletions docs/source/overview/sim/sensors/contact_sensor.md
Original file line number Diff line number Diff line change
Expand Up @@ -142,3 +142,23 @@ may be left over from an earlier update.
Newton MuJoCo-Warp exposes contact forces, so both impulse fields are available. Other supported Newton rigid solvers currently expose contact geometry with zero-valued impulse fields. MuJoCo CPU mode does not expose device contact buffers, and DexUni does not currently publish rigid contacts through `ContactQuery`; those modes are therefore unsupported by this sensor.

Default Direct GPU reports static counterparts with actor ID `-1` because its raw contact buffer does not expose their object identity. To monitor a dynamic body or articulation link against arbitrary static geometry, select the dynamic/link object and set `filter_need_both_actor=False`. Default CPU and Newton can identify registered static shapes.


## CUDA substep sampling

When a contact history is registered, `BaseEnv` samples the sensor after every
physics substep. On CUDA, the sensor captures dense-row scattering, overflow
accumulation and all history reductions in one graph. The first sample warms
kernels; capture on the following sample does not advance history, and replay
advances it exactly once. Selected-row reset writes into the same live buffers.
Changing the sampling interval, history threshold or counterpart option, or
registering another history rebuilds this graph.

Contact fetching and actor metadata refresh still run before sampling. DexSim
owns its snapshot and query graphs; the sensor consumes the current query
buffer and its device-resident count. Results are ordered onto the caller's
Torch stream. CPU sampling uses the same reductions without capture.

The physics loop, simulation clock and rendering callbacks remain outside the
sensor graph. Window, offscreen and browser visualization retain their existing
update order.
36 changes: 36 additions & 0 deletions docs/source/overview/sim/sim_articulation.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,34 @@ Articulations are configured using the {class}`~cfg.ArticulationCfg` dataclass.
At runtime, call `articulation.set_gravity(...)` to change gravity for every
environment or for a selected set of environment indices.

### Combined state writes

Use `Articulation.set_state` (also inherited by `Robot`) to reset selected
articulations in one Scene batch operation:

```python
robot.set_state(
env_ids=env_ids,
joint_ids=policy_joint_ids,
root_pose=pose, # (N, 7): environment-local xyz + xyzw
qpos=joint_position, # (N, J), clipped to the selected joint limits
target_qpos=joint_position,
clear_dynamics=True,
)
```

All supplied joint fields share the selected row and joint order. Optional
`qvel`, `target_qvel`, `qf`, and `root_velocity` fields override reset defaults.
`clear_dynamics=True` clears velocities, forces, external wrenches and solver
history, and holds the final joint positions where no target is supplied.
Newton/MJWarp requires all articulations in each affected solver world when
clearing solver history. Other environments retain their state.

The Scene batch validates shapes and selections before writes, then propagates
final joint kinematics once. Native execution errors are reported to the caller.
Single-field joint setters use this same writer; target-only writes do not
recompute kinematics. State changes follow the existing render-frame publication lifecycle.

### Root velocity writes

`Articulation.set_root_velocity(velocity, env_ids=None)` writes root-link velocities
Expand Down Expand Up @@ -142,6 +170,14 @@ Inspect the resolved properties and test the response after changing solvers.

### Joint Position Limits

When binding a finalized Scene, `ArticulationData` initializes position,
velocity and effort limits through one DexSim `ArticulationBatch` property
read. Each environment retains its own values in public DOF order. The
`joint_stiffness`, `joint_damping`, `joint_friction` and `joint_armature`
properties also use this batch interface and return independent snapshots of
the current model. Retained descriptor edits that require a Scene rebuild
become visible after that rebuild.

Use `qpos_limits` to override the limits defined in the asset file. This is the
articulation's effective physical limit in simulation, so it is also the range
used when `set_qpos(...)` clamps requested joint positions.
Expand Down
12 changes: 11 additions & 1 deletion docs/source/overview/sim/sim_manager/rendering/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ ray-tracing sample count, tone mapping, and DLSS settings used by

| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `renderer` | `str` | `"auto"` | Renderer backend: `auto`, `hybrid`, `fast-rt`, or `rt`. |
| `renderer` | `str` | `"auto"` | Renderer backend: `auto`, `no-render`, `hybrid`, `fast-rt`, or `rt`. |
| `spp` | `int` | `1` | Samples per pixel for ray-traced rendering. Must be at least `1`. |
| `tone_mapping_enabled` | `bool` | `False` | Apply modified Reinhard tone mapping to RGB output. |
| `tone_mapping_exposure` | `float` | `1.0` | Fixed linear exposure multiplier used before tone mapping. |
Expand All @@ -20,6 +20,16 @@ unchanged.

## Renderer selection

For state-based RL, use `RenderCfg(renderer="no-render")` with
`SimulationManagerCfg(headless=True)`. This selects DexSim's `Renderer.NORENDER`
backend. EmbodiChain skips background, light and visual-material creation,
retains physical ground, and skips Newton render-state publication.
Native cameras and windows require `hybrid`, `fast-rt`, or `rt`.
`headless=True` with a native renderer keeps offscreen rendering available.

Policy evaluation accepts `--renderer no-render`. For visual evaluation of a
NoRender checkpoint, `--viewer` selects Hybrid by default.

With `renderer="auto"`, EmbodiChain selects a backend from the GPU detected at
the configured `gpu_id` when the simulation manager is constructed:

Expand Down
2 changes: 1 addition & 1 deletion embodichain/cli/sim.py
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ def add_sim_args_to_parser(parser: argparse.ArgumentParser) -> None:
parser.add_argument(
"--renderer",
type=str,
choices=["auto", "hybrid", "fast-rt", "rt"],
choices=["auto", "no-render", "hybrid", "fast-rt", "rt"],
default="auto",
help="Renderer backend; omission preserves the launcher/config default.",
)
Expand Down
25 changes: 22 additions & 3 deletions embodichain/lab/gym/envs/base_env.py
Original file line number Diff line number Diff line change
Expand Up @@ -727,8 +727,18 @@ def get_obs(self, **kwargs) -> EnvObs:
"""

with self._profiler.section("proprio"):
# A Python index list makes each proprioception field upload its
# own indices. Reuse the device index until the selection changes.
selection = (tuple(self.active_joint_ids), self.device)
cached = getattr(self, "_proprio_joint_index_cache", None)
if cached is None or cached[0] != selection:
cached = (
selection,
torch.tensor(selection[0], dtype=torch.long, device=self.device),
)
self._proprio_joint_index_cache = cached
obs = TensorDict(
dict(robot=self.robot.get_proprioception()[:, self.active_joint_ids]),
dict(robot=self.robot.get_proprioception()[:, cached[1]]),
batch_size=[self.num_envs],
device=self.device,
)
Expand Down Expand Up @@ -930,7 +940,13 @@ def reset(

with self._profiler.section("reset_objects_state"):
self.sim.reset_objects_state(
env_ids=reset_ids, excluded_uids=self._detached_uids_for_reset
env_ids=reset_ids,
# Environment sensors are reset below, including sensors
# supplied by custom _setup_sensors implementations.
excluded_uids=[
*self._detached_uids_for_reset,
*(sensor.uid for sensor in self.sensors.values()),
],
)

for sensor in self.sensors.values():
Expand All @@ -940,7 +956,10 @@ def reset(
with self._profiler.section("initialize_episode"):
self._initialize_episode(reset_ids, **options)
self._reset_physical_objective(reset_ids)
self._elapsed_steps[reset_ids] = 0
elapsed_ids = torch.as_tensor(reset_ids, dtype=torch.long).to(
device=self._elapsed_steps.device, non_blocking=True
)
self._elapsed_steps.index_fill_(0, elapsed_ids, 0)

with self.sim.render_frame(force_visualization=True):
with self._profiler.section("get_obs"):
Expand Down
14 changes: 10 additions & 4 deletions embodichain/lab/gym/envs/embodied_env.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@
from math import log
from functools import wraps
from datetime import datetime
import logging
import os
import threading
import copy
Expand Down Expand Up @@ -898,7 +899,9 @@ def _update_episode_success_status(
dtype=torch.bool,
device=self.episode_success_status.device,
)
self.episode_success_status[update_mask] = success[update_mask]
self.episode_success_status.copy_(
torch.where(update_mask, success, self.episode_success_status)
)

def _extend_obs(self, obs: EnvObs, **kwargs) -> EnvObs:
if self.observation_manager:
Expand Down Expand Up @@ -986,7 +989,10 @@ def _reset_physical_objective(self, env_ids: Sequence[int] | torch.Tensor) -> No
def _initialize_episode(
self, env_ids: Sequence[int] | None = None, **kwargs
) -> None:
logger.log_debug(f"Initializing episode for env_ids: {env_ids}", color="blue")
if logger.logger.isEnabledFor(logging.DEBUG):
logger.log_debug(
f"Initializing episode for env_ids: {env_ids}", color="blue"
)
save_data = kwargs.get("save_data", True)

# Determine which environments to process
Expand Down Expand Up @@ -1086,7 +1092,7 @@ def _initialize_episode(

_traj_steps = getattr(self, "_traj_steps", None)
if _traj_steps is not None:
_traj_steps[env_ids_to_process] = 0
_traj_steps.index_fill_(0, env_ids_to_process, 0)

# Clear episode buffers only after every recorder has consumed them.
if self.rollout_buffer is not None and self._rollout_buffer_mode != "rl":
Expand All @@ -1111,7 +1117,7 @@ def _initialize_episode(
self._demo_active_rollout_start_steps[demo_ids] = 0
self._demo_steps[demo_ids] = 0

self.episode_success_status[env_ids_to_process] = False
self.episode_success_status.index_fill_(0, env_ids_to_process, False)

# Stateful managers reset selected rows before reset-mode events run.
action_manager = getattr(self, "action_manager", None)
Expand Down
20 changes: 19 additions & 1 deletion embodichain/lab/gym/envs/managers/actions.py
Original file line number Diff line number Diff line change
Expand Up @@ -429,7 +429,25 @@ def apply_actions(self) -> None:
)

def reset(self, env_ids: list[int] | torch.Tensor | None = None) -> None:
super().reset(env_ids)
"""Clear selected action history and commands without changing bias.

Args:
env_ids: Rows to reset. None selects all rows.
"""
buffers = (
self._raw_actions,
self._previous_raw_actions,
self._processed_actions,
)
if env_ids is None:
for buffer in buffers:
buffer.zero_()
else:
ids = torch.as_tensor(env_ids, dtype=torch.long).to(
device=self.device, non_blocking=True
)
for buffer in buffers:
buffer.index_fill_(0, ids, 0)


class EefPoseAction(ActionTerm):
Expand Down
12 changes: 12 additions & 0 deletions embodichain/lab/gym/envs/managers/event_manager.py
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,7 @@ def __init__(self, cfg: object, env: EmbodiedEnv):

# call the base class (this will parse the functors config)
super().__init__(cfg, env)
self._all_env_ids = torch.arange(env.num_envs, device=env.device)

def __str__(self) -> str:
"""Returns: A string representation for event manager."""
Expand Down Expand Up @@ -310,6 +311,17 @@ def apply(
self._call_event_functor(
mode, functor_name, functor_cfg, self._env, None
)
elif not functor_cfg.is_global and functor_cfg.interval_step == 1:
# Every row is due, including rows reset on the last step.
# Reuse the known selection instead of synchronizing a
# CUDA nonzero merely to recover all environment indices.
self._call_event_functor(
mode,
functor_name,
functor_cfg,
self._env,
self._all_env_ids,
)
else:
valid_env_ids = (
(
Expand Down
Loading
Loading