Skip to content

Sync Task Engine with main - #729

Open
yuecideng wants to merge 151 commits into
mainfrom
codex/action-engine-v36-main-sync
Open

yuecideng wants to merge 151 commits into
mainfrom
codex/action-engine-v36-main-sync

Conversation

@yuecideng

@yuecideng yuecideng commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Description

Synchronize the current Task Engine v38 line, including E6 recovery and standalone E9 pressing, with main at 7e56e321. This resolves the conflicts introduced when PR #729 was updated to the newer Task Engine branch.

  • Preserve the GenSim invocation-scoped motion/grasp policies, bounded E6 recovery, bottom-up drawer-set ordering, per-handle arm preference, and simultaneous terminal drawer checks.
  • Integrate main's plan-transform API, rigidized-articulation object bindings, non-mutating TCP inverse transform, planner configuration, and whole-scene USD preview/export flow. Retain GenSim's VisACD and authored revolute-limit handling in the relocated scene loader.
  • Retain drawer ownership validation, scope close-arm rewrites to the relevant drawer routes, and attempt cleanup of every cuRobo payload attachment even if one detach fails. Do not reinstate the older blanket ban on mixed drawer/ordinary tasks: v38 already owns route-local services.
  • Restore the shared PyTorch solver's 0.5 mm position default. Generated GenSim configurations explicitly retain their previously qualified 5 mm position tolerance, while preserving caller overrides and the existing rotation tolerance.
  • Document render_scene_visual_evidence, make_visual_grounding_caller, and visual_grounding_available, including their evidence limits and failure behavior. Repair short RST headings and gate the visual rendering test on isolated EGL availability.

E9 remains a standalone, single-button, single-environment contact-press workflow. Its existing generated unit-scale/calibration policy, 10 mm extra commanded travel, contact-supported 0.05 mm event criterion, and retreat checks are unchanged. This does not certify device activation, self-latching, or success in the original unadapted scene.

Refs #531

Dependencies: the GenSim optional extra includes shapely, yourdfpy>=0.0.60, and python-fcl>=0.7.0.11. Documentation was built with the project's pinned documentation requirements in an isolated Python 3.11 environment.

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • Enhancement
  • New feature
  • Breaking change
  • Documentation update

Validation

  • black . with Black 26.3.1: 1297 files unchanged after formatting the targeted edits.
  • git diff --cached --check and agent context checks: passed.
  • python docs/scripts/check_api_docs.py: 2695/2695 exports documented.
  • API documentation checker tests: 8 passed; context tooling tests: 40 passed.
  • GenSim, Task Program, configured-integration and segment-policy regression: 1911 passed, 1 skipped, 1 deselected; 1 external asset-download failure described below.
  • Shared atomic-action, math, topology and IK tests, excluding native simulation/GPU cases: 1042 passed, 87 deselected.
  • Additional documentation, EGL, drawer-scope, cleanup and local-IK-default regressions: 122 passed.
  • Sphinx 7.4.7 dummy build against an isolated clean snapshot of the staged tree: succeeded. It is not warning-free: 776 warnings remain, including existing cross-reference ambiguities and optional-import warnings. The newly documented visual evidence/grounding functions have no function-specific warnings in that build.
  • Fixed local scene replays through the normal bundle runner, seed 0 and one environment: task1154 succeeded (2601 steps), task1155 succeeded (1994 steps), task1107/E9 succeeded (815 steps). E6 simultaneous terminal checks passed. E9 recorded approximately 5.959 mm contact-supported travel, verified contact release and retreat, and had zero invalid samples among 2153 recorded substeps. activation_verified remains false.

The physical replays used existing local scenes and deterministic/provider-free preparation; no external language-model or scene-generation API was called. Source scene and asset hashes were unchanged. These results do not claim all-scene or multi-seed qualification.

Remaining environment limitation

test_all_examples_register_plain_embodied_env_under_config_selected_ids[open_drawer] cannot complete because the configured mirror returns a Drawer.zip with an invalid MD5. The asset checksum was not changed and the check was not disabled to hide this failure. Other affected integration tests passed; live cuRobo and the full native/GPU test matrix were not rerun.

Screenshots

No UI change. Local simulation videos, execution reports, contact traces, and trajectory comparisons were retained with the validation artifacts; generated scene assets and recordings are not included in this PR.

Checklist

  • I have run black . to format the code base.
  • I reviewed affected documentation and agent context and updated the relevant contracts.
  • Public API changes are reflected in the API docs.
  • Focused regression tests cover the compatibility and safety-boundary fixes.
  • Dependency declarations reflect the GenSim runtime requirements.

i
1. modify cli/start logic
2. add xy position info in scene object data sturcture
3. add some test files
4. modify scene export and scene import, thus the scene graph will be export and import automatically
…libration in image-conditioned scene engine pipeline (but only calibrated the upright bottle-like assets currently)
@yuecideng yuecideng added bug Something isn't working refactor task A task written in openai gym format for imitation learning or reinforcement learning atomic action atomic action related functionality solver Robot kinematics solver rendering Things related to rendering (eg, performace, efficiency, bug) labels Sep 30, 2026
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 3/5

[High risk] Adds task engine module with new runtime code paths.

The PR does not appear safe to merge while successful drawer planning can fail during cleanup and articulation proxies can render at incorrect positions.

Fix All in CodexFindings

  1. P1 Cleanup discards successful plans ▶
  2. P1 Articulation proxies render misplaced ▶
  3. P1 Visual test loses EGL gate ▶
  4. P2 Repair loses error details ▶
Fix with agent prompt
### Issue 1
embodichain/gen_sim/task_engine/_task_program/drawer_curobo.py:312-313
If an attachment manager fails to detach after drawer transport planning succeeds, `close()` raises from the `finally` block in `DrawerMotionGenerator.generate()`. That replaces the usable plan with an exception and fails the drawer call, even though attachment cleanup is meant to be best-effort.

### Issue 2
embodichain/gen_sim/task_engine/orchestration/visual_evidence.py:106-117
When a scene has an articulation proxy and a nonzero scene rotation or XY translation, scene preparation transforms `init_pos` and `init_rot` but leaves `proxy_init_pos` unchanged. This renderer combines the unchanged proxy position with the transformed rotation, so the articulation's label and mask appear away from its prepared scene position. Visual grounding can then use misleading evidence even for a task that does not operate the articulation.

### Issue 3
tests/gen_sim/task_engine/orchestration/test_visual_evidence.py:undefined-67
If a CPU-only test host has no usable EGL renderer, this test now runs its EGL-dependent rendering subprocess unconditionally. Rendering fails and the test fails instead of skipping as it did with the removed capability check.

### Issue 4
embodichain/gen_sim/task_engine/orchestration/grounding.py:513-515
When a grounding response selects an unknown UID or omits a reference, the one repair prompt now says only `ValueError`. It no longer tells the model which binding failed validation, making that retry less useful and increasing the chance that an otherwise groundable candidate is rejected.

```suggestion
                "UIDs from the supplied candidate inventory. Validation error: "
                f"{first_error}"
```

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR synchronizes the Task Engine and scene integration with main. Since the previous review, it aligns robot-profile defaults, changes grounding repair feedback, copies geometry inputs before proxy fitting, and adds focused default-profile tests.

Reviews (4) · Last reviewed commit: "fix(task-engine): align phase-one defaul..."

Comment on lines +312 to +313
if first_error is not None:
raise first_error

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Cleanup discards successful plans

If an attachment manager fails to detach after drawer transport planning succeeds, close() raises from the finally block in DrawerMotionGenerator.generate(). That replaces the usable plan with an exception and fails the drawer call, even though attachment cleanup is meant to be best-effort.

Prompt To Fix With AI
This is a comment left during a code review.
Path: embodichain/gen_sim/task_engine/_task_program/drawer_curobo.py
Line: 312-313

Comment:
**Cleanup discards successful plans**

If an attachment manager fails to detach after drawer transport planning succeeds, `close()` raises from the `finally` block in `DrawerMotionGenerator.generate()`. That replaces the usable plan with an exception and fails the drawer call, even though attachment cleanup is meant to be best-effort.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code

Comment on lines +92 to +103
def _pose(config: dict[str, Any], *, proxy: bool) -> np.ndarray:
if not proxy and config.get("init_local_pose") is not None:
return np.asarray(config["init_local_pose"], dtype=float)
matrix = np.eye(4)
matrix[:3, :3] = Rotation.from_euler(
"XYZ", config.get("init_rot", [0.0, 0.0, 0.0]), degrees=True
).as_matrix()
matrix[:3, 3] = np.asarray(
config.get("proxy_init_pos" if proxy else "init_pos", config.get("init_pos")),
dtype=float,
)
return matrix

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Articulation proxies render misplaced

When a scene has an articulation proxy and a nonzero scene rotation or XY translation, scene preparation transforms init_pos and init_rot but leaves proxy_init_pos unchanged. This renderer combines the unchanged proxy position with the transformed rotation, so the articulation's label and mask appear away from its prepared scene position. Visual grounding can then use misleading evidence even for a task that does not operate the articulation.

Knowledge Base Used: Generative simulation pipelines

Prompt To Fix With AI
This is a comment left during a code review.
Path: embodichain/gen_sim/task_engine/orchestration/visual_evidence.py
Line: 92-103

Comment:
**Articulation proxies render misplaced**

When a scene has an articulation proxy and a nonzero scene rotation or XY translation, scene preparation transforms `init_pos` and `init_rot` but leaves `proxy_init_pos` unchanged. This renderer combines the unchanged proxy position with the transformed rotation, so the articulation's label and mask appear away from its prepared scene position. Visual grounding can then use misleading evidence even for a task that does not operate the articulation.

**Knowledge Base Used:** [Generative simulation pipelines](https://app.greptile.com/dexforce/-/custom-context/knowledge-base/dexforce/embodichain/-/docs/generative-simulation.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code

…refactor_v38

Integrate standalone contact-verified E9 pressing while retaining v38 invocation policies and E6 recovery and terminal validation. Align container landing tests with the confirmed 35 mm wall margin.

Validation: 1100 tests passed, 1 skipped; task1154/task1155 trajectories matched their baselines, and task1107 passed the contact-press contract. Existing API documentation gaps are deferred.
@skywhite1024
skywhite1024 force-pushed the codex/action-engine-v36-main-sync branch from cfea1d6 to 58e5690 Compare September 30, 2026 04:13
Comment thread embodichain/gen_sim/task_engine/_task_program/drawer_binding.py Outdated
return root


def test_visual_evidence_renders_current_assets_with_uid_masks(tmp_path: Path) -> None:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Visual test loses EGL gate

If a CPU-only test host has no usable EGL renderer, this test now runs its EGL-dependent rendering subprocess unconditionally. Rendering fails and the test fails instead of skipping as it did with the removed capability check.

Prompt To Fix With AI
This is a comment left during a code review.
Path: tests/gen_sim/task_engine/orchestration/test_visual_evidence.py
Line: 67

Comment:
**Visual test loses EGL gate**

If a CPU-only test host has no usable EGL renderer, this test now runs its EGL-dependent rendering subprocess unconditionally. Rendering fails and the test fails instead of skipping as it did with the removed capability check.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code

Preserve v38 E6/E9 policies while integrating main planner, plan-transform, rigidized object, and scene USD APIs. Restore scoped drawer cleanup and arm-routing fixes, keep IK tolerances GenSim-local, and document visual grounding exports.

Validated API documentation coverage, a clean-snapshot Sphinx build, affected CPU tests, and three fixed E6/E9 physical replays. The configured Drawer asset mirror still fails checksum validation.
@skywhite1024 skywhite1024 added the docs Improvements or additions to documentation label Sep 30, 2026
@skywhite1024 skywhite1024 changed the title sync action engine v36 with main and harden drawer execution Sync Task Engine v38 with main and fix API documentation Sep 30, 2026
@skywhite1024 skywhite1024 changed the title Sync Task Engine v38 with main and fix API documentation Sync Task Engine with main Sep 30, 2026
Comment on lines +513 to +515
"UIDs from the supplied candidate inventory. The previous "
f"validation failed with {type(first_error).__name__}; return a "
"fresh object that satisfies the schema."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Repair loses error details

When a grounding response selects an unknown UID or omits a reference, the one repair prompt now says only ValueError. It no longer tells the model which binding failed validation, making that retry less useful and increasing the chance that an otherwise groundable candidate is rejected.

Suggested change
"UIDs from the supplied candidate inventory. The previous "
f"validation failed with {type(first_error).__name__}; return a "
"fresh object that satisfies the schema."
"UIDs from the supplied candidate inventory. Validation error: "
f"{first_error}"
Prompt To Fix With AI
This is a comment left during a code review.
Path: embodichain/gen_sim/task_engine/orchestration/grounding.py
Line: 513-515

Comment:
**Repair loses error details**

When a grounding response selects an unknown UID or omits a reference, the one repair prompt now says only `ValueError`. It no longer tells the model which binding failed validation, making that retry less useful and increasing the chance that an otherwise groundable candidate is rejected.

```suggestion
                "UIDs from the supplied candidate inventory. Validation error: "
                f"{first_error}"
```

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Codex Fix in Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

atomic action atomic action related functionality bug Something isn't working docs Improvements or additions to documentation refactor rendering Things related to rendering (eg, performace, efficiency, bug) solver Robot kinematics solver task A task written in openai gym format for imitation learning or reinforcement learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants