Skip to content

About

Predictive Action Diffusion for Steerable Onboard Humanoid Control

Resources

Stars

2 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

32 Commits

Folders and files

Repository files navigation

PredActor: predictive action diffusion for steerable onboard humanoid control. Internal future states, direct actions, CG and CFG guidance, and 50 Hz onboard execution.

PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control

Project website arXiv paper Demo videos Evaluation code available Hugging Face artifacts MIT license

Harbin Institute of Technology, Shanghai Innovation Institute, RoboParty Lab, Tsinghua University, Shanghai Jiao Tong University, HexLab, and SFTR

Important

The evaluation code is available now, with browser-based MuJoCo evaluation of the published PDP051 checkpoint. Training, data collection, and robot deployment releases are coming soon.

Quick evaluation

Install uv, then run these commands on Linux:

git clone https://github.com/MasterYip/PredActor.git
cd PredActor
uv sync --locked
uv run --locked python scripts/hf_download.py --filter checkpoints
uv run --locked predactor-eval

The download command retrieves the released checkpoints from MasterYip/PredActor_Artifacts.

The last command opens http://127.0.0.1:8765/. It uses CUDA when available and otherwise falls back to CPU. The locked environment targets Python 3.10 and includes the evaluator, MuJoCo, and the browser interface; Isaac Sim and the training repositories are not required.

For a headless startup, use uv run --locked predactor-eval --headless --web-no-browser. Run uv run --locked predactor-eval --help for optional checkpoint, config, device, and Web UI overrides.

Overview

PredActor augments action diffusion with an internal future-state trajectory for look-ahead guidance while retaining direct action execution. Within a joint state-action formulation, it predicts future states and actions from proprioceptive history and optional task context. Unlike a generator-tracker pipeline, predicted states remain inside the policy rather than becoming motion references for a separate tracker. Unlike action-only diffusion, the policy provides an explicit future-state trajectory for guidance.

This representation supports two complementary steering mechanisms: classifier guidance (CG) applies test-time objectives to predicted states, and classifier-free guidance (CFG) strengthens learned behavior conditions, including text commands. The selected action is sent directly to the joint controller; policy observations require only proprioception, not externally estimated full-body states.

For onboard execution, rolling denoising and computation-preserving runtime optimizations support 50 Hz control on a Unitree G1's Jetson Orin NX. The measured complete callback takes 16.790 ms median and 19.383 ms p95, both below the 20 ms control period. Across simulation and physical-robot evaluation, demonstrations cover text commands, disturbance response, joystick steering, and semantic interpolation.

Project preview

PredActor humanoid control project preview

At a glance

  • Predictive direct control: jointly predicts future states and actions, keeps states internal, and executes the selected action without a separate motion-reference tracker.
  • Proprioceptive inputs: conditions on onboard proprioceptive history without requiring externally estimated full-body states as policy inputs.
  • CG and CFG steering: combines test-time objectives on predicted states with learned behavior conditioning.
  • Rolling onboard inference: reuses the denoising horizon across control ticks and optimizes runtime for 50 Hz operation on Jetson Orin NX.
  • Simulation and hardware evidence: demonstrates commands, transitions, steering, and disturbance response.

Release checklist

  • Evaluation: browser-based MuJoCo evaluation, G1 assets, and published checkpoints.
  • Data collection and labeling: motion collection, preprocessing, and dataset labeling tools.
  • BC training: behavior-cloning training code, configurations, and reproducibility assets.
  • DAgger: interactive data aggregation and policy refinement pipeline.
  • Deployment: onboard deployment and hardware-control tools.

Demos

Each result is independently playable below. Visit the project website for the complete gallery.

Hardware

Text control
Walk, squat down, and return to walking.
predactor-text-walk-squat-walk.mp4
Behavior transitions
Walk, accelerate to a jog, and transition into a squat.
predactor-text-walk-jog-squat.mp4
Physical interaction
Walk and stand commands under external interference.
predactor-behavioral-response.mp4
Outdoor pathway
Outdoor locomotion on the physical G1.
predactor-outdoor-pathway.mp4

Simulation

Joystick steering
Directional steering with text-selected locomotion modes.
predactor-joystick-steering.mp4
Text and joystick
Text-selected behavior with simultaneous directional control.
predactor-sim-text-joystick.mp4
Text control
Behavior selection and transitions from text commands.
predactor-sim-text-control.mp4
Semantic interpolation
Continuous control between semantic motion endpoints.
predactor-sim-semantic-interpolation.mp4
Target tracking
Classifier-guided destination following.
predactor-sim-target-tracking.mp4
Disturbance response
Recovery behavior under external perturbations.
predactor-sim-disturbance-response.mp4

Acknowledgements

We thank the authors of the following open-source projects:

  • diffusion_implementation by WhoKnowsssss, which provides the diffusion training framework.
  • TextOp, which provides the RL tracker training framework and pretrained tracker checkpoint.
  • MotionCLIP, which provides the motion encoder foundation.

Citation

@misc{ye2026predactorpredictiveactiondiffusion,
  title={PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control},
  author={Lei Ye and Haibo Gao and Yitang Li and Peng Xu and Zetong Jing and Junhan Sun and Fanrong Dong and Ziqi Han and Xue Wang and Jianhua Sun and Cewu Lu and Hao Zhao and Liang Ding},
  year={2026},
  eprint={2609.24840},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2609.24840},
}

License

The contents of this repository are released under the MIT License, unless noted otherwise.

About

Predictive Action Diffusion for Steerable Onboard Humanoid Control

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages