PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决灵活动作生成与反馈响应控制的结合问题,提出PredActor方法,通过本体感受观察直接生成可执行动作和未来状态轨迹。
📝 Abstract
Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.
Problem

Research questions and friction points this paper is trying to address.

diffusion models
humanoid control
feedback-responsive
future-state trajectory
test-time objectives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predictive Action Diffusion
Proprioceptive Observations
Classifier-free Guidance
Rolling Denoising
Runtime Optimization
🔎 Similar Papers
2024-10-07arXiv.orgCitations: 9
2024-07-16Neural Information Processing SystemsCitations: 16
L
Lei Ye
Harbin Institute of Technology
H
Haibo Gao
Harbin Institute of Technology
Yitang Li
Yitang Li
Tsinghua University
Computer visionRobotics
P
Peng Xu
Harbin Institute of Technology
Z
Zetong Jing
RoboParty Lab
J
Junhan Sun
RoboParty Lab
F
Fanrong Dong
RoboParty Lab
Z
Ziqi Han
RoboParty Lab
X
Xue Wang
RoboParty Lab
Jianhua Sun
Jianhua Sun
Shanghai Jiao Tong University
Computer VisionRobot Learning
C
Cewu Lu
Harbin Institute of Technology
H
Hao Zhao
RoboParty Lab
Liang Ding
Liang Ding
Professor in State Key Laboratory of Robotics and System, Harbin Institute of Technology
RoboticsSpace roboticsContact mechanicsArtificial IntelligenceControl