🤖 AI Summary
This study addresses the limited generalizability and reliance on task-specific policies in humanoid physical interaction by proposing the first behavior foundation model for humanoid object manipulation. The approach integrates forward-backward representations with unsupervised reinforcement learning, acquiring coupled dynamics representations through multi-timescale joint training to enable reward-conditioned closed-loop interaction. Notably, the model supports multi-scale objectives and long-horizon task chaining without retraining. Experiments demonstrate that a single policy successfully executes complex tasks such as object carrying and achieves an 89.3% fall recovery success rate, significantly outperforming existing baselines. Furthermore, real-world deployment validates the robustness of the proposed method.
📝 Abstract
Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid-object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.