π€ AI Summary
This study addresses the inherent trade-off between locomotion performance and payload stability when quadruped robots transport unfixed stacked loads. To reconcile this challenge, we propose PAMORT, a framework grounded in multi-objective reinforcement learning and conditional policy training. Crucially, it incorporates an online preference adaptation mechanism that relies exclusively on proprioceptive feedback, enabling stable control without dedicated payload sensors. The proposed framework demonstrates zero-shot generalization capabilities across varying conditions. Extensive evaluations reveal that PAMORT consistently outperforms baseline methods in simulation. Furthermore, real-world experiments confirm its practical efficacy, achieving an average success rate of 0.85 compared to 0.675 for the baselines, thereby validating the frameworkβs superiority in balancing agile locomotion with load stability during physical deployment.
π Abstract
Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted board without dedicated payload sensors or active carrier mechanisms. To address this trade-off, we propose Payload-Adaptive Multi-Objective Reinforcement learning for Transportation (PAMORT). PAMORT trains a multi-objective base policy conditioned on a preference vector that weights locomotion and payload-stability reward groups, then trains a weight adjuster on the frozen policy to adapt this preference online from proprioception. In simulation, PAMORT achieves comparable or better overall transportation success than a corresponding single-objective baseline across different payload configurations, including an unseen three-box stack, despite training only with two boxes. Real-world experiments on a Unitree Go2 demonstrate zero-shot transfer to slopes and steps at or beyond the training difficulty, with mean success rates of 0.850 for PAMORT and 0.675 for the baseline across eight tasks. These results demonstrate robust unsecured-payload transportation with online adaptation of the locomotion--payload trade-off from proprioceptive information.