🤖 AI Summary
This study addresses the coordination of locomotion and manipulation (loco-manipulation) behaviors in humanoid robots within the context of multi-objective reinforcement learning. Leveraging the Unitree G1 platform, the authors design a curriculum learning framework comprising 13 progressively complex tasks to systematically evaluate the performance of unified versus dual-critic architectures. Their findings reveal that the choice of critic architecture is a pivotal factor—more influential than reward function design—in determining multi-objective task performance, with the dual-critic configuration effectively mitigating policy degradation. Experimental results in NVIDIA Isaac Lab demonstrate that the dual-critic approach achieves a 3.5× faster goal-completion speed, doubles throughput, and attains a 65.2% effective success rate, substantially outperforming the unified critic counterpart.
📝 Abstract
Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy. A natural design choice is whether to use a single (unified) critic that estimates the combined value of all objectives, or separate (dual) critics with disjoint reward signals. We present a controlled comparison on the Unitree G1 humanoid (23 active DoF) in NVIDIA Isaac Lab, training loco-manipulation policies through a sequential curriculum spanning 13 levels from stationary reaching to walking with variable-orientation targets. In standardized evaluation, dual-critic policies reach targets 3.5$\times$ faster (6.5 vs. 22.6 simulation steps), achieve 2$\times$ higher throughput (14.3 vs. 7.0 validated reaches per 1,000 steps), and attain higher validated reach rates (65.2% vs. 53.8%) compared to the unified-critic policy. Notably, additional anti-gaming reward mechanisms provide no further improvement beyond the architectural change alone (60.9% vs. 65.2%). These results have direct implications for the emerging paradigm of RL fine-tuning of imitation-learned policies: when refining a pre-trained manipulation policy with RL, a unified critic risks suppressing the learned behavior through competing locomotion gradients. These findings demonstrate that critic architecture is a primary - and often overlooked - design choice in multi-objective humanoid RL, with greater impact than reward engineering on reaching efficiency.