rl-based whole-body control

Design and train reinforcement-learning-based whole-body controllers—policy representations and training pipelines—that coordinate locomotion and manipulation for articulated dynamical systems, explicitly handling nonlinear coupled dynamics, contacts, and whole-body constraints. Implement and analyze objectives, architectures, and training/regression methods to improve stability, robustness to external disturbances, and coordinated loco-manipulation behavior.

rl-basedwhole-bodycontrol

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of achieving safe and compliant whole-body manipulation for legged robots in dynamic environments by proposing a hybrid architecture that integrates model-driven admittance control with a reinforcement learning–based gait policy. Safety during physical interaction is ensured through a reference governor, while a neural network–enhanced Kalman filter improves base velocity estimation accuracy. A unified whole-body force response is realized using six-degree-of-freedom force/torque sensing. Experimental validation on the Unitree Go2 platform demonstrates significant improvements in high-precision interaction tracking, compliant human–robot collaboration, and safety-critical reliability in dynamic scenarios.

compliancecontact interactionloco-manipulation

Multi-critic Learning for Whole-body End-effector Twist Tracking

Jul 11, 2025
AE
Aravind Elanjimattathil Vijayan
🏛️ ETH Zurich | ANYbotics AG

Quadrupedal robots face conflicting objectives between whole-body locomotion and manipulator operation: stable base pose is required for locomotion, whereas end-effector trajectory tracking often necessitates base tilting to expand the reachable workspace; moreover, existing reinforcement learning (RL) methods relying on pose-based task specifications struggle to achieve smooth velocity tracking. Method: We propose a multi-critic Actor-Critic framework that decouples locomotion and manipulation reward signals, and introduce twist-based end-effector velocity tracking—enabling emergent whole-body coordination without explicit dynamic modeling. Contribution/Results: The method supports both discrete pose and continuous trajectory tracking, accommodates dynamic gaits and real-time manipulation, and is validated in simulation and on a physical quadrupedal manipulator. It achieves high-precision end-effector velocity tracking during locomotion and demonstrates adaptive base tilting to extend the operational workspace through coordinated whole-body motion.

Conflicting goals in locomotion and arm motion controlDifficulty in tracking both poses and motion trajectoriesLack of direct end-effector velocity control in RL

Humanoid robots face a fundamental trade-off between control accuracy and robustness in locomanipulation tasks, stemming from high-dimensional, unstable dynamics and complex multi-contact interactions. To address this, we propose a hybrid paradigm integrating model-driven trajectory optimization with end-to-end reinforcement learning: torque-constrained differential dynamic programming (DDP) generates physically feasible, natural whole-body reference trajectories; these are then refined via PPO or SAC for policy fine-tuning and robustness enhancement. We further develop a full-body rigid-body dynamics model and embed a real-time trajectory tracking controller. Evaluated on the Digit robot, our approach achieves 2.3× faster training convergence, high-fidelity trajectory tracking, and successful real-world deployment—demonstrating robust locomanipulation capabilities including walking-while-grasping and obstacle negotiation across diverse scenarios. The method significantly improves sim-to-real transfer performance and enhances motion naturalness.

Complex contact-rich loco-manipulation task challengesHigh-dimensional unstable dynamics in humanoid robotsTrade-offs between model-based control and reinforcement learning

Learning Whole-Body Loco-Manipulation for Omni-Directional Task Space Pose Tracking With a Wheeled-Quadrupedal-Manipulator

Dec 04, 2024
KJ
Kaiwen Jiang
🏛️ Southern University of Science and Technology | LimX Dynamics | Zhejiang University-University of Illinois Urbana-Champaign Institute

To address the challenge of achieving precise six-degree-of-freedom (6-DOF) end-effector (EE) pose tracking in task space for wheeled quadrupedal manipulator robots, this paper proposes a deep reinforcement learning (DRL)-based whole-body coordinated control framework. The method directly implements closed-loop 6D pose control in task space—bypassing hierarchical planning or inverse kinematics decomposition. Its key contributions are: (1) a nonlinear Reward Fusion Module (RFM) that explicitly models the multi-stage coupling among base motion, manipulator operation, and balance maintenance; and (2) a teacher–student hierarchical RL training paradigm to mitigate motion–balance coupling under high kinematic redundancy. Extensive simulation and real-robot experiments demonstrate smooth, robust tracking performance, achieving mean position error < 5 cm and orientation error < 0.1 rad—setting a new state-of-the-art.

Achieve precise 6D pose tracking with reinforcement learningBalance redundant degrees in whole-body loco-manipulation motionCoordinate floating base and robotic arm for 6D end-effector tracking

Latest Papers

What's happening recently
View more

This study addresses the challenges of balance maintenance and command tracking in humanoid robots carrying heavy loads, where center-of-mass shifts and sustained payloads induce instability. To this end, a whole-body force-controlled mobile manipulation framework is proposed. Methodologically, model predictive control guides reinforcement learning to train separate teacher policies for dual-arm manipulation and locomotion, which are subsequently distilled into a unified policy. Furthermore, a capture-point-based control barrier function is introduced to augment the wrist-force teacher, thereby enhancing dynamic stability under heavy loads. Experimental results demonstrate that the proposed strategy achieves minimal velocity tracking error under a 10 kg payload, reduces the divergent component of motion (DCM) deviation by 35.7%, and successfully withstands external disturbances up to 130 N applied as torso pushes.

BalanceCommand trackingHeavy objects

This work addresses the challenges of high-dimensional control, postural instability, and real-time perceptual demands faced by humanoid robots operating on complex terrains. To this end, the authors propose a whole-body motion control framework that integrates diffusion models with reinforcement learning. The approach uniquely combines real-time terrain-aware diffusion-based motion generation with a reinforcement learning-based tracking controller, augmented by a closed-loop fine-tuning mechanism to enable online adaptation of reference motions and coordinated whole-body responses. Evaluated on the Unitree G1 platform, the system successfully executes diverse locomotion tasks—including stepping over boxes, traversing rails, climbing stairs, and navigating mixed-terrain environments—demonstrating significantly enhanced motion generalization and robustness.

humanoid locomotionmotion generationreal-time perception

Accurately inferring neuromuscular control strategies from observed kinematics in high-dimensional, redundant musculoskeletal systems remains a formidable challenge. This work proposes a large-scale parallel musculoskeletal simulation framework that integrates GPU-accelerated dynamics, adversarial reward aggregation, and value-guided flow exploration to overcome optimization bottlenecks in high-dimensional reinforcement learning for muscle control. For the first time, the method simultaneously achieves high-fidelity motion tracking—accurately reproducing joint angles and body positions—and discovers diverse neuromuscular control policies across complex dynamic tasks such as dancing, cartwheels, and backflips. The results reveal that markedly distinct muscle activation patterns can produce nearly identical external movement outcomes, highlighting the inherent redundancy and flexibility of biological motor control.

embodied learninghigh-dimensional redundancymotor control diversity

This study addresses the challenge of planning whole-body, contact-rich manipulation for humanoid robots using first-person RGB and proprioceptive inputs. The authors propose an end-to-end framework that employs a visual planner to predict torso, wrist, and foot poses, coupled with a reinforcement learning-based whole-body tracker for action execution. Notably, the approach enables direct training from human demonstrations without requiring motion retargeting and introduces a cross-error augmentation mechanism to enhance multi-module coordination robustness. Experiments on a real Unitree G1 robot demonstrate successful execution of complex tasks, including box-kicking sorting, object catching and throwing, and climbing. In simulation, the system achieves a 77% success rate for undisturbed sorting and maintains 64% under 40 N·s impulse disturbances, validating its effectiveness in dynamic, contact-intensive scenarios.

Contact-rich planningEgocentric visionHumanoid robots

This work addresses the challenges of high-dimensional action spaces, underutilized redundancy, and insufficient control accuracy in high-precision whole-body loco-manipulation for highly articulated robots. The authors propose a hierarchical control framework wherein a high-level planner leverages Kinematic Normalizing Flows to generate diverse, kinematically feasible partial reference trajectories in a latent space, effectively exploring redundant solutions. A low-level controller then employs imitation learning to accurately track these references while ensuring physical feasibility. By integrating a large-scale kinematic dataset with high-dimensional action modeling, the approach significantly outperforms existing methods in simulation. Hardware experiments across eight tasks and 24 trials demonstrate state-of-the-art performance, achieving end-effector pose errors of 4.5 cm and 0.14 rad, as well as mobile tracking errors of 0.1 m/s and 0.01 rad/s.

high-dimensional action spacehigh-DoF robotic systemskinematic redundancy

Hot Scholars

DZ

Ding Zhao

Carnegie Mellon University
Trustworthy AIAI safetyreinforcement learningautonomous vehicles
CQ

Chen Qiu

Research Scientist, Bosch Center for AI
Machine LearningDeep LearningAnomaly Detection
HE

H. Eric Tseng

Uni. of Texas at Arlington
Automotive Control
BC

Bingqing Chen

Bosch Center for Artificial Intelligence
Reinforcement LearningEnergy & Sustainability
YN

Yaru Niu

Carnegie Mellon University
Robot LearningRoboticsReinforcement LearningMulti-Agent Systems