latent-space residual reinforcement learning

Designs and implements hybrid control systems that learn a residual policy via reinforcement learning to correct or augment a base controller or model, where the residual operates in a learned latent representation rather than the raw observation or action space. This competence covers building the latent encoder/decoder, integrating the residual RL agent with the base controller for online residual exploration and local corrections near demonstrations, and training/evaluating sample-efficient residual policies that constrain perturbations to the base behavior.

latent-spaceresidualreinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Residual Off-Policy RL for Finetuning Behavior Cloning Policies

Sep 23, 2025
LA
Lars Ankile
🏛️ Amazon FAR (Frontier AI & Robotics) | Stanford University | Carnegie Mellon University | UC Berkeley

High-degree-of-freedom robots face low sample efficiency, difficulty optimizing under sparse rewards, and insufficient safety guarantees during real-world reinforcement learning (RL) training. Method: We propose a residual offline RL fine-tuning framework: a behavior cloning (BC) policy serves as a fixed base, and only lightweight per-step residual corrections are learned—requiring neither dense reward signals nor online interaction, but only sparse binary rewards. Contribution/Results: This is the first work to achieve end-to-end RL training for embodied dexterous humanoid hands in real-world settings, significantly alleviating bottlenecks in sample efficiency and long-horizon task learning. The method attains state-of-the-art performance on both simulated and real-world visuomotor control tasks, demonstrating its effectiveness in high-dimensional systems and feasibility for practical deployment.

Enabling effective RL training on high-degree-of-freedom real-world systemsImproving behavior cloning policies limited by human demonstration qualityOvercoming RL challenges like sample inefficiency and safety concerns

Controlling hybrid dynamical systems—such as legged robots and autonomous vehicles—driven by latent-variable-induced mode switches remains challenging due to the tight coupling between continuous dynamics and unobservable discrete events; conventional model-based methods neglect uncertainty, while model-free reinforcement learning suffers from poor generalization across modes. To address this, we propose SAC-MoE: a Soft Actor-Critic architecture augmented with a Mixture-of-Experts (MoE) structure, where a learnable router dynamically selects specialized policy experts conditioned on inferred latent dynamic modes. We further introduce a challenge-oriented curriculum learning strategy to enhance cross-mode transferability. To our knowledge, SAC-MoE is the first framework to enable latent-aware adaptive policy routing within SAC. Empirical evaluation on hybrid autonomous driving and legged locomotion tasks demonstrates up to 6× improvement in zero-shot generalization performance over prior methods.

Address poor generalization of standard RL methods during abrupt mode transitionsControl hybrid dynamical systems with unobservable latent parameters and mode switchesImprove robustness to unseen environments and switching locations through curriculum learning

This work addresses the limited execution precision of Vision-Language-Action (VLA) models in contact-rich tasks, alongside the high cost and safety risks of real-world robotic reinforcement learning. We propose VLaRL, a framework that freezes a pretrained VLA and leverages its internal latent representations as control conditions and a sim-to-real transfer interface, thereby circumventing pixel-level alignment. By training a residual policy in simulation and employing a lightweight mapper to align cross-domain latent feature distributions, VLaRL achieves efficient transfer and zero-shot online deployment without adaptation. Experiments across four contact-rich tasks and two VLA backbones demonstrate that VLaRL significantly improves real-world success rates, validating the effectiveness of the latent conditioning and feature alignment mechanisms.

Contact-rich manipulationResidual reinforcement learningSim-to-real transfer

Accelerating Residual Reinforcement Learning with Uncertainty Estimation

Jun 20, 2025
LD
Lakshita Dodeja
🏛️ Brown University | RAI Institute

Existing residual reinforcement learning (RL) suffers from low sample efficiency under sparse rewards and struggles to accommodate stochastic base policies (e.g., Gaussian or diffusion-based policies). To address this, we propose an uncertainty-guided off-policy residual RL framework. First, we estimate base policy uncertainty via Bayesian or ensemble methods to dynamically guide exploration. Second, we introduce an off-policy residual Q-learning mechanism with observable base actions—enabling stable training for the first time with stochastic base policies. Our method seamlessly integrates Gaussian policy optimization and diffusion-based policy modeling. Evaluated on multi-task benchmarks (Robosuite and D4RL), it significantly outperforms fine-tuning, imitation-augmented, and prior residual RL approaches. Moreover, it achieves zero-shot sim-to-real transfer and robust execution on real robots. Key contributions include: (1) uncertainty-driven exploration grounded in base policy estimation, and (2) the first off-policy residual RL framework supporting stochastic base policies.

Addressing sparse rewards in Residual Reinforcement LearningEnhancing sample efficiency in Residual RL for stochastic policiesImproving robustness for sim-to-real transfer in RL

Latest Papers

What's happening recently
View more

Traditional residual reinforcement learning adjusts actions only through additive corrections, which cannot alter the shape, scale, or state-dependent structure of the base policy’s action distribution, thereby limiting adaptability under dynamic changes. This work proposes Warp RL, the first approach to incorporate invertible flow models into policy adaptation by employing state-conditioned monotonic rational-quadratic spline flows to construct invertible transformations of action distributions. Initialized as the identity map, this method enables full distribution reshaping—beyond mere translation—strictly generalizing additive residual approaches while remaining compatible with both policy gradient and gradient-free optimization. In ManiSkill3 tasks requiring distributional reshaping, Warp RL significantly outperforms residual methods; on a real-world peg-insertion task, it achieves comparable success rates but completes episodes 30% faster.

action distributiondistribution reshapingdynamics shift

This work addresses the longstanding methodological, objective, and cultural divide between reinforcement learning and control theory by proposing a novel paradigm that integrates adaptive control with actor-critic reinforcement learning. The resulting framework enables data-driven optimization of controllers by unifying dynamic programming and online learning mechanisms, thereby reconciling modeling and optimization perspectives from both fields within classical motion control tasks. Theoretical analysis elucidates fundamental differences between the two approaches, while empirical results demonstrate the efficacy of the integrated strategy. This synthesis offers a solution for controlling systems with unknown dynamics that simultaneously guarantees stability and retains strong learning capabilities, fostering interoperability and synergistic development across disciplinary boundaries.

Actor-Critic AlgorithmsAdaptive ControlControl Theory

Hot Scholars

GS

Guanya Shi

Assistant Professor, CMU RI | Amazon Scholar, FAR (Frontier AI & Robotics)
RoboticsRobot LearningReinforcement LearningControl
HQ

Haozhi Qi

UC Berkeley
RoboticsDeep LearningComputer Vision
CL

Chenran Li

PhD, University of California, Berkeley
reinforcement learningmotion planningsimulationautonomous driving
DS

Dorsa Sadigh

Stanford University
RoboticsHuman-Robot InteractionMachine LearningArtificial Intelligence
KS

Koushil Sreenath

Mechanical Engineering, UC Berkeley
ControlRoboticsLearning