Score
Designs and implements hybrid control systems that learn a residual policy via reinforcement learning to correct or augment a base controller or model, where the residual operates in a learned latent representation rather than the raw observation or action space. This competence covers building the latent encoder/decoder, integrating the residual RL agent with the base controller for online residual exploration and local corrections near demonstrations, and training/evaluating sample-efficient residual policies that constrain perturbations to the base behavior.
This paper addresses the challenge of deeply integrating model predictive control (MPC) and reinforcement learning (RL), stemming from their fundamentally divergent model usage paradigms. To resolve this, we propose the first unified taxonomy for MPC–RL fusion, centered on *how models are used*, categorizing approaches into three paradigms: MPC-augmented RL, RL-augmented MPC, and co-designed architectures. Leveraging a unified Actor–Critic modeling framework, we systematically analyze how MPC’s online optimization enhances RL’s closed-loop performance and establish a performance-gain-oriented evaluation perspective grounded in closed-loop metrics. The survey comprehensively covers six application domains—including robotics, energy systems, and autonomous driving—and synthesizes cross-cutting modeling techniques bridging control theory and RL. Our work provides a scalable methodology and principled design guidelines for hybrid intelligent control systems.
High-degree-of-freedom robots face low sample efficiency, difficulty optimizing under sparse rewards, and insufficient safety guarantees during real-world reinforcement learning (RL) training. Method: We propose a residual offline RL fine-tuning framework: a behavior cloning (BC) policy serves as a fixed base, and only lightweight per-step residual corrections are learned—requiring neither dense reward signals nor online interaction, but only sparse binary rewards. Contribution/Results: This is the first work to achieve end-to-end RL training for embodied dexterous humanoid hands in real-world settings, significantly alleviating bottlenecks in sample efficiency and long-horizon task learning. The method attains state-of-the-art performance on both simulated and real-world visuomotor control tasks, demonstrating its effectiveness in high-dimensional systems and feasibility for practical deployment.
Controlling hybrid dynamical systems—such as legged robots and autonomous vehicles—driven by latent-variable-induced mode switches remains challenging due to the tight coupling between continuous dynamics and unobservable discrete events; conventional model-based methods neglect uncertainty, while model-free reinforcement learning suffers from poor generalization across modes. To address this, we propose SAC-MoE: a Soft Actor-Critic architecture augmented with a Mixture-of-Experts (MoE) structure, where a learnable router dynamically selects specialized policy experts conditioned on inferred latent dynamic modes. We further introduce a challenge-oriented curriculum learning strategy to enhance cross-mode transferability. To our knowledge, SAC-MoE is the first framework to enable latent-aware adaptive policy routing within SAC. Empirical evaluation on hybrid autonomous driving and legged locomotion tasks demonstrates up to 6× improvement in zero-shot generalization performance over prior methods.
本文通过条件流匹配方法将控制潜变量解码为四旋翼模型分布,以实现固定策略的在线预测调优和鲁棒性分析。
This work addresses the limited execution precision of Vision-Language-Action (VLA) models in contact-rich tasks, alongside the high cost and safety risks of real-world robotic reinforcement learning. We propose VLaRL, a framework that freezes a pretrained VLA and leverages its internal latent representations as control conditions and a sim-to-real transfer interface, thereby circumventing pixel-level alignment. By training a residual policy in simulation and employing a lightweight mapper to align cross-domain latent feature distributions, VLaRL achieves efficient transfer and zero-shot online deployment without adaptation. Experiments across four contact-rich tasks and two VLA backbones demonstrate that VLaRL significantly improves real-world success rates, validating the effectiveness of the latent conditioning and feature alignment mechanisms.
Existing residual reinforcement learning (RL) suffers from low sample efficiency under sparse rewards and struggles to accommodate stochastic base policies (e.g., Gaussian or diffusion-based policies). To address this, we propose an uncertainty-guided off-policy residual RL framework. First, we estimate base policy uncertainty via Bayesian or ensemble methods to dynamically guide exploration. Second, we introduce an off-policy residual Q-learning mechanism with observable base actions—enabling stable training for the first time with stochastic base policies. Our method seamlessly integrates Gaussian policy optimization and diffusion-based policy modeling. Evaluated on multi-task benchmarks (Robosuite and D4RL), it significantly outperforms fine-tuning, imitation-augmented, and prior residual RL approaches. Moreover, it achieves zero-shot sim-to-real transfer and robust execution on real robots. Key contributions include: (1) uncertainty-driven exploration grounded in base policy estimation, and (2) the first off-policy residual RL framework supporting stochastic base policies.
Traditional residual reinforcement learning adjusts actions only through additive corrections, which cannot alter the shape, scale, or state-dependent structure of the base policy’s action distribution, thereby limiting adaptability under dynamic changes. This work proposes Warp RL, the first approach to incorporate invertible flow models into policy adaptation by employing state-conditioned monotonic rational-quadratic spline flows to construct invertible transformations of action distributions. Initialized as the identity map, this method enables full distribution reshaping—beyond mere translation—strictly generalizing additive residual approaches while remaining compatible with both policy gradient and gradient-free optimization. In ManiSkill3 tasks requiring distributional reshaping, Warp RL significantly outperforms residual methods; on a real-world peg-insertion task, it achieves comparable success rates but completes episodes 30% faster.
This work addresses the longstanding methodological, objective, and cultural divide between reinforcement learning and control theory by proposing a novel paradigm that integrates adaptive control with actor-critic reinforcement learning. The resulting framework enables data-driven optimization of controllers by unifying dynamic programming and online learning mechanisms, thereby reconciling modeling and optimization perspectives from both fields within classical motion control tasks. Theoretical analysis elucidates fundamental differences between the two approaches, while empirical results demonstrate the efficacy of the integrated strategy. This synthesis offers a solution for controlling systems with unknown dynamics that simultaneously guarantees stability and retains strong learning capabilities, fostering interoperability and synergistic development across disciplinary boundaries.
本文提出一种延迟感知框架ARLI,通过状态增强和中继观察解决因推理延迟导致的机器人策略改进问题。
为降低强化学习入门难度,开发了RLLBC-Lib代码库,通过实现表格型和深度RL方法来阐明理论基础,并支持自动评分的编程作业。
本文针对生成式机器人策略中仅能控制噪声空间而无法调节动作表示的问题,提出了一种双潜在空间强化学习框架,通过控制初始噪声和中间动作表示来提高性能。