An Analysis of Streaming Deep Reinforcement Learning for Adaptive Continual Learning in Robotics

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of robotic policies in adapting to unseen environments or morphological changes, as well as the inability of traditional offline batch learning to perform real-time updates. To this end, it explores streaming deep reinforcement learning for continuous robotic adaptation. By integrating specialized optimizers with plasticity loss mitigation techniques, the proposed framework dynamically adjusts pretrained policies using real-time experience streams, providing the first systematic validation of its feasibility. Experimental results demonstrate that this approach overcomes offline update constraints to achieve timestep-level adaptation. In quadruped locomotion tasks, it rapidly accommodates diverse unknown perturbations, improving success rates by up to 90% over pretrained baselines and significantly outperforming batch online methods. However, stability in manipulation tasks requires further improvement.
📝 Abstract
Over the course of a lifetime, robots may encounter novel scenarios unaccounted for in its original training that result in performance degradation. One common approach to mitigating this issue is to further grow the offline training dataset in hopes of producing a policy robust to these changes. In contrast, biological learning occurs moment-to-moment via a stream of experience, unlike the predominantly batch-based and offline nature of deep learning. Although recent works show the feasibility of stream-based deep reinforcement learning, where updates use only the latest experience, none have shown it to be a viable continual learning framework for adapting robotic policies to unseen changes. In this paper, we present the first analysis of streaming deep reinforcement learning for adaptive continual learning in robotics. In particular, we show that, following an initial pretraining phase, streaming deep RL can enable a robot to successfully adapt to unforeseen changes to itself, its environment, or goals. Our primary experiments within quadruped locomotion demonstrate that a deep neural network robotic policy with certain optimizers and plasticity loss mitigation techniques can successfully leverage domain task knowledge from its pretraining to quickly adapt online to diverse changes via stream learning, outperforming batch-based on-policy methods and improving task success rates by up to 90% over the pretrained policy. Furthermore, we perform additional evaluations on robotic manipulation tasks to determine if our previous observations extend to different robotic morphologies and scenarios. Our results show that the successes observed in quadruped locomotion can be partially realized in manipulation with stability and performance limitations. We conclude with a discussion on the limitations of our work and its implications for the future of continual robot learning.
Problem

Research questions and friction points this paper is trying to address.

streaming deep reinforcement learning
continual learning
robotics
adaptive policy
online adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming Deep Reinforcement Learning
Adaptive Continual Learning
Plasticity Loss Mitigation
Online Adaptation
Robotic Locomotion and Manipulation
💼 Related Jobs
No related jobs found.
T
Teeratham Vitchutripop
Department of Computer Science, Yale University, New Haven, CT 06520, USA
A
Alyssa Quarles
Department of Computer Science, Yale University, New Haven, CT 06520, USA
W
Wenhe Zhang
Department of Computer Science, Yale University, New Haven, CT 06520, USA
R
Richard Xue
Department of Computer Science, Yale University, New Haven, CT 06520, USA
Daniel Rakita
Daniel Rakita
Yale University
roboticsmotion planningoptimizationmachine learning