🤖 AI Summary
To address low online adaptation efficiency and catastrophic forgetting in reinforcement learning agents facing abrupt environmental changes during dynamic deployment, this paper proposes an online meta-reinforcement learning framework integrating change-awareness and knowledge preservation. The method introduces: (1) a priority-based exploration sampling mechanism for rapid detection and response to environmental shifts; (2) a decomposable policy representation with parameter-isolated updates to selectively retain knowledge from prior tasks; and (3) a structured knowledge-preserving representation that disentangles shared and task-specific policy modules. Evaluated across diverse simulated environments featuring sudden dynamics shifts and goal drift, the approach achieves a 3.1× improvement in sample efficiency and reduces forgetting by 62% compared to state-of-the-art online adaptation baselines.
📝 Abstract
Real-world autonomous decision-making systems, from robots to recommendation engines, must operate in environments that change over time. While deep reinforcement learning (RL) has shown an impressive ability to learn optimal policies in stationary environments, most methods are data intensive and assume a world that does not change between training and test time. As a result, conventional RL methods struggle to adapt when conditions change. This poses a fundamental challenge: how can RL agents efficiently adapt their behavior when encountering novel environmental changes during deployment without catastrophically forgetting useful prior knowledge? This dissertation demonstrates that efficient online adaptation requires two key capabilities: (1) prioritized exploration and sampling strategies that help identify and learn from relevant experiences, and (2) selective preservation of prior knowledge through structured representations that can be updated without disruption to reusable components.