How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
In lifelong reinforcement learning, transferring a single historical policy to novel tasks is often ineffective. To address this limitation, this work proposes AMSC, a method that dynamically selects and combines multiple historical policies weighted by online experience to accelerate learning. Specifically, AMSC employs non-parametric Wasserstein embeddings to quantify task similarity, integrating Z-score normalization with the Sparsemax algorithm to achieve adaptive sparse policy combination over variable-length support sets. Experimental evaluations on CT-graph and MiniGrid benchmarks demonstrate that AMSC significantly improves average performance and forward transfer capabilities while effectively mitigating catastrophic forgetting.
πŸ“ Abstract
In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and Composition (AMSC) is designed to estimate similarity from online experience via non-parametric Wasserstein task embeddings from state-action-reward samples. The z-score-normalized sparsemax of the similarity scores are used to derive a variable-size support to periodically choose and weight policies to form a prior when learning a new task. On CT-graph and MiniGrid, AMSC achieves higher mean performance and forward transfer than the evaluated modular composition baselines while exhibiting no forgetting. Results on Continual World suggest that identifying relevant prior knowledge and determining its layer-specific composition may require additional layer-specific tuning. Ablations show that selecting relevant sources and determining how strongly to reuse them are central to these gains. Independently measured pairwise transfer is also positively associated with task-embedding similarity. These results indicate that task similarity can be an effective criterion to select and weight specific knowledge for reuse in lifelong reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

Lifelong Reinforcement Learning
Policy Reuse
Task Similarity
Continual Adaptation
Knowledge Transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lifelong Reinforcement Learning
Adaptive Mask Selection and Composition
Wasserstein Task Embeddings
Policy Reuse
Continual Adaptation