Score
Designs and evaluates algorithms and training procedures that continually update and optimize control policies across deployments to improve task performance and data efficiency without relying on full resets. This includes methods for continual policy optimization (CPO), behavioral-KL or prior-task policy regularization, and learning from recovery trajectories to maintain stability and improve learning under reset and timestep budgets.
This work addresses the challenge that AI agents with frozen weights after deployment struggle to learn continuously from experience, often failing on repeated tasks. The authors propose a continual learning mechanism leveraging external memory, which distills minimal feedback—either a single-bit outcome or natural language corrections—from each interaction into retrievable rules. Integrated with retrieval-augmented generation (RAG) and frozen large language models (e.g., Mistral Large, Claude Sonnet 5), this approach enables performance improvement without fine-tuning. The method demonstrates, for the first time, that extremely sparse feedback alone can drive sustained enhancement in frozen models and supports memory transfer across models. On the τ-bench banking tasks, it achieves success rates 1.6× (outcome-only feedback) and 2.6× (with corrections) higher than baseline, resolving 22 out of 84 tasks on which the baseline completely fails.
This work addresses the performance degradation of real-world reinforcement learning systems under the common “train-then-deploy” paradigm, which fails to adapt to dynamic environmental changes. To overcome this limitation, the paper introduces a novel “deploy-and-continuously-learn” paradigm that treats deployed agents as lifelong reinforcement learning systems. It identifies four key sources of non-stationarity encountered post-deployment and integrates online adaptation, explicit non-stationarity modeling, and lifelong learning mechanisms to maintain robust performance. Through analysis of real-world deployment scenarios, the study demonstrates the advantages of this paradigm in sustaining long-term effectiveness. Furthermore, it proposes new evaluation metrics tailored to continuous learning settings, aiming to shift the research community’s focus from static models toward lifelong learning agents capable of enduring environmental dynamics.
This work addresses the fundamental challenge of jointly ensuring safety and enabling continual learning in nonlinear, nonstationary systems—such as those subject to unknown faults or abrupt constraint changes—where conventional safe reinforcement learning (Safe RL) and continual RL methods fail to coexist. We identify and analyze the intrinsic mechanism by which continual learning erodes safety constraints. To resolve this, we propose a safety-prioritized, elastic reward shaping framework that online integrates Elastic Weight Consolidation (EWC) with Constrained Policy Optimization (CPO), thereby achieving synergistic optimization of safety constraint satisfaction and task performance stability. Extensive evaluation on MuJoCo (HalfCheetah, Ant) under diverse nonstationary fault scenarios—including joint failures and sudden velocity constraint shifts—demonstrates that our method achieves over 92% safety constraint satisfaction, reduces performance degradation by 67%, and significantly mitigates catastrophic forgetting. To the best of our knowledge, this is the first approach to provably reconcile Safe RL and continual RL in dynamic nonlinear systems.
This work addresses a critical limitation in existing reinforcement learning post-training methods—the absence of mechanisms to validate the efficacy of policy updates, which often leads to optimization drift or collapse. To mitigate this, the authors propose the PIRL framework, which reframes the optimization objective from immediate reward maximization to cumulative policy improvement across training iterations. Central to this framework is the PIPO algorithm, which introduces, for the first time, a policy improvement feedback mechanism. By employing a sliding window to retrospectively validate historical baselines, PIPO establishes a self-correcting closed-loop optimization process that guarantees each policy update positively contributes to final performance. Empirical evaluations on mathematical reasoning benchmarks demonstrate that PIPO achieves superior stability and performance compared to GRPO and its variants.
This work addresses the challenge of simultaneously ensuring behavioral preference satisfaction and monotonic policy improvement while maintaining efficient exploration in policy optimization. We propose the ε-retraining framework, which introduces (i) an iterative retraining region construction mechanism that dynamically identifies preference-violating regions in the state space via behavior bias localization; (ii) a decaying ε-scheduling strategy to jointly balance global exploration and local correction; and (iii) neural network formal verification—using ReLU partitioning and linear programming—to quantify preference adherence. Evaluated across motion control, navigation, and power grid dispatch tasks with over one hundred random seeds, our method achieves a 37.2% increase in preference compliance rate and accelerates convergence by 2.1×, significantly improving sample efficiency and policy reliability.
Traditional continual learning is constrained by a parameter-centric paradigm, limiting its capacity to meet system-level adaptation demands in dynamic environments. This work proposes a “Tri-Axis Framework” (When, How, Where), offering a unified perspective that reorients continual learning beyond mere parameter updates toward external architectures and inference-time adaptation. By integrating off-policy/on-policy learning, test-time training, external memory systems, and skill repositories, the framework transcends the limitations of static parameter spaces and gradient-based optimization. A systematic review elucidates the field’s evolutionary trajectory and highlights pivotal challenges and future directions inherent in this paradigm shift.
This work addresses catastrophic forgetting in continual learning under non-stationary data streams by proposing the COLD framework, which introduces, for the first time, the Drift-Plus-Penalty stochastic optimization method from control theory into this domain. COLD formulates forgetting as a controlled dynamic process, employing virtual queues to track performance deviations on historical tasks and jointly minimizing the current task loss and queue drift at each optimization step. This mechanism explicitly governs the stability-plasticity trade-off. The framework provides theoretical guarantees on stability and convergence, and achieves significantly superior performance over state-of-the-art methods on standard benchmarks, enabling controllable and efficient suppression of catastrophic forgetting.
Policy updates in reinforcement learning are highly sensitive to distributional shifts, a problem exacerbated in large-scale settings where discrepancies in numerical precision and sampling between training and inference introduce further instability. Existing approaches often rely on fixed hyperparameters, limiting their adaptability to variations in tasks, model scales, or data distributions. This work proposes a batch-adaptive policy optimization objective that dynamically modulates update intensity based on the effective sample size of policy ratios within each batch. By replacing fixed clipping with an adaptive mechanism grounded in the empirical distribution of ratios, the method jointly addresses trust-region constraints and off-policy data reliability without introducing additional hyperparameters. Empirical results demonstrate that the proposed approach matches or surpasses carefully tuned baselines across diverse settings, significantly enhancing algorithmic robustness and generalization.
This work addresses the challenge that autonomous driving policies struggle to continuously learn from their own mistakes in long-tail traffic scenarios while preserving previously acquired capabilities. To this end, the paper introduces the R²LPL framework, which uniquely integrates lifelong learning with an error-driven correction mechanism. By leveraging rollout-based error detection, recoverable state filtering, and corrective target retrieval, the method transforms sparse, unsupervised signals from closed-loop failures into compact supervisory knowledge for efficient policy updates. Requiring only a small number of rollouts and learning iterations, R²LPL elevates a moderately performing planner to state-of-the-art performance on the large-scale nuPlan closed-loop benchmark, demonstrating particularly significant improvements over existing approaches in the challenging Test14-hard long-tail scenarios.
This study investigates whether introducing a learned command adapter onto a frozen locomotion policy yields observable and recoverable performance gains. To this end, the authors propose an adapter necessity auditing framework that integrates closed-loop system identification, counterfactual reasoning, cluster refitting, and constraint violation analysis to disentangle deployment gain, state allocation gain, and global operational gain, thereby supporting GO/NO-GO/ABSTAIN deployment decisions. Experiments on the Go2 platform reveal only 0.55% recoverable allocation gain; direct querying yields a NO-GO verdict, while VGCC and MPC-based queries result in ABSTAIN, indicating that the value of an adapter must be grounded in observable evidence rather than prior assumptions.