Score
Designs and evaluates model-update mechanisms that combine experience replay with knowledge-distillation losses applied to replayed samples, e.g., matching logits, intermediate features, or representations from a stored teacher to the current student. Implements and analyses the replay buffer, distillation targets, loss weighting schedules, and training procedures used to limit representation drift and preserve previously learned features during incremental updates.
This work systematically investigates the underutilized potential of experience replay in reinforcement learning-based post-training of large language models, challenging the prevailing assumption that fresh online data generation is indispensable. By carefully balancing data staleness, sample diversity, and computational cost, the authors design an efficient replay buffer mechanism that effectively substitutes strict online sampling. Their approach demonstrates, for the first time, that experience replay can significantly reduce inference-time computational overhead while maintaining or even improving model performance and effectively preserving policy entropy. This finding offers a compelling alternative to costly online data collection, suggesting that strategic reuse of historical interactions can sustain training efficacy without compromising behavioral diversity or learning stability.
While knowledge distillation often preserves task performance in student models, it frequently overlooks degradation in critical capabilities such as uncertainty calibration, boundary behavior, and safety. This work reframes distillation as a lossy projection of teacher behavior and introduces the “Distillation Loss Statement” framework, which integrates context-specific capability preservation objectives to establish a measurable and accountable evaluation paradigm. Through behavioral fidelity analysis, a capability taxonomy, and multidimensional assessment, the study systematically identifies and quantifies non-task-related capability losses. The resulting reproducible taxonomy of distillation-induced losses advances the field beyond mere performance retention toward practical standards that jointly prioritize reliability and responsibility in distilled models.
In continual learning, deep neural networks suffer from “plasticity loss”—a severe degradation in adaptability to new tasks—after multi-task backpropagation training. This work demonstrates that memory-based experience replay fundamentally mitigates this phenomenon and, for the first time, shows that replay alone—when integrated with standard Transformer architectures—fully restores continual plasticity. Crucially, no modifications to backpropagation, activation functions, or regularization are required; instead, replay triggers the Transformer’s inherent in-context learning capability. Evaluated across diverse continual learning benchmarks—including regression, classification, and policy evaluation—the method eliminates catastrophic forgetting entirely, sustaining long-term performance at the level of the initial network. These results establish memory-based replay coupled with off-the-shelf Transformers as an effective, non-invasive paradigm for continual adaptation.
In replay-based continual learning, limited buffer capacity and heuristic sample selection exacerbate catastrophic forgetting. To address this, we propose a data distillation framework tailored for continual learning. Our core innovation is a learnable soft-label distillation mechanism: instead of parameterizing the entire buffer, we decouple global knowledge distillation into a lightweight, trainable label generation module—substantially reducing computational overhead. This mechanism jointly optimizes distillation efficiency and generalization capability while enabling dynamic memory content updates. Evaluated on multiple standard benchmarks, our method significantly mitigates forgetting, achieves performance competitive with state-of-the-art replay approaches, and incurs lower memory footprint and computational cost.
To address catastrophic forgetting and new-knowledge suppression arising from the imbalance between memory consolidation and plasticity in continual learning, this paper proposes an enhanced experience replay framework. Methodologically, it extends Dark Experience Replay (DER) with a weight-adaptive mechanism, an erroneous-sample filtering module, and a historical-output correction strategy; concurrently, it refines Reservoir Sampling (RS) via a generalized acceptance probability function, a multi-level hierarchical buffer structure, and an active redundancy removal mechanism. Comprehensive evaluation across diverse continual learning benchmarks—including classification, regression, and reinforcement learning—demonstrates that the proposed approach significantly mitigates forgetting, accelerates adaptation to novel tasks, and improves both long-term performance stability and cross-task generalization robustness.
Online platform user behavior continuously evolves, causing behavioral analysis models to degrade due to data drift and catastrophic forgetting. To address this, we propose a knowledge-enhanced continual learning framework: (1) it integrates external knowledge bases to guide data augmentation, thereby overcoming the capacity limitations of conventional replay buffers; and (2) it introduces a synergistic mechanism combining multi-strategy augmentation with knowledge distillation to achieve dynamic balance between old and new knowledge. Evaluated on three anomaly behavior classification datasets, our method significantly outperforms classical replay-based baselines, achieving an average +4.2% improvement in F1-score. It effectively mitigates catastrophic forgetting and enhances long-term model stability. The framework provides a scalable solution for robust behavioral modeling in dynamic network environments, advancing continual learning for real-world behavioral analytics.
This work addresses the problem of model collapse in large language models caused by training on data contaminated with their own generated content. From a learning-theoretic perspective, the authors propose a replay adversary framework that injects historical model outputs into the training stream, offering the first fine-grained theoretical characterization of model collapse within the language generation limit paradigm. By integrating language generation limit theory, replay adversarial modeling, and non-uniform generation analysis, they reveal that replay is benign under strong consistency generation but induces performance degradation under weak generation paradigms. This study not only elucidates the effectiveness and limitations of existing mitigation strategies—such as data filtering, watermarking, and output curation—but also provides rigorous theoretical grounding and delineates their failure boundaries.
This work addresses a critical gap in continual learning research: while most existing methods focus on mitigating catastrophic forgetting, they largely overlook the conditions under which forward transfer—where knowledge from past tasks benefits new ones—can be effectively realized. The paper introduces, for the first time, a systematic three-condition framework that characterizes when forward transfer is feasible and proposes Transfer-Selective Replay (TSR), a novel method that leverages a zero-training-overhead task signature mechanism to automatically identify and replay only those historical samples beneficial to the current task. TSR integrates knowledge distillation to preserve performance on previous tasks while explicitly promoting forward transfer as a first-class objective. Experiments demonstrate that TSR significantly enhances forward transfer across both homogeneous and heterogeneous task sequences and consistently outperforms existing replay-based baselines, especially under limited replay budgets.