Score
Designs training protocols that partition optimization into discrete phases in which selected model parameters or components (for example specific layers, heads, or reward-weight modules) are held fixed while others are updated; this includes choosing which parts to freeze, determining phase durations and transitions, and implementing frozen-phase shaping. It also involves measuring and analyzing the effects of those schedules on learning dynamics and stationarity—e.g., preventing replay-buffer contamination from stale labels, maintaining off-policy learning stationarity, and trading off stability versus adaptability.
Learning rate scheduling in large language model training lacks rigorous theoretical foundations, leading to heuristic designs and suboptimal convergence. Method: This paper establishes, for the first time, a quantitative alignment between practical schedulers (e.g., linear decay) and tight non-smooth convex optimization lower bounds—eliminating spurious logarithmic factors in prior analyses and enabling principled cross-scheduler optimal learning rate transfer. We integrate convex optimization theory, scheduler modeling, and empirical validation, conducting systematic evaluations on 124M- and 210M-parameter Llama models. Results: Theory-guided scheduler design yields faster convergence and improved stability, empirically validating optimization theory’s practical relevance for large-model training. Core contribution: bridging the gap between theoretical performance bounds and engineering schedulers by providing a transferable, interpretable, and theoretically grounded framework for learning rate tuning.
This study addresses the unclear applicability boundaries between single-policy and multi-policy approaches in non-stationary reinforcement learning by proposing a phase-structure-based duration decomposition method. By uncovering the root causes underlying the divergence between theoretical equivalence and practical performance, we establish a policy selection criterion centered on the ratio of transient to quasi-stationary durations, validated through state augmentation techniques and numerical simulations. The research confirms key hypotheses, demonstrating that prolonged phases favor multi-policy methods while increased environmental heterogeneity imposes greater burdens on single-policy approaches. These findings provide a rigorous theoretical foundation for determining optimal policy types in non-stationary environments, thereby effectively guiding efficient policy design.
This work addresses the instability and task-agnostic collapse commonly observed in self-play policy distillation, which often stem from ill-timed updates of the teacher policy. Through a systematic analysis of the temporal coupling between the teacher’s freezing interval (quarantine period) and the student’s learning dynamics, the study identifies clock-driven teacher refreshes as a primary cause of collapse. To mitigate this, the authors propose Consolidation-Gated Teacher Refresh (CGTR), an adaptive gating mechanism that triggers teacher updates only when jointly validated by improvements in reward and safe trajectory length. Requiring no task-specific hyperparameter tuning, CGTR achieves zero collapse across four diverse tasks—Chemistry, Biology, Physics, and ToolUse—while attaining state-of-the-art performance under a unified hyperparameter configuration and automatically adjusting the teacher refresh frequency per task.
This work addresses the challenge that AI agents with frozen weights after deployment struggle to learn continuously from experience, often failing on repeated tasks. The authors propose a continual learning mechanism leveraging external memory, which distills minimal feedback—either a single-bit outcome or natural language corrections—from each interaction into retrievable rules. Integrated with retrieval-augmented generation (RAG) and frozen large language models (e.g., Mistral Large, Claude Sonnet 5), this approach enables performance improvement without fine-tuning. The method demonstrates, for the first time, that extremely sparse feedback alone can drive sustained enhancement in frozen models and supports memory transfer across models. On the τ-bench banking tasks, it achieves success rates 1.6× (outcome-only feedback) and 2.6× (with corrections) higher than baseline, resolving 22 out of 84 tasks on which the baseline completely fails.
In continual learning, artificial neural networks suffer from catastrophic forgetting—performance on previously learned tasks degrades significantly upon training on new tasks. Existing approaches rely on heuristic task-scheduling protocols lacking theoretical guarantees of optimality. This paper bridges statistical physics and optimal control theory to establish, for the first time, an analytically tractable and provably optimal framework for task selection dynamics. Leveraging a teacher–student model, we derive exact training dynamics via dynamic mean-field analysis and obtain a closed-form optimal scheduling protocol that explicitly incorporates task similarity as a key regulator of forgetting. Empirical evaluation on synthetic data and real-world benchmarks (e.g., CIFAR-100) demonstrates substantial reduction in forgetting rates. Crucially, theoretical predictions align closely with experimental results, validating the framework’s strong interpretability, formal optimality guarantee, and cross-dataset generalizability.
This study addresses the challenge of continual learning for deployed agents, where offline replay is impractical and online replay risks interfering with ongoing computation. To overcome this, we propose a silent degree-of-freedom replay method based on a local sleep mechanism. By employing k-WTA hidden layers and refractory period rules, the approach leverages input-unused degrees of freedom, while synaptic masking techniques strictly confine memory replay to silent synapses. This enables interference-free continual learning without requiring an offline phase. Experimental results demonstrate that the proposed method surpasses multiple baselines on Split-MNIST and outperforms DER++ in single-pass streaming scenarios, effectively balancing computational efficiency with model accuracy.
This work addresses the instability in multi-agent reinforcement learning caused by dynamic reward weights generated by large language models (LLMs), which violate the stationarity assumption of potential-based reward shaping (PBRS), contaminate experience replay, and destabilize training. The study is the first to identify three distinct failure modes arising from LLM-induced reward dynamics that disrupt mechanistic dependencies in the learning process. To mitigate these issues, the authors propose two stabilization strategies: training-phase-dependent weight freezing and exponential moving average (EMA) smoothing of reward signals. Integrated with QMIX and VDN architectures, the approach achieves success rates of 86.7%, 95.9%, and 99.9% on the Simple Spread, Level-Based Foraging, and SMAC 3m benchmarks, respectively, substantially outperforming baseline methods. These results establish reward signal stationarity as a critical design constraint for LLM-augmented multi-agent systems.
This study addresses the aging of production models caused by stale training data snapshots and the limited scheduling efficacy of existing global refresh strategies. It reveals the equivalence limitations of single-age triggers and proposes an optimal refresh budget allocation mechanism based on differentiated data segmentation. Through stochastic process modeling and Cauchy-Schwarz inequality optimization, this work proves that a global staleness budget yields no additional gains, establishes the theoretical foundation for segment-independent age metrics, and derives closed-form optimal solutions. Experimental results demonstrate that, under equivalent budgets, the proposed strategy reduces weighted staleness exposure by 8%–29% compared to uniform timers while maintaining significant robustness in noisy environments.
Traditional continual learning is constrained by a parameter-centric paradigm, limiting its capacity to meet system-level adaptation demands in dynamic environments. This work proposes a “Tri-Axis Framework” (When, How, Where), offering a unified perspective that reorients continual learning beyond mere parameter updates toward external architectures and inference-time adaptation. By integrating off-policy/on-policy learning, test-time training, external memory systems, and skill repositories, the framework transcends the limitations of static parameter spaces and gradient-based optimization. A systematic review elucidates the field’s evolutionary trajectory and highlights pivotal challenges and future directions inherent in this paradigm shift.
研究通过层次潜因器探讨了高级控制器何时应保留或修改低级过程策略的问题,发现状态依赖并不等同于决策价值。