Score
Design and implement training regimes that maintain a teacher network by updating its parameters as an exponential moving average (EMA) of a student network so the teacher produces stable targets for the student; this includes building the EMA update mechanism, siamese-style student setups, and the target-generation and loss-coupling components. Analyze and tune EMA decay rates, update schedules, and their interaction with optimization to stabilize training and improve sample and compute efficiency of learned representations.
In language model fine-tuning, stochasticity induced by small-batch training causes severe fluctuations in generation quality; while standard Exponential Moving Average (EMA) of weights improves training smoothness, it introduces optimization lag due to historical weight accumulation. To address this, we propose Bias-Corrected EMA (BEMA), a theoretically grounded variant that eliminates iteration lag while preserving EMA’s variance-reduction capability via an explicit bias-correction mechanism. We establish a convergence analysis framework for BEMA and prove its superior theoretical convergence rate over both standard EMA and vanilla SGD. Empirically, BEMA consistently enhances training stability, accelerates convergence, and achieves higher final performance across multiple mainstream language model fine-tuning benchmarks—demonstrating both effectiveness and practical utility.
This work addresses the challenges faced by learning-based systems in heterogeneous, dynamic, and long-running environments, where environmental shifts often lead to high retraining costs, substantial labeling overhead, performance degradation, and sluggish adaptation. To tackle these issues, the paper introduces EMA—a lightweight, adaptive framework that supports diverse system and model architectures through a system-driven, data-centric approach. EMA employs a state transformer to align representations between old and new environments, enabling warm-start model adaptation, and intelligently prioritizes high-utility data samples for labeling based on their expected contribution to performance. Experimental evaluation across eight representative systems demonstrates that EMA reduces adaptation costs—such as GPU training time—by 14.9% to 42.4% while simultaneously improving system performance metrics, including network throughput, by 6.9% to 31.3%.
In continual learning, artificial neural networks suffer from catastrophic forgetting—performance on previously learned tasks degrades significantly upon training on new tasks. Existing approaches rely on heuristic task-scheduling protocols lacking theoretical guarantees of optimality. This paper bridges statistical physics and optimal control theory to establish, for the first time, an analytically tractable and provably optimal framework for task selection dynamics. Leveraging a teacher–student model, we derive exact training dynamics via dynamic mean-field analysis and obtain a closed-form optimal scheduling protocol that explicitly incorporates task similarity as a key regulator of forgetting. Empirical evaluation on synthetic data and real-world benchmarks (e.g., CIFAR-100) demonstrates substantial reduction in forgetting rates. Crucially, theoretical predictions align closely with experimental results, validating the framework’s strong interpretability, formal optimality guarantee, and cross-dataset generalizability.
This work investigates the learning mechanism of restricted Boltzmann machines (RBMs) for structured data within a teacher–student framework, specifically addressing whether a student RBM can recover the true latent representation from data generated by a teacher RBM exhibiting correlations among hidden variables. Using statistical-physics-based mean-field analysis and temperature-regularized inference, we systematically characterize how structural strength—quantified by the number of latent patterns and inter-pattern correlation in weight rows—affects the critical sample size required for successful learning. We find that enhanced structure drastically reduces the necessary sample complexity; in the absence of correlations, performance is independent of pattern count; and excessively low inference temperatures suppress pattern acquisition, leading to learning failure. Crucially, we establish for the first time that the student can achieve exact one-to-one or one-to-many pattern matching—surpassing the conventional two-hidden-unit limitation. These results provide the first analytically tractable generative-model foundation for the “lottery ticket hypothesis.”
Tuning weight decay (WD) in AdamW for large-scale models and datasets remains challenging due to its opaque interaction with model and data scale. Method: We introduce the exponential moving average (EMA) timescale—measured in epochs—as a physically interpretable core variable, decoupling and reformulating WD accordingly. Contribution/Results: Through theoretical analysis and extensive experiments across architectures (ResNet-18, ViT, NanoGPT) and datasets (CIFAR-10, ImageNet, OpenWebText), we find that the optimal EMA timescale is approximately invariant in epoch units. This yields an inverse scaling law: WD should increase with model size but decrease with dataset size. This law coheres with muP learning rate scaling; neglecting WD scaling while only scaling learning rate causes significant performance degradation in mid-to-late training. Our work establishes the first interpretable, principled scaling framework for AdamW weight decay, enhancing training stability and generalization in large-model regimes.
研究通过调整教师网络的几何形状,解决了教师-学生网络中学习能力差异问题,采用分析损失景观和调整学习率的方法提高学习成功率。
研究重新思考了教师-学生框架在测试时适应中的应用,通过使用不更新权重的顽固教师来解决长期稳定性问题,从而提高性能和鲁棒性。
This study investigates how effective learning rate drift induced by normalized updates affects training acceleration, stability, and resource scaling. Leveraging the random feature model and dynamical mean field theory (DMFT), the work elucidates the acceleration mechanisms and late-stage instabilities of normalized SGD, with theoretical predictions validated through linearized ResNet experiments. The primary contribution is the first solvable theoretical framework connecting normalization-induced acceleration, marginal stability, and width-batch allocation. Furthermore, this research quantifies the computational efficiency trade-offs between batch size and network width across distinct scaling regimes, empirically corroborating the predicted trends on CIFAR-5M.
This work addresses the instability and task-agnostic collapse commonly observed in self-play policy distillation, which often stem from ill-timed updates of the teacher policy. Through a systematic analysis of the temporal coupling between the teacher’s freezing interval (quarantine period) and the student’s learning dynamics, the study identifies clock-driven teacher refreshes as a primary cause of collapse. To mitigate this, the authors propose Consolidation-Gated Teacher Refresh (CGTR), an adaptive gating mechanism that triggers teacher updates only when jointly validated by improvements in reward and safe trajectory length. Requiring no task-specific hyperparameter tuning, CGTR achieves zero collapse across four diverse tasks—Chemistry, Biology, Physics, and ToolUse—while attaining state-of-the-art performance under a unified hyperparameter configuration and automatically adjusting the teacher refresh frequency per task.
研究探讨了权重衰减和学习率调度在尺度不变优化中的互动机制,通过精确的离散时间定律揭示了稳定与不稳定的边界,并提供了一个统一的框架来解释不同优化器的行为。