train with ema teacher

Design and implement training regimes that maintain a teacher network by updating its parameters as an exponential moving average (EMA) of a student network so the teacher produces stable targets for the student; this includes building the EMA update mechanism, siamese-style student setups, and the target-generation and loss-coupling components. Analyze and tune EMA decay rates, update schedules, and their interaction with optimization to stabilize training and improve sample and compute efficiency of learned representations.

trainwithemateacher

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes

Jul 31, 2025
AB
Adam Block
🏛️ Columbia University | OpenAI

In language model fine-tuning, stochasticity induced by small-batch training causes severe fluctuations in generation quality; while standard Exponential Moving Average (EMA) of weights improves training smoothness, it introduces optimization lag due to historical weight accumulation. To address this, we propose Bias-Corrected EMA (BEMA), a theoretically grounded variant that eliminates iteration lag while preserving EMA’s variance-reduction capability via an explicit bias-correction mechanism. We establish a convergence analysis framework for BEMA and prove its superior theoretical convergence rate over both standard EMA and vanilla SGD. Empirically, BEMA consistently enhances training stability, accelerates convergence, and achieves higher final performance across multiple mainstream language model fine-tuning benchmarks—demonstrating both effectiveness and practical utility.

Improves convergence rates in language model fine-tuningMitigates training instability from small batch sizesReduces bias lag in exponential moving averages

This work addresses the challenges faced by learning-based systems in heterogeneous, dynamic, and long-running environments, where environmental shifts often lead to high retraining costs, substantial labeling overhead, performance degradation, and sluggish adaptation. To tackle these issues, the paper introduces EMA—a lightweight, adaptive framework that supports diverse system and model architectures through a system-driven, data-centric approach. EMA employs a state transformer to align representations between old and new environments, enabling warm-start model adaptation, and intelligently prioritizes high-utility data samples for labeling based on their expected contribution to performance. Experimental evaluation across eight representative systems demonstrates that EMA reduces adaptation costs—such as GPU training time—by 14.9% to 42.4% while simultaneously improving system performance metrics, including network throughput, by 6.9% to 31.3%.

data labelingdynamic environmentslearning-based systems

Optimal Protocols for Continual Learning via Statistical Physics and Control Theory

Sep 26, 2024
FM
Francesco Mori
🏛️ University of Oxford | Chalmers University of Technology | University of Gothenburg | City University of New York | Princeton University

In continual learning, artificial neural networks suffer from catastrophic forgetting—performance on previously learned tasks degrades significantly upon training on new tasks. Existing approaches rely on heuristic task-scheduling protocols lacking theoretical guarantees of optimality. This paper bridges statistical physics and optimal control theory to establish, for the first time, an analytically tractable and provably optimal framework for task selection dynamics. Leveraging a teacher–student model, we derive exact training dynamics via dynamic mean-field analysis and obtain a closed-form optimal scheduling protocol that explicitly incorporates task similarity as a key regulator of forgetting. Empirical evaluation on synthetic data and real-world benchmarks (e.g., CIFAR-100) demonstrates substantial reduction in forgetting rates. Crucially, theoretical predictions align closely with experimental results, validating the framework’s strong interpretability, formal optimality guarantee, and cross-dataset generalizability.

Addresses catastrophic forgetting in neural networks during sequential task learning.Develops optimal task-selection protocols using statistical physics and control theory.Validates theoretical strategies for minimizing forgetting on real-world data.

Modelling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting

Oct 21, 2024
RT
Robin Th'eriault
🏛️ Scuola Normale Superiore di Pisa | The University of British Columbia | University of Bologna

This work investigates the learning mechanism of restricted Boltzmann machines (RBMs) for structured data within a teacher–student framework, specifically addressing whether a student RBM can recover the true latent representation from data generated by a teacher RBM exhibiting correlations among hidden variables. Using statistical-physics-based mean-field analysis and temperature-regularized inference, we systematically characterize how structural strength—quantified by the number of latent patterns and inter-pattern correlation in weight rows—affects the critical sample size required for successful learning. We find that enhanced structure drastically reduces the necessary sample complexity; in the absence of correlations, performance is independent of pattern count; and excessively low inference temperatures suppress pattern acquisition, leading to learning failure. Crucially, we establish for the first time that the student can achieve exact one-to-one or one-to-many pattern matching—surpassing the conventional two-hidden-unit limitation. These results provide the first analytically tractable generative-model foundation for the “lottery ticket hypothesis.”

Analyzing impact of pattern correlations on critical data requirementsExploring temperature effects on teacher pattern learnabilityStudying RBM learning of structured data in teacher-student setting

How to set AdamW's weight decay as you scale model and dataset size

May 22, 2024
XW
Xi Wang
🏛️ University of Massachusetts Amherst | University of Bristol

Tuning weight decay (WD) in AdamW for large-scale models and datasets remains challenging due to its opaque interaction with model and data scale. Method: We introduce the exponential moving average (EMA) timescale—measured in epochs—as a physically interpretable core variable, decoupling and reformulating WD accordingly. Contribution/Results: Through theoretical analysis and extensive experiments across architectures (ResNet-18, ViT, NanoGPT) and datasets (CIFAR-10, ImageNet, OpenWebText), we find that the optimal EMA timescale is approximately invariant in epoch units. This yields an inverse scaling law: WD should increase with model size but decrease with dataset size. This law coheres with muP learning rate scaling; neglecting WD scaling while only scaling learning rate causes significant performance degradation in mid-to-late training. Our work establishes the first interpretable, principled scaling framework for AdamW weight decay, enhancing training stability and generalization in large-model regimes.

AdamW OptimizationModel ScalingWeight Decay Parameter

Latest Papers

What's happening recently
View more

This study investigates how effective learning rate drift induced by normalized updates affects training acceleration, stability, and resource scaling. Leveraging the random feature model and dynamical mean field theory (DMFT), the work elucidates the acceleration mechanisms and late-stage instabilities of normalized SGD, with theoretical predictions validated through linearized ResNet experiments. The primary contribution is the first solvable theoretical framework connecting normalization-induced acceleration, marginal stability, and width-batch allocation. Furthermore, this research quantifies the computational efficiency trade-offs between batch size and network width across distinct scaling regimes, empirically corroborating the predicted trends on CIFAR-5M.

adaptive learning rateedge of stabilitygradient normalization

This work addresses the instability and task-agnostic collapse commonly observed in self-play policy distillation, which often stem from ill-timed updates of the teacher policy. Through a systematic analysis of the temporal coupling between the teacher’s freezing interval (quarantine period) and the student’s learning dynamics, the study identifies clock-driven teacher refreshes as a primary cause of collapse. To mitigate this, the authors propose Consolidation-Gated Teacher Refresh (CGTR), an adaptive gating mechanism that triggers teacher updates only when jointly validated by improvements in reward and safe trajectory length. Requiring no task-specific hyperparameter tuning, CGTR achieves zero collapse across four diverse tasks—Chemistry, Biology, Physics, and ToolUse—while attaining state-of-the-art performance under a unified hyperparameter configuration and automatically adjusting the teacher refresh frequency per task.

isolation periodsself-distillationstate-oblivious collapse

Hot Scholars

SL

Shijian Lu

College of Computing and Data Science, NTU
Image and video analyticscomputer visionmachine learning
HL

Haodong Li

UC San Diego. Prev: HKUST, ZJU, Tencent.
3DVGenerative ModelsAgents
LC

Lijun Chen

University of Colorado at Boulder
Optimization and control of networked systemsComputer networksPower networksOptimization
YW

Yu-Wing Tai

Dartmouth College
Computer VisionDeep LearningMulti-modalities Generative AI
IM

Ian McLoughlin

Professor Singapore Institute of Technology (Singapore) and USTC (China)
AI for speech & audiosignal processingembedded systemscomputer architecture