kl annealing schedule

Designs and implements parameter schedules that gradually change regularization, divergence, truncation, or temperature-like terms during training or inference — e.g., forward-to-reverse KL transitions, competence-weighted KL schedules, progressive or shrinking truncation times, reverse- and simulated-annealing variants, and related pause/time choices — and provides tuning procedures to select schedule shape, anneal time, pauses, and truncation horizons. Analyzes convergence and freeze-out locations, quantifies effects on compounding errors and optimization dynamics, and tunes schedules to balance exploration–exploitation and acquisition trade-offs such as uncertainty versus novelty/diversity under annotation or sampling budgets.

klannealingschedule

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Schedulers for Schedule-free: Theoretically inspired hyperparameters

Nov 11, 2025
YP
Yuen-Man Pun
🏛️ Australian National University | University of British Columbia | Flatiron Institute

This work addresses theoretical limitations of schedule-free optimization methods by proposing a unified analytical framework compatible with arbitrary learning rate schedules. Methodologically, it extends classical last-iterate convergence theory—previously restricted to fixed or decaying learning rates—to general schedules (e.g., warm-up–stable–decay), introduces a dynamic averaging mechanism for model parameters that rigorously adapts to time-varying learning rates, and designs an adaptive Polyak step-size rule achieving optimal anytime convergence rate $O(1/sqrt{T})$. The theoretical analysis is rigorously established under convexity assumptions. Empirically, the method significantly outperforms SGD, Adam, and existing schedule-free baselines on black-box model distillation tasks, demonstrating strong predictive power of the theory for practical training performance.

Developing adaptive Polyak learning rate for optimal anytime convergenceExtending schedule-free theory to support arbitrary learning rate schedulersValidating theoretical convergence predictions through deep network experiments

This work addresses the limitations of conventional learning rate warmup strategies, which rely on heuristic hyperparameter tuning and lack theoretical grounding—particularly exhibiting instability under norm-constrained optimizers such as Muon and Lion. Building upon a generalized smoothness assumption that links local curvature to the suboptimality gap, the paper derives, for the first time, a learning rate schedule that naturally integrates both warmup and decay directly from convergence analysis. The resulting method is fully adaptive, requiring no additional hyperparameters and automatically adjusting warmup duration. Evaluated on LLaMA large language model pretraining, it consistently matches or surpasses the performance of manually tuned baselines across all experimental settings, significantly enhancing training efficiency and robustness.

adaptive schedulinglarge language modelslearning rate

Learning rate scheduling in large language model training lacks rigorous theoretical foundations, leading to heuristic designs and suboptimal convergence. Method: This paper establishes, for the first time, a quantitative alignment between practical schedulers (e.g., linear decay) and tight non-smooth convex optimization lower bounds—eliminating spurious logarithmic factors in prior analyses and enabling principled cross-scheduler optimal learning rate transfer. We integrate convex optimization theory, scheduler modeling, and empirical validation, conducting systematic evaluations on 124M- and 210M-parameter Llama models. Results: Theory-guided scheduler design yields faster convergence and improved stability, empirically validating optimization theory’s practical relevance for large-model training. Core contribution: bridging the gap between theoretical performance bounds and engineering schedulers by providing a transferable, interpretable, and theoretically grounded framework for learning rate tuning.

Large Model TrainingLearning Rate AdjustmentTraining Efficiency

HyperbolicLR: Epoch insensitive learning rate scheduler

Jul 21, 2024
TK
Tae-Geun Kim
🏛️ Yonsei University

Traditional learning rate schedulers suffer from sensitivity to training duration (number of epochs), poor cross-task transferability, and insufficient robustness under resource constraints. To address these issues, this work proposes two novel analytical schedulers—HyperbolicLR and ExpHyperbolicLR—that leverage the asymptotic properties of hyperbolic functions to construct epoch-agnostic scheduling mechanisms, enabling seamless hyperparameter reuse across varying training lengths. We further introduce a two-phase strategy: rapid hyperparameter tuning in early epochs followed by fixed-rate scheduling in later epochs—balancing convergence speed and optimization stability. Extensive experiments on image classification, time-series forecasting, and operator learning demonstrate that our methods significantly outperform baselines—including StepLR and CosineAnnealing—in long-horizon training, achieving superior performance stability and generalization. Results confirm enhanced efficiency and robustness under computational constraints.

Deep Learning OptimizationLearning Rate SchedulingResource-limited Learning

Latest Papers

What's happening recently
View more

This work addresses the challenge of entropy coefficient selection in real-world reinforcement learning, where environmental non-stationarity often leads to insufficient or excessive exploration under a fixed entropy weight, compromising both efficiency during stable periods and adaptability following abrupt changes. To this end, the paper introduces Adaptive Entropy Scheduling (AES), the first method to frame entropy scheduling under non-stationary environments as a one-dimensional online trade-off problem. AES dynamically adjusts the temperature coefficient using a lightweight online proxy metric for distributional drift, without altering the underlying algorithm architecture. Evaluated across four algorithms, twelve tasks, and four distinct drift patterns, AES consistently reduces performance degradation and accelerates policy recovery after environmental shifts.

entropy schedulingenvironment driftexploration-exploitation trade-off

This work addresses the lack of theoretical grounding in noise scheduling for diffusion models, which has hindered the understanding of empirically effective scheduling strategies. The authors formulate the problem for the first time as an optimal control problem, where the Fisher information serves as the state variable and the noise schedule acts as the control input, with the objective of minimizing an upper bound on the KL sampling error. Within this framework, they derive sufficient conditions for achieving near-optimal sampling error and obtain a tunable closed-form expression for the noise schedule. The proposed method unifies and generalizes exponential and sigmoidal schedules, and, after parameter tuning, achieves improved Fréchet Inception Distance (FID) scores on standard image generation benchmarks.

diffusion modelsFisher informationnoise schedule

This work addresses the optimization of noise scheduling and time discretization in diffusion models under a limited number of sampling steps to minimize the discrepancy between the generated and target distributions. By constructing a simplified diffusion model with a Gaussian source distribution, the authors derive a closed-form solution for the reverse process and analyze discretization error using KL divergence and the Euler–Maclaurin expansion. Leveraging variational calculus, they propose a “tangent law” for noise scheduling, whose parameters are analytically determined by the eigen-spectrum of the source covariance matrix. This approach provides an optimal time discretization criterion for pre-trained models without requiring retraining and consistently outperforms existing baselines across multiple datasets and architectures, particularly excelling under extremely low sampling budgets.

generative diffusion modelsKL divergencenoise scheduling

This work addresses the prevailing reliance on heuristic-based learning rate scheduling strategies, which lack a systematic understanding of optimal schedule shapes. The authors propose a method that decouples the base learning rate from the schedule shape, enabling automatic discovery of near-optimal learning rate schedules within a parameterized family tailored to specific tasks. Experiments across image classification, language modeling, and linear regression reveal that warmup followed by decay constitutes a robust characteristic of high-performing schedules, that commonly used schedule families are often suboptimal, and that weight decay significantly influences the optimal schedule shape. This study provides the first systematic characterization of universal properties underlying near-optimal learning rate schedules and establishes a clear connection between schedule morphology and optimization hyperparameters.

hyperparameterslearning rate scheduleneural network training

Dynamic Learning Rate Scheduling based on Loss Changes Leads to Faster Convergence

Dec 16, 2025
SS
Shreyas Subramanian
🏛️ Amazon Web Services

Existing learning rate schedulers (e.g., cosine decay) rely on predefined annealing curves and lack real-time responsiveness to training dynamics. This work proposes GreedyLR—a gradient-sign- and magnitude-aware greedy adaptive scheduler that requires no hyperparameters, incurs zero additional computational overhead, and integrates seamlessly with mainstream deep learning frameworks. Its core contributions are threefold: (1) a novel loss-difference-driven dynamic step-size scaling mechanism; (2) theoretical convergence guarantees, along with derivation of the optimal scaling factor that maximizes convergence rate; and (3) built-in robustness to gradient noise. Extensive experiments across NLP, CV, and 7B-scale model pretraining and fine-tuning demonstrate that GreedyLR achieves an average 1.8× speedup in convergence and consistently surpasses state-of-the-art schedulers—including cosine annealing and linear decay—in final accuracy.

Demonstrates improved accuracy, speed, and convergence over existing schedulersProposes GreedyLR scheduler for adaptive learning rate adjustmentValidates scheduler on NLP, CV, and LLM tasks up to 7B parameters

Hot Scholars

MO

Masayuki Ohzeki

Graduate School of Information Sciences, Tohoku University
Statistical MechanicsMachine LearningSpin GlassPhase transition
LW

Liangliang Wang

Simon Fraser University
Bayesian statisticsMonte Carlo methodsComputational biologyFunctional Data Analysis
FI

Francesca Iacopi

Fellow @Imec USA and Adj Professor @University of Technology Sydney and @Purdue University
Electronic Materials2D MaterialsMetamaterialsNanoelectronics
MH

Mohamed Hibat-Allah

Assistant Professor, University of Waterloo
Quantum PhysicsStatistical PhysicsMachine Learning