Score
Designs and implements loss terms, filtering modules, and training constraints that enforce smooth, coherent behavior over time by penalizing rapid changes in model outputs or latent trajectories (e.g., finite-difference penalties, temporal convolutional smoothing, or frame-to-frame consistency losses). These regularizers and filters are applied during model training or post-processing to suppress temporal artifacts and stabilize sequence or time-series outputs and their latent representations.
Traditional regularization methods require auxiliary penalty terms and struggle to jointly optimize accuracy, smoothness, and robustness. To address this, we propose Lai loss—a novel paradigm that intrinsically embeds gradient regularization into the geometric structure of the loss function. Unlike conventional approaches, Lai loss explicitly constrains both the magnitude and direction of input gradients in an end-to-end differentiable manner, directly modulating model sensitivity without introducing separate regularization terms. We further design a gradient self-constraining training mechanism to ensure optimization stability and convergence. Extensive experiments on Kaggle multi-task benchmarks demonstrate that Lai loss maintains predictive accuracy while significantly improving model smoothness and noise robustness—particularly enhancing invariance to perturbed features. This work provides a generalization-driven framework for loss function design, advancing the integration of robustness and smoothness as inherent geometric properties of the loss landscape.
This study addresses the problem of non-smooth action oscillations that arise when deploying deep reinforcement learning policies on physical robots. To this end, we propose CATS, a method that introduces temporal regularization to constrain discrepancies between consecutive actions. We theoretically demonstrate that temporal penalties effectively bound spatial smoothness. Furthermore, a linearly increasing scheduling mechanism is designed to optimize the training process, balancing policy exploration with convergence stability. Experimental results indicate that CATS significantly suppresses action oscillations in both simulated and real-world environments while maintaining high task returns with minimal computational overhead. This work thus provides an efficient solution for the safe deployment of reinforcement learning on physical robotic systems.
Existing methods often rely on soft penalties to approximate sample-level constraints, which struggle to strictly enforce hard requirements. This work proposes the first sample-wise constrained learning framework based on the sequential penalty method, enabling strict satisfaction of per-sample constraints within deep learning while providing convergence guarantees. By systematically integrating sequential penalty mechanisms into end-to-end training, the approach balances theoretical rigor with practical feasibility. Experiments on image processing tasks demonstrate that the proposed framework not only ensures strict adherence to constraints but also maintains competitive model performance.
This work addresses the lack of adaptive regularization in neural network loss functions by proposing a meta-learning framework for loss function optimization, termed TaylorGLO. Methodologically, it integrates Taylor-expansion-driven meta-optimization, learning rule decomposition, and dynamical systems analysis. Theoretically, it establishes for the first time that this paradigm intrinsically induces a phase-wise regularization mechanism: suppressing parameter oscillations in early training, preserving gradient flow dynamical invariance during mid-training to accelerate meta-convergence, and tightening generalization bounds in late training. Experiments demonstrate significant improvements in model generalization, training speed, few-shot data efficiency, and adversarial robustness. This work introduces the first theoretically grounded paradigm for adaptive loss-function regularization in meta-learning, providing formal guarantees on both regularization behavior and meta-optimization dynamics.
This paper addresses the lack of a unified early-stopping mechanism for implicit regularization in iterative learning. To bridge this gap, the authors propose a theory-driven early-stopping framework. They develop EarlyStopping, an open-source Python toolkit that—uniquely—systematically integrates truncated SVD, Landweber iteration, conjugate gradient, L2-boosting, and regression trees. The toolkit supports user-defined data generation and enables real-time monitoring of theoretical regularization strength, including effective degrees of freedom and bias–variance trade-offs. Implemented in NumPy/SciPy, it provides sequential risk estimation and analytically derived stopping boundaries. Experiments reproduce key theoretical results on implicit regularization, demonstrating that principled early stopping effectively suppresses noise propagation, constrains generalization error growth, and significantly enhances algorithmic robustness and interpretability—thereby narrowing the gap between theoretical analysis and practical deployment.
This work addresses the discrepancy in diffusion models between the denoising score matching objective and the Fokker-Planck (FP) equation that governs the true data evolution. While existing strong FP regularization methods are computationally expensive and do not consistently improve generation quality, this study systematically evaluates lightweight regularization terms targeting FP residuals. The authors propose a low-overhead weak FP regularization strategy that substantially reduces computational burden while preserving high-fidelity image generation. Empirical results demonstrate that such lightweight regularization effectively balances efficiency and performance, confirming its feasibility and advantages in practical diffusion model training.
Implicit regularization in deep learning—such as the bias introduced by early stopping or Dropout—is notoriously difficult to characterize analytically, particularly under complex training strategies where general estimation methods are lacking. This work proposes an empirical framework based on gradient matching that quantifies the discrepancy between weight updates and loss gradients, enabling estimation of implicit regularization in arbitrary deep networks without requiring analytical derivations. The approach provides the first scalable and general-purpose analysis of biases induced by sophisticated training mechanisms. It not only recovers the classical result that early stopping is equivalent to ℓ² regularization but also reveals that Dropout similarly introduces a significant implicit ℓ² regularization effect in deep networks.
This study addresses the limited generalization performance of standard norm-based regularization in neural networks when dealing with high-dimensional or feature-correlated settings, where conventional methods inadequately control model complexity. To overcome this, the authors propose two covariance-aware adaptive regularization techniques: first, incorporating the input feature covariance structure into ℓ₂ weight decay to refine ridge-type penalties; second, combining ℓ₁ sparsity with covariance-informed ℓ₂ regularization to achieve structured sparsity. By innovatively embedding feature covariance information into classical Lasso and ridge regression frameworks, the approach enables more precise complexity control. Extensive experiments on Monte Carlo simulations and real-world datasets—including building cooling load prediction and leukemia cell classification—demonstrate that the proposed methods significantly outperform traditional regularization strategies and substantially enhance model generalization.
This work addresses the degradation in generation quality of diffusion models under few-step sampling, which stems from insufficient cross-temporal consistency in the denoising trajectory. The authors propose the first formulation of the diffusion process as a Markov reward process, reframing denoising as a policy evaluation problem in reinforcement learning. To enforce consistency across multi-step denoising paths, they introduce a temporal difference (TD) objective that penalizes trajectory inconsistencies and incorporate a sample reweighting strategy to stabilize training. The resulting method is universally applicable to both discrete- and continuous-time diffusion models and achieves substantial improvements in generation quality—measured by FID—under limited sampling steps, demonstrating particularly pronounced advantages in low-compute regimes.
本文提出了一种模型无关的高维时间序列降噪框架,通过估计低维动态子空间和最优投影来恢复被观测白噪声污染的低维潜在动态成分。