Score
Designs and analyzes training objectives and procedures that mix discriminative and generative loss terms to create tunable interpolations between classification-style and likelihood-style objectives. Engineers loss weightings, hybrid training schedules, or regularizers and evaluates their effects on model behavior (e.g., categorical accuracy versus input-structure sensitivity) to optimize tradeoffs such as human alignment, robustness, or benchmark performance.
This work systematically investigates the impact of loss weighting strategies and output parameterizations on model performance in flow matching. Through numerical experiments on both synthetic data with controllable geometric structures and real-world images, the study disentangles their interaction effects across varying data manifold dimensions, model architectures, and dataset scales, using PSNR and FID as evaluation metrics. The analysis reveals, for the first time, how the optimal choice of loss weighting and parameterization depends critically on the intrinsic structure of the data. Building on these insights, the authors formulate practical design principles that substantially improve denoising accuracy and generation quality.
This work addresses the fundamental trade-off between standard accuracy and adversarial robustness in supervised learning. Methodologically, it introduces the first architecture-level accuracy–robustness trade-off curve, quantifying the inverse relationship between these objectives across diverse neural network architectures; defines a sensitivity influence function to theoretically characterize the stability of optimal solutions under adversarial perturbations; and reveals—via theoretical analysis of overparameterized linear models—that adversarial training implicitly regularizes model dynamics, interpolating between L₁ (LASSO) and L₂ (ridge regression) behaviors. The approach integrates rigorous theoretical analysis, influence-function-based modeling, and extensive empirical evaluation across fully connected, deep, and varying-width networks. Results consistently validate the existence and structure of the trade-off, providing an interpretable, predictive theoretical foundation for principled neural architecture selection.
This work addresses the lack of adaptive regularization in neural network loss functions by proposing a meta-learning framework for loss function optimization, termed TaylorGLO. Methodologically, it integrates Taylor-expansion-driven meta-optimization, learning rule decomposition, and dynamical systems analysis. Theoretically, it establishes for the first time that this paradigm intrinsically induces a phase-wise regularization mechanism: suppressing parameter oscillations in early training, preserving gradient flow dynamical invariance during mid-training to accelerate meta-convergence, and tightening generalization bounds in late training. Experiments demonstrate significant improvements in model generalization, training speed, few-shot data efficiency, and adversarial robustness. This work introduces the first theoretically grounded paradigm for adaptive loss-function regularization in meta-learning, providing formal guarantees on both regularization behavior and meta-optimization dynamics.
To bridge the gap between data-hungry deep models and human-like few-shot learning efficiency, this paper investigates the long-overlooked loss function component within meta-learning frameworks and proposes a novel dynamic adaptive loss learning paradigm. Methodologically, it introduces (1) EvoMAL—an interpretable symbolic loss evolution method that integrates symbolic regression with evolutionary algorithms to generate lightweight, task-adaptive, and interpretable loss functions; (2) Sparse Label Smoothing Regularization (SparseLSR), a new regularization technique for mitigating label noise in low-data regimes; and (3) NPBML—a unified framework jointly optimizing meta-initialization, meta-optimizer, and loss function. Experiments across multiple few-shot benchmarks demonstrate state-of-the-art performance: classification accuracy improves significantly, while memory overhead for loss learning is reduced by over 80%.
This work addresses the challenges of jointly optimizing numerous loss terms and managing high memory and computational overhead in multi-objective deep learning. Methodologically, we propose a hierarchical output-feedback control framework that eliminates explicit Lagrange multipliers by introducing time-varying multipliers, dynamically reshaping the loss landscape at the epoch level. We further introduce a novel hypervolume-based likelihood probabilistic graphical model that jointly captures the co-evolution of model parameters and multipliers, decomposing multi-objective optimization into a sequence of Pareto-adaptive constrained hierarchical optimal control subproblems. Evaluated on the PACS domain generalization benchmark—featuring a six-loss-term variational autoencoder—we demonstrate that our approach significantly outperforms existing multiplier-scheduling methods in both accuracy and robustness, while substantially reducing memory footprint and computational cost. Moreover, the framework supports modular extension for diverse multi-objective architectures.
This work addresses the reliance on empirical design and poor generalizability of hand-crafted optimizers in gradient-based learning. Methodologically, it formulates the optimizer as a learnable functional mapping from gradients to parameter updates, and—novelty—systematically recasts this as a sequence of analytically solvable convex optimization problems. This unified framework yields closed-form derivations of mainstream optimizers (e.g., SGD, Adam) along with their theoretically optimal hyperparameters. Furthermore, it incorporates a runtime gradient statistics mechanism enabling dynamic, adaptive tuning during training. Experiments demonstrate substantial improvements in convergence speed and training stability, while preserving theoretical rigor and practical deployability.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This work investigates the implicit regularization mechanism induced by the Deep Linear Discriminant Analysis (Deep LDA) objective during optimization, addressing a theoretical gap in understanding implicit biases in metric learning. By analyzing the gradient flow of the Deep LDA loss over an L-layer diagonal linear network, we uncover how, under balanced initialization, additive gradient updates are effectively transformed into multiplicative weight updates. We theoretically establish that this process inherently preserves a (2/L)-quasinorm conservation law, thereby forging the first explicit link between Deep LDA’s implicit regularization, network architecture, and the underlying optimization geometry. This insight offers a novel perspective on the generalization behavior of metric learning objectives.
This work investigates how fine-tuning large language models—even with benign data—can inadvertently degrade alignment and adversarial robustness, with the influence of fine-tuning objectives remaining unclear. Under controlled conditions fixing data, domain, architecture, and optimization settings, the study systematically compares six fine-tuning objectives: supervised fine-tuning (SFT), direct preference optimization (DPO), conditional fine-tuning, inoculation prompting, odds ratio preference optimization (ORPO), and KL regularization. Results show that while all methods perform similarly under limited training scales, ORPO and KL regularization substantially enhance adversarial robustness and mitigate role drift at larger scales, highlighting the critical role of constrained optimization objectives in balancing safety and capability.
This work addresses the challenge in hybrid modeling where the high flexibility of machine learning components often marginalizes scientific models, thereby compromising interpretability and physical consistency. To mitigate this issue, the study introduces Sharpness-Aware Minimization (SAM) into the hybrid modeling framework for the first time. By optimizing the flatness of minima in the loss landscape, SAM implicitly enforces a simplicity bias without requiring explicit regularization terms tailored to specific model architectures or domain knowledge. Experimental results across diverse model architectures and datasets demonstrate that this approach significantly enhances the robustness and interpretability of hybrid models while promoting more effective integration of scientific models into the learning process.