Score
Design, implement, and analyze objective functions and their constituent terms—including robust, noise- and imbalance-aware penalties; auxiliary, surrogate, hybrid, and normalization terms; geometry-, topology-, and spectral-aware losses; permutation-invariant and photometric terms; and specialized regression, reconstruction and ranking losses—so they impose desired invariances, sensitivities, and regularization on model outputs. Evaluate and shape their optimization- and gradient-level behavior (e.g., stability, surrogate tightness, activation distributions, and class-difficulty weighting), and construct loss combinations and normalization schemes that preserve training dynamics and intended performance trade-offs.
This paper addresses the lack of systematic guidance for loss function selection and design in deep learning. We propose the first cross-task loss taxonomy—covering vision, time-series, and tabular data—and unifying discriminative and generative paradigms. Through comprehensive survey analysis, rigorous mathematical modeling, and multi-scenario empirical evaluation, we characterize the applicability boundaries and failure modes of 12 mainstream losses, identifying three fundamental challenges: low computational efficiency, gradient instability, and poor adaptability to real-world constraints. Building on these insights, we formulate next-generation loss design principles centered on robustness, interpretability, and adaptivity. Furthermore, we deliver a practical, industry-deployment-oriented loss selection guide—grounded in empirical evidence and operational feasibility—to bridge the gap between theoretical design and practical application.
This paper addresses the interpretability deficit and inflexibility of scalarization in multi-objective optimization (MOO) and hyperparameter optimization (HPO). To this end, it proposes a unified optimization framework grounded in utility theory. Methodologically, it presents the first systematic Python implementation of Kuhn’s utility theory, integrated into the SPOT platform; it supports direct optimization, surrogate-assisted sequential optimization (e.g., via Gaussian processes), and ML hyperparameter search—all enabled by configurable, interpretable utility modeling for principled scalarization. Key contributions include: (1) an open-source, production-ready package—spotdesirability; (2) empirical validation across three representative scenarios, demonstrating significant improvements in optimization efficiency, robustness, and decision transparency; and (3) the first scalable, reproducible, utility-theoretic unification of MOO and HPO.
This work addresses the lack of adaptive regularization in neural network loss functions by proposing a meta-learning framework for loss function optimization, termed TaylorGLO. Methodologically, it integrates Taylor-expansion-driven meta-optimization, learning rule decomposition, and dynamical systems analysis. Theoretically, it establishes for the first time that this paradigm intrinsically induces a phase-wise regularization mechanism: suppressing parameter oscillations in early training, preserving gradient flow dynamical invariance during mid-training to accelerate meta-convergence, and tightening generalization bounds in late training. Experiments demonstrate significant improvements in model generalization, training speed, few-shot data efficiency, and adversarial robustness. This work introduces the first theoretically grounded paradigm for adaptive loss-function regularization in meta-learning, providing formal guarantees on both regularization behavior and meta-optimization dynamics.
In offline optimization, surrogate models suffer from poor calibration in out-of-distribution regions; existing conditional methods exhibit weak generalization and strong model dependency. This paper proposes a model-agnostic gradient norm regularization that explicitly constrains the local sharpness of surrogate models during training. We are the first to extend sharpness-based generalization theory—from prediction loss to the gradient level—establishing a theoretical bound linking training-set gradient sharpness to worst-case gradient sharpness on unseen data. The proposed regularization is architecture-agnostic and seamlessly integrates into arbitrary surrogate models (e.g., Gaussian processes, neural networks) without structural modification. Empirical evaluation on multi-objective black-box optimization tasks demonstrates an average performance improvement of 9.6%, with significant gains in generalization and robustness. The implementation is publicly available.
This paper addresses the challenge of jointly compressing high-dimensional input and output spaces in goal-oriented applications—such as sensor placement and sensitivity analysis—where conventional dimensionality reduction methods treat inputs and outputs independently. We propose an input–output co-dimensional reduction framework that jointly optimizes coupled input and output subspaces. Crucially, it reformulates the NP-hard combinatorial selection problem into a differentiable optimization over diagonal entries of a diagnostic matrix, obviating costly evaluations of the objective function. By integrating gradient-based upper-bound optimization, expected information gain, Sobol’ sensitivity indices, and spectral analysis of the diagnostic matrix, our method achieves substantial computational efficiency gains. Experiments demonstrate its effectiveness and scalability in sensor layout optimization and parameter importance ranking. The approach establishes a novel paradigm for high-dimensional, goal-oriented experimental design and global sensitivity analysis.
Offline optimization of expensive black-box functions in materials engineering suffers from poor robustness due to the high sensitivity of surrogate models to parameter perturbations. Method: We propose, for the first time, an optimizable surrogate sensitivity metric and design a sensitivity-aware regularization method orthogonal to existing frameworks. This approach integrates gradient-based sensitivity analysis with deep-learning-based surrogate modeling and is compatible with mainstream paradigms such as offline Bayesian optimization. Contribution/Results: Evaluated on multiple materials design benchmarks, our method significantly improves optimization success rate (average gain of +23.6%) and solution quality (objective value improvement up to 17.4%). Empirical results demonstrate that explicit sensitivity control delivers critical performance gains for offline optimization of expensive black-box functions in materials engineering.
This work addresses the inefficiency of global optimization when standard neural network surrogates are embedded into mixed-integer linear programs (MILPs), a challenge stemming from the lack of control over their structural properties. The authors propose a novel differentiable regularizer that, for the first time, approximates the full gradient of the LP relaxation gap with respect to network parameters, enabling direct optimization of key structural attributes such as big-M constants, the number of unstable neurons, and the LP relaxation gap itself. Built upon ReLU networks and MILP formulations, the method leverages gradients from LP dual variables and requires no custom automatic differentiation. Experiments demonstrate up to four orders of magnitude reduction in MILP solve time on nonconvex benchmark functions and two-stage stochastic programming problems, all while preserving predictive accuracy.
This study addresses the impact of objective scale disparity on the definition and approximation of regions of interest (ROIs) in preference-driven evolutionary multi-objective optimization. It systematically investigates whether ROIs should be defined in the normalized or original objective space, conducting comparative experiments using an evolutionary algorithm that incorporates estimates of both ideal and extreme points. The work reveals, for the first time, the fundamental reason why ROIs defined in normalized space are inherently difficult to approximate accurately. It demonstrates that defining ROIs in the original objective space yields significantly better approximation quality, particularly when objective scales are heterogeneous. These findings provide a theoretical foundation and practical guidance for selecting the appropriate objective space in preference-guided optimization frameworks.
This work addresses the lack of provable generalization guarantees for multidimensional hyperparameter tuning in complex, non-smooth spaces. It proposes the first data-driven framework for such settings, establishing generalization bounds under mild assumptions by integrating structured loss with validation loss. Leveraging tools from real algebraic geometry, the analysis characterizes the complexity of semi-algebraic function classes, yielding tighter and more broadly applicable generalization bounds. The framework’s effectiveness and learnability are demonstrated on models such as weighted group Lasso and weighted fused Lasso, offering both theoretical foundations and practical methodologies for multidimensional hyperparameter optimization.
This work proposes a unified functional analytic framework that interprets both supervised and unsupervised learning as variational optimization problems within a function space induced by the data distribution. The central insight is that the fundamental distinction between these learning paradigms arises from the choice of the functional being optimized, rather than from differences in the underlying function space itself. Data structure is characterized via operators induced by the distribution, and target functions are estimated in the eigenbasis of these operators. This framework systematically integrates classical algorithms—including kernel methods, spectral clustering, and manifold learning—revealing their intrinsic coherence and underscoring the foundational role of function spaces and associated operators in modern machine learning.