Score
Designs and implements optimization procedures that treat hyperparameter selection as an outer (validation) problem whose objective depends on model parameters produced by an inner training optimization, including deriving or approximating the validation hypergradients via implicit differentiation, truncated/unrolled optimization, or adjoint methods. Builds and analyzes algorithms to jointly update hyperparameters and model weights, and to evaluate trade-offs in generalization, computational cost, and memory when tuning continuous hyperparameters (e.g., regularization) using validation-based gradients.
This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.
本文针对流数据中超参数优化问题,提出四种边界约束处理策略,通过实验验证其优于现有方法。
本文提出了一种通过自动微分优化拓扑结构和超参数的方法,解决了拓扑优化中超参数调优的问题,且该方法可扩展至数千个超参数。
Cross-validation (CV) for hyperparameter optimization in smoothing-penalized regression suffers from high computational cost and systematic underestimation of smoothing parameters due to unmodeled short-range autocorrelation. Method: We propose Neighborhood CV—a novel CV variant that combines QR decomposition and matrix calculus to compute CV gradients analytically and efficiently; it employs a neighborhood hold-out strategy with robust variance estimation, reducing hyperparameter optimization complexity to O(n), equivalent to a single model fit. Contribution/Results: This is the first method enabling efficient hyperparameter optimization and uncertainty quantification via neighborhood CV under short-range autocorrelation. Experiments demonstrate substantial improvements in accuracy and robustness of smoothing parameter estimation across smoothing quantile regression and GAMLSS models. Neighborhood CV establishes a scalable, theoretically grounded, and practical paradigm for hyperparameter selection in high-dimensional nonparametric modeling.
This work investigates the theoretical hardness and algorithmic efficiency of finding stationary points of the hyperobjective in nonconvex–convex and nonconvex–nonconvex bilevel optimization, under the Polyak–Łojasiewicz (PL) condition—rather than strong convexity—on the lower-level objective. We first establish an impossibility result: for zero-respecting algorithms, computing a hyperstationary point is fundamentally intractable in the nonconvex–convex setting. Under the PL condition, we break the reliance on strong convexity and derive tighter hypergradient convergence complexity bounds. We propose a novel analytical framework unifying implicit function differentiation, hypergradient estimation, and first-order optimization. This yields complexity guarantees of $ ilde{mathcal{O}}(varepsilon^{-2})$, $ ilde{mathcal{O}}(varepsilon^{-4})$, and $ ilde{mathcal{O}}(varepsilon^{-6})$ for deterministic, partially stochastic, and fully stochastic settings, respectively—substantially improving upon existing nonconvex bilevel optimization methods.
This work addresses the lack of statistical reliability guarantees in existing hyperparameter selection methods—such as grid search and Bayesian optimization—with respect to critical metrics like risk and safety. Building upon the learn-then-test (LTT) paradigm, the paper introduces a unified statistical framework that formulates hyperparameter selection as a multiple hypothesis testing problem, accommodating user-specified constraints on average risk, quantile risk, or information-theoretic measures. Leveraging tools from statistical inference—including p-values, e-values, and concentration inequalities—the method derives explicit, finite-sample bounds on error probabilities from first principles. This approach enables theoretically grounded validation and selection of hyperparameters, substantially enhancing the reliability and safety of AI systems in real-world deployment scenarios.
This work addresses a critical limitation in existing gradient-based hyperparameter optimization methods, which often neglect the impact of variance in hypergradient estimation, leading to overfitting on the validation set. For the first time, the authors present a complete bias-variance decomposition of hypergradient estimation error, systematically revealing the pivotal role of variance in generalization performance. Building on this insight, they propose an ensemble hypergradient strategy that effectively reduces estimation variance by integrating bilevel optimization with ensemble learning. The method significantly improves hypergradient quality across diverse tasks—including regularized learning, data cleaning, and few-shot learning—thereby mitigating overfitting and enhancing model generalization.
This work addresses the lack of provable generalization guarantees for multidimensional hyperparameter tuning in complex, non-smooth spaces. It proposes the first data-driven framework for such settings, establishing generalization bounds under mild assumptions by integrating structured loss with validation loss. Leveraging tools from real algebraic geometry, the analysis characterizes the complexity of semi-algebraic function classes, yielding tighter and more broadly applicable generalization bounds. The framework’s effectiveness and learnability are demonstrated on models such as weighted group Lasso and weighted fused Lasso, offering both theoretical foundations and practical methodologies for multidimensional hyperparameter optimization.
Bayesian hyperparameter optimization (HPO) in large-scale machine learning suffers from high computational cost and difficulty in posterior sampling. To address this, we propose a generative hyperparameter tuning framework: first, we employ weighted Bayesian bootstrapping to efficiently approximate the hyperparameter posterior distribution; second, we learn a transport mapping from hyperparameters to optimizers, yielding a lookup-table-based generative estimator. This work is the first to introduce generative modeling into Bayesian HPO, enabling amortized optimization over the hyperparameter space. The method supports millisecond-scale evaluation and uncertainty quantification for both continuous and discrete hyperparameter grids. Experiments demonstrate a 10–100× speedup in search time over conventional Bayesian optimization, while achieving superior generalization performance and well-calibrated uncertainty estimates across multiple benchmarks.
This work addresses the lack of interpretability in existing methods regarding the influence of hyperparameters in multi-objective optimization. The authors propose a novel game-theoretic framework that, for the first time, integrates Shapley effects with the Pareto front to enable objective-aware global sensitivity analysis, thereby uncovering key hyperparameters and their interactions under different optimization objectives. This approach not only identifies efficient hyperparameter configurations and substantially reduces the search space but also facilitates early-stage model performance estimation. The effectiveness and generalizability of the framework are empirically validated across three distinct neural network architectures and tasks.