select smoothing parameters

Designs and implements procedures to choose or estimate smoothing/scale parameters (for example penalty weights, bandwidths, interpolation smoothing constants) used in models, estimators and algorithms. This includes automatic and locally adaptive selection methods (e.g., cross‑validation, marginal likelihood, scale‑space or local Lipschitz estimation) that trade off fidelity versus smoothness and avoid instability from relying on high‑order derivative estimates.

selectsmoothingparameters

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

On Neighbourhood Cross Validation

Apr 25, 2024
SN
Simon N. Wood
🏛️ University of Edinburgh

Cross-validation (CV) for hyperparameter optimization in smoothing-penalized regression suffers from high computational cost and systematic underestimation of smoothing parameters due to unmodeled short-range autocorrelation. Method: We propose Neighborhood CV—a novel CV variant that combines QR decomposition and matrix calculus to compute CV gradients analytically and efficiently; it employs a neighborhood hold-out strategy with robust variance estimation, reducing hyperparameter optimization complexity to O(n), equivalent to a single model fit. Contribution/Results: This is the first method enabling efficient hyperparameter optimization and uncertainty quantification via neighborhood CV under short-range autocorrelation. Experiments demonstrate substantial improvements in accuracy and robustness of smoothing parameter estimation across smoothing quantile regression and GAMLSS models. Neighborhood CV establishes a scalable, theoretically grounded, and practical paradigm for hyperparameter selection in high-dimensional nonparametric modeling.

Address un-modelled short range autocorrelation in smoothing parametersEfficiently compute cross validation for penalized regression hyperparametersReduce computational cost to that of a single model fit

Optimal smoothing parameter in Eilers-Wittaker smoother

Oct 02, 2025
RB
Roberto Bernal-Arencibia
🏛️ Universidad de La Habana | Universidad de las Ciencias Informáticas (UCI)

To address the challenge of robustly selecting the regularization parameter in Eilers–Whittaker smoothing under large-scale data and serially correlated noise, this paper proposes an automatic optimization method based on residual spectral entropy. The core innovation lies in modeling the sigmoidal relationship between residual spectral entropy and the regularization parameter, with the optimal parameter identified at the absolute maximum of this S-shaped curve. Unlike cross-validation and the V-curve method, the proposed approach is insensitive to noise structure, computationally efficient, and highly robust. Leveraging spectral entropy analysis, S-curve modeling, and Euclidean distance–assisted decision-making, extensive validation on diverse synthetic and real-world time-series datasets demonstrates substantial improvements in both parameter selection stability and smoothing accuracy. The method establishes a new, interpretable, reproducible, and scalable paradigm for automated hyperparameter tuning of Eilers–Whittaker smoothers.

Addressing poor performance of cross-validation with serially correlated noiseDetermining optimal regularization parameter for Eilers-Whittaker smoother automaticallyDeveloping spectral entropy-based method for robust parameter selection

This work addresses the computational challenges of Bayesian marginal likelihood optimization in generalized smooth models, where traditional Laplace approximation relies on high-order derivatives that are both costly to compute and difficult to implement. The authors propose a quasi-Newton extended Fellner–Schall (qEFS) method that employs a structured limited-memory secant approximation to the Hessian of the log-likelihood, requiring only first-order derivatives to efficiently approximate critical subblocks. The approach further allows selective incorporation of exact Hessian columns to enhance accuracy. Balancing computational efficiency with estimation precision, qEFS demonstrates robustness under broader conditions. Empirical results show that qEFS converges to the classical EFS in simulations, substantially simplifies implementation in hidden Markov and Tweedie models, and achieves near-nominal performance in confidence interval coverage and model selection tasks.

general smooth modelsHessian matrixLaplace approximation

This work addresses the critical dependence of B-spline regression performance in generalized additive models (GAMs) on knot placement, a challenge wherein conventional approaches struggle to balance fitting accuracy and model parsimony. The authors propose an explicit, automated knot selection method that integrates the knot placement mechanism of adaptive splines (A-splines) with a tailored Fellner–Schall smoothing parameter optimization strategy, enabling efficient sparse modeling. By reintroducing explicit knot selection—an aspect often overlooked—into the GAM framework, the method achieves predictive performance comparable to P-splines and state-of-the-art alternatives while substantially reducing the number of basis functions, thereby enhancing both model interpretability and computational efficiency.

B-spline regressiongeneralized additive modelsknot selection

LASER: A new method for locally adaptive nonparametric regression

Dec 27, 2024
SC
Sabyasachi Chatterjee
🏛️ University of Illinois Urbana–Champaign | Tata Institute of Fundamental Research | Indian Statistical Institute

To address insufficient local adaptivity in nonparametric regression, this paper proposes LASER: a variable-bandwidth local polynomial regression framework. Methodologically, LASER achieves pointwise optimal local adaptation—matching the local Hölder regularity of the regression function at each domain point—under a single global tuning parameter, without requiring prespecified model structure. Theoretically, LASER attains the local minimax optimal convergence rate, providing rigorous statistical optimality guarantees. Algorithmically, it integrates local polynomial fitting, data-driven bandwidth selection, and adaptive estimation of the local Hölder exponent. Numerical experiments demonstrate that LASER significantly outperforms existing locally adaptive methods across diverse smoothness-heterogeneous settings, while maintaining computational efficiency. By unifying strong theoretical foundations with practical implementability, LASER fills a critical gap in the literature on locally adaptive nonparametric regression.

Flexible RegressionLocal Data AdaptationModel-free Approach

Latest Papers

What's happening recently
View more

This work addresses the high computational cost of traditional spline regression, which relies on grid search and cross-validation to select resolution hyperparameters. By leveraging approximation theory and ANOVA decomposition, the authors derive—for the first time—a closed-form analytical solution for the optimal resolution, thereby eliminating costly hyperparameter tuning. The proposed method, KORE, estimates bias and noise via two pilot fits and rapidly computes leave-one-out error using the PRESS identity. KORE reveals a Kolmogorov-optimal scaling law dependent solely on effective density and independent of input dimensionality, and it integrates multiple model selection criteria—including GCV, Cp, AIC, and BIC—for efficient modeling. On additive and sparse pairwise functions up to 80 dimensions, KORE achieves the accuracy of exhaustive cross-validation with only about one-eighth of the fitting effort; across 36 real-world tabular datasets, it attains the highest accuracy per unit computation among 21 competing methods.

hyperparameter tuningKolmogorov n-widthresolution selection

This study addresses the O(h²) bias bottleneck inherent in local linear derivative estimation by proposing an iterative data sharpening method. By constructing sharpened estimators based on residual operators, this approach reduces the bias order to O(h^{2l+2}) while preserving the simplicity of local fitting. Furthermore, a closed-form single-bandwidth expression is derived specifically for Gaussian kernels. Simulation results demonstrate that the proposed algorithm significantly mitigates estimation bias and clearly elucidates the bias-variance trade-off mechanism. Consequently, this work establishes a novel paradigm for nonparametric derivative estimation that effectively balances high precision with computational efficiency, offering a robust solution to longstanding limitations in local polynomial regression techniques.

Bias reductionDerivative estimationLocal polynomial smoothing

This work addresses the inadequate characterization of uncertainty in existing variance estimation methods for policy coefficients, which fail to fully exploit the asymptotic normality of the adaptive Lasso, particularly in sparse settings. The paper proposes a novel variance estimator that, for the first time, seamlessly integrates the asymptotic normality theorem of the adaptive Lasso with its variable selection consistency, explicitly modeling how the selection process influences uncertainty quantification. By doing so, the method maintains theoretical rigor while substantially improving the accuracy of variance estimation under sparsity. Empirically, it enables more reliable uncertainty visualization in clinical policy learning, thereby enhancing the safety and robustness of data-driven decision-making.

adaptive lassopolicy learningrelative sparsity

This work addresses the computational burden and non-differentiability of traditional information criteria, which rely on discrete search procedures. We propose the Smooth Information Criterion (SIC), which employs an ε-scaling continuation strategy to continuously relax and progressively sharpen the model dimensionality penalty term. This reformulates discrete variable selection as a differentiable optimization problem, enabling coefficient-level selection in generalized linear models without requiring data-driven regularization parameters. Experimental results demonstrate that SIC replicates the model selection performance of exhaustive BIC search while outperforming stepwise regression and LASSO. Furthermore, it substantially reduces computational overhead in high-dimensional settings.

BICGeneralised Linear ModelsInformation Criterion

Hot Scholars

GA

Gilberto A. Paula

Professor of Statistics, Universidade de São Paulo
Modelos de RegressãoRegression ModelsGeneralized Linear Models
AB

Anders Bjorholm Dahl

Professor, Image Analysis, Technical University of Denmark
Image analysis
SP

Snigdha Panigrahi

University of Michigan
Selective InferenceCausal InferenceRandomizationMachine Learning
RW

Rebecca Willett

University of Chicago, Professor of Statistics and Computer Science
Machine learningData scienceSignal processingInformation theory
EC

Erjia Cui

Division of Biostatistics and Health Data Science, University of Minnesota
Biostatistics