Score
Procedures for choosing smoothing parameters (bandwidths) in kernel-based estimators to trade bias and variance and capture local structure. Used to reduce variance in inverse-propensity estimators, estimate spatially varying intensities, and enable semiparametric estimation with Nadaraya–Watson kernels.
Bandwidth selection in kernel density estimation is often hindered by either insufficient flexibility of nonparametric methods or overly restrictive normality assumptions. This work proposes a semiparametric bandwidth selection framework that builds upon the normal reference rule and incorporates a small number of data-driven correction coefficients via Hermite series expansion, enabling flexible adjustment of the bandwidth. The approach preserves the structural benefits of the normal prior while significantly improving adaptability to the actual data distribution. Theoretical analysis demonstrates that the proposed correction form enjoys favorable asymptotic properties, thereby providing a solid foundation for subsequent algorithmic implementation and numerical validation.
This study addresses the challenge of enhancing density estimation performance by integrating the structural advantages of parametric models while preserving the flexibility of nonparametric methods. The authors propose a locally parametrized nonparametric density estimator that, for each point \(x\), estimates an optimal local parameter \(\hat{\theta}(x)\) via kernel-smoothed likelihood, yielding an estimator of the form \(f(x, \hat{\theta}(x))\). This approach achieves near-full-likelihood efficiency under correct model specification and retains nonparametric robustness under misspecification, effectively serving as a semiparametric realization of higher-order kernel methods. Theoretical analysis and empirical experiments demonstrate that the proposed estimator exhibits variance comparable to classical kernel density estimation but substantially reduced bias, leading to significantly improved accuracy in neighborhoods of well-specified parametric models.
This work addresses the long-standing challenge of efficient and stable bandwidth selection in Beta kernel density estimation, where existing iterative optimization methods are computationally expensive and prone to instability. The authors propose the first closed-form bandwidth selector, derived by approximating the unweighted asymptotic mean integrated squared error (AMISE) via the method of moments, reducing bandwidth computation to constant time complexity. To handle integrability issues near boundaries—particularly for U-shaped and J-shaped distributions—they introduce tailored boundary-aware heuristics. In Monte Carlo simulations, the proposed method achieves accuracy comparable to numerical optimization while accelerating computation by over 35,000-fold. Applied to real-world socioeconomic data, it effectively mitigates the boundary bias commonly observed with Gaussian kernels, substantially improving both estimation efficiency and stability.
This paper challenges the conventional wisdom that kernel choice is inconsequential in local polynomial density (LPD) estimation at boundary points, demonstrating its decisive impact on statistical efficiency. We show that commonly used compactly supported kernels—e.g., the triangular kernel—induce severe variance inflation, elevated mean squared error, overly wide confidence intervals, and even variance divergence in small samples, critically undermining the power of manipulation tests in regression discontinuity design (RDD). To address this, we provide the first systematic theoretical and empirical analysis establishing the critical role of kernel selection in boundary LPD estimation. We propose a novel unbounded-support kernel—the Laplace spline kernel—and rigorously analyze its properties. Both asymptotic theory and Monte Carlo simulations confirm that this kernel reduces boundary estimation variance by 30–50%, substantially enhancing the statistical power of manipulation tests—particularly under small-sample and weak-discontinuity regimes.
This paper addresses the nonparametric estimation of conditional probability distributions for locally stationary processes (LSPs), aiming to accurately capture smoothly time-varying statistical structures—such as mean and variance—in time series. We propose a Nadaraya–Watson kernel-smoothed estimator for the conditional distribution function. Theoretically, we establish, for the first time, the minimax optimal convergence rate of this estimator under the Wasserstein distance, extending the result from univariate to multivariate settings. To balance computational feasibility with theoretical rigor, we adopt the sliced Wasserstein distance as a tractable surrogate metric. Numerical experiments on synthetic data and real-world financial and biological time series demonstrate that the proposed method achieves high estimation accuracy and robustness.
This study addresses the challenge of data-driven bandwidth selection for spectral density estimation near zero frequency by proposing a local-to-zero cross-validated log-likelihood criterion, denoted CVLL_c, which selects the bandwidth by summing over the frequency interval $[1, n^c]$ with $0 < c < 1$. It is shown for the first time that when $4/5 < c < 1$, the expectation of the key term in CVLL_c converges to the asymptotic mean squared error of the zero-frequency spectral estimator, thereby overcoming limitations inherent in conventional global CVLL approaches. The theoretical foundation for applying this criterion to heteroskedasticity- and autocorrelation-consistent (HAC) standard error estimation is established through Fourier frequency truncation, Taylor expansion, and asymptotic analysis.
This work addresses the challenges of bandwidth selection in kernel density estimation and its poor performance on multimodal, spiky, or discrete data by proposing a spectral method operating in the characteristic function domain. Treating bandwidth as a spectral truncation point, the approach integrates a data-driven noise-floor estimator, an adaptive Wiener tapering kernel, and a reconstruction architecture that combines a smooth basis with a band-limited residual. This framework unifies spectral decomposition, Wiener filtering, and Gaussian mixture modeling to enable automatic bandwidth selection and high-fidelity density estimation. Evaluated on the Marron–Wand benchmark suite (n = 5000), the method achieves the lowest integrated mean squared error and demonstrates superior robustness and reproducibility across six diverse real-world and synthetic datasets.
In additive noise models, regression functions estimated by machine learning often induce spurious dependence between residuals and covariates, compromising the validity of downstream inference. This work proposes the first semiparametrically efficient inference method tailored to kernel-based heteroskedasticity, constructing a Hilbert space–valued one-step estimator for the kernel covariance operator between covariates and residuals. Coupled with a bootstrap calibration procedure, the approach enables valid tests for residual independence and model goodness-of-fit. The method accommodates settings with additional covariates, supports efficient inference on heterogeneity in residual noise distributions across treatment groups, and yields asymptotically valid confidence intervals. Simulations demonstrate that, compared to naive plug-in residual methods, the proposed approach achieves substantially improved calibration and statistical power.
This study addresses the significant negative bias of existing spectral variance estimators under positively correlated data and the absence of an optimal bandwidth selection rule for lugsail kernels. Focusing on stationary data with unknown serial correlation, the work establishes, for the first time, an inference-optimal bandwidth rule for the lugsail long-run variance estimator based on a nonstandard fixed-smoothing asymptotic distribution. This rule achieves zero asymptotic bias under arbitrary dependence structures while balancing bias correction and estimation variability. Theoretical analysis and simulation experiments demonstrate that the proposed method substantially reduces estimation bias and enhances the robustness and accuracy of statistical inference.
This study addresses the challenge of specification testing in conditional moment models under high-dimensional nuisance parameters, where conventional approaches relying on asymptotically linear estimators struggle to accommodate modern machine learning methods. The authors propose a kernel-based locally robust testing framework that uniquely integrates Neyman-orthogonal moments, cross-fitting, and reproducing kernel Hilbert space techniques to achieve first-order insensitivity to estimation errors in nuisance parameters. Under mild convergence rates for nuisance estimators, the method ensures oracle equivalence between feasible and infeasible tests and precisely characterizes local power. Employing a fast multiplier bootstrap, the framework demonstrates excellent finite-sample performance across diverse applications—including specification tests for high-dimensional linear and logistic regression, significance testing in machine learning regression, and tests for constant conditional treatment effects—as validated by Monte Carlo simulations and empirical analysis.