Score
Designs and implements bootstrap-based goodness-of-fit tests using the Kernel Stein Discrepancy (KSD), including constructing bootstrap resampling procedures and KSD test statistics and proving properties such as preservation of asymptotic test level, local consistency, and test power. Builds and analyzes scalable approximations (e.g., Nyström estimators) and evaluates how these computational reductions affect test validity and runtime.
This work addresses the computational inefficiency of traditional kernel Stein discrepancy (KSD) tests, which suffer from quadratic time complexity due to their reliance on U- or V-statistics and require computationally intensive bootstrap procedures to approximate the null distribution. To overcome these limitations, the authors propose an accelerated KSD test based on the Nyström approximation. They provide the first theoretical guarantee that this approach preserves asymptotic type-I error control and local consistency within a bootstrap framework, while substantially reducing computational cost. Empirical evaluations on spherical and functional data demonstrate that the accelerated method achieves statistical performance comparable to the original KSD test but with significantly improved computational efficiency, thereby enabling scalable and theoretically sound nonparametric goodness-of-fit testing.
To address the $O(n^2)$ computational bottleneck of kernelized Stein discrepancy (KSD) under large-scale data—arising from its reliance on U- or V-statistics—this paper introduces, for the first time, the Nyström low-rank kernel approximation into KSD estimation, yielding a scalable and accelerated KSD estimator. The proposed method reduces time complexity to $O(mn + m^3)$, where $m ll n$, and establishes $sqrt{n}$-consistency under sub-Gaussian assumptions. Theoretical analysis is grounded in the Stein operator and reproducing kernel Hilbert space (RKHS) framework, balancing statistical efficiency with computational tractability. Extensive benchmark experiments demonstrate that the new estimator retains statistical power comparable to the original KSD while substantially enhancing practicality for large-scale goodness-of-fit testing. This work provides an efficient, theoretically sound tool for high-dimensional distribution fitting and hypothesis testing.
Existing kernel-based goodness-of-fit (GoF) tests fail under both qualitative and quantitative robustness, and robustification strategies—such as tilted kernels—cannot simultaneously satisfy both criteria in GoF testing. Method: Addressing the practical question “Is the model sufficiently accurate?”, we propose the first robust GoF testing framework based on a kernel Stein discrepancy (KSD) ball. This framework rigorously formalizes robust GoF testing and theoretically establishes that conventional kernel tests—and their tilted-kernel variants—lack dual robustness. Contribution/Results: Our test achieves both stability and statistical power under diverse contamination models—including Huber contamination and density bands—enabling unified robust modeling. Empirical evaluation confirms its effectiveness in finite-sample settings, resolving a long-standing challenge in designing robust kernel-based GoF tests.
This paper addresses the goodness-of-fit testing problem under composite hypotheses: “Does a model belong to a given parametric family?” We propose a unified, data-splitting-free, kernel-based testing framework. For the first time, it enables parameter estimation and hypothesis testing to share the same dataset while rigorously controlling the Type-I error rate. The method accommodates unnormalized densities and simulator-based models without requiring explicit density evaluation. By integrating maximum mean discrepancy (MMD), kernel Stein discrepancy (KSD), and minimum distance estimation—within a theoretically grounded composite null testing framework—it substantially improves statistical power and broadens applicability. Experiments on unnormalized density models and biological cell-network simulators demonstrate its effectiveness. Theoretically, the test level is precisely controllable, and the framework extends the scope of goodness-of-fit testing to complex, intractable models.
This paper addresses the multivariate two-sample and $k$-sample goodness-of-fit testing problem by introducing the first unified kernelized quadratic distance (KQD) framework. Methodologically, it integrates both two-sample and $k$-sample tests within a single statistical model grounded in matrix-valued distances and reproducing kernel Hilbert space (RKHS) theory; derives the asymptotic null distribution of the test statistic rigorously; and enables finite-sample inference via Monte Carlo or permutation procedures. Key contributions include: (i) establishing the first theoretical equivalence between KQD-based tests and maximum mean discrepancy (MMD) tests; (ii) proposing a scalable, statistically rigorous paradigm for multi-group testing; and (iii) releasing QuadratiK, an open-source software package supporting both R and Python. Extensive simulations and real-data analyses demonstrate that the method maintains accurate Type-I error control while achieving substantial gains in statistical power.
Traditional goodness-of-fit tests struggle to distinguish between “no significant difference” and “practical equivalence,” as failure to reject the null hypothesis may merely reflect insufficient test power. This work proposes the first kernel-based framework for full-distribution equivalence testing, leveraging Kernel Stein Discrepancy (KSD) and Maximum Mean Discrepancy (MMD) to quantify the distance between distributions while incorporating a prespecified minimum equivalence margin. By employing asymptotic normal approximations and bootstrap procedures to compute critical values, the method overcomes the limitations of existing equivalence tests, which are typically confined to parametric models or specific moments. Numerical experiments demonstrate that the proposed approach reliably assesses whether two distributions are equivalent within the specified margin while effectively controlling both Type I and Type II error rates.
For models with intractable likelihoods but tractable normalizing constants, existing goodness-of-fit tests lack theoretical guarantees or computational feasibility. Method: We propose the semiparametric kernelized Stein discrepancy (SKSD) test, unifying score-based and distance-based frameworks. SKSD leverages exponential tilting models, integral probability metrics, a kernelized Stein function class, the Stein identity, and parametric bootstrap. Contribution/Results: SKSD is the first test proven to be universally consistent, Pitman-optimal, and robust to nuisance parameters. Crucially, it reveals that classical distance-based tests—including Kolmogorov–Smirnov, Wasserstein-1, and MMD—are special cases of score-based constructions under specific Stein operators. Empirically, SKSD achieves competitive power against specialized normality tests (e.g., Anderson–Darling, Lilliefors) on kernel exponential families and conditional Gaussian models. It provides the first general-purpose goodness-of-fit test for complex latent-variable models that simultaneously satisfies strong theoretical guarantees and practical computability.
This work addresses the problem of nonparametric goodness-of-fit testing for the target distribution under covariate shift, where labels are available only from the source distribution. The authors propose a novel approach based on truncated importance-weighted kernel ridge regression combined with a multiplier bootstrap procedure. By introducing a truncation mechanism to stabilize importance weights under heavy-tailed density ratios and employing bootstrap calibration to construct confidence sets for the regression function, the method achieves both theoretical rigor and practical efficacy. Under an operator compatibility condition, explicit non-asymptotic coverage error bounds are established when the density ratio satisfies either bounded moment or sub-exponential tail assumptions. Numerical experiments demonstrate the superior performance of the proposed method.
This paper investigates how grid resolution affects coverage accuracy when constructing uniform confidence bands for functions via the multiplier bootstrap on a finite evaluation grid. Existing approaches fail to disentangle discretization error from high-dimensional bootstrap approximation error. Method: We propose the first decoupled analytical framework that separately quantifies these two error sources, deriving an explicit upper bound on the overall coverage error. Our approach integrates extreme-value statistics, high-dimensional approximation theory, and kernel density estimation techniques, ensuring computational feasibility while rigorously controlling total coverage bias. Contribution/Results: Based on the theoretical bound, we establish a practical, operationally feasible rule for selecting grid size. The theoretical results are validated in kernel density estimation, demonstrating that our criterion significantly improves the empirical coverage probability of confidence bands. The method thus bridges theoretical rigor with practical utility for uncertainty quantification in nonparametric function estimation.