analyze u-statistic estimator

Designs and analyzes estimators built from symmetric kernel-based U-statistics by specifying U-statistic constructions, deriving finite-sample expectations, bias and variance expressions, and performing bias correction when needed. Uses theoretical tools such as the Hoeffding decomposition and limit theorems to establish consistency, asymptotic normality (or other limit laws), rates of convergence, and asymptotic lower bounds on bias and variance under degeneracy and regularity conditions.

analyzeu-statisticestimator

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Degenerate U-statistics face two key challenges: nonstandard asymptotic distributions and high computational cost. This paper departs from the classical Hoeffding decomposition framework and develops a unified analytical approach grounded in hypergraph theory and combinatorial design. It establishes, for the first time, Berry–Esseen-type convergence rate bounds for degenerate U-statistics whose order diverges, along with rigorous theoretical guarantees for normal approximation. We further propose an efficient algorithm for constructing replicated incomplete U-statistics, circumventing permutation testing entirely. The method applies to arbitrary deterministic designs and substantially improves computational efficiency. Empirical validation is conducted within nonparametric two-sample and independence testing using maximum mean discrepancy (MMD) and Hilbert–Schmidt independence criterion (HSIC). On CIFAR-10, our approach enables permutation-free MMD testing, drastically reducing computational overhead while strictly controlling Type-I error and preserving statistical power.

Addressing computational cost and degeneracy in U-statistics estimationCreating efficient algorithms for equireplicate designs in kernel testsDeveloping Berry-Esseen bounds for incomplete U-statistics with deterministic designs

Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics

Dec 15, 2025
AW
Aaron Wei
🏛️ University of British Columbia | Independent researcher | Amii

Conventional maximum mean discrepancy (MMD) two-sample tests require equal sample sizes, forcing data discarding in practice and reducing statistical power. Method: We establish, for the first time, the asymptotic normality of the MMD estimator under unequal sample sizes—breaking the long-standing reliance on balanced sampling. We propose a power-optimization criterion based on generalized U-statistics to enable high-power testing using all available data. Contributions: We identify a novel phenomenon: MMD estimator degeneracy does not necessarily imply zero population MMD, correcting a common misconception. We derive a more concise and precise variance characterization and rigorously prove statistical consistency and computational feasibility under non-proportional sampling. The proposed framework combines theoretical rigor with practical robustness, substantially enhancing MMD’s applicability and efficacy in real-world scenarios.

Develops new asymptotic distribution theory for MMD with unequal samples.Extends MMD testing to handle unequal sample sizes without discarding data.Provides a criterion to optimize test power in unequal sample scenarios.

This work addresses the problem of robust and efficient estimation of expectations of symmetric kernel functions under the presence of outliers or heavy-tailed distributions. The authors propose the Median-of-Incomplete U-statistics (MIU), which combines subsampling with median aggregation to simultaneously ensure computational efficiency and enhanced robustness. For the first time, non-asymptotic error bounds and concentration rates for MIU are established under finite-sample settings, demonstrating that the estimator achieves both statistical consistency and strong robustness in high-dimensional and non-ideal data environments. This provides a novel and theoretically grounded tool for robust U-statistic inference.

expectation estimationfinite-sample concentration rateMedian-of-Incomplete-U-Statistics

Distance and Kernel-Based Measures for Global and Local Two-Sample Conditional Distribution Testing

Oct 15, 2022
JY
Jian Yan
🏛️ Cornell University | Texas A&M University

This paper addresses the critical yet underexplored problem of testing equivalence between two conditional distributions—a fundamental task in transfer learning and causal inference. We propose the first unified framework for both global and local two-sample conditional distribution testing. Our method introduces: (1) novel distance and kernel-based metrics that characterize conditional distribution homogeneity; (2) an estimation theory grounded in conditional U-statistics, enabling integrated modeling for both global and local tests; and (3) a principled combination of RKHS embeddings and localized bootstrap resampling, yielding convergence rates and asymptotic null/alternative distributions of the estimators. Theoretical analysis guarantees strong statistical power, while empirical evaluations on synthetic and real-world datasets demonstrate high detection accuracy and robustness.

Developing distance and kernel-based measures for distribution homogeneityProposing consistent estimators and hypothesis tests for conditional distributionsTesting equality of two conditional distributions globally and locally

Latest Papers

What's happening recently
View more

This study addresses the problem of constructing anytime-valid confidence sequences for second-order U-statistics under continuous monitoring, proposing distinct approaches for non-degenerate and degenerate cases. In the non-degenerate setting, the authors leverage Hoeffding’s projection to reduce the problem to a time-uniform central limit theorem for first-order partial sums and employ a leave-one-out jackknife estimator to yield a data-driven confidence sequence. For the more challenging degenerate case, they introduce the first anytime-valid confidence sequence by developing the SAGE boundary and a computationally feasible scheme based on truncated spectral estimation. The resulting sequences achieve near-optimal width rates of √(log log n / n) and log log n / n in the non-degenerate and degenerate regimes, respectively, with numerical experiments confirming their empirical validity and efficiency.

anytime-valid inferenceconfidence sequencesdegenerate case

This study addresses the problem of testing equality of estimable parameters—such as variance, correlation coefficients, and Gini indices—across multiple independent populations using U-statistics. It develops a unified inferential framework that, for the first time, accommodates both fixed-dimensional and increasing-dimensional asymptotic settings within a single approach. The proposed method delivers an asymptotically exact distribution-free inference procedure, combining Wald-type and ANOVA-type statistics with a weighted bootstrap approximation. Theoretical analysis and extensive simulations demonstrate that the method achieves high statistical power and computational efficiency even in finite samples. Its practical utility is further corroborated through real-data applications, confirming its effectiveness in real-world scenarios.

equality testingestimable parametersmultiple populations

In additive noise models, regression functions estimated by machine learning often induce spurious dependence between residuals and covariates, compromising the validity of downstream inference. This work proposes the first semiparametrically efficient inference method tailored to kernel-based heteroskedasticity, constructing a Hilbert space–valued one-step estimator for the kernel covariance operator between covariates and residuals. Coupled with a bootstrap calibration procedure, the approach enables valid tests for residual independence and model goodness-of-fit. The method accommodates settings with additional covariates, supports efficient inference on heterogeneity in residual noise distributions across treatment groups, and yields asymptotically valid confidence intervals. Simulations demonstrate that, compared to naive plug-in residual methods, the proposed approach achieves substantially improved calibration and statistical power.

additive noise modelskernel covariance operatornoise heterogeneity

This study addresses the problem of efficiently computing unbiased estimators for high-order U-statistics, such as distance covariance and HSIC. By establishing the equivalence between U-centering and the residual structure in ANOVA, the authors propose a generalized high-order U-centering framework: symmetric hollow arrays are interpreted as least-squares residuals after fitting additive endpoint effects, and this perspective is extended to r-tuple subset indexing to eliminate lower-order effects involving fewer than r sample labels. This approach unifies high-order Hoeffding decompositions with variance component estimation, yielding an unbiased estimator with O(n^r) computational complexity. The method accurately computes inner products of r-th order Hoeffding components and expresses the highest-order variance component as a non-negative mean squared residual, substantially enhancing both computational efficiency and theoretical clarity.

ANOVA residualizationhigher-order interactionsHoeffding decomposition

This study investigates the asymptotic admissibility of Double Machine Learning (DML) estimators for quadratic functionals and integral functionals of densities under structural agnosticism. By integrating higher-order influence functions (HOIF), U-statistic theory, and a structure-free modeling framework, the authors establish—for the first time—that DML is asymptotically inadmissible for these two classes of functionals and construct a second-order influence function estimator that asymptotically dominates DML. For a third class of functionals, both DML and HOIF estimators achieve minimax optimality but neither dominates the other. These findings reveal fundamental limitations of DML under weak structural assumptions and provide superior alternatives grounded in higher-order influence functions.

Asymptotic InadmissibilityDouble Machine LearningHigher-Order Influence Functions

Hot Scholars

MY

Mingao Yuan

Department of Mathematical Sciences, The University of Texas at El Paso
Network Data AnalysisStatistical InferenceInformation Geometry
FL

Fan Li

Department of Statistical Science, Duke University
statisticscausal inferencecomparative effectiveness researchmissing data
HS

Helton Saulo

Assistant Professor of Statistics, University of Brasilia
EconometricsStatistical Learning
XH

Xiaoming Huo

Professor, Georgia Institute of Technology
statisticsdata sciencemachine learning
CH

Christian H. Weiß

Professor, Department of Mathematics and Statistics, Helmut Schmidt University, Hamburg
time series analysisstatistical process controlcomputational statistics