uniform empirical-process analysis

Designs and carries out uniform empirical-process analyses that establish uniform probabilistic bounds and convergence rates for estimators and stochastic processes indexed by function classes, including accommodation of clustered or subject-specific sampling schemes. Uses VC-type complexity conditions for intrinsic kernel classes (including non‑Lipschitz kernels) to control suprema of empirical processes and to derive spectral perturbation bounds for covariance/operator estimators and their eigenvalues and eigenfunctions.

uniformempirical-processanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates how to effectively transfer the supremum-norm convergence rates and weak convergence properties of covariance kernel estimators to functional principal components (FPCs). By integrating $L_2$ perturbation theory with the Nyström method and minimax lower bound analysis, the work establishes, for the first time, optimal supremum-norm convergence rates and asymptotic normality for FPC estimators. The analysis uncovers a novel phenomenon in sparse observation regimes, where eigenvalue estimation is dominated by discretization effects. These theoretical results are further extended to cross-covariance, long-run covariance, and derivative process covariance kernels. Empirical validation on temperature curves demonstrates the method’s effectiveness across the spectrum from sparse to dense observation designs.

covariance kernel estimationdiscrete observational modelfunctional principal components

This paper addresses the optimal nonparametric estimation of the covariance kernel under supremum-norm loss for synchronously sampled functional data. We propose a kernel-based estimator constructed directly from discrete synchronous observations, without requiring prior estimation or smoothness assumptions on the mean function; it accommodates high-order smoothness off the diagonal while allowing low regularity on the diagonal. Theoretically, we establish, for the first time under dense sampling, a $sqrt{n}$-rate convergence without logarithmic penalties; under sparse sampling, the rate degrades to that of one-dimensional mean estimation, breaking the two-dimensional structural barrier. We derive an information-theoretic lower bound and achieve matching upper bounds, yielding tight minimax-optimal rates. Moreover, in the dense regime, we prove a supremum-norm central limit theorem, enabling construction of uniform confidence sets. Simulation studies and real-data analysis demonstrate the method’s effectiveness and robustness.

Addressing transition from dense to sparse design scenariosEstimating covariance kernel from synchronously sampled functional dataObtaining minimax-optimal convergence rates in supremum norm

This paper addresses the problem of controlling the supremum of empirical processes under general mixing random processes, aiming to establish a unified maximal inequality for function classes under arbitrary mixing rates—both fast and slow. To overcome the limitation of classical approaches, which rely on specific mixing assumptions, the authors introduce a novel complexity measure that integrates functional-analytic techniques, probabilistic inequalities, and empirical process theory. This enables, for the first time, a universal characterization of the supremum of sample averages under general mixing conditions. The analysis reveals a phase-transition phenomenon: the mixing rate critically determines the concentration speed. Based on this insight, the paper derives new Glivenko–Cantelli-type uniform convergence results and a strong approximation theorem, substantially extending the scope of Donsker-type functional central limit theorems to weakly dependent sequences.

Analyzing concentration rates under arbitrary mixing conditionsBounding supremum of sample averages for mixing processesEstablishing strong approximations for general mixing processes

This work addresses the lack of theoretical guarantees for nonparametric regression in reproducing kernel Hilbert spaces under model misspecification, high-dimensional settings, and nonconvex losses. It establishes a unified theoretical framework for regularized M-estimators encompassing a broad class of both convex and nonconvex loss functions. By introducing a novel complexity measure, the analysis achieves an explicit bias–variance decomposition. Leveraging tools from functional analysis and empirical process theory, the study proves the existence, measurability, and asymptotic linearity of the estimator without requiring closed-form solutions or global Lipschitz assumptions. Notably, within tensor-product Sobolev spaces, the framework reveals a mechanism to circumvent the curse of dimensionality, yielding minimax-optimal convergence rates that depend on mixed smoothness of the underlying function. The variance component is shown to be robust to model misspecification, and numerical experiments in C++ corroborate the theoretical findings.

convergence ratescurse of dimensionalityM-estimation

Gaussian processes (GPs) struggle to rigorously incorporate uncountably infinite-dimensional functional prior information—such as boundary conditions or global physical constraints satisfied by PDE solutions. Method: This paper proposes a unified modeling framework grounded in reproducing kernel Hilbert spaces (RKHS), establishing for the first time a rigorous equivalence between the GP conditional expectation and orthogonal projection in RKHS. This enables direct embedding of functional constraints (e.g., Dirichlet or Neumann boundary conditions) into the GP prior, bypassing conventional pseudo-point approximations. Contribution/Results: We provide theoretical guarantees on existence, uniqueness, and convergence of the constrained GP posterior. Computationally, we design a practical numerical approximation algorithm. Experiments on PDE inverse problems demonstrate substantial improvements in uncertainty quantification accuracy and posterior consistency. The framework delivers a rigorous, general, and computationally tractable paradigm for integrating domain knowledge into Bayesian modeling.

Addressing boundary value problems without pseudo-training pointsModeling Gaussian processes with uncountable functional informationUnifying finite data and uncountable information via kernel methods

Latest Papers

What's happening recently
View more

This work investigates when the conditional expectation operator (CEO) defines a bounded and Hilbert–Schmidt map from a function space into a reproducing kernel Hilbert space (RKHS). By analyzing the regularity of the Radon–Nikodym density of the conditional distribution, the authors establish a unified and verifiable sufficient condition that integrates probabilistic regularity, operator theory, and kernel methods. The criterion is validated across three distinct settings—nonparametric regression, Bayesian inverse problems, and Koopman operator theory—providing a direct pathway to verify the well-definedness and error bounds of conditional mean embeddings. This result establishes a cohesive theoretical framework applicable across diverse domains.

Conditional Expectation OperatorsConditional Mean EmbeddingsRegularity Criterion

This work investigates the optimal sampling complexity required to achieve dimension-independent $(1\pm\varepsilon)$ relative error guarantees in regularized classification tasks. Focusing on Lipschitz continuous loss functions—such as logistic, hinge, and ReLU losses—combined with $\ell_1/k$, $\ell_2/k$, and $\ell_2^2/k$ regularization, the paper introduces a novel analytical framework based on higher-order moments and empirical processes, circumventing the overcounting issues inherent in traditional VC-dimension and sensitivity-based approaches. Matching upper and lower bounds are established for the first time for all three regularization schemes: tight bounds of $k^2/\varepsilon^2$ and $k/\varepsilon^2$ are achieved for $\ell_2/k$ and $\ell_1/k$, respectively, while for $\ell_2^2/k$ regularization under the condition $g'(0)=0$, a linear-in-$k$ sampling complexity is attained, substantially improving upon the prior $k^3/\varepsilon^2$ bound. The proposed algorithms employ uniform or (squared) norm-based sampling and demonstrate strong theoretical and practical performance.

dimension-free samplingLipschitz loss functionsoptimal bounds

This work addresses kernel regression under non-Gaussian noise—including sub-Gaussian, bounded, sub-exponential, and moment-bounded types—by establishing a unified probabilistic uniform error bound within a non-asymptotic framework. It overcomes the prevailing limitation of existing approaches that assume conditionally independent sub-Gaussian noise, and for the first time simultaneously accommodates multiple non-Gaussian noise distributions and dependent noise structures. By integrating concentration inequalities with reproducing kernel Hilbert space theory, the derived non-conservative error bounds substantially enhance the reliability of uncertainty quantification. This leads to significantly tighter confidence regions in safety-critical control tasks, outperforming current state-of-the-art methods in both theoretical rigor and practical performance.

kernel regressionnon-asymptotic boundsnon-Gaussian noise

This study investigates the robust stability of statistical estimators under η-fraction data corruption. To this end, it introduces “empirical sensitivity” as a novel metric quantifying an estimator’s susceptibility to data perturbations, with a focus on Gaussian mean estimation. Theoretical analysis establishes, for the first time, a lower bound of Ω(η + √(ηd/n)) on empirical sensitivity and demonstrates its tightness up to logarithmic factors. By integrating the Efron–Stein inequality with recent advances in robust mean estimation algorithms, the authors construct an estimator achieving a matching upper bound, thereby confirming the attainability of this limit. This work provides a new theoretical framework and quantitative tool for analyzing the stability of robust estimators.

data perturbationempirical sensitivityGaussian mean estimation

Hot Scholars

MK

Masahiro Kato

Mizuho-DL Financial Technology Co., Ltd. / The University of Tokyo
Economics
SS

Sergey Samsonov

HSE university, Moscow
high-dimensional probabilityMarkov ChainsMCMC
AR

Aaditya Ramdas

Associate Professor (with tenure), Carnegie Mellon University
Machine LearningStatistics
EM

Eric Moulines

Professeur, Ecole Polytechnique, Membre de l'Académie des Sciences
StatisticsMachine learningSignal Processing
GX

Gongjun Xu

University of Michigan
StatisticsMachine LearningPsychometrics