Score
Designs and carries out rigorous derivations of probabilistic limit theorems and asymptotic properties for estimators and test statistics, including proofs of consistency, convergence rates, limit distributions, leading-order expansions, and asymptotic comparisons using weak convergence, functional central limit theorems, and empirical-process methods. Uses these derivations to compute asymptotic variances and standard errors and to justify and construct asymptotic inference procedures (confidence regions, Wald tests, hypothesis tests), and to analyze second-order risk behavior and specialized estimator inference (e.g., HTC or half-trek estimators).
Existing sequential mean testing methods struggle to characterize the higher-order asymptotic behavior of stopping times, particularly lacking second-order information-theoretic optimality analyses under bounded distributions. This work proposes a sequential test based on the KL_inf statistic, which not only achieves first-order asymptotic optimality but also establishes, for the first time, a central limit theorem for this statistic. Consequently, the stopping time—after appropriate centering and scaling by √log(1/α)—converges in distribution to a Gaussian limit. This result provides a refined second-order characterization of sequential test performance under bounded distributions, with theoretical proof showing convergence to a normal distribution having an explicit variance. Numerical experiments corroborate the accuracy of this second-order asymptotic analysis.
This study develops a statistical inference framework for topological features derived from persistent homology and applies it to goodness-of-fit testing for spatial point patterns. By introducing stabilization techniques—previously unexplored in persistent homology analysis—the authors establish a central limit theorem for bounded functionals of persistence diagrams generated by germ–grain random set models, particularly those exhibiting exponential decay of correlations, as the observation window expands. Building on this asymptotic theory, they construct effective goodness-of-fit tests by combining rectangular partitions of persistence diagrams with functional summaries such as accumulated persistence functions (APFs) and support functions of lift zonoids. The proposed methodology accurately distinguishes between clustered and repulsive spatial structures and successfully detects spatial interaction patterns in breast histology images, demonstrating its practical utility.
This paper identifies the systematic failure of classical statistical methods—including mean-based inference, principal component analysis (PCA), and asymptotic normality assumptions—under heavy-tailed distributions, particularly in medium-sample-size (medium-*n*) real-world settings where they are routinely misapplied. Methodologically, it challenges the uncritical adoption of Gaussian and stable-distribution assumptions and introduces the “Median Law” theoretical framework, which formalizes fundamental limitations under heavy tails: unreliable sample means, distorted empirical distributions, and degenerate principal components. The approach integrates extreme value theory, generalized stable distribution modeling, robust parametric estimation, and pre-asymptotic analysis. Empirical validation draws on counterexamples from finance, economics, and psychology, supplemented by cross-disciplinary case studies. The core contribution is a foundational rethinking of uncertainty quantification and causal inference: it demonstrates that many canonical “cognitive biases” are, in fact, rational inferences under heavy-tailed probability structures—thereby advocating a paradigm shift in statistical practice from idealized asymptotics to empirically grounded probabilistic modeling.
This paper addresses statistical inference for the averaged estimator in temporal-difference (TD) learning under a Markov chain setting, establishing the first non-asymptotic central limit theorem (CLT). Methodologically, it pioneers the integration of Stein’s method with the Poisson equation to derive a non-asymptotic CLT for vector-valued martingale difference sequences, extended to functionals of ergodic Markov chains. Key contributions include: (1) an $O(1/sqrt{n})$ convergence rate for the TD averaged estimator with explicit, non-asymptotic error bounds; (2) the first non-asymptotic characterization of normality for TD estimators; and (3) a rigorous statistical foundation for constructing confidence intervals and conducting hypothesis tests in reinforcement learning. By unifying stochastic approximation, Markov ergodic theory, and Stein’s method, the work significantly advances the interpretability and reliability analysis of RL algorithms.
This study addresses volatility estimation in high-frequency financial data contaminated by market frictions or microstructure noise. The authors propose an adaptive subsampling approach that directly infers the asymptotic (conditional) covariance matrix of volatility estimators without explicitly modeling the noise structure. By employing time-rescaled statistics over local intervals to assess sampling variability, the method automatically selects tuning parameters while ensuring the resulting covariance matrix is positive semidefinite. Theoretical analysis, Monte Carlo simulations, and empirical applications demonstrate that the proposed estimator is consistent, exhibits strong finite-sample performance, and enables robust and feasible statistical inference.
This study addresses the estimation of parameters of the form θ₀ = E[F_Y⁻¹∘F_Z(X)] in the “changes-in-changes” model, for which existing methods lack theoretical guarantees when variables are unbounded. The authors construct a plug-in estimator based on empirical quantiles and establish its √n-consistency and asymptotic normality under assumptions weaker than those in the current literature. They further propose a novel consistent estimator for the asymptotic variance. The theoretical analysis leverages empirical process theory and plug-in methods for quantile functions. Monte Carlo simulations demonstrate that the proposed variance estimator substantially outperforms existing alternatives, leading to markedly improved inference accuracy.
This study addresses the failure of standard influence function–based inference in finite samples under “near-boundary” settings of semiparametric models, where second-order remainder terms non-negligibly contribute to sampling variance. The authors propose a finite-sample variance decomposition framework that separates influence function variance from remainder-induced variance and establish necessary and sufficient conditions for the consistency of sandwich variance estimators. Building on this framework, they develop two robust variance estimators—the leave-one-unit-out jackknife and a paired-cluster bootstrap—and derive an analytical expression for the interaction between remainder terms and within-cluster correlation in clustered data. Their approach enables valid confidence interval construction near the boundary: the jackknife Wald interval is numerically equivalent to a bias-corrected sandwich estimator and accurately captures the mechanism by which clustering amplifies variance estimation bias.
Classical empirical processes fail under heavy-tailed or skewed distributions where moments (e.g., mean or variance) do not exist. To address this, we propose the trimmed functional empirical process (TFEP) framework, which relaxes the conventional finite-variance requirement. Under mild regularity conditions, TFEP converges weakly to a Gaussian process, enabling novel asymptotic theory for one-sample and two-sample hypothesis tests as well as confidence intervals. Our methodology integrates extreme-order-statistic trimming, functional empirical process analysis, and Monte Carlo simulation, supporting robust inference for trimmed means, variances, and their ratios. Empirical evaluations demonstrate that TFEP substantially outperforms classical functional empirical processes and normal approximation methods under Pareto and Cauchy distributions—yielding more accurate confidence interval coverage. We further validate its practical utility through application to real-world income data analysis.
This work investigates the joint asymptotic distribution of multiple test statistics—such as those employed in NIST STS and TestU01—for assessing randomness, under the null hypothesis $H_0$ (i.i.d. sequence) and a local alternative $H_1$ (sequences asymptotically approaching $H_0$ as sample size grows). Leveraging empirical process theory, the multivariate central limit theorem, and characteristic function techniques, we establish, for the first time, the joint asymptotic normality of such composite test statistics. We derive an explicit Berry–Esseen-type bound on the convergence rate, with computable constants, and identify sufficient conditions for asymptotic independence among the statistics. These theoretical advances provide a rigorous foundation for high-dimensional joint testing of pseudorandom sequences, substantially enhancing both detection sensitivity and statistical reliability for subtle deviations from ideal randomness.
This study addresses the challenge of accurately controlling the family-wise error rate (FWER) under weak dependence among normal test statistics—a setting where existing methods often fail, particularly in high-dimensional applications such as genomics. The authors provide the first rigorous proof that, as the number of hypotheses grows to infinity, suitably adjusted Bonferroni and Šidák procedures achieve asymptotically exact FWER control under weak correlation structures. Furthermore, they characterize the limiting behavior of these procedures with respect to generalized FWER and statistical power. Combining asymptotic analysis, probabilistic limit theory, and a multiple testing framework—supported by numerical simulations—the work fills a critical theoretical gap and demonstrates the validity and robustness of these classical correction methods in high-dimensional, weakly dependent settings.