Score
Deriving probabilistic limit theorems and finite-sample characterizations for multivariate/high-dimensional random variables and statistics (e.g., multivariate CLTs, extreme-value limits, Taylor expansions and state-evolution analyses), including establishing probabilistic thresholds and dimensionality regimes (relations between p and n) under which asymptotic or finite-sample guarantees hold.
Classical Taylor’s theorem—developed in a deterministic setting—lacks measurability guarantees for the intermediate point when applied to random functions (e.g., likelihoods), undermining the rigorous probabilistic interpretation of asymptotic expansions in statistics. This foundational issue has long been overlooked. Method: Under mild measurability assumptions, we establish the first multivariate Taylor and mean-value theorems applicable to random vectors and random functions, explicitly ensuring the measurability of the intermediate point. Our approach integrates stochastic processes, measure theory, multivariate differential calculus, and statistical asymptotics. Contribution/Results: The results bridge a critical theoretical gap between stochastic analysis and statistical inference. They provide a strictly measurable foundation for asymptotic expansions of estimators—including maximum likelihood and M-estimators—thereby strengthening the mathematical rigor of statistical inference and enabling valid probabilistic reasoning in high-order approximations.
This paper addresses the conditional distribution of Banach space-valued jointly Gaussian random variables. It establishes that the conditional distribution remains Gaussian and develops a finite-dimensional approximation scheme based on Banach space-valued martingales to compute the conditional mean and covariance operator exactly. Methodologically, it unifies nuclear norm convergence and weak convergence analyses—yielding, for the first time, a rigorously convergent Gaussian conditioning theory in general Banach spaces. The framework applies broadly, including to reproducing kernel Hilbert spaces (RKHS) and spaces of continuous functions. For continuous Gaussian process paths, it guarantees uniform convergence of the conditional mean and covariance functions, as well as weak convergence of the conditional probability measures. These results provide a rigorous, general mathematical foundation for infinite-dimensional statistical inference and Bayesian inverse problems involving Gaussian processes.
This paper investigates the theoretical foundations of generative capability, formalizing necessary and sufficient conditions for decidability of generation—unified, non-uniform, and prompt-based—and elucidating its fundamental relationship with predictability. Building on statistical learning theory, we introduce *closure dimension*—a novel combinatorial measure—as the core criterion characterizing generative capacity, thereby establishing a general framework for generability. We rigorously prove that generability and predictability (in both PAC and online learning models) are mutually exclusive. We provide complete, precise decidability characterizations—necessary and sufficient conditions—for all three generation paradigms. Furthermore, we unify and extend the framework of Kleinberg & Mullainathan (2024), modeling generation as a formal language closure problem over binary hypothesis classes. Our results bridge foundational learning theory with modern generative modeling, offering rigorous criteria to distinguish generative from predictive computation.
This paper addresses uncertainty quantification for risk-minimizing estimators in machine learning, overcoming limitations of classical approaches that rely on restrictive distributional assumptions and asymptotic theory. We propose the first general-purpose, finite-sample, distribution-free, and frequentist-valid inference framework applicable to *any* risk minimizer. Our method is grounded in the generalized likelihood ratio test, integrated with empirical process analysis and data-driven tuning, and inherently supports anytime-valid inference. Theoretically, it guarantees exact coverage of confidence sets for *all* finite sample sizes—without asymptotic approximations. Empirically, it consistently outperforms classical asymptotic methods across diverse tasks, demonstrating both high accuracy and strong robustness. This work establishes a new paradigm for model-agnostic statistical inference.
This work addresses posterior functional computation in high-dimensional generalized linear models under non-log-concave likelihoods. To overcome the lack of non-asymptotic theoretical guarantees for conventional MCMC methods under non-convex posteriors, we propose a gradient-driven MCMC algorithm that achieves statistically optimal posterior sampling in polynomial time, requiring only local likelihood regularity and a suitable initial point. Our contribution is the first non-asymptotic convergence theory for posterior sampling that operates outside the M-estimation framework—free from asymptotic assumptions and scalable to high dimensions—applicable to density estimation, nonparametric regression, and PDE inverse problems. Crucially, the theoretical guarantees hold rigorously under non-log-concave likelihoods. Empirical evaluations confirm the method’s effectiveness and computational efficiency in both generalized linear models and PDE inverse problems.
This work addresses the absence of a unified theoretical framework for the Central Limit Theorem (CLT), which has led to fragmented proofs and obscured insights into its intrinsic structure. To resolve this, the authors propose an abstract framework grounded in enriched category theory with expansive seminorms, uniquely integrating categorical methods with expansion structures to systematically characterize and derive diverse CLTs. This approach not only recovers and strengthens classical results—including the standard CLT and the Law of Large Numbers—but also extends them to statistical mechanical contexts, yielding a novel CLT for observables on symplectic manifolds. The framework thus establishes a categorical foundation for probabilistic limit theorems and opens avenues for broad theoretical and applied generalizations.
This work addresses the efficient construction and utilization of Stein methods for distributional metrics and optimization in probabilistic reasoning and learning. By establishing a general framework for constructing Stein discrepancies, it systematically analyzes their computability, separating properties, and capacity to detect and control convergence of probability measures. The study further elucidates the theoretical connection between Stein operators and Stein Variational Gradient Descent (SVGD). Beyond clarifying the foundational theoretical properties of Stein discrepancies, this research develops a unified framework that bridges variational inference and nonparametric distribution approximation, thereby providing a rigorous theoretical foundation and practical pathway for designing efficient and verifiable probabilistic learning algorithms.
This work investigates the minimax detection boundary for uniformity testing of multinomial distributions under ℓₚ deviations, focusing on the intermediate regime where the sample size satisfies \( N = o(n^2) \) and the signal-to-noise ratio converges to a finite positive constant. By introducing a Poissonized model, constructing a Poisson mixture prior, and applying a conditional central limit theorem for weighted sums, the authors establish—for the first time—a matching lower bound on the minimax risk in this regime. This lower bound precisely coincides with the known upper bound, thereby fully characterizing the asymptotic behavior of the minimax risk at the level of sharp constants: the risk converges to the nontrivial constant \( 2\Phi(-u^*/2) \).
This study investigates the large-sample asymptotic behavior of the frequency spectrum in Pitman–Yor random partitions, with particular emphasis on the limiting distribution of the sum of frequency counts over intervals of the form ∑_{j=⌊λn⌋}^{⌊μn⌋} M_{jn}. By leveraging the theory of Gibbs-type partitions, asymptotic analysis, and methods from combinatorial stochastic structures, the work establishes, for the first time, a limit theorem for the Pitman–Yor frequency spectrum that holds uniformly across broad ranges of such intervals, and further explores its functional limit form. The results uncover a profound connection between the frequency spectrum and the limiting shape of associated random combinatorial structures, thereby providing a rigorous theoretical foundation for applications in population genetics, particularly in the analysis of allele frequency spectra.
This study investigates the asymptotic normality of pattern counts in random planar maps. By directly analyzing the bivariate coefficient asymptotics of functional equations involving a single catalytic variable, the authors establish a central limit theorem for pattern occurrences without relying on prior assumptions about face-degree distributions. The approach leverages techniques from analytic combinatorics—specifically, functional equations and bivariate asymptotic analysis—to yield a more streamlined proof while extending the result to arbitrary boundary conditions and broader classes of maps. This work substantially broadens the scope under which asymptotic normality holds, enhancing both the generality and technical efficiency of the underlying theory.