Score
Computing and using statistical moments, moment-generating functions, and moment-closure approximations to characterize distributions, derive critical thresholds or timescales, and obtain tractable approximations for dynamics or discrepancy measures while balancing bias and variance.
The empirical statistical behavior of sample skewness and kurtosis under small-sample regimes remains poorly understood. Method: Leveraging asymptotic analysis, heavy-tailed distribution modeling, and experiments on both synthetic and real-world datasets, we derive rigorous theoretical bounds and examine moment estimation properties. Contribution/Results: We establish, for the first time, a strict asymptotic lower bound on sample kurtosis as a function of sample size and skewness. Extending the Taylor power law to higher-order moments, we reveal that the skewness–kurtosis 4/3 scaling law arises intrinsically from the asymptotic structure of heavy-tailed distributions. This scaling is robust only for heavy-tailed data and sufficiently large samples (n ≳ 100); under small samples, kurtosis exhibits a pronounced negative bias with a nontrivial lower bound, rendering classical moment estimators severely biased. Our results provide a theoretically grounded criterion for the validity domain of moment-based estimation and correct the misconception of universal power-law applicability.
This paper addresses the challenge of unknown moment properties for self-normalized statistics—such as ratios and the Gini coefficient—under nonnegative data. We propose a general framework for computing exact moments based solely on sample means. For the first time, we derive closed-form moment expressions whose computational complexity remains constant with increasing sample size, overcoming scalability limitations of conventional methods in high dimensions or large samples. Theoretical analysis, supported by extensive numerical experiments, reveals novel patterns in bias and variance dependence on distributional skewness and sample size—enabling the design of a broadly applicable debiasing algorithm. The derived closed-form solutions apply to diverse ratio-based metrics across economics, ecology, and machine learning. Empirical evaluation on both synthetic and real-world datasets demonstrates substantial improvements in estimation accuracy, with average bias reduction exceeding 60%, while maintaining computational scalability and practical utility.
This paper addresses risk assessment of rare events in nonstationary complex systems, focusing on modeling heavy-tailed multivariate distributions of interdependent variables and elucidating how time-varying dependence amplifies tail risk. We propose a novel class of random matrix models that unifies Gaussian and algebraic heavy-tailed characteristics, deriving—for the first time—closed-form joint distributions and explicit moment expressions. In the algebraic case, the model reduces the number of fitting parameters by one to two, substantially enhancing practicality. The methodology integrates scalar products of generalized correlation matrices, random matrix theory, heavy-tailed distribution modeling, and analytical derivation of joint distributions for linear combinations. Empirical validation on financial data demonstrates high accuracy. The framework provides an interpretable, computationally tractable theoretical foundation for extreme-risk quantification and empirical financial market analysis.
Conventional moment computation relies on analytic derivatives of probability density functions (PDFs) or moment-generating functions (MGFs), rendering it inapplicable to fractional, complex-order, and central/non-central moments when PDFs are unavailable or MGF derivatives are intractable. Method: This paper proposes the Complex-extended Moment Generating Function (CMGF) integration method—a novel framework that requires only integrability of the MGF over a contour in the complex plane. By leveraging analytic continuation and numerical contour integration, CMGF directly computes arbitrary real-order, complex-order, absolute, central, and non-central moments without invoking PDFs or MGF derivatives. Contribution/Results: As the first general-purpose MGF-based moment computation framework relying on integration rather than differentiation, CMGF is validated across three canonical scenarios where closed-form MGFs exist but PDFs or their derivatives do not. It significantly simplifies high-order and non-integer moment evaluation, enhancing efficiency in statistical inference and stochastic modeling.
Kernel Stein Discrepancy (KSD) lacks theoretical guarantees for moment convergence and is only valid under weak convergence, limiting its applicability in distribution approximation tasks requiring higher-order moment control. Method: We propose Diffusion KSD—a novel discrepancy measure integrating the Stein operator, reproducing kernel Hilbert space (RKHS), and diffusion process theory—to rigorously characterize *q*-Wasserstein convergence. Contribution/Results: Diffusion KSD is the first KSD variant that simultaneously ensures weak convergence and convergence of all moments up to order *q*. We establish sufficient conditions under which KSD controls moment convergence, thereby extending the theoretical convergence boundary of classical KSD. The method provides a computationally tractable yet theoretically rigorous tool for MCMC diagnostic analysis and fitting unnormalized models. Its principled foundation in optimal transport and Stein methodology enables reliable assessment of distributional approximation quality, with broad implications for Bayesian inference, generative modeling, and variational inference.
This work addresses the lack of a unified and efficient inference framework for models where exact likelihoods are intractable. It proposes the first general modeling and inference framework based on saddlepoint approximation, which preserves access to the moment-generating function through high-level operations to automatically construct the cumulant-generating function, its saddlepoint, and associated gradients. By integrating automatic differentiation, the framework optimizes the saddlepoint likelihood efficiently. It supports flexible modeling with complex distributional compositions and introduces diagnostic metrics that assess approximation error without requiring the true likelihood. Empirical results demonstrate that the method enables accurate parameter estimation, standard error computation, and error diagnosis with high computational efficiency and remarkable flexibility across multiple case studies.
This work addresses the challenge of estimating statistical moments—such as means and covariances—for continuous-time Markov chains, which typically lack closed-form solutions and incur prohibitive computational costs when using conventional Monte Carlo methods in multi-parameter settings. The authors propose a simulation-based surrogate modeling framework that employs neural networks to learn the mapping from model parameters to statistical moments using Monte Carlo–noisy simulation data, thereby enabling efficient approximation across the entire parameter space. A key contribution lies in uncovering the heterogeneous impact of simulation noise on moment estimation, which informs a tailored allocation of computational resources and a robust learning algorithm. Under a fixed computational budget, the method yields accurate population-level moment estimates and demonstrates significant improvements over baseline approaches in downstream tasks such as whitening.
This work addresses the lack of a unified computational framework for analyzing non-ergodicity, modeling heavy-tailed dynamics, and studying decision-making under uncertainty in stochastic processes. To this end, we introduce an open-source Python library that, for the first time, integrates non-ergodicity diagnostics, simulation of heavy-tailed processes—such as multiplicative Lévy growth and memory-dependent mean-reverting dynamics—and agent-based experimentation within a single platform. Built upon the scientific Python ecosystem (NumPy/SciPy), the library supports end-to-end workflows including stochastic process definition, simulation, parameter inference, and partial solution of stochastic differential equations. Through several reproducible examples—ranging from heavy-tailed ensemble diffusion to pre-asymptotic fluctuation analysis—it substantially reduces boilerplate code and enhances both reproducibility and development efficiency in the study of time-averaged behaviors of complex stochastic systems.
This work addresses the challenge of parameter estimation in dynamic models where the likelihood function is intractable, and existing likelihood-free methods either rely on handcrafted summary statistics or computationally expensive neural networks. To overcome these limitations, the authors propose a simulation-based inference approach leveraging random features. Their key innovation lies in introducing embedding theory from nonlinear dynamical systems into simulation-based inference, enabling identification of a p-dimensional parameter model by matching only a small number (2p+1) of random features between observed and simulated data. The method applies to both stationary and non-stationary processes and, under mild regularity conditions, yields consistent estimators. This framework establishes a new paradigm for dynamic system parameter estimation that is efficient, broadly applicable, and theoretically grounded.
This study addresses the computational challenges and inefficiency associated with Bayesian inference for stochastic volatility mean models under heavy-tailed distributions. To overcome these limitations, the authors extend the model to the class of scale mixtures of normal distributions and develop a fast approximate Bayesian inference framework based on hidden Markov models. By leveraging special functions to circumvent numerical integration and incorporating parallel computing strategies, the proposed approach substantially enhances both algorithmic stability and computational efficiency. The method achieves inference accuracy comparable to that of conventional Markov chain Monte Carlo (MCMC) techniques while accelerating computation by approximately an order of magnitude, thereby offering an efficient and practical Bayesian solution for modeling high-dimensional financial time series.