Score
Designs and analyzes methods for computing distributions of sums or linear combinations of random variables via probability convolution, using analytical, transform-based, or numerical convolution techniques. Builds and derives coupled noise-generation and distribution-transformation procedures and their sampling algorithms so that joint or parameterized distributions meet specified constraints or invariants (for example matched marginals or other statistical properties).
This study addresses the challenge of modeling predictive distributions for nonlinear, multivariate time series by proposing a general generative representation framework grounded in measure-theoretic probability, which is, to the authors’ knowledge, the first to be integrated with conditional generative adversarial networks (CGANs). Under a mild temporal dependence assumption, the method establishes estimation consistency in the Hausdorff metric and enables efficient simulation and computation of conditional means, variances, and risk measures. Empirical results demonstrate that the model achieves strong predictive performance on tasks involving stock returns, realized variances, and covariances, delivering high accuracy with remarkable computational efficiency—requiring only about one minute for a single training run.
This study addresses the fundamental statistical problem of conditional density estimation by systematically comparing classical nonparametric approaches—such as single-index models, basis expansion methods (e.g., FlexCode), and DeepCDE—with modern generative models, including conditional GANs and conditional denoising diffusion probabilistic models. For the first time, these methods are evaluated within a unified and reproducible framework using metrics like mean squared error and Wasserstein distance to assess their accuracy, flexibility, and computational cost in estimating conditional means and standard deviations. The analysis clarifies the performance boundaries and practical applicability of each approach, offering both theoretical insights and actionable guidance for selecting appropriate methods in predictive modeling, uncertainty quantification, and probabilistic inference tasks.
This paper addresses risk assessment of rare events in nonstationary complex systems, focusing on modeling heavy-tailed multivariate distributions of interdependent variables and elucidating how time-varying dependence amplifies tail risk. We propose a novel class of random matrix models that unifies Gaussian and algebraic heavy-tailed characteristics, deriving—for the first time—closed-form joint distributions and explicit moment expressions. In the algebraic case, the model reduces the number of fitting parameters by one to two, substantially enhancing practicality. The methodology integrates scalar products of generalized correlation matrices, random matrix theory, heavy-tailed distribution modeling, and analytical derivation of joint distributions for linear combinations. Empirical validation on financial data demonstrates high accuracy. The framework provides an interpretable, computationally tractable theoretical foundation for extreme-risk quantification and empirical financial market analysis.
This work addresses the longstanding limitation in conditional density estimation—namely, the absence of closed-form solutions for multivariate conditional densities under non-Gaussian assumptions. We propose a generative conditional density estimation framework grounded in copula modeling and analytic conditionalization in latent space. Methodologically, we first establish the inheritability of “conditional stability” under mixture and transformation operations, thereby extending analytically tractable conditional families to non-Gaussian, nonlinear, and cross-dimensional settings. The core components include a Gaussian Mixture Copula Model (GMCM), an explicit latent-space conditionalization mechanism, and joint copula modeling. Experiments on synthetic and real-world datasets demonstrate substantial improvements in conditional density estimation accuracy and robustness to missing data imputation. Crucially, our approach enables efficient, differentiable, and sampling-free deterministic conditional inference.
This paper challenges the reliance of machine learning research in the social sciences on abstract data-generating distributions, arguing that such assumptions lack empirical grounding in finite-population settings and engender interpretability and reproducibility issues. Method: The authors advocate replacing distributional assumptions with finite-population modeling, systematically advancing five core arguments grounded in statistical foundations, philosophical epistemology, and ML empirical analysis. They reconstruct the premises of learning theory by explicitly identifying the implicit assumptions and boundary conditions underlying distributional modeling. Contribution/Results: The proposed framework enhances theoretical coherence, modeling transparency, causal traceability, and practical applicability. It provides a novel paradigm and methodological foundation for sampling design, bias correction in evaluation, and reproducibility research—thereby addressing critical limitations of conventional distribution-based approaches in social-science ML applications.
This work proposes the first fast and universal random number generation algorithm that covers the entire parameter space of the Pearson Type IV distribution, which has long lacked an efficient sampling method due to its complex parameter structure. By employing a carefully designed transformation combined with an adaptive rejection sampling strategy, the algorithm achieves uniformly efficient sampling across all admissible shape parameters. The method not only fills a critical gap in existing computational techniques for this distribution but also demonstrates practical utility in Bayesian inference tasks, confirming its effectiveness in real-world statistical modeling. As such, it provides a key computational tool for implementing complex probabilistic models involving the Pearson Type IV distribution.
This study addresses the challenge of statistical inference in the presence of heteroscedastic measurement errors, where conventional methods often fail due to either neglecting the noise structure or incurring prohibitive computational costs. The authors propose a convolutional Maximum Mean Discrepancy (convMMD) framework, which extends MMD to noisy settings for the first time by convolving observed samples with the known noise distribution, thereby enabling robust nonparametric testing and estimation. Theoretical contributions include finite-sample bias bounds, an equivalence between noise-convolved MMD tests and kernel smoothing, and proofs of consistency and asymptotic normality of the resulting estimators. Empirical evaluations demonstrate that the method achieves both computational efficiency and practical utility across simulations and real-world applications in astronomy and social sciences.
This paper addresses the efficient discrete approximation of Gaussian mixture models (GMMs) under the Wasserstein distance, motivated by dual requirements of quantization accuracy and computational scalability in control and cyber-physical system verification. We propose an enhanced quantization framework that integrates sigma-point sampling with adaptive clustering, enabling robust handling of high-dimensional, large-scale, and degenerate GMMs. A rigorous upper bound on the Wasserstein approximation error is derived, and a modular interface is provided to support customizable approximation schemes. Experiments demonstrate that our method achieves sublinear error convergence while significantly reducing computational overhead—outperforming state-of-the-art approaches in accuracy. The core contribution lies in the first systematic integration of sigma-point mechanisms with Wasserstein quantization theory, thereby unifying theoretical guarantees with practical deployability.
This study addresses accuracy and consistency issues in the numerical construction of the Karhunen–Loève expansion (KLE) arising from discretization, quadrature rules, and finite sample sizes. It establishes an algebraic equivalence between the spectral decomposition of the Fredholm integral equation and the singular value decomposition (SVD) of a weighted sample covariance matrix, thereby unifying model-driven and data-driven KLE frameworks. The work innovatively constructs the covariance function on a non-simply-connected three-dimensional toroidal domain using the shortest interior path distance, and implements the approach numerically with unstructured meshes and Gaussian quadrature. Experiments demonstrate that, in a one-dimensional benchmark problem, SVD-based eigenvalue estimates and empirical KL coefficients converge to the theoretical 𝒩(0,1) distribution. In two-dimensional irregular and three-dimensional toroidal domains, the study systematically quantifies the combined influence of discretization strategy, quadrature accuracy, and sample size on KLE reconstruction error.