Score
Designs and carries out mathematical and computational analyses of ensembles of random matrices and matrix-valued random processes, including computing spectra, eigenvalue/eigenvector and singular-value distributions, and quantifying finite-sample bias and concentration. Builds and analyzes models of random matrix products and their dynamics—deriving limiting laws, stability and convergence conditions, and analytic or probabilistic techniques for predicting the behavior of matrix-valued random systems.
Spectral concentration of high-order random matrices—termed “matrix chaos”—remains a fundamental challenge in theoretical computer science, particularly in average-case analysis of the Sum-of-Squares (SoS) hierarchy. Existing techniques rely heavily on restrictive linear or low-degree assumptions, failing to handle general polynomial matrix ensembles. Method: We develop the first general spectral bound theory for arbitrary polynomial random matrices, bypassing traditional linearity constraints. Our approach introduces a unified matrix concentration inequality based on coefficient tensor flattenings, and identifies a combinatorially structured class of matrix chaos whose parameters are mechanically computable via tensor decomposition and noncommutative probability tools. Contribution/Results: The framework yields automated, order-optimal spectral bounds with improved dimension dependence for graph matrices, Khatri–Rao products, and SoS average-case analysis. It provides a systematic theoretical toolkit for efficient and precise spectral analysis of complex random combinatorial systems.
This paper addresses risk assessment of rare events in nonstationary complex systems, focusing on modeling heavy-tailed multivariate distributions of interdependent variables and elucidating how time-varying dependence amplifies tail risk. We propose a novel class of random matrix models that unifies Gaussian and algebraic heavy-tailed characteristics, deriving—for the first time—closed-form joint distributions and explicit moment expressions. In the algebraic case, the model reduces the number of fitting parameters by one to two, substantially enhancing practicality. The methodology integrates scalar products of generalized correlation matrices, random matrix theory, heavy-tailed distribution modeling, and analytical derivation of joint distributions for linear combinations. Empirical validation on financial data demonstrates high accuracy. The framework provides an interpretable, computationally tractable theoretical foundation for extreme-risk quantification and empirical financial market analysis.
This study investigates the stationary distribution and stability of stochastic differential equation systems with multimodal uncertain parameters exhibiting superposition effects, using the nonlinear Rosenzweig–MacArthur predator–prey model as a case study. For the first time, multimodal mixture-distributed parameters are incorporated into the stationary analysis of stochastic dynamical systems. System stability is quantified through the eigenvalue distribution of the Jacobian matrix, and posterior stationary density estimates are obtained via the Monte Carlo method proposed by Hoegele (2026). The results reveal that under multimodal parameter uncertainty, the system exhibits a multimodal stationary distribution, accurately delineating regions of stability. This demonstrates the effectiveness and novelty of the proposed framework for uncertainty quantification in complex ecological dynamics.
This work addresses the limited understanding of how disorder influences training and generalization in high-dimensional dynamical systems relevant to machine learning. By integrating dynamical mean-field theory (DMFT), random matrix theory, the cavity method, and path integrals, the authors reduce complex high-dimensional coupled systems to effective single-site stochastic processes driven by non-Hermitian random matrices. They uncover a novel mechanism—rooted in non-Hermitian structure—that leads to non-monotonic training loss dynamics in settings such as gradient flow, random feature models, and deep linear networks. A DMFT-based bias–variance decomposition is introduced via ensemble averaging over noise realizations, and the emergent spiked random matrix structure underlying feature learning in deep linear networks is characterized. Finally, asymptotic dynamics of both training and test losses are derived for high-dimensional random data, providing a theoretical foundation for quantifying strategies like ensemble learning.
This work investigates the asymptotic behavior of randomly initialized neural networks in the wide-limit regime, addressing three core problems: (i) the convergence rate to a Gaussian process, (ii) the Frobenius-norm convergence rate of the neural tangent kernel (NTK) to its deterministic limit, and (iii) higher-order moments of the limiting spectral distribution of the Jacobian matrix. Methodologically, we introduce a unified analytical framework: (i) extending the Faà di Bruno formula to multivariate composition for linearizing activation function effects; (ii) pioneering the application of genus expansion to neural network initialization analysis; and (iii) developing a graph-indexed stochastic multilinear mapping expansion. We rigorously establish Gaussian process convergence as width tends to infinity; quantify the Frobenius convergence rate of the NTK; and derive, for the first time, closed-form expressions for arbitrary-order moments of the limiting Jacobian spectrum—generalized to sparse and non-Gaussian weight settings.
This study investigates the low-degree polynomial approximation of the leading eigenpair of random symmetric matrices. Focusing on the Spiked Gaussian Orthogonal Ensemble (GOE) and standard GOE models, it integrates exact spectral methods with the low-degree algorithmic framework. By leveraging the extremal properties of Chebyshev polynomials and random matrix theory, this work rectifies prevailing misconceptions regarding the required number of iterations in classical power methods. It establishes a critical degree threshold for approximating the leading eigenpair and derives an exact expression for the asymptotic overlap. The resulting theoretical predictions significantly improve upon existing bounds, offering new insights into the fundamental limits of polynomial-based algorithms for random matrix computations.
This study addresses a fundamental challenge in system identification: distinguishing spurious eigenvalues arising from limited data from those genuinely reflecting the underlying system dynamics. To this end, the paper introduces—for the first time—the probabilistic sampling pseudospectrum \( P(\lambda) \) and its computationally efficient estimator \( \hat{P}(\lambda) \). By leveraging resampling and statistical inference, this framework quantifies the uncertainty of eigenvalues across the complex plane. The proposed approach provides a general and rigorous statistical criterion for data-driven methods such as Dynamic Mode Decomposition and subspace identification, substantially enhancing the reliability of identifying true dynamical modes from noisy, finite-length observations.
This project addresses efficient randomized algorithms for three fundamental matrix computations under resource constraints: (1) low-rank positive semidefinite (PSD) matrix approximation, aiming for exact low-rank reconstruction with minimal element accesses; (2) estimation of implicit matrix properties—trace, diagonal entries, and row norms—using only matrix-vector products; and (3) numerically robust floating-point solutions to overdetermined linear least-squares problems. Methodologically, we propose a randomized pivoted Cholesky decomposition to drastically reduce element accesses in PSD approximation; introduce a novel leave-one-out framework for property estimation, achieving optimal sample complexity; and design the first randomized least-squares solver that simultaneously attains asymptotic acceleration—O(nd log d) time for n × d matrices—and numerical backward stability, matching the accuracy of direct methods. These contributions advance both theoretical guarantees and practical efficiency in large-scale, memory- or query-constrained numerical linear algebra.
This study addresses accuracy and consistency issues in the numerical construction of the Karhunen–Loève expansion (KLE) arising from discretization, quadrature rules, and finite sample sizes. It establishes an algebraic equivalence between the spectral decomposition of the Fredholm integral equation and the singular value decomposition (SVD) of a weighted sample covariance matrix, thereby unifying model-driven and data-driven KLE frameworks. The work innovatively constructs the covariance function on a non-simply-connected three-dimensional toroidal domain using the shortest interior path distance, and implements the approach numerically with unstructured meshes and Gaussian quadrature. Experiments demonstrate that, in a one-dimensional benchmark problem, SVD-based eigenvalue estimates and empirical KL coefficients converge to the theoretical 𝒩(0,1) distribution. In two-dimensional irregular and three-dimensional toroidal domains, the study systematically quantifies the combined influence of discretization strategy, quadrature accuracy, and sample size on KLE reconstruction error.
This work addresses the gap between algorithmic prototypes and efficient implementations in scientific research by proposing a lightweight approach to translate statistical and machine learning algorithms—such as kernel ridge regression and stochastic gradient descent matrix factorization—from mathematical formulations into readable, high-performance C++ code. Leveraging the Eigen template library for core linear algebra operations—including kernel matrix construction, regularized solvers, and vectorized updates—the implementation seamlessly integrates into the Python ecosystem via pybind11, enabling efficient interoperability with NumPy arrays. The project provides concise, reproducible code examples that encapsulate common computational patterns in research, significantly lowering the barrier for researchers to adopt C++ for high-performance development while balancing performance, readability, and usability.