Score
Designs, builds, and analyzes estimators, decompositions, approximations, and computational procedures for covariance objects and their parameters — e.g., covariance matrix estimation, diagonal-plus-low-rank and tower decompositions, covariance-aware shrinkage and regularization, and methods for efficient covariance computation and approximation — and derives their statistical properties such as finite-sample risk bounds and data-driven shrinkage sizes. Also develops and integrates covariate procedures (covariate adjustment, selection, balance diagnostics, control, and covariate-shift correction) that use or correct covariance structure and propagate covariance estimates through inference pipelines.
This work addresses the high statistical risk inherent in high-dimensional covariance estimation by reframing covariance shrinkage as a parameterized empirical risk minimization problem based on stochastic interpolation between source and target distributions. The proposed approach extends the theoretical foundations of classical shrinkage estimation through a synergistic integration of optimal transport couplings, eigenvector regularization induced by nonlinear flow maps, an early-stopping mechanism grounded in vector field regression, and an adaptive scheduling strategy. By unifying stochastic interpolation, neural estimators, and quadratic risk upper-bound analysis, the method demonstrates strong empirical performance on synthetic data and achieves superior regularization efficacy and estimation accuracy on real neuroimaging datasets.
Testing structural assumptions on covariance and correlation matrices—such as sphericity, diagonality, bandedness, or factor structure—remains a fundamental yet challenging task in multivariate statistics, especially under non-normality or limited distributional assumptions. Method: This paper proposes a general semi-parametric testing framework that relaxes traditional normality and parametric modeling requirements. It accommodates a broad class of structured null hypotheses and employs bootstrap-based p-value calibration to ensure accurate size control and high statistical power, even in small-sample settings. Contribution/Results: We develop the open-source R package CovCorTest, featuring a modular, extensible API. Extensive simulations and real-data applications across finance, genomics, and psychology demonstrate that the method substantially improves robustness, applicability, and practical usability for covariance structure inference. It provides a reliable foundational tool for high-dimensional dependence modeling.
Conventional single-target shrinkage estimators for high-dimensional covariance matrices suffer from poor adaptability and limited robustness. Method: We propose a multi-target linear shrinkage estimator that combines the sample covariance matrix with multiple constant target matrices—such as the identity and diagonal matrices—via data-driven weights. Contribution/Results: Our approach is the first to construct both theoretically grounded oracle and feasible estimators, rigorously establishing their consistency and convergence under the Kolmogorov asymptotic regime. By relaxing the restrictive single-target constraint, it enables flexible, data-adaptive selection of target combinations. Leveraging random matrix theory and empirical risk minimization, the estimator achieves computational efficiency alongside strong theoretical guarantees. Extensive experiments demonstrate its significant superiority over classical single-target methods—including Ledoit–Wolf—across diverse high-dimensional settings, delivering enhanced robustness, adaptability, and empirical performance.
This paper addresses covariance and precision matrix estimation from weighted sample covariance matrices—particularly exponentially weighted ones—and proposes an asymptotically optimal nonlinear shrinkage method while characterizing the joint distribution of sample and population eigenvector overlaps. Leveraging random matrix theory and the Ledoit–Péché asymptotic framework, we derive, for the first time, explicit asymptotic formulas for nonlinear shrinkage under general weighting schemes; establish a rigorous theory for the joint distribution of eigenvector overlaps; and develop a computationally tractable, robust numerical algorithm. Experiments demonstrate substantial improvements over conventional linear shrinkage estimators. Crucially, both theoretical guarantees and algorithmic performance remain valid under heavy-tailed data distributions. Our core contributions are: (1) the first explicit asymptotic solution for weighted nonlinear shrinkage; (2) a precise characterization of the joint distribution of sample–population eigenvector overlaps; and (3) a unified estimation framework that achieves theoretical optimality while maintaining practical robustness.
Existing spectral shrinkage methods for high-dimensional covariance estimation rely either on heuristic objectives or strong distributional assumptions. To address this, we propose a geometrically unified framework grounded in distributionally robust optimization (DRO). Our approach defines geometric divergence measures—including KL, Fisher–Rao, and Wasserstein divergences—on the covariance manifold, constructs data-driven ambiguity sets, and automatically derives spectral shrinkage estimators by minimizing the worst-case Frobenius error. Crucially, we establish for the first time that the functional form of shrinkage is intrinsically determined by the geometric structure of the chosen divergence, thereby eliminating dependence on prior distributional assumptions. We prove asymptotic consistency of the proposed estimator and derive an $O(1/n)$ finite-sample error bound. Moreover, we obtain three explicit robust estimators, which match state-of-the-art performance on both synthetic and real-world datasets.
This work addresses the trade-off between bias control and learning efficiency in multi-source heterogeneous data settings by proposing a parameter-free, covariance-aware sequential shrinkage framework. The method adaptively constructs shrinkage directions based on the covariance structure of source data and integrates multiple sources sequentially to maximize risk reduction under limited sample sizes. It achieves, for the first time, a fully data-driven, tuning-free multi-source shrinkage estimator, establishes non-asymptotic risk bounds, and proves that the proposed strategy asymptotically attains the oracle risk. Both theoretical analysis and empirical experiments demonstrate that, particularly in highly heterogeneous scenarios, the approach significantly outperforms existing methods in terms of estimation accuracy and computational efficiency.
This study addresses the lack of theoretical justification for covariate-adaptive randomization in high-dimensional settings where the number of covariates grows with the sample size. The authors develop a unified theoretical framework to analyze the imbalance properties of two classes of adaptive randomization procedures, both for specified and unspecified covariates. Leveraging high-dimensional probabilistic limit theory and imbalance analysis, they establish—for the first time—the convergence rates of covariate imbalance under this asymptotic regime and prove the asymptotic normality of the average treatment effect estimator. Furthermore, they derive valid confidence intervals based on these results. Numerical experiments corroborate the practical relevance of the theoretical findings, thereby providing a rigorous statistical foundation for high-dimensional causal inference.
This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.