Score
Designs and analyzes estimators and algorithms that estimate sparse precision (inverse covariance) matrices separately for predefined or learned blocks of variables, recovering per-block graphical structures such as class- or group-specific dependence graphs. Builds computationally scalable procedures that exploit a shared block partition and provides finite-sample statistical error bounds for the within-block precision-matrix estimates.
This paper addresses the challenging problem of jointly estimating high-dimensional sparse partial correlation and inverse covariance matrices. We propose a two-stage joint partial regression method: within a per-variable linear regression framework, we first embed both sparsity-inducing regularization and positive definite cone projection into a unified optimization objective, thereby simultaneously ensuring the positive definiteness and sparsity of the estimated matrix. The method is efficiently implemented via a proximal splitting algorithm. We establish theoretical guarantees showing that the inverse covariance estimator achieves the optimal convergence rate, while the partial correlation estimator attains a strictly tighter error bound than existing state-of-the-art methods. Extensive experiments on synthetic and real-world high-dimensional datasets demonstrate that our approach significantly outperforms mainstream competitors—including Graphical Lasso and CLIME—in both estimation accuracy and computational efficiency. This work provides a new paradigm for high-dimensional graphical model learning that combines strong theoretical foundations with practical effectiveness.
This work addresses the joint estimation of multiple precision matrices in high-dimensional settings where they share a common sparsity pattern yet exhibit heterogeneous edge strengths. The authors propose Multiplicative Graphical Lasso (Mglasso), which decomposes each precision matrix as the Schur–Hadamard product of a shared structural matrix and a group-specific strength matrix. This decomposition uniquely disentangles graph topology from edge intensities, enabling simultaneous learning of structural consistency and strength heterogeneity within a unified framework that combines ℓ₁ and Frobenius norm penalties. The resulting penalized log-likelihood is efficiently optimized via an alternating direction method of multipliers (ADMM) coupled with gradient descent. Theoretical analysis establishes high-dimensional consistency and exact support recovery, while experiments demonstrate that Mglasso significantly outperforms Group Graphical Lasso in small-sample regimes, achieving superior model selection consistency and practical utility.
This work addresses the problem of community detection in the two-community stochastic block model with multiple independent graph samples. It proposes averaging the adjacency matrices across samples followed by spectral partitioning to suppress noise and enhance recovery accuracy. Theoretically, the study establishes a spectral norm bound for the aggregated noise matrix under multiple samples and proves that the probability of community recovery error decays exponentially with the number of samples, providing the first rigorous guarantee for graph data augmentation in this setting. The analysis combines matrix averaging, a simplified spectral algorithm, and tools from the Davis–Kahan theorem and random matrix theory. Experiments on graphs with up to 1,000 nodes and as few as 2–3 samples demonstrate the tightness of the theoretical bounds and confirm that even a small number of samples significantly improves detection performance.
This paper addresses joint testing of the mean vector and covariance matrix under a uniform block structure in high-dimensional data with missing observations. Method: We develop the first unified statistical inference framework accommodating missing data, introducing a novel block-wise Hadamard product representation for uniformly structured block matrices. This enables closed-form expressions for the likelihood ratio and information statistics, along with their exact null distributions. We further propose an FDP-controlled simultaneous marginal mean testing procedure. Contribution/Results: Theoretical analysis and extensive simulations demonstrate accurate distributional characterization of test statistics, robust and reliable FDP control, and strong robustness against perturbations in the covariance structure and arbitrary missingness mechanisms. The method is successfully applied to hypothesis testing in high-dimensional neuroimaging data, substantially broadening the practical applicability of block-structured covariance models in high-dimensional inference.
Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.
Block-structured latent variable models are widely employed in psychology, education, economics, and genetics, yet their identifiability and estimation performance have long lacked a systematic theoretical foundation. This work establishes, for the first time, identifiability conditions for such models under various block designs and introduces a Lagrangian-type nonconvex optimization framework based on constrained maximum likelihood estimation. The study derives both non-asymptotic error bounds and asymptotic distributions for the resulting estimators. The proposed algorithm enjoys strong theoretical guarantees and, as demonstrated through extensive simulations and empirical analyses, efficiently and accurately estimates latent variable models across diverse block structures.
This study addresses the challenges in inferring conditional dependence structures of high-dimensional stationary time series in the frequency domain, which are hindered by truncation and smoothing biases arising from finite-sample discrete Fourier transforms (DFTs) and the difficulty of estimating complex-valued spectral precision matrices in high dimensions. To overcome these issues, the authors propose a debiased complex-valued graphical Lasso estimator based on the full likelihood across neighboring frequency points. The method enables high-dimensional inference of sparse spectral precision matrices at fixed frequencies and achieves entrywise consistent covariance estimation through cross-frequency information aggregation. It constitutes the first approach to perform full-likelihood inference directly on the DFT, effectively controlling regularization, truncation, and smoothing biases. Simulations demonstrate accurate confidence interval coverage across all non-zero frequencies, higher statistical power than existing methods, and false discovery rates close to nominal levels.
This work addresses the computational inefficiency in estimating precision matrices for high-dimensional fully positive Gaussian graphical models by proposing a generalized Bayesian framework based on the D-trace loss, which circumvents costly log-determinant computations. To induce sparsity, a spike-and-slab prior is incorporated. Three key techniques are introduced to enhance scalability and sampling efficiency in high dimensions: conditional independence–driven data augmentation, a fast matrix-normal sampler leveraging the Gram structure of the sample covariance, and an intertwining strategy that combines augmented and direct updates. Experimental results demonstrate that the proposed method achieves substantial computational acceleration on both synthetic and financial datasets while maintaining high estimation accuracy and superior graph structure recovery.
This study addresses the poor performance of precision matrix estimation in high-dimensional graphical models when edge weights exhibit clustering structures and data are heavy-tailed. To overcome this limitation, the authors propose Graphical SLOPE, a method that employs SLOPE regularization to simultaneously recover sparsity and cluster edges with similar strengths. For the first time, they establish asymptotic theory for estimation error and pattern recovery under both Gaussian and elliptical distributions, elucidating its clustering selection mechanism. Furthermore, they introduce TSLOPE, which integrates a multivariate t-loss function to robustly handle heavy-tailed data. Theoretical analysis and empirical experiments demonstrate that Graphical SLOPE yields more accurate estimates under structured edge patterns, while TSLOPE substantially outperforms conventional Gaussian-based approaches and effectively identifies economically meaningful dependency clusters.
This work addresses the limitations of high-dimensional covariance matrix estimation, which often stems from restrictive invariance assumptions or a lack of theoretical optimality. The authors propose a hierarchical Bayesian framework leveraging the orthogonal group $O(p)$-equivariant structure and establish, for the first time, that the minimum risk estimator within the $O(p)$-equivariant class coincides with the Oracle Bayes rule under the Haar measure. They flexibly model the unknown eigenvalue distribution using a finite Pólya tree prior and implement posterior inference via Gibbs sampling. The resulting shrinkage estimator asymptotically approaches theoretical optimality under multiple loss functions. Simulations demonstrate that the method accurately recovers the spectral shape of eigenvalues and substantially outperforms classical approaches such as Ledoit–Wolf and Haff estimators, achieving performance close to that of the Oracle.