blockwise sparse precision estimation

Designs and analyzes estimators and algorithms that estimate sparse precision (inverse covariance) matrices separately for predefined or learned blocks of variables, recovering per-block graphical structures such as class- or group-specific dependence graphs. Builds computationally scalable procedures that exploit a shared block partition and provides finite-sample statistical error bounds for the within-block precision-matrix estimates.

blockwisesparseprecisionestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Sparse Estimation of Inverse Covariance and Partial Correlation Matrices via Joint Partial Regression

Feb 12, 2025
SE
Samuel Erickson
🏛️ KTH Royal Institute of Technology | Lynx Asset Management AB

This paper addresses the challenging problem of jointly estimating high-dimensional sparse partial correlation and inverse covariance matrices. We propose a two-stage joint partial regression method: within a per-variable linear regression framework, we first embed both sparsity-inducing regularization and positive definite cone projection into a unified optimization objective, thereby simultaneously ensuring the positive definiteness and sparsity of the estimated matrix. The method is efficiently implemented via a proximal splitting algorithm. We establish theoretical guarantees showing that the inverse covariance estimator achieves the optimal convergence rate, while the partial correlation estimator attains a strictly tighter error bound than existing state-of-the-art methods. Extensive experiments on synthetic and real-world high-dimensional datasets demonstrate that our approach significantly outperforms mainstream competitors—including Graphical Lasso and CLIME—in both estimation accuracy and computational efficiency. This work provides a new paradigm for high-dimensional graphical model learning that combines strong theoretical foundations with practical effectiveness.

Enhances partial correlation matrix estimationEstimates sparse inverse covariance matricesProposes efficient proximal splitting algorithm

This work addresses the joint estimation of multiple precision matrices in high-dimensional settings where they share a common sparsity pattern yet exhibit heterogeneous edge strengths. The authors propose Multiplicative Graphical Lasso (Mglasso), which decomposes each precision matrix as the Schur–Hadamard product of a shared structural matrix and a group-specific strength matrix. This decomposition uniquely disentangles graph topology from edge intensities, enabling simultaneous learning of structural consistency and strength heterogeneity within a unified framework that combines ℓ₁ and Frobenius norm penalties. The resulting penalized log-likelihood is efficiently optimized via an alternating direction method of multipliers (ADMM) coupled with gradient descent. Theoretical analysis establishes high-dimensional consistency and exact support recovery, while experiments demonstrate that Mglasso significantly outperforms Group Graphical Lasso in small-sample regimes, achieving superior model selection consistency and practical utility.

Gaussian graphical modelheterogeneous edge strengthshigh-dimensional estimation

This work addresses the problem of community detection in the two-community stochastic block model with multiple independent graph samples. It proposes averaging the adjacency matrices across samples followed by spectral partitioning to suppress noise and enhance recovery accuracy. Theoretically, the study establishes a spectral norm bound for the aggregated noise matrix under multiple samples and proves that the probability of community recovery error decays exponentially with the number of samples, providing the first rigorous guarantee for graph data augmentation in this setting. The analysis combines matrix averaging, a simplified spectral algorithm, and tools from the Davis–Kahan theorem and random matrix theory. Experiments on graphs with up to 1,000 nodes and as few as 2–3 samples demonstrate the tightness of the theoretical bounds and confirm that even a small number of samples significantly improves detection performance.

community detectiongraph data augmentationmultiple graph samples

Hypothesis testing under uniform-block covariance structures

Apr 17, 2023
YY
Yifan Yang
🏛️ Case Western Reserve University | University of Maryland

This paper addresses joint testing of the mean vector and covariance matrix under a uniform block structure in high-dimensional data with missing observations. Method: We develop the first unified statistical inference framework accommodating missing data, introducing a novel block-wise Hadamard product representation for uniformly structured block matrices. This enables closed-form expressions for the likelihood ratio and information statistics, along with their exact null distributions. We further propose an FDP-controlled simultaneous marginal mean testing procedure. Contribution/Results: Theoretical analysis and extensive simulations demonstrate accurate distributional characterization of test statistics, robust and reliable FDP control, and strong robustness against perturbations in the covariance structure and arbitrary missingness mechanisms. The method is successfully applied to hypothesis testing in high-dimensional neuroimaging data, substantially broadening the practical applicability of block-structured covariance models in high-dimensional inference.

Developing joint tests for covariance and mean structures with block matricesHypothesis testing for uniform-block covariance structures in high-dimensional dataValidating test statistics and FDP control in simulations and real data

Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing

Aug 28, 2023
YZ
Yihan Zhang
🏛️ Institute of Science and Technology Austria | University of Cambridge

Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.

Characterizing spectral estimators for correlated Gaussian designsEstimating parameters in high-dimensional generalized linear modelsIdentifying optimal preprocessing for efficient parameter estimation

Latest Papers

What's happening recently
View more

Block-structured latent variable models are widely employed in psychology, education, economics, and genetics, yet their identifiability and estimation performance have long lacked a systematic theoretical foundation. This work establishes, for the first time, identifiability conditions for such models under various block designs and introduces a Lagrangian-type nonconvex optimization framework based on constrained maximum likelihood estimation. The study derives both non-asymptotic error bounds and asymptotic distributions for the resulting estimators. The proposed algorithm enjoys strong theoretical guarantees and, as demonstrated through extensive simulations and empirical analyses, efficiently and accurately estimates latent variable models across diverse block structures.

block structured latent variable modelsidentifiabilitymaximum likelihood estimation

This study addresses the challenges in inferring conditional dependence structures of high-dimensional stationary time series in the frequency domain, which are hindered by truncation and smoothing biases arising from finite-sample discrete Fourier transforms (DFTs) and the difficulty of estimating complex-valued spectral precision matrices in high dimensions. To overcome these issues, the authors propose a debiased complex-valued graphical Lasso estimator based on the full likelihood across neighboring frequency points. The method enables high-dimensional inference of sparse spectral precision matrices at fixed frequencies and achieves entrywise consistent covariance estimation through cross-frequency information aggregation. It constitutes the first approach to perform full-likelihood inference directly on the DFT, effectively controlling regularization, truncation, and smoothing biases. Simulations demonstrate accurate confidence interval coverage across all non-zero frequencies, higher statistical power than existing methods, and false discovery rates close to nominal levels.

conditional dependenceGaussian graphical modelshigh-dimensional inference

This work addresses the computational inefficiency in estimating precision matrices for high-dimensional fully positive Gaussian graphical models by proposing a generalized Bayesian framework based on the D-trace loss, which circumvents costly log-determinant computations. To induce sparsity, a spike-and-slab prior is incorporated. Three key techniques are introduced to enhance scalability and sampling efficiency in high dimensions: conditional independence–driven data augmentation, a fast matrix-normal sampler leveraging the Gram structure of the sample covariance, and an intertwining strategy that combines augmented and direct updates. Experimental results demonstrate that the proposed method achieves substantial computational acceleration on both synthetic and financial datasets while maintaining high estimation accuracy and superior graph structure recovery.

Bayesian inferenceGaussian graphical modelshigh-dimensional statistics

This study addresses the poor performance of precision matrix estimation in high-dimensional graphical models when edge weights exhibit clustering structures and data are heavy-tailed. To overcome this limitation, the authors propose Graphical SLOPE, a method that employs SLOPE regularization to simultaneously recover sparsity and cluster edges with similar strengths. For the first time, they establish asymptotic theory for estimation error and pattern recovery under both Gaussian and elliptical distributions, elucidating its clustering selection mechanism. Furthermore, they introduce TSLOPE, which integrates a multivariate t-loss function to robustly handle heavy-tailed data. Theoretical analysis and empirical experiments demonstrate that Graphical SLOPE yields more accurate estimates under structured edge patterns, while TSLOPE substantially outperforms conventional Gaussian-based approaches and effectively identifies economically meaningful dependency clusters.

edge clusteringheavy-tailed distributionsnon-Gaussian data

This work addresses the limitations of high-dimensional covariance matrix estimation, which often stems from restrictive invariance assumptions or a lack of theoretical optimality. The authors propose a hierarchical Bayesian framework leveraging the orthogonal group $O(p)$-equivariant structure and establish, for the first time, that the minimum risk estimator within the $O(p)$-equivariant class coincides with the Oracle Bayes rule under the Haar measure. They flexibly model the unknown eigenvalue distribution using a finite Pólya tree prior and implement posterior inference via Gibbs sampling. The resulting shrinkage estimator asymptotically approaches theoretical optimality under multiple loss functions. Simulations demonstrate that the method accurately recovers the spectral shape of eigenvalues and substantially outperforms classical approaches such as Ledoit–Wolf and Haff estimators, achieving performance close to that of the Oracle.

covariance matrix estimationeigenvalue distributionequivariance

Hot Scholars

YX

Yanxun Xu

Johns Hopkins University
BayesianClinical trial DesignElectronic Health Record DataNetwork Data
LZ

Lunbin Zeng

Huazhong University of Science and Technology
compute vision
PZ

Pengyu Zhao

Peking University
Neural Architecture SearchRecommender System360-degree Video
XL

Xunhao Lai

Peking University
Machine LearningNatural Language ProcessingLarge language model
BF

Bernardo Flores

PhD Candidate, University of Texas at Austin
Bayesian statisticsBayesian nonparametricsbiostatistics