Score
Designs and implements low‑rank latent‑variable models that perform principal‑component‑style dimension reduction for mixed‑type data by modeling each observed variable with an exponential‑family distribution whose natural parameter is a linear function of shared Gaussian latent variables. Builds estimation procedures (e.g., method‑of‑moments or likelihood‑based algorithms) to recover the latent covariance and principal components and to produce sparse component loadings for interpretability.
This work addresses the challenge of performing principal component analysis on mixed-type data comprising continuous, binary, integer, and positive continuous variables. The authors propose a unified probabilistic latent variable framework in which observed variables are assumed to be generated from exponential family distributions driven by shared Gaussian latent factors. The covariance matrix of these latent variables is estimated via the method of moments, and sparsity constraints are imposed on the loading matrix to yield interpretable sparse principal components. This approach extends classical sparse PCA theory to heterogeneous data settings, seamlessly integrating principal component score estimation with sparse loading recovery. Experiments on both synthetic data and the real-world Zoo dataset demonstrate that the proposed method effectively extracts sparse principal components with clear structure and strong interpretability.
This work proposes a novel approach that integrates principal component–guided sparse regularization (pcLasso) into the reduced-rank regression framework, addressing a key limitation of existing methods which struggle to simultaneously exploit the principal component structure and group structure of predictors while effectively biasing regression coefficients toward high-variance principal component directions. By explicitly incorporating predictor principal component orientations, group information, and inter-response correlations, the proposed method overcomes constraints inherent in conventional models, enhancing both predictive accuracy and model interpretability. Extensive numerical simulations and real-data analyses demonstrate the substantial advantages of this approach in terms of prediction performance and explanatory power.
This work addresses the tractability of exact inference and learning in exponential-family latent variable models (LVMs), seeking to characterize the precise boundary of models admitting closed-form analytical solutions without approximation. Method: We derive necessary and sufficient conditions for prior–posterior conjugacy in exponential-family LVMs, providing the first systematic characterization of exact solvability. We further propose a composable graphical model construction framework that preserves structural flexibility while guaranteeing analytic tractability throughout. A general-purpose exact Bayesian inference and parameter learning algorithm is developed, accompanied by an open-source implementation supporting empirical validation across diverse models. Contribution/Results: Our results substantially broaden the class of LVMs amenable to exact inference—bypassing variational approximations or Monte Carlo sampling—and establish a rigorous theoretical foundation and practical toolkit for interpretable, high-precision latent-variable modeling.
This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.
Existing exponential family factor analysis (EFFA) frameworks for non-Gaussian, missing, and heteroscedastic matrix data suffer from restrictive distributional assumptions and asymptotic bias in simulation-based maximum likelihood (SML) estimation. Method: We propose the first quasi-likelihood-based EFFA model, explicitly incorporating dispersion parameters and element-wise weights to enhance robustness against heteroscedasticity and arbitrary missingness mechanisms. We further design an EM-SGD hybrid algorithm that eliminates SML’s asymptotic bias, achieving a theoretical error bound of O(1/p) and enabling scalable inference. Results: Extensive experiments on synthetic data and three real-world modalities—count, binary, and skewed continuous matrices—demonstrate substantial improvements in low-rank covariance structure recovery and missing value imputation accuracy over state-of-the-art baselines.
This study addresses the challenges posed by high-dimensional health monitoring data, which often exhibit multicollinearity among predictors, mixed-type responses (Gaussian, Bernoulli, and negative binomial), and latent heterogeneity, making simultaneous clustering, dimensionality reduction, and interpretable modeling difficult. To tackle this, we propose a Bayesian latent class low-rank regression model that represents the response as a finite mixture of regression surfaces, each with its own mean shift and low-rank coefficient matrix. The approach innovatively integrates latent classes, an adaptive low-rank structure, and a mixed exponential family likelihood, employing a multiplicative gamma process to automatically shrink ineffective ranks per class. Model selection for both the number of classes and the maximum rank is performed jointly via WAIC. Theoretical guarantees are provided for posterior consistency of both regression surfaces and predictor-side singular subspaces. Empirical results demonstrate accurate recovery of true cluster numbers across diverse response types, outperforming benchmarks such as K-means and mclust, and revealing interpretable individual patterns and regional clusters in three real-world health applications.
Block-structured latent variable models are widely employed in psychology, education, economics, and genetics, yet their identifiability and estimation performance have long lacked a systematic theoretical foundation. This work establishes, for the first time, identifiability conditions for such models under various block designs and introduces a Lagrangian-type nonconvex optimization framework based on constrained maximum likelihood estimation. The study derives both non-asymptotic error bounds and asymptotic distributions for the resulting estimators. The proposed algorithm enjoys strong theoretical guarantees and, as demonstrated through extensive simulations and empirical analyses, efficiently and accurately estimates latent variable models across diverse block structures.
This work proposes a novel method—nuclear-norm-penalized principal covariates regression (PcovR-nnp)—to address the challenges of high-dimensional regression, where dimensionality reduction and regularized coefficient estimation are typically performed sequentially in an ad hoc order, compromising model stability. By introducing the nuclear norm into the principal covariates regression framework for the first time, PcovR-nnp jointly optimizes matrix decomposition and regularized regression, thereby simultaneously achieving dimension selection and coefficient estimation. This integrated approach eliminates the need for subjective, stepwise procedural choices inherent in conventional pipelines and substantially enhances both the accuracy and robustness of high-dimensional regression modeling.
This work addresses the challenge of high-dimensional output regression under scarce training data, where conventional multi-output Gaussian processes (GPs) suffer from poor scalability and existing compression-prediction approaches—such as PCA-GP—rely on fixed, task-agnostic bases. To overcome these limitations, we propose Gaussian Process Latent Factor Regression (GPLFR), which models high-dimensional outputs as linear-Gaussian decodings of low-dimensional latent states governed by GP priors. By analytically marginalizing decoding weights, GPLFR enables end-to-end joint optimization of dimensionality reduction and prediction. Combining variational inference with scalable kernel methods, GPLFR substantially outperforms baseline methods like PCA-GP in global climate modeling of rocky exoplanets, achieving efficient, spatially resolved predictions in low-data, high-dimensional settings for the first time.
This study addresses the substantial bias in existing proportion of variance explained (FVE) estimators—such as those from GWAS or LMM-REML—when predictors exhibit strong correlations in high-dimensional linear models. To mitigate this issue, the authors propose a two-component FVE estimation framework based on principal component decomposition, which partitions covariates into a low-dimensional subspace of strongly correlated variables and a high-dimensional complement of weakly correlated ones, each handled with tailored estimation strategies. The resulting estimator effectively reduces bias induced by high-dimensional strong correlation structures and enjoys desirable asymptotic consistency. Extensive simulations and real-data analysis using the ABCD neuroimaging cohort demonstrate that the proposed method significantly improves FVE estimation accuracy and more reliably captures heritability signals underlying cognitive phenotypes.