Score
Designs, fits, and evaluates statistical models that represent multivariate observed variables as combinations of a small set of latent factors, estimating factor loadings, factor scores, and the number of factors. Builds and analyzes compact, interpretable low-dimensional representations to reduce input dimensionality, interpret underlying constructs, and supply features for downstream models or analyses.
Existing multidimensional factor models struggle to accommodate the typical 3–5 dimensional latent constructs in psychometrics and lack a unified framework for modeling multiparameter moderation effects. This paper proposes a scalable penalized maximum likelihood estimation method applicable to arbitrarily many factors, enabling— for the first time—the joint estimation of linear and nonlinear moderation effects within high-dimensional models. By incorporating ridge, lasso, and alignment penalties, the approach simultaneously stabilizes parameter estimation, detects partial measurement noninvariance, and enhances interpretability. Leveraging closed-form analytical gradients, the method avoids computationally intensive numerical integration and MCMC sampling, substantially improving computational efficiency. Simulation and empirical studies demonstrate accurate recovery of complex moderation patterns. The proposed method provides a scalable, efficient, and robust new tool for measurement invariance research involving multidimensional constructs.
This paper addresses the challenges of covariate effect inference and future outcome prediction in high-dimensional multivariate longitudinal data—characterized by complex dependencies among variables and over time, mixed-type outcomes (continuous and discrete), and irregular observation patterns (missingness or right-censoring). We propose a novel latent-variable modeling framework that (i) unifies the treatment of mixed-type responses and irregular temporal observations for the first time; (ii) introduces an information criterion tailored to high-dimensional longitudinal settings for automatic selection of the latent factor dimension; and (iii) establishes a rigorous central limit theorem for regression coefficient estimators, ensuring valid statistical inference. Evaluated on a customer shopping behavior prediction task, our method significantly improves long-term trend modeling accuracy and robustness of personalized forecasting, demonstrating both practical utility and theoretical soundness in real-world high-dimensional longitudinal applications.
This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.
Traditional factor analysis struggles to distinguish latent factors shared across multiple studies from study- or subgroup-specific sources of variation. To address this, we propose an adaptive multi-study joint factor model that employs a novel hierarchical shrinkage prior to induce sparsity and structural adaptivity in factor loadings. This is the first Bayesian framework to rigorously ensure identifiability of multi-study factor loadings while enabling unbiased estimation of subgroup-specific factors. The method flexibly infers hierarchical factor structures—from globally shared factors, to cross-study subgroups, down to fine-grained factors nested within individual studies. Simulation studies demonstrate estimation accuracy comparable to state-of-the-art methods, with substantially improved interpretability. Applied to avian co-occurrence and ovarian cancer gene expression datasets, the model successfully identifies robust cross-cohort biological signals and subgroup-specific driver factors.
Bayesian inference for high-dimensional covariance matrices is computationally prohibitive with conventional MCMC methods (e.g., Gibbs sampling), limiting scalability. Method: We propose FABLE, an efficient pseudo-posterior construction framework that bypasses MCMC entirely. Leveraging the “blessing of dimensionality”—a newly identified phenomenon wherein spectral structure stabilizes in high dimensions—we integrate singular value decomposition (SVD) with joint conjugate priors to construct theoretically justified, high-accuracy pseudo-posteriors. FABLE models low-rank structure via factor analysis, combines SVD-based dimension reduction with closed-form conjugate updates, and calibrates Bayesian credible intervals. Contribution/Results: We establish Wasserstein distance convergence guarantees for the pseudo-posterior. Empirical evaluation on simulated data and real gene expression datasets shows estimation accuracy comparable to MCMC, with 10–100× speedup. FABLE thus achieves both statistical reliability and computational scalability for high-dimensional covariance inference.
Block-structured latent variable models are widely employed in psychology, education, economics, and genetics, yet their identifiability and estimation performance have long lacked a systematic theoretical foundation. This work establishes, for the first time, identifiability conditions for such models under various block designs and introduces a Lagrangian-type nonconvex optimization framework based on constrained maximum likelihood estimation. The study derives both non-asymptotic error bounds and asymptotic distributions for the resulting estimators. The proposed algorithm enjoys strong theoretical guarantees and, as demonstrated through extensive simulations and empirical analyses, efficiently and accurately estimates latent variable models across diverse block structures.
研究高维多元回归模型中的结构变化检测问题,通过构建基于最小二乘的Wald统计量并扫描候选断点段落,提出适用于单个和多个断点的方法。
本文提出一种设计辅助回归框架,通过利用协变量分布信息来稳定弱设计方向和修正潜在效应扭曲,从而改进估计性能。
This study addresses the challenge of regression with high-dimensional responses and covariates when the responses are influenced by both observed covariates and unobserved latent variables, a setting where conventional multivariate regression methods fail to provide effective modeling. The authors propose a generalized latent variable model that accommodates mixed-type high-dimensional responses and allows for flexible dependence structures between covariates and latent factors. By decomposing the non-convex estimation problem into a sequence of convex subproblems through alternating optimization, and integrating debiased estimation with asymptotic normality analysis, the work achieves, for the first time, valid statistical inference on covariate effects within a high-dimensional generalized latent variable framework. The proposed estimator enjoys statistical consistency and guaranteed error bounds, while the debiased version exhibits asymptotic normality, as demonstrated empirically in an application to PISA data for assessing educational equity.
This study addresses the challenges in exploratory factor analysis arising from unknown factor structures and indeterminate numbers of latent factors, which often hinder model identification and evaluation. The authors propose a variational Bayesian variable selection framework that employs spike-and-slab priors to recover the underlying factor structure and introduces a post-selection model fit assessment system. By recasting hard and soft selection strategies as covariance models, the approach facilitates diagnostic evaluation and determination of the number of factors. A novel dimensionless gain rule, combined with multidimensional fit indices—including RMSEA, SRMR, CFI, TLI, AIC, BIC, and ELBO—is introduced to effectively prevent misidentification of factor count. Simulations demonstrate that absolute fit indices sensitively track loading recovery and detect underfactoring, while the gain rule accurately recovers the true dimensionality, with the ELBO variant exhibiting the greatest robustness. Applied to the 100-item PID-5 dataset, the method significantly outperforms a prespecified 25-factor confirmatory model.