Score
Designs and implements probabilistic latent-factor models that represent high-dimensional observed variables using a compact set of latent factors, covering static factorization, factor-graph formulations, and dynamic low-rank state‑space models. This competence includes constructing factor graphs and potentials, parameterizing factor loadings and priors, specifying temporal dynamics (e.g., AR processes), decomposing factors into components, and building inference procedures (message passing, marginalization) to link latent factors to observed outcomes.
This work addresses the absence of a unified and reproducible platform for Bayesian factor models, which has hindered fair comparisons and efficient implementation in large-scale, complex datasets. To bridge this gap, the study introduces a standardized, modular computational framework that enables seamless integration and direct comparison of diverse modern Bayesian factor models, including those with random effects. Built upon this framework, the authors develop and publicly release the factorverse R package, substantially enhancing model deployment efficiency and reproducibility. The proposed framework not only facilitates rigorous empirical evaluation across modeling approaches but also offers practical utility for both applied research and educational purposes.
This study addresses the instability of principal component-based estimation in high-dimensional approximate factor models, where the number of variables far exceeds the sample size, leading to severe distortion in the eigenstructure of the sample covariance matrix. The paper presents the first systematic application of a Bayesian framework to this problem, enabling joint posterior inference of factor loadings and latent factors. Theoretically, the proposed approach achieves a posterior contraction rate comparable to the benchmark established for high-dimensional spiked covariance models. Simulation studies demonstrate that the method more accurately recovers the underlying factor structure. In empirical applications to macro-financial data, the estimated factors exhibit clear economic interpretability and significantly outperform existing methods in predictive tasks.
This study addresses the challenges of estimating the number of latent factors and learning the mapping structure in nonlinear latent factor models. Methodologically, it constructs a graph representation based on pairwise dependencies among observed variables and proposes a dependency thresholding algorithm to jointly identify the number of latent factors and the nonlinear mapping structure. A structure-constrained neural network is further designed for efficient optimization. The core contribution lies in overcoming traditional linearity assumptions by establishing an identifiability and consistency theoretical framework grounded in general dependence measures. Simulation experiments demonstrate that the proposed algorithm achieves superior accuracy and robustness in high-dimensional settings, effectively recovering the underlying nonlinear functions.
This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.
This paper addresses the challenges of covariate effect inference and future outcome prediction in high-dimensional multivariate longitudinal data—characterized by complex dependencies among variables and over time, mixed-type outcomes (continuous and discrete), and irregular observation patterns (missingness or right-censoring). We propose a novel latent-variable modeling framework that (i) unifies the treatment of mixed-type responses and irregular temporal observations for the first time; (ii) introduces an information criterion tailored to high-dimensional longitudinal settings for automatic selection of the latent factor dimension; and (iii) establishes a rigorous central limit theorem for regression coefficient estimators, ensuring valid statistical inference. Evaluated on a customer shopping behavior prediction task, our method significantly improves long-term trend modeling accuracy and robustness of personalized forecasting, demonstrating both practical utility and theoretical soundness in real-world high-dimensional longitudinal applications.
Bayesian inference for high-dimensional covariance matrices is computationally prohibitive with conventional MCMC methods (e.g., Gibbs sampling), limiting scalability. Method: We propose FABLE, an efficient pseudo-posterior construction framework that bypasses MCMC entirely. Leveraging the “blessing of dimensionality”—a newly identified phenomenon wherein spectral structure stabilizes in high dimensions—we integrate singular value decomposition (SVD) with joint conjugate priors to construct theoretically justified, high-accuracy pseudo-posteriors. FABLE models low-rank structure via factor analysis, combines SVD-based dimension reduction with closed-form conjugate updates, and calibrates Bayesian credible intervals. Contribution/Results: We establish Wasserstein distance convergence guarantees for the pseudo-posterior. Empirical evaluation on simulated data and real gene expression datasets shows estimation accuracy comparable to MCMC, with 10–100× speedup. FABLE thus achieves both statistical reliability and computational scalability for high-dimensional covariance inference.
This study addresses the challenge of identifying direct causal effects among observed variables in densely confounded linear structural equation models with latent variables, where conventional methods often fail. The authors propose a novel identification criterion that explicitly models latent variables, employs a recursive identification strategy, and systematically handles unidentified causal parents. By transforming the combinatorial search problem into an efficient network flow computation, the method substantially enhances the identifiability of direct causal effects in dense confounding settings. Accompanied by an open-source algorithmic implementation, this approach combines theoretical rigor with practical utility for causal inference in complex observational data.
This study addresses the poor interpretability of multivariate factor models caused by rotational invariance, as well as the complex Bayesian implementation and challenging prior specification associated with generalized lower triangular structures. To overcome these limitations, this work proposes a generalized lower triangular process prior that supports dual sparsity at both global and within-component levels. The proposed approach substantially simplifies prior hyperparameter specification and establishes a theoretical connection to the Indian Buffet Process. Furthermore, an efficient Gibbs sampler incorporating Metropolis-Hastings steps is designed for posterior inference. Simulation studies and an empirical analysis of Big Five personality data demonstrate that the method automatically infers the number of factors while discovering sparse structures, thereby significantly enhancing model interpretability and practical utility.
本文开发了一类针对矩阵时间序列数据的线性状态空间模型,通过矩阵版本的卡尔曼滤波等方法估计潜在状态矩阵和参数,适用于处理混合频率、异方差性和异常值问题。
本文针对存在潜在混淆变量时因果关系发现的问题,提出了一种通过恢复观测变量的精度矩阵为稀疏加低秩矩阵的方法来重建观测变量间的有向无环图。
This study addresses the challenges in exploratory factor analysis arising from unknown factor structures and indeterminate numbers of latent factors, which often hinder model identification and evaluation. The authors propose a variational Bayesian variable selection framework that employs spike-and-slab priors to recover the underlying factor structure and introduces a post-selection model fit assessment system. By recasting hard and soft selection strategies as covariance models, the approach facilitates diagnostic evaluation and determination of the number of factors. A novel dimensionless gain rule, combined with multidimensional fit indices—including RMSEA, SRMR, CFI, TLI, AIC, BIC, and ELBO—is introduced to effectively prevent misidentification of factor count. Simulations demonstrate that absolute fit indices sensitively track loading recovery and detect underfactoring, while the gain rule accurately recovers the true dimensionality, with the ELBO variant exhibiting the greatest robustness. Applied to the 100-item PID-5 dataset, the method significantly outperforms a prespecified 25-factor confirmatory model.