build latent factor models

Designs and implements probabilistic latent-factor models that represent high-dimensional observed variables using a compact set of latent factors, covering static factorization, factor-graph formulations, and dynamic low-rank state‑space models. This competence includes constructing factor graphs and potentials, parameterizing factor loadings and priors, specifying temporal dynamics (e.g., AR processes), decomposing factors into components, and building inference procedures (message passing, marginalization) to link latent factors to observed outcomes.

buildlatentfactormodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$191K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the instability of principal component-based estimation in high-dimensional approximate factor models, where the number of variables far exceeds the sample size, leading to severe distortion in the eigenstructure of the sample covariance matrix. The paper presents the first systematic application of a Bayesian framework to this problem, enabling joint posterior inference of factor loadings and latent factors. Theoretically, the proposed approach achieves a posterior contraction rate comparable to the benchmark established for high-dimensional spiked covariance models. Simulation studies demonstrate that the method more accurately recovers the underlying factor structure. In empirical applications to macro-financial data, the estimated factors exhibit clear economic interpretability and significantly outperform existing methods in predictive tasks.

approximate factor modelseigenstructurehigh-dimensional

This study addresses the challenges of estimating the number of latent factors and learning the mapping structure in nonlinear latent factor models. Methodologically, it constructs a graph representation based on pairwise dependencies among observed variables and proposes a dependency thresholding algorithm to jointly identify the number of latent factors and the nonlinear mapping structure. A structure-constrained neural network is further designed for efficient optimization. The core contribution lies in overcoming traditional linearity assumptions by establishing an identifiability and consistency theoretical framework grounded in general dependence measures. Simulation experiments demonstrate that the proposed algorithm achieves superior accuracy and robustness in high-dimensional settings, effectively recovering the underlying nonlinear functions.

dependence measuresidentifiabilitylatent factor models

Optimal discriminant analysis in high-dimensional latent factor models

Oct 23, 2022
XB
Xin Bing
🏛️ University of Toronto | Cornell University

This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.

Analyze convergence rates of excess risk in classificationDevelop efficient classifier for high-dimensional latent factor modelsSelect optimal principal components in data-driven projection

A Latent Variable Approach to Learning High-dimensional Multivariate longitudinal Data

May 23, 2024
SM
Sze Ming Lee
🏛️ London School of Economics and Political Science | The Chinese University of Hong Kong

This paper addresses the challenges of covariate effect inference and future outcome prediction in high-dimensional multivariate longitudinal data—characterized by complex dependencies among variables and over time, mixed-type outcomes (continuous and discrete), and irregular observation patterns (missingness or right-censoring). We propose a novel latent-variable modeling framework that (i) unifies the treatment of mixed-type responses and irregular temporal observations for the first time; (ii) introduces an information criterion tailored to high-dimensional longitudinal settings for automatic selection of the latent factor dimension; and (iii) establishes a rigorous central limit theorem for regression coefficient estimators, ensuring valid statistical inference. Evaluated on a customer shopping behavior prediction task, our method significantly improves long-term trend modeling accuracy and robustness of personalized forecasting, demonstrating both practical utility and theoretical soundness in real-world high-dimensional longitudinal applications.

Analyzing covariate effects using latent variable approachModeling high-dimensional multivariate longitudinal data dependenciesPredicting future outcomes with mixed-type and missing data

Bayesian inference for high-dimensional covariance matrices is computationally prohibitive with conventional MCMC methods (e.g., Gibbs sampling), limiting scalability. Method: We propose FABLE, an efficient pseudo-posterior construction framework that bypasses MCMC entirely. Leveraging the “blessing of dimensionality”—a newly identified phenomenon wherein spectral structure stabilizes in high dimensions—we integrate singular value decomposition (SVD) with joint conjugate priors to construct theoretically justified, high-accuracy pseudo-posteriors. FABLE models low-rank structure via factor analysis, combines SVD-based dimension reduction with closed-form conjugate updates, and calibrates Bayesian credible intervals. Contribution/Results: We establish Wasserstein distance convergence guarantees for the pseudo-posterior. Empirical evaluation on simulated data and real gene expression datasets shows estimation accuracy comparable to MCMC, with 10–100× speedup. FABLE thus achieves both statistical reliability and computational scalability for high-dimensional covariance inference.

Develops accurate covariance estimation without Gibbs samplingEnables efficient high-dimensional factor analysis computationImproves pseudo-posterior accuracy with higher dimensionality

Latest Papers

What's happening recently
View more

This study addresses the challenge of identifying direct causal effects among observed variables in densely confounded linear structural equation models with latent variables, where conventional methods often fail. The authors propose a novel identification criterion that explicitly models latent variables, employs a recursive identification strategy, and systematically handles unidentified causal parents. By transforming the combinatorial search problem into an efficient network flow computation, the method substantially enhances the identifiability of direct causal effects in dense confounding settings. Accompanied by an open-source algorithmic implementation, this approach combines theoretical rigor with practical utility for causal inference in complex observational data.

causal identificationconfoundingdirect causal effects

This study addresses the poor interpretability of multivariate factor models caused by rotational invariance, as well as the complex Bayesian implementation and challenging prior specification associated with generalized lower triangular structures. To overcome these limitations, this work proposes a generalized lower triangular process prior that supports dual sparsity at both global and within-component levels. The proposed approach substantially simplifies prior hyperparameter specification and establishes a theoretical connection to the Indian Buffet Process. Furthermore, an efficient Gibbs sampler incorporating Metropolis-Hastings steps is designed for posterior inference. Simulation studies and an empirical analysis of Big Five personality data demonstrate that the method automatically infers the number of factors while discovering sparse structures, thereby significantly enhancing model interpretability and practical utility.

Bayesian implementationfactor modelsgeneralized lower triangular structure

本文开发了一类针对矩阵时间序列数据的线性状态空间模型,通过矩阵版本的卡尔曼滤波等方法估计潜在状态矩阵和参数,适用于处理混合频率、异方差性和异常值问题。

heteroskedasticitylatent matrix normal processmatrix-valued time series

This study addresses the challenges in exploratory factor analysis arising from unknown factor structures and indeterminate numbers of latent factors, which often hinder model identification and evaluation. The authors propose a variational Bayesian variable selection framework that employs spike-and-slab priors to recover the underlying factor structure and introduces a post-selection model fit assessment system. By recasting hard and soft selection strategies as covariance models, the approach facilitates diagnostic evaluation and determination of the number of factors. A novel dimensionless gain rule, combined with multidimensional fit indices—including RMSEA, SRMR, CFI, TLI, AIC, BIC, and ELBO—is introduced to effectively prevent misidentification of factor count. Simulations demonstrate that absolute fit indices sensitively track loading recovery and detect underfactoring, while the gain rule accurately recovers the true dimensionality, with the ELBO variant exhibiting the greatest robustness. Applied to the 100-item PID-5 dataset, the method significantly outperforms a prespecified 25-factor confirmatory model.

Bayesian variable selectionfactor number selectionlatent structure recovery

Hot Scholars

HC

Haeran Cho

University of Bristol
change-point detectionnonstationary time series analysishigh-dimensional data analysisenergy data modelling
HV

Holger Voos

University of Luxembourg, SnT Automation & Robotics Research Group
Control EngineeringAutomationMobile Robotics
JB

Jushan Bai

Columbia University
EconometricsFactor ModelsInteractive EffectsStructural Changes
MW

Martin Weidner

Department of Economics & Nuffield College, University of Oxford
Econometrics
YY

Yanrong Yang

Australian National University
High Dimensional Statistics