latent variable regression

Designs and estimates regression models that represent and use unobserved (latent) variables: build a measurement component that maps observed indicators (continuous, ordinal, or categorical) onto latent traits and a structural component that regresses outcomes or predictors on those latent constructs. Produces estimated latent-variable scores and interpretable regression coefficients quantifying relationships between latent constructs and observed variables.

latentvariableregression

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Two-step estimation of latent trait models

Mar 28, 2023
JK
J. Kuha
🏛️ London School of Economics and Political Science | Leiden University

To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.

Evaluating performance compared to one-step and three-step methodsExamining properties through simulation studies and applicationsTwo-step estimation for latent trait models

This study addresses a critical yet previously unrecognized issue in observational causal inference: measurement-induced confounding, wherein latent variables—such as motivation or self-efficacy—are imperfectly measured, leading to biased estimates of adjusted causal effects. The authors formally identify and name this problem, moving beyond conventional two-stage adjustment approaches. They propose a novel Bayesian joint estimation framework that simultaneously models the latent variable’s measurement structure, the treatment assignment mechanism, and the potential outcomes model. This integrated approach effectively corrects bias in average treatment effect estimation and restores the nominal coverage of uncertainty intervals, thereby substantially enhancing the reliability of causal inferences drawn from observational data with error-prone proxies for unobserved confounders.

average treatment effectcausal inferencelatent confounding

A Latent Variable Approach to Learning High-dimensional Multivariate longitudinal Data

May 23, 2024
SM
Sze Ming Lee
🏛️ London School of Economics and Political Science | The Chinese University of Hong Kong

This paper addresses the challenges of covariate effect inference and future outcome prediction in high-dimensional multivariate longitudinal data—characterized by complex dependencies among variables and over time, mixed-type outcomes (continuous and discrete), and irregular observation patterns (missingness or right-censoring). We propose a novel latent-variable modeling framework that (i) unifies the treatment of mixed-type responses and irregular temporal observations for the first time; (ii) introduces an information criterion tailored to high-dimensional longitudinal settings for automatic selection of the latent factor dimension; and (iii) establishes a rigorous central limit theorem for regression coefficient estimators, ensuring valid statistical inference. Evaluated on a customer shopping behavior prediction task, our method significantly improves long-term trend modeling accuracy and robustness of personalized forecasting, demonstrating both practical utility and theoretical soundness in real-world high-dimensional longitudinal applications.

Analyzing covariate effects using latent variable approachModeling high-dimensional multivariate longitudinal data dependenciesPredicting future outcomes with mixed-type and missing data

Attenuation Bias with Latent Predictors

Jul 29, 2025
CT
Connor T. Jerzak
🏛️ University of Texas at Austin

In social science research, measurement error in latent variables induces attenuation bias—regression coefficients biased toward zero—yet existing correction methods often ignore interactions between such errors and identification constraints on latent variables, sometimes exacerbating bias. This paper identifies the underlying mechanism and proposes a novel coefficient correction method that jointly models latent-variable identification constraints and measurement-error structure, enabling simultaneous adjustment of regression coefficients during estimation. The approach imposes no strong distributional or functional-form assumptions and is compatible with diverse latent-variable estimation strategies (e.g., CFA, SEM, Bayesian latent-variable modeling). Empirical evaluations demonstrate that corrected coefficients increase by 30–50% on average relative to naive OLS estimates, substantially outperforming both uncorrected regression and mainstream error-correction techniques (e.g., regression calibration, SIMEX). The method effectively recovers true effect magnitudes and enhances the validity of causal inference in latent-variable contexts.

Addressing attenuation bias in latent variable regression modelsCorrecting measurement error impact on latent predictor coefficientsImproving accuracy of estimates for unobservable theoretical constructs

Estimating the variance-covariance matrix of two-step estimates of latent variable models: A general simulation-based approach

Jul 22, 2025
RD
Roberto Di Mari
🏛️ University of Catania | London School of Economics and Political Science

This paper addresses the challenge of analytically computing the variance–covariance matrix of structural parameters in two-step estimation of latent variable models. We propose a general, asymptotically consistent simulation-based estimator. The method repeatedly draws simulated values from the sampling distribution of measurement parameters estimated in the first step and substitutes them into the second-step estimation to directly quantify the variability of structural parameters—thereby avoiding error-prone analytical differentiation of cross-derivative matrices required by conventional approaches. It is particularly well-suited for latent variable models with categorical observed indicators. Simulation studies and empirical analyses of two distinct model classes demonstrate that the proposed method exhibits excellent finite-sample statistical properties, computational efficiency, and robustness. Consequently, it substantially enhances the feasibility and reliability of statistical inference in two-step estimation frameworks.

Applying method to categorical and continuous latent variablesAvoiding complex cross-derivative matrix evaluationEstimating variance-covariance matrix for latent variable models

Latest Papers

What's happening recently
View more

This study investigates the mechanism through which academic performance—an ordinal variable—influences self-efficacy, a continuous outcome. To this end, the authors propose a conditional Bayesian modeling framework that introduces a latent academic achievement variable and integrates Gaussian copula regression with Bayesian variable selection to identify key covariates specific to each outcome type. Methodologically, they develop a tailored partially collapsed Gibbs sampler that substantially enhances computational efficiency in estimating integrated regression coefficients and improves the accuracy of variable selection. Simulation studies demonstrate that the proposed approach markedly outperforms existing joint modeling strategies in both sampling efficiency and variable selection performance. Application to data from the Longitudinal Study of Australian Children reveals distinct association pathways between academic achievement and self-efficacy, along with markedly different covariate structures for the two outcomes.

academic performancecontinuous scalelatent academic achievement

This study addresses the challenge of identifying direct causal effects among observed variables in densely confounded linear structural equation models with latent variables, where conventional methods often fail. The authors propose a novel identification criterion that explicitly models latent variables, employs a recursive identification strategy, and systematically handles unidentified causal parents. By transforming the combinatorial search problem into an efficient network flow computation, the method substantially enhances the identifiability of direct causal effects in dense confounding settings. Accompanied by an open-source algorithmic implementation, this approach combines theoretical rigor with practical utility for causal inference in complex observational data.

causal identificationconfoundingdirect causal effects

This study addresses identification bias in latent variable regression coefficients arising from multi-source nonlinear measurement error—exemplified by divergent measures of occupational exposure to artificial intelligence—by proposing a partial identification approach based on curvature constraints. Assuming a linear consensus measurement function and bounding heterogeneity in the curvature of individual measurement sources relative to the slope, the method constructs closed-form identification intervals that are invariant to unknown measurement loadings. These intervals exhibit sharpness, with half-widths that are second-order small relative to the curvature bounds. The curvature bounds are estimated via split-sample instrumental variable techniques, and inference with uniform coverage is achieved by combining Imbens–Manski confidence intervals with Stoye critical values. Applied to 8.88 million person-years of U.S. community survey data, the approach yields a consensus coefficient of −0.239 across five AI exposure measures, with a partial identification half-width amounting to only 1.23% of the point estimate.

latent regressormeasurement errornonlinear measurements

This study addresses a critical limitation in conventional two-stage approaches that link individual-level distributional characteristics—such as variability and skewness—to downstream outcomes, which ignore estimation error in the first stage and consequently yield biased estimates and inflated Type I error rates. To overcome this, the authors propose the Distributional Feature Latent Variable Model (DFLVM), which, for the first time, integrates distributional features into a latent variable framework. DFLVM captures between-individual heterogeneity through random intercepts and jointly models both the distributional features and their effects on outcomes within a single-step maximum likelihood estimation procedure. This unified approach circumvents the inherent bias of two-stage methods. Simulation studies and empirical analyses demonstrate that DFLVM substantially reduces estimation bias and false positive rates while enhancing inferential accuracy.

distributional featuresestimation errorlatent variable models

本文提出一种设计辅助回归框架,通过利用协变量分布信息来稳定弱设计方向和修正潜在效应扭曲,从而改进估计性能。

covariatesdesign informationestimation

Hot Scholars

YC

Yunxiao Chen

Department of Statistics, London School of Economics and Political Science
Multivariate StatisticsPsychometrics
BH

Biwei Huang

UCSD
CausalityMachine LearningComputational Science
YZ

Yujia Zheng

Carnegie Mellon University
Machine LearningCausal Discovery and InferenceLatent Variable ModelsGenerative Models
YL

Yuhang Liu

The University of Adelaide
Representation LearningLLMsLatent Variable ModelsResponsible AI
GX

Gongjun Xu

University of Michigan
StatisticsMachine LearningPsychometrics