Score
Designs and estimates regression models that represent and use unobserved (latent) variables: build a measurement component that maps observed indicators (continuous, ordinal, or categorical) onto latent traits and a structural component that regresses outcomes or predictors on those latent constructs. Produces estimated latent-variable scores and interpretable regression coefficients quantifying relationships between latent constructs and observed variables.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
This study addresses a critical yet previously unrecognized issue in observational causal inference: measurement-induced confounding, wherein latent variables—such as motivation or self-efficacy—are imperfectly measured, leading to biased estimates of adjusted causal effects. The authors formally identify and name this problem, moving beyond conventional two-stage adjustment approaches. They propose a novel Bayesian joint estimation framework that simultaneously models the latent variable’s measurement structure, the treatment assignment mechanism, and the potential outcomes model. This integrated approach effectively corrects bias in average treatment effect estimation and restores the nominal coverage of uncertainty intervals, thereby substantially enhancing the reliability of causal inferences drawn from observational data with error-prone proxies for unobserved confounders.
This paper addresses the challenges of covariate effect inference and future outcome prediction in high-dimensional multivariate longitudinal data—characterized by complex dependencies among variables and over time, mixed-type outcomes (continuous and discrete), and irregular observation patterns (missingness or right-censoring). We propose a novel latent-variable modeling framework that (i) unifies the treatment of mixed-type responses and irregular temporal observations for the first time; (ii) introduces an information criterion tailored to high-dimensional longitudinal settings for automatic selection of the latent factor dimension; and (iii) establishes a rigorous central limit theorem for regression coefficient estimators, ensuring valid statistical inference. Evaluated on a customer shopping behavior prediction task, our method significantly improves long-term trend modeling accuracy and robustness of personalized forecasting, demonstrating both practical utility and theoretical soundness in real-world high-dimensional longitudinal applications.
In social science research, measurement error in latent variables induces attenuation bias—regression coefficients biased toward zero—yet existing correction methods often ignore interactions between such errors and identification constraints on latent variables, sometimes exacerbating bias. This paper identifies the underlying mechanism and proposes a novel coefficient correction method that jointly models latent-variable identification constraints and measurement-error structure, enabling simultaneous adjustment of regression coefficients during estimation. The approach imposes no strong distributional or functional-form assumptions and is compatible with diverse latent-variable estimation strategies (e.g., CFA, SEM, Bayesian latent-variable modeling). Empirical evaluations demonstrate that corrected coefficients increase by 30–50% on average relative to naive OLS estimates, substantially outperforming both uncorrected regression and mainstream error-correction techniques (e.g., regression calibration, SIMEX). The method effectively recovers true effect magnitudes and enhances the validity of causal inference in latent-variable contexts.
This paper addresses the challenge of analytically computing the variance–covariance matrix of structural parameters in two-step estimation of latent variable models. We propose a general, asymptotically consistent simulation-based estimator. The method repeatedly draws simulated values from the sampling distribution of measurement parameters estimated in the first step and substitutes them into the second-step estimation to directly quantify the variability of structural parameters—thereby avoiding error-prone analytical differentiation of cross-derivative matrices required by conventional approaches. It is particularly well-suited for latent variable models with categorical observed indicators. Simulation studies and empirical analyses of two distinct model classes demonstrate that the proposed method exhibits excellent finite-sample statistical properties, computational efficiency, and robustness. Consequently, it substantially enhances the feasibility and reliability of statistical inference in two-step estimation frameworks.
This study investigates the mechanism through which academic performance—an ordinal variable—influences self-efficacy, a continuous outcome. To this end, the authors propose a conditional Bayesian modeling framework that introduces a latent academic achievement variable and integrates Gaussian copula regression with Bayesian variable selection to identify key covariates specific to each outcome type. Methodologically, they develop a tailored partially collapsed Gibbs sampler that substantially enhances computational efficiency in estimating integrated regression coefficients and improves the accuracy of variable selection. Simulation studies demonstrate that the proposed approach markedly outperforms existing joint modeling strategies in both sampling efficiency and variable selection performance. Application to data from the Longitudinal Study of Australian Children reveals distinct association pathways between academic achievement and self-efficacy, along with markedly different covariate structures for the two outcomes.
This study addresses the challenge of identifying direct causal effects among observed variables in densely confounded linear structural equation models with latent variables, where conventional methods often fail. The authors propose a novel identification criterion that explicitly models latent variables, employs a recursive identification strategy, and systematically handles unidentified causal parents. By transforming the combinatorial search problem into an efficient network flow computation, the method substantially enhances the identifiability of direct causal effects in dense confounding settings. Accompanied by an open-source algorithmic implementation, this approach combines theoretical rigor with practical utility for causal inference in complex observational data.
This study addresses identification bias in latent variable regression coefficients arising from multi-source nonlinear measurement error—exemplified by divergent measures of occupational exposure to artificial intelligence—by proposing a partial identification approach based on curvature constraints. Assuming a linear consensus measurement function and bounding heterogeneity in the curvature of individual measurement sources relative to the slope, the method constructs closed-form identification intervals that are invariant to unknown measurement loadings. These intervals exhibit sharpness, with half-widths that are second-order small relative to the curvature bounds. The curvature bounds are estimated via split-sample instrumental variable techniques, and inference with uniform coverage is achieved by combining Imbens–Manski confidence intervals with Stoye critical values. Applied to 8.88 million person-years of U.S. community survey data, the approach yields a consensus coefficient of −0.239 across five AI exposure measures, with a partial identification half-width amounting to only 1.23% of the point estimate.
This study addresses a critical limitation in conventional two-stage approaches that link individual-level distributional characteristics—such as variability and skewness—to downstream outcomes, which ignore estimation error in the first stage and consequently yield biased estimates and inflated Type I error rates. To overcome this, the authors propose the Distributional Feature Latent Variable Model (DFLVM), which, for the first time, integrates distributional features into a latent variable framework. DFLVM captures between-individual heterogeneity through random intercepts and jointly models both the distributional features and their effects on outcomes within a single-step maximum likelihood estimation procedure. This unified approach circumvents the inherent bias of two-stage methods. Simulation studies and empirical analyses demonstrate that DFLVM substantially reduces estimation bias and false positive rates while enhancing inferential accuracy.
本文提出一种设计辅助回归框架,通过利用协变量分布信息来稳定弱设计方向和修正潜在效应扭曲,从而改进估计性能。