Score
Specifying, estimating, and interpreting latent‑variable measurement models (exploratory and confirmatory factor analysis) to determine whether observed responses reflect intended constructs, control identifiability, and identify influential components.
Traditional structural equation modeling (SEM) relies predominantly on reflective latent variable specifications, limiting its capacity to flexibly represent composite constructs—linear combinations of observed indicators. Existing compositional modeling approaches either compromise core SEM functionalities (e.g., overall model fit assessment, missing data handling, multi-group comparison) or inflate model complexity via auxiliary latent variables. Method: We propose the first SEM framework that unifies composites and latent variables within a single covariance structure model, leveraging maximum likelihood and generalized least squares estimation to directly specify the implied covariance matrix incorporating composites. Contribution/Results: Our approach eliminates the need for redundant latent variables while fully preserving SEM’s diagnostic and inferential capabilities—including fit evaluation, standard error estimation, and hypothesis testing. It significantly enhances expressive power and analytical flexibility for hybrid constructs (reflective + formative) and extends SEM’s applicability to more complex theoretical models.
This study addresses a critical yet previously unrecognized issue in observational causal inference: measurement-induced confounding, wherein latent variables—such as motivation or self-efficacy—are imperfectly measured, leading to biased estimates of adjusted causal effects. The authors formally identify and name this problem, moving beyond conventional two-stage adjustment approaches. They propose a novel Bayesian joint estimation framework that simultaneously models the latent variable’s measurement structure, the treatment assignment mechanism, and the potential outcomes model. This integrated approach effectively corrects bias in average treatment effect estimation and restores the nominal coverage of uncertainty intervals, thereby substantially enhancing the reliability of causal inferences drawn from observational data with error-prone proxies for unobserved confounders.
Existing multidimensional factor models struggle to accommodate the typical 3–5 dimensional latent constructs in psychometrics and lack a unified framework for modeling multiparameter moderation effects. This paper proposes a scalable penalized maximum likelihood estimation method applicable to arbitrarily many factors, enabling— for the first time—the joint estimation of linear and nonlinear moderation effects within high-dimensional models. By incorporating ridge, lasso, and alignment penalties, the approach simultaneously stabilizes parameter estimation, detects partial measurement noninvariance, and enhances interpretability. Leveraging closed-form analytical gradients, the method avoids computationally intensive numerical integration and MCMC sampling, substantially improving computational efficiency. Simulation and empirical studies demonstrate accurate recovery of complex moderation patterns. The proposed method provides a scalable, efficient, and robust new tool for measurement invariance research involving multidimensional constructs.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
Existing structural equation modeling (SEM) frameworks struggle to model latent variable variances that depend on other latent variables, thereby limiting the characterization of latent heteroscedasticity—such as in psychological constructs like personality or creativity. To address this, we propose Bayesian Gaussian Distributional SEM, the first SEM extension integrating distributional regression into the SEM framework to jointly model both the mean and variance of latent variables. Leveraging Bayesian inference and MCMC sampling, our approach flexibly specifies latent variances as arbitrary functions of other latent variables. Simulation studies demonstrate high statistical reliability and computational efficiency. Empirical analysis of personality data reveals that emotional stability significantly moderates the variability of neuroticism—a finding inaccessible under conventional SEM. This work introduces a novel theoretical tool and methodological paradigm for modeling latent heteroscedasticity, advancing both substantive theory testing and statistical methodology in behavioral and social sciences.
This paper addresses the identification challenge in multidimensional continuous measurement error models, where all observed variables are contaminated by latent, mutually dependent errors and no injective mapping is available. Methodologically, we construct a third-order cross-moment tensor and integrate Kruskal decomposition, the Kotlarski identity, and integral operator theory; we introduce the novel concept of “signal rank” and generalize Kruskal rank, achieving global identification of the joint distribution of latent variables and measurement errors for the first time under non-injective, multidimensional nonlinear settings. Our contribution lies in breaking free from conventional reliance on clean measurements or injectivity assumptions, providing empirically testable identification conditions, and substantially broadening the applicability of latent variable models. The proposed framework delivers a robust and general theoretical foundation for modeling noisy data in economics, psychology, and related disciplines.
Existing research frequently suffers from model misspecification of formative constructs, and the absence of a consensus-based validation methodology leads scholars to erroneously apply reflective measurement frameworks, thereby compromising construct validity. Method: This paper introduces the first dedicated, multi-stage validation framework for formative constructs, integrating systematic literature review, descriptive statistics, multicollinearity diagnostics, and formative-model-specific tests to rigorously distinguish formative (causal) from reflective (effect) measurement logic. Contribution/Results: The framework ensures both theoretical rigor and practical feasibility, substantially enhancing the psychometric soundness and statistical integrity of formative indicators. It provides a reproducible, defensible methodological pathway for scale development and construct validation, directly addressing longstanding measurement challenges in behavioral and social science research.
Latent class analysis (LCA) often suffers from limited interpretability due to the complexity of item response probability matrices. This work proposes a post-estimation sparsification method that enhances traditional LCA by incorporating item-wise pseudolikelihood with sparse regularization, which penalizes the number of non-zero response probability levels per item. This approach automatically merges redundant response levels, thereby improving the interpretability of latent classes. The method is computationally efficient and enjoys theoretical consistency guarantees for accurately recovering the true sparse structure. Combined with Bayesian Information Criterion (BIC) for selecting the number of latent classes, the proposed framework demonstrates strong performance in both simulation studies and an empirical analysis of social role performance survey data, yielding concise and interpretable latent class characterizations. The implementation code is publicly available.
This study addresses the challenges in exploratory factor analysis arising from unknown factor structures and indeterminate numbers of latent factors, which often hinder model identification and evaluation. The authors propose a variational Bayesian variable selection framework that employs spike-and-slab priors to recover the underlying factor structure and introduces a post-selection model fit assessment system. By recasting hard and soft selection strategies as covariance models, the approach facilitates diagnostic evaluation and determination of the number of factors. A novel dimensionless gain rule, combined with multidimensional fit indices—including RMSEA, SRMR, CFI, TLI, AIC, BIC, and ELBO—is introduced to effectively prevent misidentification of factor count. Simulations demonstrate that absolute fit indices sensitively track loading recovery and detect underfactoring, while the gain rule accurately recovers the true dimensionality, with the ELBO variant exhibiting the greatest robustness. Applied to the 100-item PID-5 dataset, the method significantly outperforms a prespecified 25-factor confirmatory model.
This study addresses the challenge of regression with high-dimensional responses and covariates when the responses are influenced by both observed covariates and unobserved latent variables, a setting where conventional multivariate regression methods fail to provide effective modeling. The authors propose a generalized latent variable model that accommodates mixed-type high-dimensional responses and allows for flexible dependence structures between covariates and latent factors. By decomposing the non-convex estimation problem into a sequence of convex subproblems through alternating optimization, and integrating debiased estimation with asymptotic normality analysis, the work achieves, for the first time, valid statistical inference on covariate effects within a high-dimensional generalized latent variable framework. The proposed estimator enjoys statistical consistency and guaranteed error bounds, while the debiased version exhibits asymptotic normality, as demonstrated empirically in an application to PISA data for assessing educational equity.