Score
Designs, estimates, and evaluates structural equation models that represent causal paths among observed and latent variables and implements confirmatory factor analysis to impose and test explicit measurement models. Specifies and fits models to estimate direct and indirect effects, test mediation and moderation, assess model fit and validity, and resolve identification and parameter‑estimation issues.
This study addresses construct type misspecification in structural equation modeling (SEM): specifically, the consequences of mismatching the assumed construct type—latent variable, causal-formative construct, or composite indicator—with its true underlying nature. Using large-scale Monte Carlo simulations, we systematically disentangle the independent effects of construct misspecification from those of estimation methods (e.g., maximum likelihood, PLS), examining bias patterns across six realistic–hypothetical construct-type pairings. Results demonstrate that construct misspecification alone induces substantial and systematic bias in path coefficient estimates. Moreover, conventional fit indices—including χ², CFI, and RMSEA—fail to reliably distinguish the correct construct type. The findings underscore that theoretically grounded construct definition must take precedence over statistical fit criteria. This work provides critical methodological guidance for SEM practice, cautioning against reliance on global fit metrics to justify construct typology and highlighting the necessity of substantive theory in model specification.
Traditional structural equation modeling (SEM) relies predominantly on reflective latent variable specifications, limiting its capacity to flexibly represent composite constructs—linear combinations of observed indicators. Existing compositional modeling approaches either compromise core SEM functionalities (e.g., overall model fit assessment, missing data handling, multi-group comparison) or inflate model complexity via auxiliary latent variables. Method: We propose the first SEM framework that unifies composites and latent variables within a single covariance structure model, leveraging maximum likelihood and generalized least squares estimation to directly specify the implied covariance matrix incorporating composites. Contribution/Results: Our approach eliminates the need for redundant latent variables while fully preserving SEM’s diagnostic and inferential capabilities—including fit evaluation, standard error estimation, and hypothesis testing. It significantly enhances expressive power and analytical flexibility for hybrid constructs (reflective + formative) and extends SEM’s applicability to more complex theoretical models.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
Existing structural equation modeling (SEM) approaches face multiple limitations in modeling composite variables—such as indices, formative constructs, and bundle variables—including inability to represent their construction process, fixed (non-estimated) weights, difficulty in specifying them as endogenous, and lack of capacity to test full mediation or component-variable influence. This paper introduces two novel methods grounded in the H-O specification, integrating phantom variables with pseudo-indicator techniques. For the first time, these methods treat inverse or full weights as freely estimated parameters, enabling flexible inclusion of composite variables at any model position—including endogenous—and rigorous testing of effect transmission. By unifying linear combination weight estimation with the SEM framework, the methods successfully estimate weights, model composites endogenously, and distinguish mediation from formative mechanisms in empirical data. This significantly enhances modeling accuracy, interpretability, and guidance for model selection in multivariate behavioral research.
Existing structural equation modeling (SEM) frameworks struggle to model latent variable variances that depend on other latent variables, thereby limiting the characterization of latent heteroscedasticity—such as in psychological constructs like personality or creativity. To address this, we propose Bayesian Gaussian Distributional SEM, the first SEM extension integrating distributional regression into the SEM framework to jointly model both the mean and variance of latent variables. Leveraging Bayesian inference and MCMC sampling, our approach flexibly specifies latent variances as arbitrary functions of other latent variables. Simulation studies demonstrate high statistical reliability and computational efficiency. Empirical analysis of personality data reveals that emotional stability significantly moderates the variability of neuroticism—a finding inaccessible under conventional SEM. This work introduces a novel theoretical tool and methodological paradigm for modeling latent heteroscedasticity, advancing both substantive theory testing and statistical methodology in behavioral and social sciences.
This study addresses the challenge of fitting covariance-based structural equation models when the number of variables exceeds the sample size ($p > n$), a setting in which traditional factor-based approaches fail due to singularity of the sample covariance matrix. The authors propose a novel method that decomposes the covariance structure into auto-covariance and cross-covariance components, integrating a likelihood-based feasible set with relative error constraints to achieve stable estimation in small-sample regimes. This approach enables, for the first time, covariance-based structural equation modeling in $p > n$ scenarios, substantially improving parameter estimation stability and accurately recovering the signs and directions of structural parameters. Empirical evaluations on both synthetic and real-world datasets demonstrate its superior performance, highlighting its practical utility for decision-making applications.
This study addresses the challenge of identifying direct causal effects among observed variables in densely confounded linear structural equation models with latent variables, where conventional methods often fail. The authors propose a novel identification criterion that explicitly models latent variables, employs a recursive identification strategy, and systematically handles unidentified causal parents. By transforming the combinatorial search problem into an efficient network flow computation, the method substantially enhances the identifiability of direct causal effects in dense confounding settings. Accompanied by an open-source algorithmic implementation, this approach combines theoretical rigor with practical utility for causal inference in complex observational data.
This study addresses a critical yet previously unrecognized issue in observational causal inference: measurement-induced confounding, wherein latent variables—such as motivation or self-efficacy—are imperfectly measured, leading to biased estimates of adjusted causal effects. The authors formally identify and name this problem, moving beyond conventional two-stage adjustment approaches. They propose a novel Bayesian joint estimation framework that simultaneously models the latent variable’s measurement structure, the treatment assignment mechanism, and the potential outcomes model. This integrated approach effectively corrects bias in average treatment effect estimation and restores the nominal coverage of uncertainty intervals, thereby substantially enhancing the reliability of causal inferences drawn from observational data with error-prone proxies for unobserved confounders.
This study addresses the limitation of existing information criteria in structural equation modeling, which fail to explicitly leverage latent variable structures, thereby constraining model selection performance. To overcome this, the authors propose two novel information criteria based on the complete-data likelihood, uniquely integrating complete-data likelihood with importance sampling to explicitly incorporate latent variable structure into criterion design. Within the Gaussian structural equation modeling framework, the approach accurately estimates latent variables, thereby effectively recovering dependencies among observed variables. Empirical evaluations demonstrate that the proposed criteria exhibit robust performance across diverse latent structures and sample conditions, significantly outperforming existing methods—particularly when latent variable estimation is accurate.
This study addresses the problem of testing whether a treatment effect operates entirely through observed mediators and identifying causal mechanisms under control for covariates. The authors propose a statistical test based on double machine learning, extending— for the first time—the joint evaluation of full mediation and causal mechanism identification to non-randomized treatment settings. By integrating conditional independence testing, the method achieves root-n consistent and asymptotically normal inference even in the presence of high-dimensional covariates. Simulation studies demonstrate favorable finite-sample performance, and the approach is successfully applied to two randomized experiments examining maternal mental health and social norms.