Score
Designs and conducts statistical analyses to evaluate whether a measurement instrument (e.g., scale, test, or latent construct) functions equivalently across groups or occasions; this involves specifying and comparing nested measurement models (e.g., configural, metric, scalar, strict invariance) using multi‑group confirmatory factor analysis or item response theory and imposing equality constraints on loadings, thresholds/intercepts, residuals, or item parameters. Interprets fit indices and significance tests, detects noninvariant items (DIF), and recommends model modifications, item removal, or score adjustments so that comparisons of observed or latent scores across groups or time are valid.
This study addresses a critical limitation in cross-cultural research: single-item measures lack latent variable proxies and multiple indicators, rendering conventional tests for differential item functioning (DIF) or measurement invariance (MI) infeasible. To overcome this challenge, the authors propose a penalized heteroskedastic ordered probit model that integrates regularization techniques to resolve identification and estimation issues inherent in single-item data. This approach enables, for the first time, a viable framework for conducting DIF and MI analyses with single-item responses, thereby circumventing the traditional reliance on multi-item scales or multiple indicators. The method offers a novel analytical tool for cross-cultural comparisons in resource-constrained settings or contexts where only single-item measures are available, significantly expanding the scope of valid cross-cultural inference under practical constraints.
In large-scale international educational assessments (e.g., PISA), linguistic, cultural, and curricular differences frequently induce differential item functioning (DIF), severely biasing group ability distributions and rankings estimated via conventional item response theory (IRT). Existing DIF detection and calibration methods rely on unrealistic assumptions—such as the existence of a well-defined reference group or invariant anchor items—compromising statistical consistency. Method: We propose a reference-group- and anchor-item-free multi-group DIF-robust calibration framework, integrating high-dimensional statistical inference with nonconvex optimization to jointly estimate item and group parameters across populations. Contribution/Results: Our approach is the first to theoretically guarantee asymptotically consistent recovery of cross-group ability ordering under arbitrary DIF. Rigorous asymptotic analysis establishes its statistical validity. Empirical validation on PISA 2022 mathematics, science, and reading data demonstrates substantial correction of ability distribution bias and yields more reliable national rankings.
This study addresses the challenge of detecting differential item functioning (DIF) in ordinal-scale items when known group labels or anchor items are unavailable. The authors propose a hybrid latent class item response model that employs a proportional odds framework to model ordered responses, probabilistically assigning individuals to latent classes. Within this framework, uniform and nonuniform DIF are captured through class-specific intercept and slope deviations, respectively. The method requires no prespecified grouping variables or anchor items; instead, it leverages sparsity assumptions and L1 regularization to automatically identify DIF effects. A tailored EM algorithm is developed to optimize the L1-penalized marginal likelihood. Simulation results demonstrate accurate parameter recovery and effective DIF detection, while empirical analysis of a personality inventory reveals latent subgroups with heterogeneous response patterns and potentially biased items.
This paper addresses the distortion of effect size estimates in educational and psychological intervention research due to differential item functioning (DIF). Moving beyond conventional differential test functioning (DTF) analyses that rely on total-score differences, we propose a novel causal robustness framework grounded in item response theory (IRT). We formally define “impact” as between-group differences in the latent trait distribution and develop a Hausman-type test that integrates DIF modeling directly into causal effect identification—thereby disentangling true construct-level impact from item-specific bias. Methodologically, we introduce a DIF-robust doubly robust estimator and a testable framework for effect generalizability inference. Empirical validation across item-level data from 34 randomized trials shows that DIF correction substantially reduces discrepancies between effect estimates derived from researcher-developed versus independent measures, thereby enhancing construct validity and cross-measure comparability of effect interpretations.
To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.
This study addresses the challenge of jointly modeling latent group effects and differential item functioning (DIF) in measurement invariance assessment when group membership is unobserved and anchor items are unavailable. The authors propose a novel approach grounded in asymmetric item response theory (IRT), integrating a mixture IRT model with an ℓ₁-regularized estimator. By introducing latent classes to capture population heterogeneity and item-specific shifts to represent DIF effects, the method simultaneously identifies latent groups and DIF items without requiring known group labels or pre-specified anchor items. This work represents the first effort to achieve joint estimation of latent impact and DIF within an asymmetric IRT framework under completely unsupervised conditions, thereby overcoming limitations imposed by traditional symmetric link functions and reliance on anchor items. Simulation and empirical analyses demonstrate that the proposed method accurately recovers underlying parameter structures and successfully distinguishes between pure latent impact and pronounced DIF in educational assessments.
Existing effect size measures for differential item functioning (DIF) suffer from inconsistent classification schemes, systematic underestimation, and sensitivity to design factors, and lack a unified, cross-method standard for practical significance. This study systematically reviews current effect size indices and classification criteria, and evaluates their performance under Mantel-Haenszel, SIBTEST, and model-based approaches through large-scale simulation studies and real-data analyses. It innovatively introduces an improved form of area-based effect sizes and proposes unified cutoff values with clearly defined applicability boundaries, revising classification thresholds and usage guidelines accordingly. The resulting framework is implemented in R, substantially enhancing the consistency, accuracy, and practical interpretability of DIF effect sizes, thereby advancing the standardization of DIF analysis.
This study addresses the common practice in confirmatory factor analysis of accepting standardized factor loadings as low as 0.50, which leads to elevated measurement error, compromised construct validity, and unstable factor solutions. Building on the logic of average variance extracted (AVE) and communality, the authors propose and justify a uniform item-level threshold of λ ≥ 0.70, aligning it with construct-level validity requirements. Through theoretical derivation, Monte Carlo simulations, and structural equation modeling, the research systematically evaluates the impact of weak loadings on measurement quality, factor score determinacy, and model fit. Findings demonstrate that retaining indicators with λ < 0.70 significantly undermines model accuracy and robustness, whereas enforcing the λ ≥ 0.70 criterion enhances the explanatory power and overall quality of latent variable models.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.
This study addresses a critical yet previously unrecognized issue in observational causal inference: measurement-induced confounding, wherein latent variables—such as motivation or self-efficacy—are imperfectly measured, leading to biased estimates of adjusted causal effects. The authors formally identify and name this problem, moving beyond conventional two-stage adjustment approaches. They propose a novel Bayesian joint estimation framework that simultaneously models the latent variable’s measurement structure, the treatment assignment mechanism, and the potential outcomes model. This integrated approach effectively corrects bias in average treatment effect estimation and restores the nominal coverage of uncertainty intervals, thereby substantially enhancing the reliability of causal inferences drawn from observational data with error-prone proxies for unobserved confounders.