🤖 AI Summary
Existing effect size measures for differential item functioning (DIF) suffer from inconsistent classification schemes, systematic underestimation, and sensitivity to design factors, and lack a unified, cross-method standard for practical significance. This study systematically reviews current effect size indices and classification criteria, and evaluates their performance under Mantel-Haenszel, SIBTEST, and model-based approaches through large-scale simulation studies and real-data analyses. It innovatively introduces an improved form of area-based effect sizes and proposes unified cutoff values with clearly defined applicability boundaries, revising classification thresholds and usage guidelines accordingly. The resulting framework is implemented in R, substantially enhancing the consistency, accuracy, and practical interpretability of DIF effect sizes, thereby advancing the standardization of DIF analysis.
📝 Abstract
Differential Item Functioning (DIF) analysis is used to identify potentially biased items in multi-item measurements. In addition to testing the statistical significance, it is essential to evaluate the practical significance of DIF through effect-size measures. We review existing DIF effect-size measures and cut-off values used to classify the effect-size magnitudes for the Mantel-Haenszel test, SIBTEST, and model-based methods for binary items, and introduce a refinement of area-based effect-size measures. A simulation study is conducted to investigate the properties of these effect-size measures and existing classification guidelines, and to assess their comparative performance. The results indicate that some commonly used effect-size measures exhibit undesirable properties, including inconsistent classifications, systematic underestimation of the magnitude of the underlying DIF, and strong dependence on design factors. To address these issues, we introduce usage restrictions for some effect-size measures, revise cut-off values that unify results across different methods, and propose new cut-off values for area-based effect-size measures. The methods are demonstrated using two real data examples. Implementation is provided in the R software.