differential item functioning analysis

Designs and executes statistical analyses to detect test or questionnaire items that function differently across defined subgroups; builds and applies item-level models and tests to identify uniform and nonuniform differential item functioning, estimate effect sizes, and evaluate the impact of flagged items on scale validity and fairness.

differentialitemfunctioninganalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the distortion of effect size estimates in educational and psychological intervention research due to differential item functioning (DIF). Moving beyond conventional differential test functioning (DTF) analyses that rely on total-score differences, we propose a novel causal robustness framework grounded in item response theory (IRT). We formally define “impact” as between-group differences in the latent trait distribution and develop a Hausman-type test that integrates DIF modeling directly into causal effect identification—thereby disentangling true construct-level impact from item-specific bias. Methodologically, we introduce a DIF-robust doubly robust estimator and a testable framework for effect generalizability inference. Empirical validation across item-level data from 34 randomized trials shows that DIF correction substantially reduces discrepancies between effect estimates derived from researcher-developed versus independent measures, thereby enhancing construct validity and cross-measure comparability of effect interpretations.

Compares latent trait distribution differences between respondent groupsDevelops robust scaling method for consistent impact estimationProposes effect size for DIF's impact on group comparisons

QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation

Mar 07, 2025
BN
Bang Nguyen
🏛️ University of Notre Dame | University of Wisconsin-Madison | University of Illinois at Urbana-Champaign

Existing question generation (QG) evaluation methods lack alignment with psychometric metrics, failing to reflect true item quality across dimensions such as topic coverage, difficulty, discrimination, and distractor efficiency. Method: This paper introduces Classical Test Theory (CTT) into QG evaluation for the first time. We construct item pairs exhibiting significant quality differences and propose QG-SMS—a large language model (LLM)-based student modeling and simulation framework for interpretable, multi-dimensional automatic assessment. QG-SMS integrates LLM-driven student behavioral modeling, simulated response generation, CTT-based metric computation, and human validation. Contribution/Results: Experiments demonstrate that QG-SMS substantially improves the discriminative accuracy and robustness of QG systems in evaluating educational item quality. Its assessments strongly correlate with actual student performance and outperform conventional automated metrics.

Evaluating test item quality in educational assessmentsIdentifying shortcomings in existing QG evaluation methodsProposing QG-SMS for enhanced test item analysis

This work addresses the complex challenges of continuously monitoring item pool quality and health in large-scale AI-driven assessments. It proposes AQuAP, a dashboard system integrated with an item factory framework that leverages operational data analytics to support item generation and pool management. The system introduces novel metrics such as Effective Bank Size (EBS), which combines exposure rates and usage frequency to holistically evaluate the security, diversity, and efficiency of the item pool. By integrating psychometric indicators, exposure control algorithms, and advanced visualization techniques, AQuAP enables real-time monitoring of item pool vitality. The system has been successfully deployed in the Duolingo English Test, significantly enhancing the intelligence and responsiveness of item pool management.

AI-driven testingeducational assessmentitem bank health

This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.

analytical workflowcomputer-based assessmentsconsistency check

Enhancing Psychometric Analysis with Interactive ShinyItemAnalysis Modules

Jul 10, 2024
PM
Patrícia Martinková
🏛️ Institute of Computer Science of the Czech Academy of Sciences | Charles University

To address limitations in scalability of psychometric tools, high barriers to integrating novel methods, and insufficient support for reproducibility and pedagogy, this study introduces ShinyItemAnalysis (SIA), a modular extensibility framework built in R and Shiny. SIA pioneers a module mechanism supporting embedded datasets, object-oriented design, and compiled code—enabling researchers to develop, package, and share interactive psychometric methods (e.g., IRT, EFA, DIF detection) as standalone R packages while seamlessly leveraging SIA’s integrated data processing, visualization, and core analytical capabilities. The framework has been instantiated with multiple open-source example modules. These advances substantially enhance method accessibility, computational reproducibility, and instructional utility, thereby fostering an open, interactive psychometric tool ecosystem across psychology, education, and the social sciences.

Enhancing psychometric analysis with interactive SIA modulesExtending ShinyItemAnalysis for broader methodological applicationsFacilitating interactive psychometric software for research dissemination

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing LLM-as-a-Judge evaluation methods, which predominantly focus on output quality and lack a systematic framework for assessing the reliability of large language models (LLMs) as measurement instruments. To this end, the study introduces item response theory (IRT)—specifically the graded response model (GRM)—into this domain, proposing a two-stage diagnostic framework that evaluates LLM judges along two interpretable dimensions: internal consistency and human alignment. By integrating prompt perturbations with human rating data, the approach generates interpretable diagnostic signals that effectively identify unreliable LLM judgments. Empirical results demonstrate the method’s capacity to validate the reliability of LLM-as-a-Judge systems, offering both theoretical grounding and practical guidance for their trustworthy deployment in evaluation tasks.

automated evaluationItem Response TheoryLLM-as-a-Judge

Existing effect size measures for differential item functioning (DIF) suffer from inconsistent classification schemes, systematic underestimation, and sensitivity to design factors, and lack a unified, cross-method standard for practical significance. This study systematically reviews current effect size indices and classification criteria, and evaluates their performance under Mantel-Haenszel, SIBTEST, and model-based approaches through large-scale simulation studies and real-data analyses. It innovatively introduces an improved form of area-based effect sizes and proposes unified cutoff values with clearly defined applicability boundaries, revising classification thresholds and usage guidelines accordingly. The resulting framework is implemented in R, substantially enhancing the consistency, accuracy, and practical interpretability of DIF effect sizes, thereby advancing the standardization of DIF analysis.

classification guidelinesDifferential Item Functioningeffect-size measures

This study addresses the challenge of jointly modeling latent group effects and differential item functioning (DIF) in measurement invariance assessment when group membership is unobserved and anchor items are unavailable. The authors propose a novel approach grounded in asymmetric item response theory (IRT), integrating a mixture IRT model with an ℓ₁-regularized estimator. By introducing latent classes to capture population heterogeneity and item-specific shifts to represent DIF effects, the method simultaneously identifies latent groups and DIF items without requiring known group labels or pre-specified anchor items. This work represents the first effort to achieve joint estimation of latent impact and DIF within an asymmetric IRT framework under completely unsupervised conditions, thereby overcoming limitations imposed by traditional symmetric link functions and reliance on anchor items. Simulation and empirical analyses demonstrate that the proposed method accurately recovers underlying parameter structures and successfully distinguishes between pure latent impact and pronounced DIF in educational assessments.

asymmetric IRT modelsdifferential item functioninglatent impact

This study addresses the challenge of detecting differential item functioning (DIF) in ordinal-scale items when known group labels or anchor items are unavailable. The authors propose a hybrid latent class item response model that employs a proportional odds framework to model ordered responses, probabilistically assigning individuals to latent classes. Within this framework, uniform and nonuniform DIF are captured through class-specific intercept and slope deviations, respectively. The method requires no prespecified grouping variables or anchor items; instead, it leverages sparsity assumptions and L1 regularization to automatically identify DIF effects. A tailored EM algorithm is developed to optimize the L1-penalized marginal likelihood. Simulation results demonstrate accurate parameter recovery and effective DIF detection, while empirical analysis of a personality inventory reveals latent subgroups with heterogeneous response patterns and potentially biased items.

differential item functioninggroup comparisonslatent subgroups

This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.

adaptive assessmentexplanatory IRTitem calibration

Hot Scholars

SG

Shalini Ghosh

Tech Lead on AI Safety (Principal / Director level) in Google K&I
ML and Multimodal AI (LanguageSpeech/AudioComputer Vision)Physics
HS

Hong Shen

Assistant Professor, Carnegie Mellon University
human-computer interactionsocial computingcommunicationspublic policy
DD

Daniel Dan

Assistant Professor, Modul University, Vienna
Artificial IntelligenceApplied Data ScienceMarketingTourism
JP

Junyeong Park

M.S. Student at KAIST, School of Computing
NLPLLMs
AW

Anna Wróblewska

Warsaw University of Technology
machine learningnatural language processingimage processingmultimodal learning