item response model application

Designs, estimates, and applies statistical models that link individuals' latent traits to their item-level responses on assessments or questionnaires; this includes specifying IRT model families (e.g., dichotomous and polytomous models), calibrating item parameters, and producing ability/trait scores. Builds and evaluates scoring procedures, information and reliability functions, model-fit diagnostics, and tests for item functioning (e.g., differential item functioning) to support measurement decisions.

itemresponsemodelapplication

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.

adaptive assessmentexplanatory IRTitem calibration

Two-step estimation of latent trait models

Mar 28, 2023
JK
J. Kuha
🏛️ London School of Economics and Political Science | Leiden University

To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.

Evaluating performance compared to one-step and three-step methodsExamining properties through simulation studies and applicationsTwo-step estimation for latent trait models

This study addresses the challenge of jointly modeling latent group effects and differential item functioning (DIF) in measurement invariance assessment when group membership is unobserved and anchor items are unavailable. The authors propose a novel approach grounded in asymmetric item response theory (IRT), integrating a mixture IRT model with an ℓ₁-regularized estimator. By introducing latent classes to capture population heterogeneity and item-specific shifts to represent DIF effects, the method simultaneously identifies latent groups and DIF items without requiring known group labels or pre-specified anchor items. This work represents the first effort to achieve joint estimation of latent impact and DIF within an asymmetric IRT framework under completely unsupervised conditions, thereby overcoming limitations imposed by traditional symmetric link functions and reliance on anchor items. Simulation and empirical analyses demonstrate that the proposed method accurately recovers underlying parameter structures and successfully distinguishes between pure latent impact and pronounced DIF in educational assessments.

asymmetric IRT modelsdifferential item functioninglatent impact

A longstanding issue in IRT simulation—“reliability omission”—treats reliability as an implicit byproduct rather than an explicit, controllable design parameter, resulting in ambiguous signal-to-noise ratios. This paper formally defines the IRT inverse-design problem and introduces the first simulation framework enabling precise, user-specified control of marginal reliability—elevating it from an output metric to an explicit input parameter. We innovatively distinguish and calibrate two reliability types—equivalent-class (EQC) and stochastic-approximation (SAC)—yielding two deterministic and stochastic algorithms, respectively: EQC achieves near-exact calibration, while SAC ensures unbiased estimation under non-normal latent traits and realistic item pools. Leveraging Jensen’s inequality for theoretical analysis, we validate the framework across 960 experimental conditions. We publicly release the R package *IRTsimrel*, enabling standardized, reliability-aware IRT simulation.

Addresses the reliability omission in IRT simulations by making marginal reliability an explicit design factor.Provides algorithms and tools to standardize reliability as a controlled input in simulation studies.Solves the inverse design problem to achieve pre-specified target reliability through global discrimination scaling.

This paper addresses the distortion of effect size estimates in educational and psychological intervention research due to differential item functioning (DIF). Moving beyond conventional differential test functioning (DTF) analyses that rely on total-score differences, we propose a novel causal robustness framework grounded in item response theory (IRT). We formally define “impact” as between-group differences in the latent trait distribution and develop a Hausman-type test that integrates DIF modeling directly into causal effect identification—thereby disentangling true construct-level impact from item-specific bias. Methodologically, we introduce a DIF-robust doubly robust estimator and a testable framework for effect generalizability inference. Empirical validation across item-level data from 34 randomized trials shows that DIF correction substantially reduces discrepancies between effect estimates derived from researcher-developed versus independent measures, thereby enhancing construct validity and cross-measure comparability of effect interpretations.

Compares latent trait distribution differences between respondent groupsDevelops robust scaling method for consistent impact estimationProposes effect size for DIF's impact on group comparisons

Latest Papers

What's happening recently
View more

Traditional multidimensional item response theory is constrained by the assumption that latent traits follow a Gaussian distribution, which often fails to capture complex structures such as skewness, heavy tails, or multimodality, leading to biased parameter estimates. This work proposes the first integration of normalizing flows into this framework, leveraging invertible neural networks to model latent traits as flexible transformations of a simple base distribution. By combining conditional flows with variational inference, the approach jointly learns item parameters, the latent trait distribution, and its posterior. Simulation studies demonstrate that the method substantially improves the accuracy of both parameter and trait recovery under non-normal conditions. Furthermore, application to real-world personality data confirms its capacity to effectively model intricate latent distributions.

Latent Trait DistributionModel MisspecificationMultidimensional Item Response Theory

This study systematically investigates the prevalence and psychometric consequences of deviations from the normality assumption of latent trait distributions in item response theory (IRT). Analyzing 504 real-world datasets, the authors employed flexible nonparametric and semiparametric methods to estimate trait distributions and compared these against conventional normality assumptions, evaluating impacts on reliability, item parameters, predicted responses, and individual scores. The large-scale empirical analysis reveals—for the first time—that more than half of the datasets exhibit cumulative distribution discrepancies exceeding 10 percentage points (with roughly one-fifth surpassing 20 points), frequently manifesting as skewed, heavy-tailed, flat, or multimodal shapes, with substantial variation across domains. These deviations exert particularly pronounced effects when full models are refitted. The findings underscore the necessity of routinely reporting distributional sensitivity analyses in IRT applications.

distributional departureitem response theorylatent trait distribution

This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.

adaptive routingattenuation biaslocal item calibration

This study addresses the challenge that traditional models struggle to capture intra-individual variability in response times within computerized testing, where dynamic shifts in behavior may occur. To this end, we propose a novel latent variable model that incorporates individual-specific change points into log-response time modeling for the first time. Behavioral abrupt shifts are characterized through item-specific mean structure offsets, and the change point is treated as a discrete latent variable whose distribution is linked to an underlying speed factor. This unified framework enables simultaneous modeling of within-person dynamics and supports statistical inference and uncertainty quantification for both change point locations and effect parameters. Using marginal maximum likelihood estimation—integrating latent variable modeling, change point detection, and Bayesian posterior inference—simulation studies demonstrate that the model robustly and accurately recovers parameters and change point positions across varying sample sizes and test lengths.

change-pointcomputerized assessmentslatent variable model

Hot Scholars

SK

Sanmi Koyejo

Assistant Professor, Stanford University
Machine LearningHealthcare AINeuroinformatics
RR

Rachel Rudinger

Assistant Professor, Department of Computer Science, University of Maryland
YZ

Yiyun Zhou

Zhejiang University
Data MiningMultimodal LearningLarge Language Model
DC

Donato Crisostomi

Ellis Ph.D. student, Sapienza University of Rome & University of Cambridge
Deep LearningModel MergingRepresentation Alignment
ZH

Zhenya Huang

University of Science and Technology of China
Data ScienceAIKnowledge RepresentationCognitive Reasoning