item response theory calibration

Designs and fits probabilistic measurement models (Rasch and broader item response theory variants) that map observed item or response data to one or more latent proficiency traits, estimating item parameters (e.g., difficulty, discrimination, guessing) and person ability on a common scale while performing calibration, scaling/equating, and model-fit diagnostics. Produces calibrated item characteristic functions, information curves, and continuous latent proficiency scores and validated item parameters to be used as features or inputs for downstream models and adaptive assessment systems.

itemresponsetheorycalibration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Random pairing MLE for estimation of item parameters in Rasch model

Jun 20, 2024
YY
Yuepeng Yang
🏛️ University of Chicago

Estimating item parameters in the Rasch model under large-scale sparse observational data poses a fundamental challenge—balancing minimax optimality, uncertainty quantification, and computational feasibility. Method: We propose Randomized Pairwise Maximum Likelihood Estimation (RP-MLE) and its multi-replicate variant (MRP-MLE), which construct randomized pairwise comparisons among items to achieve dimensionality reduction while preserving sample independence. Contribution/Results: We establish that RP-MLE achieves finite-sample ℓ∞-norm minimax optimality and attains the information-theoretic lower bound asymptotically. Crucially, it provides the first rigorous asymptotic distributional characterization—and thus verifiable confidence intervals—for item parameter estimates under sparsity. Extensive simulations and real-data experiments demonstrate that RP-MLE and MRP-MLE significantly outperform classical estimators in both estimation accuracy and statistical inference, effectively breaking the traditional dependence of estimators on data density.

Achieving minimax optimal estimation error with uncertainty quantificationEstimating item parameters in Rasch model efficientlyHandling sparse observation data in psychometric assessments

This study addresses the challenge of calibrating explanatory item response theory (IRT) parameters under large-scale sparse data—a common scenario in adaptive testing where examinees respond to only a small fraction of items. To this end, the authors propose a Bayesian multidimensional explanatory IRT model, complemented by a tailored MCMC algorithm, a sparsity-aware data structure, and a high-performance computing engine. This integrated approach enables, for the first time, efficient and stable calibration of IRT parameters in ultra-large-scale sparse psychometric datasets. The resulting scalable SPICE calibration engine supports diverse psychometric applications and demonstrates strong performance and practical utility in real-world contexts such as computerized adaptive testing and automated item bank generation.

adaptive assessmentexplanatory IRTitem calibration

This study addresses the high cost and time demands of traditional item response theory (IRT) parameter calibration, which relies on real student response data. The authors propose a novel approach leveraging large language models: by fine-tuning the Qwen-3 dense model with LoRA, they generate synthetic responses conditioned on discrete ability descriptors to simulate students at varying proficiency levels. This enables reconstruction of item characteristic curves (ICCs) and estimation of IRT parameters without any real-world response data. To the best of the authors’ knowledge, this is the first method to implicitly model psychometric properties in a data-free setting. The approach demonstrates particularly strong performance in estimating item discrimination and achieves competitive or superior results compared to existing baselines on both sixth-grade English Language Arts items and the BEA 2024 dataset, confirming its effectiveness and practical utility.

field testingItem Characteristic Curvesitem parameter estimation

Two-step estimation of latent trait models

Mar 28, 2023
JK
J. Kuha
🏛️ London School of Economics and Political Science | Leiden University

To address the computational complexity and convergence difficulties inherent in joint estimation of measurement and structural models in item response theory (IRT), this paper proposes a two-step maximum likelihood estimation procedure: first, estimating measurement model parameters independently; second, estimating structural model parameters with measurement parameters held fixed. This work provides the first systematic theoretical justification—under settings involving continuous latent variables and categorical observed variables—of the statistical consistency, robustness, and computational efficiency of the two-step approach. Compared to conventional one-step estimation (prone to non-convergence) and three-step methods (susceptible to bias accumulation), the proposed method offers conceptual clarity, implementation simplicity, reliable standard errors, and stable convergence. Extensive simulation studies and empirical analyses validate its efficacy and generalizability across diverse latent variable models. The framework establishes a novel, general-purpose, flexible, and practical estimation paradigm for educational measurement, psychometrics, and related fields.

Evaluating performance compared to one-step and three-step methodsExamining properties through simulation studies and applicationsTwo-step estimation for latent trait models

Latest Papers

What's happening recently
View more

This study addresses the challenge of inaccurate item parameter estimation for new items in computerized adaptive testing, which arises from sparse response data and degrades scoring quality. To tackle this issue, the authors propose a content-driven compact calibration method that leverages a pretrained shallow ReLU network to map handcrafted item features into a low-dimensional embedding space. This embedding is then decoupled and integrated with a linear, interpretable three-parameter logistic (3PL) item response theory (IRT) model, eliminating the need for online parameter tuning during deployment. The approach achieves, for the first time, a disentangled joint optimization of neural embeddings and interpretable IRT parameters. Evaluated on two task types in the Duolingo English Test, the method demonstrates superior performance using only a six-dimensional embedding, outperforming larger models and confirming its efficiency and effectiveness.

cold-start problemcomputerized adaptive testingitem calibration

This study systematically investigates the prevalence and psychometric consequences of deviations from the normality assumption of latent trait distributions in item response theory (IRT). Analyzing 504 real-world datasets, the authors employed flexible nonparametric and semiparametric methods to estimate trait distributions and compared these against conventional normality assumptions, evaluating impacts on reliability, item parameters, predicted responses, and individual scores. The large-scale empirical analysis reveals—for the first time—that more than half of the datasets exhibit cumulative distribution discrepancies exceeding 10 percentage points (with roughly one-fifth surpassing 20 points), frequently manifesting as skewed, heavy-tailed, flat, or multimodal shapes, with substantial variation across domains. These deviations exert particularly pronounced effects when full models are refitted. The findings underscore the necessity of routinely reporting distributional sensitivity analyses in IRT applications.

distributional departureitem response theorylatent trait distribution

This study addresses the challenge of scale interference in operational latent variables during embedded pilot item calibration under few-shot adaptive routing. To mitigate this issue, the authors propose a Bayesian modular framework that severs feedback from pilot responses to operational latent variables, instead calibrating each pilot item independently via local logistic regression using fixed predictors derived from the posterior of operational traits. The approach innovatively incorporates an analytical debiasing mapping, enabling unbiased recovery of item parameters under a normal-ogive approximation. By integrating Firth’s penalized likelihood, multivariate delta-method covariance estimation, and Rubin’s pooling rules, the method achieves marginal likelihood calibration without per-item numerical integration while propagating parameter uncertainty into standard errors. Simulation results demonstrate substantially reduced attenuation bias and yield confidence interval coverage ranging from nominal to conservative under constrained conditions such as missing-at-random routing.

adaptive routingattenuation biaslocal item calibration

This study addresses the challenge of determining the number of latent dimensions in multidimensional graded response models by proposing an adaptive Bayesian framework for dimensionality selection. The approach introduces a cumulative ordered spike-and-slab (COSS) prior on the column variances of the item loading matrix, which automatically shrinks redundant dimensions while preserving meaningful structure. Coupled with Albert–Chib latent variable augmentation, it enables an efficient Gibbs sampler. A key innovation lies in integrating ordered shrinkage with Bayesian nonparametric principles, allowing adaptive inference of the latent dimensionality without pre-specifying a candidate set and naturally quantifying uncertainty in dimension selection. Simulation studies and analyses of real psychological assessment data demonstrate that the method accurately recovers true dimensional structures, yields more precise parameter estimates, and maintains computational efficiency.

dimension selectionlatent dimensionsmodel uncertainty

Traditional multidimensional item response theory is constrained by the assumption that latent traits follow a Gaussian distribution, which often fails to capture complex structures such as skewness, heavy tails, or multimodality, leading to biased parameter estimates. This work proposes the first integration of normalizing flows into this framework, leveraging invertible neural networks to model latent traits as flexible transformations of a simple base distribution. By combining conditional flows with variational inference, the approach jointly learns item parameters, the latent trait distribution, and its posterior. Simulation studies demonstrate that the method substantially improves the accuracy of both parameter and trait recovery under non-normal conditions. Furthermore, application to real-world personality data confirms its capacity to effectively model intricate latent distributions.

Latent Trait DistributionModel MisspecificationMultidimensional Item Response Theory

Hot Scholars

SK

Sanmi Koyejo

Assistant Professor, Stanford University
Machine LearningHealthcare AINeuroinformatics
JC

Jinsong Chen

Central China Normal University
Graph Representation LearningGraph Data MiningAI for Education
SZ

Susu Zhang

University of Illinois Urbana-Champaign
PsychologyStatistics
ZX

Ziang Xiao

Computer Science, Johns Hopkins University
AI4SocialScienceConversational AIHuman-centered EvaluationInformation Seeking
HJ

Han Jiang

Johns Hopkins University
Natural Language GenerationSocietal AIModel Evaluation